Tomaso A. Poggio

dblp:12/5544 · DBLP profile ↗
← Back
131ranked-venue papers
11as first author
7since 2021 · last 2025
0000-0002-3944-0455ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 104 · 7 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 49 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 1 first-authorDatabases, data management, data science and information retrieval · 6 · 1 first-authorTheory of computation · 2 · 1 first-authorSystems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
68 papers
Deep learning architectures and training · 23% Representation and self-supervised learning · 17% Learning theory · 12%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%

Topics — the 30 heaviest of 139, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
neural collapse
1.522025
Formation of Representations in Neural Networks · ICLR 2025
Feature learning in deep classifiers through Intermediate Neural Collapse · ICML 2023
Machine learning › Deep learning architectures and training
convolutional neural network
1.442023
Norm-based Generalization Bounds for Sparse Neural Networks · NeurIPS 2023
Biologically Inspired Mechanisms for Adversarial Robustness · NeurIPS 2020
Do Deep Neural Networks Suffer from Crowding? · NIPS 2017
Machine learning › Representation and self-supervised learning › representation matching
feature alignment
0.912025
Training the Untrainable: Introducing Inductive Bias via Representational Alignment · NeurIPS 2025
Machine learning › Learning theory
inductive bias
0.912025
Training the Untrainable: Introducing Inductive Bias via Representational Alignment · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation
knowledge transfer
0.912025
Training the Untrainable: Introducing Inductive Bias via Representational Alignment · NeurIPS 2025
Computer vision › Image recognition and object detection
object recognition
0.892025
Do Deep Neural Networks Suffer from Crowding? · NIPS 2017
Training the Untrainable: Introducing Inductive Bias via Representational Alignment · NeurIPS 2025
Why The Brain Separates Face Recognition From Object Recognition · NIPS 2011
Natural language and speech › Language models and text generation › neural language model
autoregressive language model
0.812024
On the Power of Decision Trees in Auto-Regressive Language Modeling · NeurIPS 2024
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.812024
On the Power of Decision Trees in Auto-Regressive Language Modeling · NeurIPS 2024
Machine learning › Learning theory
generalization bounds
0.722023
Norm-based Generalization Bounds for Sparse Neural Networks · NeurIPS 2023
Bounds on the Generalization Performance of Kernel Machine Ensembles · ICML 2000
Bioinformatics and computational biology
computational neuroscience
0.722023
System Identification of Neural Systems: If We Got It Right, Would We Know? · ICML 2023
Just One View: Invariances in Inferotemporal Cell Tuning · NIPS 1997
Machine learning › Representation and self-supervised learning › representation learning
representation geometry
0.712023
Feature learning in deep classifiers through Intermediate Neural Collapse · ICML 2023
Machine learning › Representation and self-supervised learning › representation analysis
representation similarity
0.712023
System Identification of Neural Systems: If We Got It Right, Would We Know? · ICML 2023
Machine learning › Efficient and distributed learning › model compression
sparse neural network
0.712023
Norm-based Generalization Bounds for Sparse Neural Networks · NeurIPS 2023
Bioinformatics and computational biology › computational neuroscience › neural response modeling
neural system identification
0.712023
System Identification of Neural Systems: If We Got It Right, Would We Know? · ICML 2023
Machine learning › Deep learning architectures and training
biologically plausible learning
0.622019
Biologically-Plausible Learning Algorithms Can Scale to Large Datasets · ICLR (Poster) 2019
How Important Is Weight Symmetry in Backpropagation? · AAAI 2016
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.412020
Biologically Inspired Mechanisms for Adversarial Robustness · NeurIPS 2020
Computer vision › Image recognition and object detection
biologically inspired vision
0.412020
Biologically Inspired Mechanisms for Adversarial Robustness · NeurIPS 2020
Machine learning › Trustworthy machine learning
robustness
0.412020
Biologically Inspired Mechanisms for Adversarial Robustness · NeurIPS 2020
Machine learning › Probabilistic and Bayesian machine learning
probabilistic inference
0.412019
Fast and Flexible Inference of Joint Distributions from their Marginals · ICML 2019
Computer vision › Face, body and person analysis
face recognition
0.392011
Why The Brain Separates Face Recognition From Object Recognition · NIPS 2011
A Component-based Framework for Face Detection and Identification · Int. J. Comput. Vis. 2007
Categorization by Learning and Combining Object Parts · NIPS 2001
Machine learning › Learning theory
approximation theory
0.332017
When and Why Are Deep Networks Better Than Shallow Ones? · AAAI 2017
Extensions of a Theory of Networks for Approximation and Learning · NIPS 1990
Networks for approximation and learning · Proc. IEEE 1990
Machine learning › Graph learning › graph neural network › expressive power
depth separation
0.312017
When and Why Are Deep Networks Better Than Shallow Ones? · AAAI 2017
Computer vision › Video understanding and tracking
action recognition
0.332013
Neural representation of action sequences: how far can a simple snippet-matching model take us? · NIPS 2013
A Biologically Inspired System for Action Recognition · ICCV 2007
HMDB: A large video database for human motion recognition · ICCV 2011
Machine learning › Deep learning architectures and training
training dynamics
0.312025
Formation of Representations in Neural Networks · ICLR 2025
Machine learning › Deep learning architectures and training
backpropagation
0.212016
How Important Is Weight Symmetry in Backpropagation? · AAAI 2016
Machine learning › Representation and self-supervised learning › representation learning › embedding learning › semantic embedding
compositional embedding
0.212016
Holographic Embeddings of Knowledge Graphs · AAAI 2016
Machine learning › Representation and self-supervised learning › representation learning › structured representation learning
relational representation learning
0.212016
Holographic Embeddings of Knowledge Graphs · AAAI 2016
Machine learning › Deep learning architectures and training › training dynamics
weight symmetry
0.212016
How Important Is Weight Symmetry in Backpropagation? · AAAI 2016
Knowledge graphs
knowledge graph embedding
0.212016
Holographic Embeddings of Knowledge Graphs · AAAI 2016
Knowledge graphs
link prediction
0.212016
Holographic Embeddings of Knowledge Graphs · AAAI 2016

Methods — techniques the papers use, named apart from their topics

linear encoding model · 1.3centered kernel alignment · 1.3power-law analysis · 0.9neural distance function · 0.9layerwise representational similarity · 0.9knowledge distillation · 0.9alignment relations · 0.9transformer representation · 0.8autoregressive decision trees · 0.8multi-scale receptive field · 0.7holographic embedding · 0.2circular correlation · 0.2snippet-matching model · 0.2linear weighted sum · 0.2relaxation · 0.1reconstruction bounds · 0.1k-flats approximation · 0.1convex analysis · 0.1
YearPublicationVenuePosition
2025 On Generalization Bounds for Neural Networks with Low Rank Layers
abstract
While previous optimization results have suggested that deep neural networks tend to favour low-rank weight matrices, the implications of this inductive bias on generalization bounds remain underexplored. In this paper, we apply a chain rule for Gaussian complexity (Maurer, 2016a) to analyze how low-rank layers in deep networks can prevent the accumulation of rank and dimensionality factors that typically multiply across layers. This approach yields generalization bounds for rank and spectral norm constrained networks. We compare our results to prior generalization bounds for deep networks, highlighting how deep networks with low-rank layers can achieve better generalization than those with full-rank layers. Additionally, we discuss how this framework provides new perspectives on the generalization capabilities of deep networks exhibiting neural collapse.
Andrea Pinto, Akshay Rangamani, Tomaso A. Poggio
ALT3
2025 Formation of Representations in Neural Networks
abstract
Understanding neural representations will help open the black box of neural networks and advance our scientific understanding of modern AI systems. However, how complex, structured, and transferable representations emerge in modern neural networks has remained a mystery. Building on previous results, we propose the Canonical Representation Hypothesis (CRH), which posits a set of six alignment relations to universally govern the formation of representations in most hidden layers of a neural network. Under the CRH, the latent representations (R), weights (W), and neuron gradients (G) become mutually aligned during training. This alignment implies that neural networks naturally learn compact representations, where neurons and weights are invariant to task-irrelevant transformations. We then show that the breaking of CRH leads to the emergence of reciprocal power-law relations between R, W, and G, which we refer to as the Polynomial Alignment Hypothesis (PAH). We present a minimal-assumption theory proving that the balance between gradient noise and regularization is crucial for the emergence of the canonical representation. The CRH and PAH lead to an exciting possibility of unifying major key deep learning phenomena, including neural collapse and the neural feature ansatz, in a single framework.
Liu Ziyin 0001, Isaac L. Chuang, Tomer Galanti, Tomaso A. Poggio
ICLR4
2025 Training the Untrainable: Introducing Inductive Bias via Representational Alignment
abstract
We demonstrate that architectures which traditionally are considered to be ill-suited for a task can be trained using inductive biases from another architecture. We call a network untrainable when it overfits, underfits, or converges to poor results even when tuning their hyperparameters. For example, fully connected networks overfit on object recognition while deep convolutional networks without residual connections underfit. The traditional answer is to change the architecture to impose some inductive bias, although the nature of that bias is unknown. We introduce guidance, where a guide network steers a target network using a neural distance function. The target minimizes its task loss plus a layerwise representational similarity against the frozen guide. If the guide is trained, this transfers over the architectural prior and knowledge of the guide to the target. If the guide is untrained, this transfers over only part of the architectural prior of the guide. We show that guidance prevents FCN overfitting on ImageNet, narrows the vanilla RNN–Transformer gap, boosts plain CNNs toward ResNet accuracy, and aids Transformers on RNN-favored tasks. We further identify that guidance-driven initialization alone can mitigate FCN overfitting. Our method provides a mathematical tool to investigate priors and architectures, and in the long term, could automate architecture design.
Vighnesh Subramaniam, David Mayo, Colin Conwell, Tomaso A. Poggio, Boris Katz, Brian Cheung, Andrei Barbu
NeurIPS4
2024 On the Power of Decision Trees in Auto-Regressive Language Modeling
abstract
Originally proposed for handling time series data, Auto-regressive Decision Trees (ARDTs) have not yet been explored for language modeling. This paper delves into both the theoretical and practical applications of ARDTs in this new context. We theoretically demonstrate that ARDTs can compute complex functions, such as simulating automata, Turing machines, and sparse circuits, by leveraging "chain-of-thought" computations. Our analysis provides bounds on the size, depth, and computational efficiency of ARDTs, highlighting their surprising computational power. Empirically, we train ARDTs on simple language generation tasks, showing that they can learn to generate coherent and grammatically correct text on par with a smaller Transformer model. Additionally, we show that ARDTs can be used on top of transformer representations to solve complex reasoning tasks. This research reveals the unique computational abilities of ARDTs, aiming to broaden the architectural diversity in language model development.
Yulu Gan, Tomer Galanti, Tomaso A. Poggio, Eran Malach
NeurIPS3
2023 System Identification of Neural Systems: If We Got It Right, Would We Know?
abstract
Artificial neural networks are being proposed as models of parts of the brain. The networks are compared to recordings of biological neurons, and good performance in reproducing neural responses is considered to support the model’s validity. A key question is how much this system identification approach tells us about brain computation. Does it validate one model architecture over another? We evaluate the most commonly used comparison techniques, such as a linear encoding model and centered kernel alignment, to correctly identify a model by replacing brain recordings with known ground truth models. System identification performance is quite variable; it also depends significantly on factors independent of the ground truth architecture, such as stimuli images. In addition, we show the limitations of using functional similarity scores in identifying higher-level architectural motifs.
Yena Han, Tomaso A. Poggio, Brian Cheung
ICML2
2023 Feature learning in deep classifiers through Intermediate Neural Collapse
abstract
In this paper, we conduct an empirical study of the feature learning process in deep classifiers. Recent research has identified a training phenomenon called Neural Collapse (NC), in which the top-layer feature embeddings of samples from the same class tend to concentrate around their means, and the top layer's weights align with those features. Our study aims to investigate if these properties extend to intermediate layers. We empirically study the evolution of the covariance and mean of representations across different layers and show that as we move deeper into a trained neural network, the within-class covariance decreases relative to the between-class covariance. Additionally, we find that in the top layers, where the between-class covariance is dominant, the subspace spanned by the class means aligns with the subspace spanned by the most significant singular vector components of the weight matrix in the corresponding layer. Finally, we discuss the relationship between NC and Associative Memories (Willshaw et. al. 1969).
Akshay Rangamani, Marius Lindegaard, Tomer Galanti, Tomaso A. Poggio
ICML4
2023 Norm-based Generalization Bounds for Sparse Neural Networks
abstract
In this paper, we derive norm-based generalization bounds for sparse ReLU neural networks, including convolutional neural networks. These bounds differ from previous ones because they consider the sparse structure of the neural network architecture and the norms of the convolutional filters, rather than the norms of the (Toeplitz) matrices associated with the convolutional layers. Theoretically, we demonstrate that these bounds are significantly tighter than standard norm-based generalization bounds. Empirically, they offer relatively tight estimations of generalization for various simple classification problems. Collectively, these findings suggest that the sparsity of the underlying target function and the model's architecture plays a crucial role in the success of deep learning.
Tomer Galanti, Mengjia Xu, Liane Galanti, Tomaso A. Poggio
NeurIPS4
2020 Approximate Inference with Wasserstein Gradient Flows
abstract
We present a novel approximate inference method for diffusion processes, based on the Wasserstein gradient flow formulation of the diffusion. In this formulation, the time-dependent density of the diffusion is derived as the limit of implicit Euler steps that follow the gradients of a particular free energy functional. Existing methods for computing Wasserstein gradient flows rely on discretization of the domain of the diffusion, prohibiting their application to domains in more than several dimensions. We propose instead a discretization-free inference method that computes the Wasserstein gradient flow directly in a space of continuous functions. We characterize approximation properties of the proposed method and evaluate it on a nonlinear filtering task, finding performance comparable to the state-of-the-art for filtering diffusions.
Charlie Frogner, Tomaso A. Poggio
AISTATS2
2020 Cross-Domain Adversarial Reprogramming of a Recurrent Neural Network
Alexandra Maria Proca, Andrzej Banburski-Fahey, Tomaso A. Poggio
CogSci3
2020 Biologically Inspired Mechanisms for Adversarial Robustness
abstract
A convolutional neural network strongly robust to adversarial perturbations at reasonable computational and performance cost has not yet been demonstrated. The primate visual ventral stream seems to be robust to small perturbations in visual stimuli but the underlying mechanisms that give rise to this robust perception are not understood. In this work, we investigate the role of two biologically plausible mechanisms in adversarial robustness. We demonstrate that the non-uniform sampling performed by the primate retina and the presence of multiple receptive fields with a range of receptive field sizes at each eccentricity improve the robustness of neural networks to small adversarial perturbations. We verify that these two mechanisms do not suffer from gradient obfuscation and study their contribution to adversarial robustness through ablation studies.
Manish Reddy Vuyyuru, Andrzej Banburski-Fahey, Nishka Pant, Tomaso A. Poggio
NeurIPS4
2020 An analysis of training and generalization errors in shallow and deep networks
Hrushikesh N. Mhaskar, Tomaso A. Poggio
Neural Networks2
2019 Fisher-Rao Metric, Geometry, and Complexity of Neural Networks
abstract
We study the relationship between geometry and capacity measures for deep neural networks from an invariance viewpoint. We introduce a new notion of capacity — the Fisher-Rao norm — that possesses desirable invariance properties and is motivated by Information Geometry. We discover an analytical characterization of the new capacity measure, through which we establish norm-comparison inequalities and further show that the new measure serves as an umbrella for several existing norm-based complexity measures. We discuss upper bounds on the generalization error induced by the proposed measure. Extensive numerical experiments on CIFAR-10 support our theoretical findings. Our theoretical analysis rests on a key structural lemma about partial derivatives of multi-layer rectifier networks.
Tengyuan Liang, Tomaso A. Poggio, Alexander Rakhlin, James Stokes
AISTATS2
2019 Biologically-Plausible Learning Algorithms Can Scale to Large Datasets
Will Xiao, Qianli Liao, Tomaso A. Poggio
ICLR (Poster)4
2019 Fast and Flexible Inference of Joint Distributions from their Marginals
abstract
Across the social sciences and elsewhere, practitioners frequently have to reason about relationships between random variables, despite lacking joint observations of the variables. This is sometimes called an "ecological" inference; given samples from the marginal distributions of the variables, one attempts to infer their joint distribution. The problem is inherently ill-posed, yet only a few models have been proposed for bringing prior information into the problem, often relying on restrictive or unrealistic assumptions and lacking a unified approach. In this paper, we treat the inference problem generally and propose a unified class of models that encompasses some of those previously proposed while including many new ones. Previous work has relied on either relaxation or approximate inference via MCMC, with the latter known to mix prohibitively slowly for this type of problem. Here we instead give a single exact inference algorithm that works for the entire model class via an efficient fixed point iteration called Dykstra’s method. We investigate empirically both the computational cost of our algorithm and the accuracy of the new models on real datasets, showing favorable performance in both cases and illustrating the impact of increased flexibility in modeling enabled by this work.
Charlie Frogner, Tomaso A. Poggio
ICML2
2019 Symmetry-adapted representation learning
Fabio Anselmi, Georgios Evangelopoulos, Lorenzo Rosasco, Tomaso A. Poggio
Pattern Recognit.4
2017 When and Why Are Deep Networks Better Than Shallow Ones?
abstract
While the universal approximation property holds both for hierarchical and shallow networks, deep networks can approximate the class of compositional functions as well as shallow networks but with exponentially lower number of training parameters and sample complexity. Compositional functions are obtained as a hierarchy of local constituent functions, where "local functions'' are functions with low dimensionality. This theorem proves an old conjecture by Bengio on the role of depth in networks, characterizing precisely the conditions under which it holds. It also suggests possible answers to the the puzzle of why high-dimensional deep networks trained on large training sets often do not seem to show overfit.
Hrushikesh N. Mhaskar, Qianli Liao, Tomaso A. Poggio
AAAI3
2017 Compression of Deep Neural Networks for Image Instance Retrieval
abstract
Image instance retrieval is the problem of retrieving images from a database which contain the same object. Convolutional Neural Network (CNN) based descriptors are becoming the dominant approach for generating global image descriptors for the instance retrieval problem. One major drawback of CNN-based global descriptors is that uncompressed deep neural network models require hundreds of megabytes of storage making them inconvenient to deploy in mobile applications or in custom hardware. In this work, we study the problem of neural network model compression focusing on the image instance retrieval task. We study quantization, coding, pruning and weight sharing techniques for reducing model size for the instance retrieval problem. We provide extensive experimental results on the trade-off between retrieval performance and model size for different types of networks on several data sets providing the most comprehensive study on this topic. We compress models to the order of a few MBs: two orders of magnitude smaller than the uncompressed models while achieving negligible loss in retrieval performance1.
Vijay Chandrasekhar 0001, Jie Lin 0001, Qianli Liao, Olivier Morère, Antoine Veillard, Ling-Yu Duan, Tomaso A. Poggio
DCC7
2017 Nested Invariance Pooling and RBM Hashing for Image Instance Retrieval
abstract
The goal of this work is the computation of very compact binary hashes for image instance retrieval. Our approach has two novel contributions. The first one is Nested Invariance Pooling (NIP), a method inspired from i-theory, a mathematical theory for computing group invariant transformations with feed-forward neural networks. NIP is able to produce compact and well-performing descriptors with visual representations extracted from convolutional neural networks. We specifically incorporate scale, translation and rotation invariances but the scheme can be extended to any arbitrary sets of transformations. We also show that using moments of increasing order throughout nesting is important. The NIP descriptors are then hashed to the target code size (32-256 bits) with a Restricted Boltzmann Machine with a novel batch-level regularization scheme specifically designed for the purpose of hashing (RBMH). A thorough empirical evaluation with state-of-the-art shows that the results obtained both with the NIP descriptors and the NIP+RBMH hashes are consistently outstanding across a wide range of datasets.
Olivier Morère, Jie Lin 0001, Antoine Veillard, Ling-Yu Duan, Vijay Chandrasekhar 0001, Tomaso A. Poggio
ICMR6
2017 Do Deep Neural Networks Suffer from Crowding?
abstract
Crowding is a visual effect suffered by humans, in which an object that can be recognized in isolation can no longer be recognized when other objects, called flankers, are placed close to it. In this work, we study the effect of crowding in artificial Deep Neural Networks (DNNs) for object recognition. We analyze both deep convolutional neural networks (DCNNs) as well as an extension of DCNNs that are multi-scale and that change the receptive field size of the convolution filters with their position in the image. The latter networks, that we call eccentricity-dependent, have been proposed for modeling the feedforward path of the primate visual cortex. Our results reveal that the eccentricity-dependent model, trained on target objects in isolation, can recognize such targets in the presence of flankers, if the targets are near the center of the image, whereas DCNNs cannot. Also, for all tested networks, when trained on targets in isolation, we find that recognition accuracy of the networks decreases the closer the flankers are to the target and the more flankers there are. We find that visual similarity between the target and flankers also plays a role and that pooling in early layers of the network leads to more crowding. Additionally, we show that incorporating flankers into the images of the training set for learning the DNNs does not lead to robustness against configurations not seen at training.
Anna Volokitin, Gemma Roig, Tomaso A. Poggio
NIPS3
2017 Invariant recognition drives neural representations of action sequences
abstract
Recognizing the actions of others from visual stimuli is a crucial aspect of human perception that allows individuals to respond to social cues. Humans are able to discriminate between similar actions despite transformations, like changes in viewpoint or actor, that substantially alter the visual appearance of a scene. This ability to generalize across complex transformations is a hallmark of human visual intelligence. Advances in understanding action recognition at the neural level have not always translated into precise accounts of the computational principles underlying what representations of action sequences are constructed by human visual cortex. Here we test the hypothesis that invariant action discrimination might fill this gap. Recently, the study of artificial systems for static object perception has produced models, Convolutional Neural Networks (CNNs), that achieve human level performance in complex discriminative tasks. Within this class, architectures that better support invariant object recognition also produce image representations that better match those implied by human and primate neural data. However, whether these models produce representations of action sequences that support recognition across complex transformations and closely follow neural representations of actions remains unknown. Here we show that spatiotemporal CNNs accurately categorize video stimuli into action classes, and that deliberate model modifications that improve performance on an invariant action recognition task lead to data representations that better match human neural recordings. Our results support our hypothesis that performance on invariant discrimination dictates the neural representations of actions computed in the brain. These results broaden the scope of the invariant recognition framework for understanding visual intelligence from perception of inanimate objects and faces in static images to the study of human perception of action sequences.
Andrea Tacchetti, Leyla Isik, Tomaso A. Poggio
PLoS Comput. Biol.3
2016 How Important Is Weight Symmetry in Backpropagation?
abstract
Gradient backpropagation (BP) requires symmetric feedforward and feedback connections — the same weights must be used for forward and backward passes. This "weight transport problem'' (Grossberg 1987) is thought to be one of the main reasons to doubt BP's biologically plausibility. Using 15 different classification datasets, we systematically investigate to what extent BP really depends on weight symmetry. In a study that turned out to be surprisingly similar in spirit to Lillicrap et al.'s demonstration (Lillicrap et al. 2014) but orthogonal in its results, our experiments indicate that: (1) the magnitudes of feedback weights do not matter to performance (2) the signs of feedback weights do matter — the more concordant signs between feedforward and their corresponding feedback connections, the better (3) with feedback weights having random magnitudes and 100% concordant signs, we were able to achieve the same or even better performance than SGD. (4) some normalizations/stabilizations are indispensable for such asymmetric BP to work, namely Batch Normalization (BN) (Ioffe and Szegedy 2015) and/or a "Batch Manhattan'' (BM) update rule.
Qianli Liao, Joel Z. Leibo, Tomaso A. Poggio
AAAI3
2016 Holographic Embeddings of Knowledge Graphs
abstract
Learning embeddings of entities and relations is an efficient and versatile method to perform machine learning on relational data such as knowledge graphs. In this work, we propose holographic embeddings (HolE) to learn compositional vector space representations of entire knowledge graphs. The proposed method is related to holographic models of associative memory in that it employs circular correlation to create compositional representations. By using correlation as the compositional operator, HolE can capture rich interactions but simultaneously remains efficient to compute, easy to train, and scalable to very large datasets. Experimentally, we show that holographic embeddings are able to outperform state-of-the-art methods for link prediction on knowledge graphs and relational learning benchmark datasets.
Maximilian Nickel, Lorenzo Rosasco, Tomaso A. Poggio
AAAI3
2016 Unsupervised learning of invariant representations
Fabio Anselmi, Joel Z. Leibo, Lorenzo Rosasco, Jim Mutch, Andrea Tacchetti, Tomaso A. Poggio
Theor. Comput. Sci.6
2015 Convex Learning of Multiple Tasks and their Structure
abstract
Reducing the amount of human supervision is a key problem in machine learning and a natural approach is that of exploiting the relations (structure) among different tasks. This is the idea at the core of multi-task learning. In this context a fundamental question is how to incorporate the tasks structure in the learning problem. We tackle this question by studying a general computational framework that allows to encode a-priori knowledge of the tasks structure in the form of a convex penalty; in this setting a variety of previously proposed methods can be recovered as special cases, including linear and non-linear approaches. Within this framework, we show that tasks and their structure can be efficiently learned considering a convex optimization problem that can be approached by means of block coordinate methods such as alternating minimization and for which we prove convergence to the global minimum.
Carlo Ciliberto, Youssef Mroueh, Tomaso A. Poggio, Lorenzo Rosasco
ICML3
2015 Discriminative template learning in group-convolutional networks for invariant speech representations
abstract
In the framework of a theory for invariant sensory signal representations, a signature which is invariant and selective for speech sounds can be obtained through projections in template signals and pooling over their transformations under a group. For locally compact groups, e.g., translations, the theory explains the resilience of convolutional neural networks with filter weight sharing and max pooling across their local translations in frequency or time. In this paper we propose a discriminative approach for learning an optimum set of templates, under a family of transformations, namely frequency transpositions and perturbations of the vocal tract length, which are among the primary sources of speech variability. Implicitly, we generalize convolutional networks to transformations other than translations, and derive data-specific templates by training a deep network with convolution-pooling layers and densely connected layers. We demonstrate that such a representation, combining group-generalized convolutions, theoretical invariance guarantees and discriminative template selection, improves frame classification performance over standard translation-CNNs and DNNs on TIMIT and Wall Street Journal datasets.
Chiyuan Zhang, Stephen Voinea, Georgios Evangelopoulos, Lorenzo Rosasco, Tomaso A. Poggio
INTERSPEECH5
2015 Learning with a Wasserstein Loss
abstract
Learning to predict multi-label outputs is challenging, but in many problems there is a natural metric on the outputs that can be used to improve predictions. In this paper we develop a loss function for multi-label learning, based on the Wasserstein distance. The Wasserstein distance provides a natural notion of dissimilarity for probability measures. Although optimizing with respect to the exact Wasserstein distance is costly, recent work has described a regularized approximation that is efficiently computed. We describe an efficient learning algorithm based on this regularization, as well as a novel extension of the Wasserstein distance from probability measures to unnormalized measures. We also describe a statistical learning bound for the loss. The Wasserstein loss can encourage smoothness of the predictions with respect to a chosen metric on the output space. We demonstrate this property on a real-data tag prediction problem, using the Yahoo Flickr Creative Commons dataset, outperforming a baseline that doesn't use the metric.
Charlie Frogner, Chiyuan Zhang, Hossein Mobahi, Mauricio Araya-Polo, Tomaso A. Poggio
NIPS5
2015 Learning with Group Invariant Features: A Kernel Perspective
abstract
We analyze in this paper a random feature map based on a theory of invariance (\emph{I-theory}) introduced in \cite{AnselmiLRMTP13}. More specifically, a group invariant signal signature is obtained through cumulative distributions of group-transformed random projections. Our analysis bridges invariant feature learning with kernel methods, as we show that this feature map defines an expected Haar-integration kernel that is invariant to the specified group action. We show how this non-linear random feature map approximates this group invariant kernel uniformly on a set of $N$ points. Moreover, we show that it defines a function space that is dense in the equivalent Invariant Reproducing Kernel Hilbert Space. Finally, we quantify error rates of the convergence of the empirical risk minimization, as well as the reduction in the sample complexity of a learning algorithm using such an invariant representation for signal classification, in a classical supervised learning setting
Youssef Mroueh, Stephen Voinea, Tomaso A. Poggio
NIPS3
2015 The Invariance Hypothesis Implies Domain-Specific Regions in Visual Cortex
abstract
Is visual cortex made up of general-purpose information processing machinery, or does it consist of a collection of specialized modules? If prior knowledge, acquired from learning a set of objects is only transferable to new objects that share properties with the old, then the recognition system's optimal organization must be one containing specialized modules for different object classes. Our analysis starts from a premise we call the invariance hypothesis: that the computational goal of the ventral stream is to compute an invariant-to-transformations and discriminative signature for recognition. The key condition enabling approximate transfer of invariance without sacrificing discriminability turns out to be that the learned and novel objects transform similarly. This implies that the optimal recognition system must contain subsystems trained only with data from similarly-transforming objects and suggests a novel interpretation of domain-specific regions like the fusiform face area (FFA). Furthermore, we can define an index of transformation-compatibility, computable from videos, that can be combined with information about the statistics of natural vision to yield predictions for which object categories ought to have domain-specific regions in agreement with the available data. The result is a unifying account linking the large literature on view-based recognition with the wealth of experimental evidence concerning domain-specific regions.
Joel Z. Leibo, Qianli Liao, Fabio Anselmi, Tomaso A. Poggio
PLoS Comput. Biol.4
2014 A deep representation for invariance and music classification
abstract
Representations in the auditory cortex might be based on mechanisms similar to the visual ventral stream; modules for building invariance to transformations and multiple layers for compositionality and selectivity. In this paper we propose the use of such computational modules for extracting invariant and discriminative audio representations. Building on a theory of invariance in hierarchical architectures, we propose a novel, mid-level representation for acoustical signals, using the empirical distributions of projections on a set of templates and their transformations. Under the assumption that, by construction, this dictionary of templates is composed from similar classes, and samples the orbit of variance-inducing signal transformations (such as shift and scale), the resulting signature is theoretically guaranteed to be unique, invariant to transformations and stable to deformations. Modules of projection and pooling can then constitute layers of deep networks, for learning composite representations. We present the main theoretical and computational aspects of a framework for unsupervised learning of invariant audio representations, empirically evaluated on music genre classification.
Chiyuan Zhang, Georgios Evangelopoulos, Stephen Voinea, Lorenzo Rosasco, Tomaso A. Poggio
ICASSP5
2014 Word-level invariant representations from acoustic waveforms
abstract
Extracting discriminant, transformation-invariant features from raw audio signals remains a serious challenge for speech recognition. The issue of speaker variability is central to this problem, as changes in accent, dialect, gender, and age alter the sound waveform of speech units at multiple levels (phonemes, words, or phrases). Approaches for dealing with this variability have typically focused on analyzing the spectral properties of speech at the level of frames, on par with frame-level acoustic modeling usually applied to speech recognition systems. In this paper, we propose a framework for representing speech at the word level and extracting features from the acoustic, temporal domain, without the need for spectral encoding or preprocessing. Leveraging recent work on unsupervised learning of invariant sensory representations, we extract a signature for a word by first projecting its raw waveform onto a set of templates and their transformations, and then forming empirical estimates of the resulting one-dimensional distributions via histograms. The representation and relevant parameters are evaluated for word classification on a series of datasets with increasing speakermismatch difficulty, and the results are compared to those of an MFCC-based representation. Index Terms: invariance, acoustic features, speech representation, word classification
Stephen Voinea, Chiyuan Zhang, Georgios Evangelopoulos, Lorenzo Rosasco, Tomaso A. Poggio
INTERSPEECH5
2014 Phone classification by a hierarchy of invariant representation layers
abstract
We propose a multi-layer feature extraction framework for speech, capable of providing invariant representations. A set of templates is generated by sampling the result of applying smooth, identity-preserving transformations (such as vocal tract length and tempo variations) to arbitrarily-selected speech signals. Templates are then stored as the weights of “neurons”. We use a cascade of such computational modules to factor out different types of transformation variability in a hierarchy, and show that it improves phone classification over baseline features. In addition, we describe empirical comparisons of a) different transformations which may be responsible for the variability in speech signals and of b) different ways of assembling template sets for training. The proposed layered system is an effort towards explaining the performance of recent deep learning networks and the principles by which the human auditory cortex might reduce the sample complexity of learning in speech recognition. Our theory and experiments suggest that invariant representations are crucial in learning from complex, real-world data like natural speech. Our model is built on basic computational primitives of cortical neurons, thus making an argument about how representations might be learned in the human auditory cortex. Index Terms: Invariance, Auditory Cortex, Phonetic Classification, Convolutional Network
Chiyuan Zhang, Stephen Voinea, Georgios Evangelopoulos, Lorenzo Rosasco, Tomaso A. Poggio
INTERSPEECH5
2013 The Computational Magic of Pattern Recognition in Cortex: A Theory of Selectivity and Invariance
Tomaso A. Poggio
ICPRAM1
2013 Learning invariant representations and applications to face verification
abstract
One approach to computer object recognition and modeling the brain's ventral stream involves unsupervised learning of representations that are invariant to common transformations. However, applications of these ideas have usually been limited to 2D affine transformations, e.g., translation and scaling, since they are easiest to solve via convolution. In accord with a recent theory of transformation-invariance, we propose a model that, while capturing other common convolutional networks as special cases, can also be used with arbitrary identity-preserving transformations. The model's wiring can be learned from videos of transforming objects---or any other grouping of images into sets by their depicted object. Through a series of successively more complex empirical tests, we study the invariance/discriminability properties of this model with respect to different transformations. First, we empirically confirm theoretical predictions for the case of 2D affine transformations. Next, we apply the model to non-affine transformations: as expected, it performs well on face verification tasks requiring invariance to the relatively smooth transformations of 3D rotation-in-depth and changes in illumination direction. Surprisingly, it can also tolerate clutter transformations'' which map an image of a face on one background to an image of the same face on a different background. Motivated by these empirical findings, we tested the same model on face verification benchmark tasks from the computer vision literature: Labeled Faces in the Wild, PubFig and a new dataset we gathered---achieving strong performance in these highly unconstrained cases as well."
Qianli Liao, Joel Z. Leibo, Tomaso A. Poggio
NIPS3
2013 Neural representation of action sequences: how far can a simple snippet-matching model take us?
abstract
The macaque Superior Temporal Sulcus (STS) is a brain area that receives and integrates inputs from both the ventral and dorsal visual processing streams (thought to specialize in form and motion processing respectively). For the processing of articulated actions, prior work has shown that even a small population of STS neurons contains sufficient information for the decoding of actor invariant to action, action invariant to actor, as well as the specific conjunction of actor and action. This paper addresses two questions. First, what are the invariance properties of individual neural representations (rather than the population representation) in STS? Second, what are the neural encoding mechanisms that can produce such individual neural representations from streams of pixel images? We find that a baseline model, one that simply computes a linear weighted sum of ventral and dorsal responses to short action “snippets”, produces surprisingly good fits to the neural data. Interestingly, even using inputs from a single stream, both actor-invariance and action-invariance can be produced simply by having different linear weights.
Cheston Tan, Jedediah M. Singer, Thomas Serre, David L. Sheinberg, Tomaso A. Poggio
NIPS5
2012 Learning Manifolds with K-Means and K-Flats
abstract
We study the problem of estimating a manifold from random samples. In particular, we consider piecewise constant and piecewise linear estimators induced by k-means and k-flats, and analyze their performance. We extend previous results for k-means in two separate directions. First, we provide new results for k-means reconstruction on manifolds and, secondly, we prove reconstruction bounds for higher-order approximation (k-flats), for which no known results were previously available. While the results for k-means are novel, some of the technical tools are well-established in the literature. In the case of k-flats, both the results and the mathematical tools are new.
Guille D. Cañas, Tomaso A. Poggio, Lorenzo Rosasco
NIPS2
2012 Multiclass Learning with Simplex Coding
abstract
In this paper we dicuss a novel framework for multiclass learning, defined by a suitable coding/decoding strategy, namely the simplex coding, that allows to generalize to multiple classes a relaxation approach commonly used in binary classification. In this framework a relaxation error analysis can be developed avoiding constraints on the considered hypotheses class. Moreover, we show that in this setting it is possible to derive the first provably consistent regularized methods with training/tuning complexity which is {\em independent} to the number of classes. Tools from convex analysis are introduced that can be used beyond the scope of this paper.
Youssef Mroueh, Tomaso A. Poggio, Lorenzo Rosasco, Jean-Jacques E. Slotine
NIPS2
2011 HMDB: A large video database for human motion recognition
abstract
With nearly one billion online videos viewed everyday, an emerging new frontier in computer vision research is recognition and search in video. While much effort has been devoted to the collection and annotation of large scalable static image datasets containing thousands of image categories, human action datasets lag far behind. Current action recognition databases contain on the order of ten different action categories collected under fairly controlled conditions. State-of-the-art performance on these datasets is now near ceiling and thus there is a need for the design and creation of new benchmarks. To address this issue we collected the largest action video database to-date with 51 action categories, which in total contain around 7,000 manually annotated clips extracted from a variety of sources ranging from digitized movies to YouTube. We use this database to evaluate the performance of two representative computer vision systems for action recognition and explore the robustness of these methods under various conditions such as camera motion, viewpoint, video quality and occlusion.
Hilde Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso A. Poggio, Thomas Serre
ICCV4
2011 Why The Brain Separates Face Recognition From Object Recognition
abstract
Many studies have uncovered evidence that visual cortex contains specialized regions involved in processing faces but not other object classes. Recent electrophysiology studies of cells in several of these specialized regions revealed that at least some of these regions are organized in a hierarchical manner with viewpoint-specific cells projecting to downstream viewpoint-invariant identity-specific cells (Freiwald and Tsao 2010). A separate computational line of reasoning leads to the claim that some transformations of visual inputs that preserve viewed object identity are class-specific. In particular, the 2D images evoked by a face undergoing a 3D rotation are not produced by the same image transformation (2D) that would produce the images evoked by an object of another class undergoing the same 3D rotation. However, within the class of faces, knowledge of the image transformation evoked by 3D rotation can be reliably transferred from previously viewed faces to help identify a novel face at a new viewpoint. We show, through computational simulations, that an architecture which applies this method of gaining invariance to class-specific transformations is effective when restricted to faces and fails spectacularly when applied across object classes. We argue here that in order to accomplish viewpoint-invariant face identification from a single example view, visual cortex must separate the circuitry involved in discounting 3D rotations of faces from the generic circuitry involved in processing other objects. The resulting model of the ventral stream of visual cortex is consistent with the recent physiology results showing the hierarchical organization of the face processing network.
Joel Z. Leibo, Jim Mutch, Tomaso A. Poggio
NIPS3
2010 Hierarchical Learning Machines and Neuroscience of Visual Cortex
Tomaso A. Poggio
ECML/PKDD (1)1
2009 On Invariance in Hierarchical Models
abstract
A goal of central importance in the study of hierarchical models for object recognition -- and indeed the visual cortex -- is that of understanding quantitatively the trade-off between invariance and selectivity, and how invariance and discrimination properties contribute towards providing an improved representation useful for learning from data. In this work we provide a general group-theoretic framework for characterizing and understanding invariance in a family of hierarchical models. We show that by taking an algebraic perspective, one can provide a concise set of conditions which must be met to establish invariance, as well as a constructive prescription for meeting those conditions. Analyses in specific cases of particular relevance to computer vision and text processing are given, yielding insight into how and when invariance can be achieved. We find that the minimal sets of transformations intrinsic to the hierarchical model needed to support a particular invariance can be clearly described, thereby encouraging efficient computational implementations.
Jake V. Bouvrie, Lorenzo Rosasco, Tomaso A. Poggio
NIPS3
2008 Localized spectro-temporal cepstral analysis of speech
abstract
Drawing on recent progress in auditory neuroscience, we present a novel speech feature analysis technique based on localized spectro- temporal cepstral analysis of speech. We proceed by extracting localized 2D patches from the spectrogram and project onto a 2D discrete cosine (2D-DCT) basis. For each time frame, a speech feature vector is then formed by concatenating low-order 2D- DCT coefficients from the set of corresponding patches. We argue that our framework has significant advantages over standard one- dimensional MFCC features. In particular, we find that our features are more robust to noise, and better capture temporal modulations important for recognizing plosive sounds. We evaluate the performance of the proposed features on a TIMIT classification task in clean, pink, and babble noise conditions, and show that our feature analysis outperforms traditional features based on MFCCs.
Jake V. Bouvrie, Tony Ezzat, Tomaso A. Poggio
ICASSP3
2008 A Canonical Neural Circuit for Cortical Nonlinear Operations
abstract
A few distinct cortical operations have been postulated over the past few years, suggested by experimental data on nonlinear neural response across different areas in the cortex. Among these, the energy model proposes the summation of quadrature pairs following a squaring nonlinearity in order to explain phase invariance of complex V1 cells. The divisive normalization model assumes a gain-controlling, divisive inhibition to explain sigmoid-like response profiles within a pool of neurons. A gaussian-like operation hypothesizes a bell-shaped response tuned to a specific, optimal pattern of activation of the presynaptic inputs. A max-like operation assumes the selection and transmission of the most active response among a set of neural inputs. We propose that these distinct neural operations can be computed by the same canonical circuitry, involving divisive normalization and polynomial nonlinearities, for different parameter values within the circuit. Hence, this canonical circuit may provide a unifying framework for several circuit models, such as the divisive normalization and the energy models. As a case in point, we consider a feedforward hierarchical model of the ventral pathway of the primate visual cortex, which is built on a combination of the gaussian-like and max-like operations. We show that when the two operations are approximated by the circuit proposed here, the model is capable of generating selective and invariant neural responses and performing object recognition, in good agreement with neurophysiological data.
Minjoon Kouh, Tomaso A. Poggio
Neural Comput.2
2007 AM-FM Demodulation of Spectrograms using Localized 2D Max-Gabor Analysis
abstract
We present a method that de-modulates a narrowband magnitude spectrogram S(f, t) into a frequency modulation term cos(Φ)(f,t)) which represents the underlying harmonic carrier, and an amplitude modulation term A(f,t) which represents the spectral envelope. Our method operates by performing a two-dimensional local patch analysis of the spectrogram, in which each patch is factored into a local carrier term and a local amplitude envelope term using a Max-Gabor analysis. We demonstrate the technique over a wide variety of speakers, and show how the spectrograms in each case may be adequately reconstructed as S(f, t) = A(f, t)cos(Φ(f, t)).
Tony Ezzat, Jake V. Bouvrie, Tomaso A. Poggio
ICASSP (4)3
2007 A Biologically Inspired System for Action Recognition
abstract
We present a biologically-motivated system for the recognition of actions from video sequences. The approach builds on recent work on object recognition based on hierarchical feedforward architectures [25, 16, 20] and extends a neurobiological model of motion processing in the visual cortex [10]. The system consists of a hierarchy of spatio-temporal feature detectors of increasing complexity: an input sequence is first analyzed by an array of motion- direction sensitive units which, through a hierarchy of processing stages, lead to position-invariant spatio-temporal feature detectors. We experiment with different types of motion-direction sensitive units as well as different system architectures. As in [16], we find that sparse features in intermediate stages outperform dense ones and that using a simple feature selection approach leads to an efficient system that performs better with far fewer features. We test the approach on different publicly available action datasets, in all cases achieving the highest results reported to date.
Hueihan Jhuang, Thomas Serre, Lior Wolf, Tomaso A. Poggio
ICCV4
2007 Spectro-temporal analysis of speech using 2-d Gabor filters
abstract
We present a 2-D spectro-temporal Gabor filterbank based on the 2-D Fast Fourier Transform, and show how it may be used to analyze localized patches of a spectrogram. We argue that the 2-D Gabor filterbank has the capacity to decompose a patch into its underlying dominant spectro-temporal components, and we illustrate the response of our filterbank to different speech phenomena such as harmonicity, formants, vertical onsets/offsets, noise, and overlapping simultaneous speakers. Index Terms: speech analysis, spectro-temporal filterbanks, 2-D Gabor
Tony Ezzat, Jake V. Bouvrie, Tomaso A. Poggio
INTERSPEECH3
2007 A Component-based Framework for Face Detection and Identification
Bernd Heisele, Thomas Serre, Tomaso A. Poggio
Int. J. Comput. Vis.3
2007 Robust Object Recognition with Cortex-Like Mechanisms
abstract
We introduce a new general framework for the recognition of complex visual scenes, which is motivated by biology: We describe a hierarchical system that closely follows the organization of visual cortex and builds an increasingly complex and invariant feature representation by alternating between a template matching and a maximum pooling operation. We demonstrate the strength of the approach on a range of recognition tasks: From invariant single object recognition in clutter to multiclass categorization problems and complex scene understanding tasks that rely on the recognition of both shape-based as well as texture-based objects. Given the biological constraints that the system had to satisfy, the approach performs surprisingly well: It has the capability of learning from only a few training examples and competes with state-of-the-art systems. We also discuss the existence of a universal, redundant dictionary of features that could handle the recognition of most object categories. In addition to its relevance for computer vision, the success of this approach suggests a plausibility proof for a class of feedforward models of object recognition in cortex.
Thomas Serre, Lior Wolf, Stanley M. Bileschi, Maximilian Riesenhuber, Tomaso A. Poggio
IEEE Trans. Pattern Anal. Mach. Intell.5
2006 Max-Gabor analysis and synthesis of spectrograms
abstract
We present a method that analyzes a two-dimensional magnitude spectrogram S(f, t) into its local constituent spectro-temporal amplitudes A(f,t), frequencies F(f, t), orientations Θ(f, t), and phases φ(f, t). The method operates by performing a twodimensional local Gabor-like analysis of the spectrogram, retaining only the parameters of the 2D-Gabor filter with maximal amplitude response within the local region. We demonstrate the technique over a wide variety of speakers, and show how the spectrograms in each case may be adequately reconstructed using the parameters of the Max-Gabor analysis. Finally, we discuss the nature of the extracted Max-Gabor parameters. Index Terms: spectrogram analysis, spectrogram reconstruction, two-dimensional Gabor, spectro-temporal frequency, spectrotemporal orientation. 1.
Tony Ezzat, Jake V. Bouvrie, Tomaso A. Poggio
INTERSPEECH3
2005 Object Recognition with Features Inspired by Visual Cortex
abstract
We introduce a novel set of features for robust object recognition. Each element of this set is a complex feature obtained by combining position- and scale-tolerant edge-detectors over neighboring positions and multiple orientations. Our system's architecture is motivated by a quantitative model of visual cortex. We show that our approach exhibits excellent recognition performance and outperforms several state-of-the-art systems on a variety of image datasets including many different object categories. We also demonstrate that our system is able to learn from very few examples. The performance of the approach constitutes a suggestive plausibility proof for a class of feedforward models of object recognition in cortex.
Thomas Serre, Lior Wolf, Tomaso A. Poggio
CVPR (2)3
2005 Learning Features of Intermediate Complexity for the Recognition of Biological Motion
Rodrigo Sigala, Thomas Serre, Tomaso A. Poggio, Martin A. Giese
ICANN (1)3
2005 Morphing spectral envelopes using audio flow
abstract
We present a method for morphing between smooth spectral magnitude envelopes of speech. An important element of our method is the notion of audio flow, which is inspired by similar notions of optical flow computed between images in computer vision applications. Audio flow defines the correspondence between two smooth spectral magnitude envelopes, and encodes the formant shifting that occurs from one sound to another. We present several algorithms for the automatic computation of audio flow from a small 20 second corpus of speech. In addition, we present an algorithm for morphing smoothly between any two spectral magnitude envelopes, given the computed audio flow between them.
Tony Ezzat, Ethan Meyers, James R. Glass, Tomaso A. Poggio
INTERSPEECH4
2003 Reanimating Faces in Images and Video
abstract
Abstract This paper presents a method for photo‐realistic animation that can be applied to any face shown in a single imageor a video. The technique does not require example data of the person's mouth movements, and the image to beanimated is not restricted in pose or illumination. Video reanimation allows for head rotations and speech in theoriginal sequence, but neither of these motions is required. In order to animate novel faces, the system transfers mouth movements and expressions across individuals, basedon a common representation of different faces and facial expressions in a vector space of 3D shapes and textures.This space is computed from 3D scans of neutral faces, and scans of facial expressions. The 3D model's versatility with respect to pose and illumination is conveyed to photo‐realistic image and videoprocessing by a framework of analysis and synthesis algorithms: The system automatically estimates 3D shape andall relevant rendering parameters, such as pose, from single images. In video, head pose and mouth movements aretracked automatically. Reanimated with new mouth movements, the 3D face is rendered into the original images. Categories and Subject Descriptors (according to ACM CCS): I.3.7 [Computer Graphics]: Animation
Volker Blanz, Curzio Basso, Tomaso A. Poggio, Thomas Vetter
Comput. Graph. Forum3
2003 Face recognition: component-based versus global approaches
Bernd Heisele, Purdy Ho, Jane Wu, Tomaso A. Poggio
Comput. Vis. Image Underst.4
2003 Hierarchical classification and feature reduction for fast face detection with support vector machines
Bernd Heisele, Thomas Serre, Sam Prentice, Tomaso A. Poggio
Pattern Recognit.4
2003 Full-body person recognition system
Chikahito Nakajima, Massimiliano Pontil, Bernd Heisele, Tomaso A. Poggio
Pattern Recognit.4
2003 Image Representations and Feature Selection for Multimedia Database Search
abstract
The success of a multimedia information system depends heavily on the way the data is represented. Although there are "natural" ways to represent numerical data, it is not clear what is a good way to represent multimedia data, such as images, video, or sound. We investigate various image representations where the quality of the representation is judged based on how well a system for searching through an image database can perform-although the same techniques and representations can be used for other types of object detection tasks or multimedia data analysis problems. The system is based on a machine learning method used to develop object detection models from example images that can subsequently be used for examples to detect-search-images of a particular object in an image database. As a base classifier for the detection task, we use support vector machines (SVM), a kernel based learning method. Within the framework of kernel classifiers, we investigate new image representations/kernels derived from probabilistic models of the class of images considered and present a new feature selection method which can be used to reduce the dimensionality of the image representation without significant losses in terms of the performance of the detection-search-system.
Theodoros Evgeniou, Massimiliano Pontil, Constantine Papageorgiou, Tomaso A. Poggio
IEEE Trans. Knowl. Data Eng.4
2002 Biophysiologically Plausible Implementations of the Maximum Operation
abstract
Visual processing in the cortex can be characterized by a predominantly hierarchical architecture, in which specialized brain regions along the processing pathways extract visual features of increasing complexity, accompanied by greater invariance in stimulus properties such as size and position. Various studies have postulated that a nonlinear pooling function such as the maximum (MAX) operation could be fundamental in achieving such selectivity and invariance. In this article, we are concerned with neurally plausible mechanisms that may be involved in realizing the MAX operation. Different canonical models are proposed, each based on neural mechanisms that have been previously discussed in the context of cortical processing. Through simulations and mathematical analysis, we compare the performance and robustness of these mechanisms. We derive experimentally verifiable predictions for each model and discuss the relevant physiological considerations.
Angela J. Yu, Martin A. Giese, Tomaso A. Poggio
Neural Comput.3
2002 Learning and vision machines
abstract
The problem of learning is arguably at the very core of the problem of intelligence, both biological and artificial. In this paper we review our approach to the problem of visual perception based on supervised learning. After a brief presentation of the theoretical background, we focus on some of the engineering applications of statistical learning to computer vision and discuss the main open problems and directions of our future research.
Bernd Heisele, Alessandro Verri, Tomaso A. Poggio
Proc. IEEE3
2002 Trainable videorealistic speech animation
abstract
We describe how to create with machine learning techniques a generative, speech animation module. A human subject is first recorded using a videocamera as he/she utters a predetermined speech corpus. After processing the corpus automatically, a visual speech module is learned from the data that is capable of synthesizing the human subject's mouth uttering entirely novel utterances that were not recorded in the original video. The synthesized utterance is re-composited onto a background sequence which contains natural head and eye movement. The final output is videorealistic in the sense that it looks like a video camera recording of the subject. At run time, the input to the system can be either real audio sequences or synthetic audio produced by a text-to-speech system, as long as they have been phonetically aligned.The two key contributions of this paper are 1) a variant of the multidimensional morphable model (MMM) to synthesize new, previously unseen mouth configurations from a small set of mouth image prototypes; and 2) a trajectory synthesis technique based on regularization, which is automatically trained from the recorded video corpus, and which is capable of synthesizing trajectories in MMM space corresponding to any desired utterance.
Tony Ezzat, Gadi Geiger, Tomaso A. Poggio
ACM Trans. Graph.3
2001 Feature Reduction and Hierarchy of Classifiers for Fast Object Detection in Video Images
abstract
We present a two-step method to speed-up object detection systems in computer vision that use Support Vector Machines (SVMs) as classifiers. In a first step we perform feature reduction by choosing relevant image features according to a measure derived from statistical learning theory. In a second step we build a hierarchy of classifiers. On the bottom level, a simple and fast classifier analyzes the whole image and rejects large parts of the background On the top level, a slower but more accurate classifier performs the final detection. Experiments with a face detection system show that combining feature reduction with hierarchical classification leads to a speed-up by a factor of 170 with similar classification performance.
Bernd Heisele, Thomas Serre, Sayan Mukherjee 0001, Tomaso A. Poggio
CVPR (2)4
2001 Component-based Face Detection
abstract
We present a component-based, trainable system for detecting frontal and near-frontal views of faces in still gray images. The system consists of a two-level hierarchy of Support Vector Machine (SVM) classifiers. On the first level, component classifiers independently detect components Of a face. On the second level, a single classifier checks if the geometrical configuration of the detected components in the image matches a geometrical model of a face. We propose a method for automatically learning components by using 3-D head models, This approach has the advantage that no manual interaction is required for choosing and extracting components. Experiments show that the component-based system is significantly more robust against rotations in depth than a comparable system trained on whole face patterns.
Bernd Heisele, Thomas Serre, Massimiliano Pontil, Tomaso A. Poggio
CVPR (1)4
2001 Face Recognition with Support Vector Machines: Global versus Component-based Approach
abstract
We present a component-based method and two global methods for face recognition and evaluate them with respect to robustness against pose changes. In the component system we first locate facial components, extract them and combine them into a single feature vector which is classified by a Support Vector Machine (SVM). The two global systems recognize faces by classifying a single feature vector consisting of the gray values of the whole face image. In the first global system we trained a single SVM classifier for each person in the database. The second system consists of sets of viewpoint-specific SVM classifiers and involves clustering during training. We performed extensive tests on a database which included faces rotated up to about 40/spl deg/ in depth. The component system clearly outperformed both global systems on all tests.
Bernd Heisele, Purdy Ho, Tomaso A. Poggio
ICCV3
2001 Categorization by Learning and Combining Object Parts
abstract
We describe an algorithm for automatically learning discriminative com- ponents of objects with SVM classifiers. It is based on growing image parts by minimizing theoretical bounds on the error probability of an SVM. Component-based face classifiers are then combined in a second stage to yield a hierarchical SVM classifier. Experimental results in face classification show considerable robustness against rotations in depth and suggest performance at significantly better level than other face detection systems. Novel aspects of our approach are: a) an algorithm to learn component-based classification experts and their combination, b) the use of 3-D morphable models for training, and c) a maximum operation on the output of each component classifier which may be relevant for bio- logical models of visual recognition.
Bernd Heisele, Thomas Serre, Massimiliano Pontil, Thomas Vetter, Tomaso A. Poggio
NIPS5
2001 Example-Based Object Detection in Images by Components
abstract
We present a general example-based framework for detecting objects in static images by components. The technique is demonstrated by developing a system that locates people in cluttered scenes. The system is structured with four distinct example-based detectors that are trained to separately find the four components of the human body: the head, legs, left arm, and right arm. After ensuring that these components are present in the proper geometric configuration, a second example-based classifier combines the results of the component detectors to classify a pattern as either a "person" or a "nonperson." We call this type of hierarchical architecture, in which learning occurs at multiple stages, an adaptive combination of classifiers (ACC). We present results that show that this system performs significantly better than a similar full-body person detector. This suggests that the improvement in performance is due to the component-based approach and the ACC data classification architecture. The algorithm is also more robust than the full-body person detection method in that it is capable of locating partially occluded views of people and people whose body parts have little contrast with the background.
Anuj Mohan, Constantine Papageorgiou, Tomaso A. Poggio
IEEE Trans. Pattern Anal. Mach. Intell.3
2000 Learning-Based Approach to Real Time Tracking and Analysis of Faces
abstract
This paper describes a trainable system capable of tracking faces and facial features like eyes and nostrils and estimating basic mouth features such as degrees of openness and smile in real time. In developing this system, we have addressed the twin issues of image representation and algorithms for learning. We have used the invariance properties of image representations based on Haar wavelets to robustly capture various facial features. Similarly, unlike previous approaches this system is entirely trained using examples and does not rely on a priori (hand-crafted) models of facial features based on an optical flow or facial musculature. The system works in several stages that begin with face detection, followed by localization of facial features and estimation of mouth parameters. Each of these stages is formulated as a problem in supervised learning from examples. We apply the new and robust technique of support vector machines (SVM) for classification in the stage of skin segmentation, face detection and eye detection. Estimation of mouth parameters is modeled as a regression from a sparse subset of coefficients (basis functions) of an overcomplete dictionary of Haar wavelets.
Vinay P. Kumar, Tomaso A. Poggio
FG2
2000 Bounds on the Generalization Performance of Kernel Machine Ensembles
Theodoros Evgeniou, Luis Pérez-Breva, Massimiliano Pontil, Tomaso A. Poggio
ICML4
2000 Object Recognition and Detection by a Combination of Support Vector Machine and Rotation Invariant Phase Only Correlation
abstract
This paper proposes an object recognition and detection method by a combination of support vector machine classifier (SVM) and rotation invariant phase-only correlation (RIPOC). SVM is a learning technique that is well founded in statistical learning theory. RIPOC is a position and rotation invariant pattern matching technique. We combined these two techniques to develop an augmented reality system. This system can recognize and detect objects from image sequences without special image marks or sensors and show information about the objects through a head-mounted display. Performance is real time.
Chikahito Nakajima, Norihiko Itoh, Massimiliano Pontil, Tomaso A. Poggio
ICPR4
2000 People Recognition and Pose Estimation in Image Sequences
abstract
Presents a system which learns from examples to automatically recognize people and estimate their poses in image sequences with the potential application to daily surveillance in indoor environments. The person in the image is represented by a set of features based on color and shape information. Recognition is carried out through a hierarchy of biclass SVM classifiers that are separately trained to recognize people and estimate their poses. The system shows a very high accuracy in people recognition and about 85% level of performance in pose estimation, outperforming in both cases k-nearest neighbors classifiers. The system works in real time.
Chikahito Nakajima, Massimiliano Pontil, Tomaso A. Poggio
IJCNN (4)3
2000 Incremental and Decremental Support Vector Machine Learning
abstract
An on-line recursive algorithm for training support vector machines, one vector at a time, is presented. Adiabatic increments retain the Kuhn(cid:173) Tucker conditions on all previously seen training data, in a number of steps each computed analytically. The incremental procedure is re(cid:173) versible, and decremental "unlearning" offers an efficient method to ex(cid:173) actly evaluate leave-one-out generalization performance. Interpretation of decremental unlearning in feature space sheds light on the relationship between generalization and geometry of the data.
Gert Cauwenberghs, Tomaso A. Poggio
NIPS2
2000 Feature Selection for SVMs
abstract
We introduce a method of feature selection for Support Vector Machines. The method is based upon finding those features which minimize bounds on the leave-one-out error. This search can be efficiently performed via gradient descent. The resulting algorithms are shown to be superior to some standard feature selection algorithms on both toy data and real-life problems of face recognition, pedestrian detection and analyzing DNA micro array data.
Jason Weston, Sayan Mukherjee 0001, Olivier Chapelle, Massimiliano Pontil, Tomaso A. Poggio, Vladimir Vapnik
NIPS5
2000 Statistical Learning Theory: A Primer
Theodoros Evgeniou, Massimiliano Pontil, Tomaso A. Poggio
Int. J. Comput. Vis.3
2000 Visual Speech Synthesis by Morphing Visemes
Tony Ezzat, Tomaso A. Poggio
Int. J. Comput. Vis.2
2000 Morphable Models for the Analysis and Synthesis of Complex Motion Patterns
Martin A. Giese, Tomaso A. Poggio
Int. J. Comput. Vis.2
2000 A Trainable System for Object Detection
Constantine Papageorgiou, Tomaso A. Poggio
Int. J. Comput. Vis.2
2000 Introduction: Learning and Vision at CBCL
Tomaso A. Poggio, Alessandro Verri
Int. J. Comput. Vis.1
1999 Sparse correlation kernel reconstruction
abstract
This paper presents a new paradigm for signal reconstruction and superresolution, correlation kernel analysis (CKA), that is based on the selection of a sparse set of bases from a large dictionary of class-specific basis functions. The basis functions that we use are the correlation functions of the class of signals we are analyzing. To choose the appropriate features from this large dictionary, we use support vector machine (SVM) regression and compare this to traditional principal component analysis (PCA) for the task of signal reconstruction. The testbed we use in this paper is a set of images of pedestrians. Based on the results presented here, we conclude that, when used with a sparse representation technique, the correlation function is an effective kernel for image reconstruction.
Constantine Papageorgiou, Federico Girosi, Tomaso A. Poggio
ICASSP3
1999 A Pattern Classification Approach to Dynamical Object Detection
abstract
Current systems for object detection in video sequences rely on explicit dynamical models like Kalman filters or hidden Markov models. There is significant overhead needed in the development of such systems as well as the a priori assumption that the object dynamics can be described with such a dynamical model. This paper describes a new pattern classification technique for object detection in video sequences that uses a rich, overcomplete dictionary of wavelet features to describe an object class. Unlike previous work where a small subset of features was selected from the dictionary, this system does no feature selection and learns the model in the full 1,326 dimensional feature space. Comparisons using different sized sets of several types of features are given. We extend this representation into the time domain without assuming any explicit model of dynamics. This data driven approach produces a model of the physical structure and short-time dynamical characteristics of people from a training set of examples; no assumptions are made about the motion of people, just that short sequences characterize their dynamics sufficiently for the purposes of detection. One of the main benefits of this approach is that transient false positives are reduced. This technique compares favorably with the static detection approach and could be applied to other object classes. We also present a real-time version of one of our static people detection systems.
Constantine Papageorgiou, Tomaso A. Poggio
ICCV2
1999 Trainable Pedestrian Detection
abstract
Robust, fast object detection systems are critical to the success of next-generation automotive vision systems. An important criteria is that the detection system be easily configurable to a new domain or environment. In this paper, we present work on a general object detection system that can be trained to detect different types of objects; we focus on the task of pedestrian detection. This paradigm of learning from examples allows us to avoid the need for a hand-crafted solution. Unlike many pedestrian detection systems, the core technique does not rely on motion information and makes no assumptions on the scene structured or the number of objects present. We discuss an extension to the system that takes advantage of dynamical information when processing video sequences to enhance accuracy. We also describe a real, real-time version of the system that has been integrated into a DaimlerChrysler test vehicle.
Constantine Papageorgiou, Tomaso A. Poggio
ICIP (4)2
1999 GEM: A Global Electronic Market System
Benny Rachlevsky-Reich, Israel Ben-Shaul, Nicholas Tung Chan, Andrew W. Lo, Tomaso A. Poggio
Inf. Syst.5
1998 MikeTalk: A Talking Facial Display Based on Morphing Visemes
abstract
We present MikeTalk, a text-to-audiovisual speech synthesizer which converts input text into an audiovisual speech stream. MikeTalk is built using visemes, which are a set of images spanning a large range of mouth shapes. The visemes are acquired from a recorded visual corpus of a human subject which is specifically designed to elicit one instantiation of each viseme. Using optical flow methods, correspondence from every viseme to every other viseme is computed automatically. By morphing along this correspondence, a smooth transition between viseme images may be generated. A complete visual utterance is constructed by concatenating viseme transitions. Finally, phoneme and timing information extracted from a text-to-speech synthesizer is exploited to determine which viseme transitions to use, and the rate at which the morphing process should occur. In this manner, we are able to synchronize the visual speech stream with the audio speech stream, and hence give the impression, of a photorealistic talking face.
Tony Ezzat, Tomaso A. Poggio
CA2
1998 Hierarchical Morphable Models
abstract
This paper presents a new technique for modelling object classes (such as faces) and matching the model to novel images from the object class. The technique can be used for a variety of image analysis applications including face recognition, object verification and facial expression analysis. The model, called a hierarchical morphable model, is "learned" from example images (partioned into components) and their correspondences. This is an extension to the work on morphable models described in previous papers. Hierarchical morphable models are shown to find good matches to novel face images and are also robust to partial occlusion.
Michael J. Jones 0001, Tomaso A. Poggio
CVPR2
1998 Multidimensional Morphable Models
abstract
We describe a flexible model for representing images of objects of a certain class, known a priori, such as faces, and introduce a new algorithm for matching it to a novel image and thereby performing image analysis. We call this model a multidimensional morphable model or just a, morphable model. The morphable model is learned from example images (called prototypes) of objects of a class. In this paper we introduce an effective stochastic gradient descent algorithm that automaticaIly matches a model to a novel image by finding the parameters that minimize the error between the image generated by the model and the novel image. Two examples demonstrate the robustness and the broad range of applicability of the matching algorithm and the underlying morphable model. Our approach can provide novel solutions to several vision tasks, including the computation of image correspondence, object verification, image synthesis and image compression.
Michael J. Jones 0001, Tomaso A. Poggio
ICCV2
1998 A General Framework for Object Detection
abstract
This paper presents a general trainable framework for object detection in static images of cluttered scenes. The detection technique we develop is based on a wavelet representation of an object class derived from a statistical analysis of the class instances. By learning an object class in terms of a subset of an overcomplete dictionary of wavelet basis functions, we derive a compact representation of an object class which is used as an input to a support vector machine classifier. This representation overcomes both the problem of in-class variability and provides a low false detection rate in unconstrained environments. We demonstrate the capabilities of the technique in two domains whose inherent information content differs significantly. The first system is face detection and the second is the domain of people which, in contrast to faces, vary greatly in color, texture, and patterns. Unlike previous approaches, this system learns from examples and does not rely on any a priori (hand-crafted) models or motion-based segmentation. The paper also presents a motion-based extension to enhance the performance of the detection algorithm over video sequences. The results presented here suggest that this architecture may well be quite general.
Constantine Papageorgiou, Michael Oren, Tomaso A. Poggio
ICCV3
1998 Multidimensional Morphable Models: A Framework for Representing and Matching Object Classes
Michael J. Jones 0001, Tomaso A. Poggio
Int. J. Comput. Vis.2
1998 A Sparse Representation For Function Approximation
abstract
We derive a new general representation for a function as a linear combination of local correlation kernels at optimal sparse locations (and scales) and characterize its relation to principal component analysis, regularization, sparsity principles, and support vector machines.
Tomaso A. Poggio, Federico Girosi
Neural Comput.1
1998 Example-Based Learning for View-Based Human Face Detection
abstract
We present an example-based learning approach for locating vertical frontal views of human faces in complex scenes. The technique models the distribution of human face patterns by means of a few view-based "face" and "nonface" model clusters. At each image location, a difference feature vector is computed between the local image pattern and the distribution-based model. A trained classifier determines, based on the difference feature vector measurements, whether or not a human face exists at the current image location. We show empirically that the distance metric we adopt for computing difference feature vectors, and the "nonface" clusters we include in our distribution-based model, are both critical for the success of our system.
Kah Kay Sung, Tomaso A. Poggio
IEEE Trans. Pattern Anal. Mach. Intell.2
1998 Incorporating prior information in machine learning by creating virtual examples
abstract
One of the key problems in supervised learning is the insufficient size of the training set. The natural way for an intelligent learner to counter this problem and successfully generalize is to exploit prior information that may be available about the domain or that can be learned from prototypical examples. We discuss the notion of using prior knowledge by creating virtual examples and thereby expanding the effective training-set size. We show that in some contexts this idea is mathematically equivalent to incorporating the prior knowledge as a regularizer, suggesting that the strategy is well motivated. The process of creating virtual examples in real-world pattern recognition tasks is highly nontrivial. We provide demonstrative examples from object recognition and speech recognition to illustrate the idea.
Partha Niyogi, Federico Girosi, Tomaso A. Poggio
Proc. IEEE3
1998 Guest Editorial Applications Of Artificial Neural Networks To Image Processing
Rama Chellappa, Kunihiko Fukushima, Aggelos K. Katsaggelos, Sun-Yuan Kung, Yann LeCun, Nasser M. Nasrabadi, Tomaso A. Poggio
IEEE Trans. Image Process.7
1997 Pedestrian Detection Using Wavelet Templates
abstract
This paper presents a trainable object detection architecture that is applied to detecting people in static images of cluttered scenes. This problem poses several challenges. People are highly non-rigid objects with a high degree of variability in size, shape, color, and texture. Unlike previous approaches, this system learns from examples and does not rely on any a priori (hand-crafted) models or on motion. The detection technique is based on the novel idea of the wavelet template that defines the shape of an object in terms of a subset of the wavelet coefficients of the image. It is invariant to changes in color and texture and can be used to robustly define a rich and complex class of objects such as people. We show how the invariant properties and computational efficiency of the wavelet template make it an effective tool for object detection.
Michael Oren, Constantine Papageorgiou, Pawan Sinha, Edgar Osuna, Tomaso A. Poggio
CVPR5
1997 A bootstrapping algorithm for learning linear models of object classes
abstract
Flexible models of object classes, based on linear combinations of prototypical images, are capable of matching novel images of the same class and have been shown to be a powerful tool to solve several fundamental vision tasks such as recognition, synthesis and correspondence. The key problem in creating a specific flexible model is the computation of pixelwise correspondence between the prototypes, a task done until now in a semiautomatic way. In this paper we describe an algorithm that automatically bootstraps the correspondence between the prototypes. The algorithm -which can be used for 2D images as well as for 3D models-is shown to synthesize successfully a flexible model of frontal face images and a flexible model of handwritten digits.
Thomas Vetter, Michael J. Jones 0001, Tomaso A. Poggio
CVPR3
1997 Just One View: Invariances in Inferotemporal Cell Tuning
Maximilian Riesenhuber, Tomaso A. Poggio
NIPS2
1997 Image-based view synthesis by combining trilinear tensors and learning techniques
abstract
We present a new method for rendering novel images of flexible 3D objects from a small number of example images in correspondence.The strength of the method is the ability to synthesize images whose viewing position is significantly far away from the viewing cone of the example images ("view extrapolation"), yet without ever modeling the 3D structure of the scene.The method relies on synthesizing a chain of "trilinear tensors" that govems the warping function from the example images to the novel image, together with a multi-dimensional interpolation function that synthesizes the non-rigidmotions of the viewed object from the virtual camera position.We show that two closely spaced example images alone are sufficient in practice to synthesize a significant viewing cone, thus demonstrating the ability of representing an object by a relatively small number of model images -for the purpose of cheap and fast viewers that can run on standard hardware.
Shai Avidan, Theodoros Evgeniou, Amnon Shashua, Tomaso A. Poggio
VRST4
1997 Linear Object Classes and Image Synthesis From a Single Example Image
abstract
The need to generate new views of a 3D object from a single real image arises in several fields, including graphics and object recognition. While the traditional approach relies on the use of 3D models, simpler techniques are applicable under restricted conditions. The approach exploits image transformations that are specific to the relevant object class, and learnable from example views of other "prototypical" objects of the same class. In this paper, we introduce such a technique by extending the notion of linear class proposed by the authors (1992). For linear object classes, it is shown that linear transformations can be learned exactly from a basis set of 2D prototypical views. We demonstrate the approach on artificial objects and then show preliminary evidence that the technique can effectively "rotate" high-resolution face images from a single 2D view.
Thomas Vetter, Tomaso A. Poggio
IEEE Trans. Pattern Anal. Mach. Intell.2
1997 Template matching: matched spatial filters and beyond
Roberto Brunelli, Tomaso A. Poggio
Pattern Recognit.2
1996 Image Synthesis from a Single Example Image
Thomas Vetter, Tomaso A. Poggio
ECCV (1)2
1996 Facial Analysis and Synthesis Using Image-Based Models
abstract
In this paper we describe image-based modeling techniques that make possible the creation of photo-realistic computer models of real human faces. The image-based model is built using example views of the face, bypassing the need for any three-dimensional computer graphics models. A learning network is trained to associate each of the example images with a set of pose and expression parameters. For a novel set of parameters, the network synthesizes a novel, intermediate view using a morphing approach. This image-based synthesis paradigm can adequately model both rigid and non-rigid facial movements. We also describe an analysis-by-synthesis algorithm, which is capable of extracting a set of high-level parameters from an image sequence involving facial movement using embedded image-based models. The parameters of the models are perturbed in a local and independent manner for each image until a correspondence-based error metric is minimized. A small sample of experimental results is presented.
Tony Ezzat, Tomaso A. Poggio
FG2
1996 3D Object Recognition: A Model of View-Tuned Neurons
Emanuela Bricolo, Tomaso A. Poggio, Nikos K. Logothetis
NIPS2
1995 Finding Human Faces with a Gaussian Mixture Distribution-Based Face Model
Tomaso A. Poggio, Kah Kay Sung
ACCV1
1995 Learning Human Face Detection in Cluttered Scenes
Kah Kay Sung, Tomaso A. Poggio
CAIP2
1995 Face Recognition from One Example View
abstract
To create a pose-invariant face recognizer, one strategy is the view-based approach, which uses a set of real example views at different poses. But what if we only have one real view available, such as a scanned passport photo-can we still recognize faces under different poses? Given one real view at a known pose, it is still possible to use the view-based approach by exploiting prior knowledge of faces to generate virtual views, or views of the face as seen from different poses. To represent prior knowledge, we use 2D example views of prototype faces under different rotations. We develop example-based techniques for applying the rotation seen in the prototypes to essentially "rotate" the single real view which is available. Next, the combined set of one real and multiple virtual views is used as example views for a view-based, pose-invariant face recognizer. Oar experiments suggest that among the techniques for expressing prior knowledge of faces, 2D example-based approaches should be considered alongside the more standard 3D modeling techniques.>
David Beymer, Tomaso A. Poggio
ICCV2
1995 Model-Based Matching of Line Drawings by Linear Combinations of Prototypes
abstract
We describe a technique for finding pixelwise correspondences between two images by using models of objects of the same class to guide the search. The object models are "learned" from example images (also called prototypes) of an object class. The models consist of a linear combination of prototypes. The flow fields giving pixelwise correspondences between a base prototype and each of the other prototypes must be given. A novel image of an object of the same class is matched to a model by minimizing an error between the novel image and the current guess for the closest model image. Currently, the algorithm applies to line drawings of objects. An extension to real grey level images is discussed.>
Michael J. Jones 0001, Tomaso A. Poggio
ICCV2
1995 Optical flow from 1-D correlation: Application to a simple time-to-crash detector
Nicola Ancona, Tomaso A. Poggio
Int. J. Comput. Vis.2
1995 Automatic person recognition by acoustic and geometric features
Roberto Brunelli, Daniele Falavigna, Tomaso A. Poggio, Luigi Stringa
Mach. Vis. Appl.3
1995 Regularization Theory and Neural Networks Architectures
abstract
We had previously shown that regularization principles lead to approximation schemes that are equivalent to networks with one layer of hidden units, called regularization networks. In particular, standard smoothness functionals lead to a subclass of regularization networks, the well known radial basis functions approximation schemes. This paper shows that regularization networks encompass a much broader range of approximation schemes, including many of the popular general additive models and some of the neural networks. In particular, we introduce new classes of smoothness functionals that lead to different classes of basis functions. Additive splines as well as some tensor product splines can be obtained from appropriate classes of smoothness functionals. Furthermore, the same generalization that extends radial basis functions (RBF) to hyper basis functions (HBF) also leads from additive models to ridge approximation models, containing as special cases Breiman's hinge functions, some forms of projection pursuit regression, and several types of neural networks. We propose to use the term generalized regularization networks for this broad class of approximation schemes that follow from an extension of regularization. In the probabilistic interpretation of regularization, the different classes of basis functions correspond to different classes of prior probabilities on the approximating function spaces, and therefore to different types of smoothness assumptions. In summary, different multilayer networks with one hidden layer, which we collectively call generalized regularization networks, correspond to different classes of priors and associated smoothness functionals in a classical regularization principle. Three broad classes are (1) radial basis functions that can be generalized to hyper basis functions, (2) some tensor product splines, and (3) additive splines that can be generalized to schemes of the type of ridge approximation, hinge functions, and several perceptron-like neural networks with one hidden layer.
Federico Girosi, Michael J. Jones 0001, Tomaso A. Poggio
Neural Comput.3
1994 Introduction to the Special Section on Learning in Computer Vision
Bir Bhanu, Tomaso A. Poggio
IEEE Trans. Pattern Anal. Mach. Intell.2
1993 Optical flow from 1D correlation: Application to a simple time-to-crash detector
abstract
The authors show that a new technique exploiting 1-D correlation of 2-D or even 1-D patches between successive frames may be sufficient to compute a satisfactory estimation of the optical flow field. The algorithm is well-suited to VLSI implementation. The sparse measurements provided by the technique can be used to compute qualitative properties of the flow for a number of different visual tasks. In particular, a method is described to combine the 1-D correlation technique with a scheme for detecting expansion or rotation (T. Poggio et al., 1991) in a simple algorithm which also suggests interesting biological implications. The algorithm provides a rough estimate of time-to-crash. It was tested on real image sequences. Its performance is discussed and the results are compared to previous approaches.>
Nicola Ancona, Tomaso A. Poggio
ICCV2
1993 Face Recognition: Features Versus Templates
abstract
Two new algorithms for computer recognition of human faces, one based on the computation of a set of geometrical features, such as nose width and length, mouth position, and chin shape, and the second based on almost-gray-level template matching, are presented. The results obtained for the testing sets show about 90% correct recognition using geometrical features and perfect recognition using template matching.>
Roberto Brunelli, Tomaso A. Poggio
IEEE Trans. Pattern Anal. Mach. Intell.2
1993 Report on Workshop on High Performance Computing and Communications for Grand Challenge Applications: Computer Vision, Speech and Natural Language Processing, and Artificial Intelligence
abstract
The findings of a workshop, the goals of which were to identify applications, research problems, and designs of high performance computing and communications (HPCC) systems for supporting applications are discussed. In computer vision, the main scientific issues are machine learning, surface reconstruction, inverse optics and integration, model acquisition, and perception and action. In speech and natural language processing (SNLP), issues were identified statistical analysis in corpus-based speech and language understanding, search strategies for language analysis, auditory and vocal-tract modeling, integration of multiple levels of speech and language analyses, and connectionist systems. In AI, important issues that need immediate attention include the development of efficient machine learning and heuristic search methods that can adapt to different architectural configurations, and the design and construction of scalable and verifiable knowledge bases, active memories, and artificial neural networks.>
Benjamin W. Wah, Thomas S. Huang, Aravind K. Joshi, Dan I. Moldovan, Yiannis Aloimonos, Ruzena Bajcsy, Dana H. Ballard, Doug DeGroot, Kenneth A. De Jong, Charles R. Dyer, Scott E. Fahlman, Ralph Grishman, Lynette Hirschman, Richard E. Korf, Stephen E. Levinson, Daniel P. Miranker, N. H. Morgan, Sergei Nirenburg, Tomaso A. Poggio, Edward M. Riseman, Craig Stanfil, Salvatore J. Stolfo, Steven L. Tanimoto, Charles C. Weems
IEEE Trans. Knowl. Data Eng.19
1992 Face Recognition through Geometrical Features
Roberto Brunelli, Tomaso A. Poggio
ECCV2
1992 Learning of visual modules from examples: A framework for understanding adaptive visual performance
Tomaso A. Poggio, Shimon Edelman, Manfred Fahle
CVGIP Image Underst.1
1992 Analog VLSI systems for image acquisition and fast early vision processing
John L. Wyatt Jr., Craig L. Keast, Mark Seidel, David L. Standley, Berthold K. P. Horn, Tom Knight, Charles G. Sodini, Hae-Seung Lee, Tomaso A. Poggio
Int. J. Comput. Vis.9
1992 Bringing the Grandmother back into the Picture: A Memory-Based View of Object Recognition
abstract
We describe experiments with a versatile pictorial prototype-based learning scheme for 3-D object recognition. The Generalized Radial Basis Function (GRBF) scheme seems to be amenable to realization in biophysical hardware because the only kind of computation it involves can be effectively carried out by combining receptive fields. Furthermore, the scheme is computationally attractive because it brings together the old notion of a "grandmother" cell and the rigorous approximation methods of regularization and splines.
Shimon Edelman, Tomaso A. Poggio
Int. J. Pattern Recognit. Artif. Intell.2
1991 HyperBF Networks for Real Object Recognition
Roberto Brunelli, Tomaso A. Poggio
IJCAI2
1990 Extensions of a Theory of Networks for Approximation and Learning
Federico Girosi, Tomaso A. Poggio, Bruno Caprile
NIPS2
1990 Networks for approximation and learning
abstract
The problem of the approximation of nonlinear mapping, (especially continuous mappings) is considered. Regularization theory and a theoretical framework for approximation (based on regularization techniques) that leads to a class of three-layer networks called regularization networks are discussed. Regularization networks are mathematically related to the radial basis functions, mainly used for strict interpolation tasks. Learning as approximation and learning as hypersurface reconstruction are discussed. Two extensions of the regularization approach are presented, along with the approach's corrections to splines, regularization, Bayes formulation, and clustering. The theory of regularization networks is generalized to a formulation that includes task-dependent clustering and dimensionality reduction. Applications of regularization networks are discussed.>
Tomaso A. Poggio, Federico Girosi
Proc. IEEE1
1989 Representation Properties of Networks: Kolmogorov's Theorem Is Irrelevant
abstract
Many neural networks can be regarded as attempting to approximate a multivariate function in terms of one-input one-output units. This note considers the problem of an exact representation of nonlinear mappings in terms of simpler functions of fewer variables. We review Kolmogorov's theorem on the representation of functions of several variables in terms of functions of one variable and show that it is irrelevant in the context of networks for learning.
Federico Girosi, Tomaso A. Poggio
Neural Comput.2
1989 Motion Field and Optical Flow: Qualitative Properties
abstract
It is shown that the motion field the 2-D vector field which is the perspective projection on the image plane of the 3-D velocity field of a moving scene, and the optical flow, defined as the estimate of the motion field which can be derived from the first-order variation of the image brightness pattern, are in general different, unless special conditions are satisfied. Therefore, dense optical flow is often ill-suited for computing structure from motion and for reconstructing the 3-D velocity field by algorithms which require a locally accurate estimate of the motion field. A different use of the optical flow is suggested. It is shown that the (smoothed) optical flow and the motion field can be interpreted as vector fields tangent to flows of planar dynamical systems. Stable qualitative properties of the motion field, which give useful informations about the 3-D velocity field and the 3-D structure of the scene, usually can be obtained from the optical flow. The idea is supported by results from the theory of structural stability of dynamical systems.>
Alessandro Verri, Tomaso A. Poggio
IEEE Trans. Pattern Anal. Mach. Intell.2
1989 Integration of vision modules and labeling of surface discontinuities
abstract
It is assumed that a major goal of the early vision modules and their integration is to deliver a cartoon of the discontinuities in the scene and to label them in terms of their physical origin. The output of each of the vision modules is noisy, possibly sparse, and sometimes not unique. The authors suggest the use of a coupled Markov random field (MRF) at the output of each module (image cues)-stereo, motion, color, and texture-to achieve two goals: first, to counteract the noise and fill in sparse data, and secondly, to integrate the image within each MRF to find the module discontinuities and align them with the intensity edges. The authors outline a theory of how to label the discontinuities in terms of depth, orientation, albedo, illumination, and specular discontinuities. They present labeling results using a simple linear classifier operating on the output of the MRF associated with each vision module and coupled to the image data. The classifier has been trained on a small set of a mixture of synthetic and real data.>
Edward B. Gamble, Davi Geiger, Tomaso A. Poggio, Daphna Weinshall
IEEE Trans. Syst. Man Cybern.3
1988 Parallel Optical Flow Using Local Voting
abstract
We describe a parallel algorithm for computing optical flow from short-range motion. Regularizing optical flow computation leads to a forruulation which minimizes matching error and, at the same time, maximises smoothness of the optical flow. We develop an approximation to the full regularization computation in which corresponding points are found by comparing local patches of the images. Selection aniong competing matches is performed using a winner-take-all scheme. The algorithm accommodates many different image transformations uniformly, with siniilar results, from brightness to edges. The optical flow computed froni different image transformations, such as edge detection and direct brightness computation, can be simply combined. The algorithm is easily implemented using local operations on a finegrained computer, and has been implemented on a Connection Machine. Experiments with natural images show that the scheme is effective and robust against noise. The algorithm leads to dense optical flow fields ; in addition, inforniation from matching facilitates segmentation.
James J. Little, Heinrich H. Bülthoff, Tomaso A. Poggio
ICCV3
1988 A Network for Image Segmentation Using Color
Anya C. Hurlbert, Tomaso A. Poggio
NIPS2
1988 A regularized solution to edge detection
Tomaso A. Poggio, Harry Voorhees, Alan L. Yuille
J. Complex.1
1988 Representation properties of multilayer feedforward networks
Tomaso A. Poggio
Neural Networks2
1988 Learning, regularization and splines
Tomaso A. Poggio
Neural Networks1
1988 Ill-posed problems in early vision
abstract
Mathematical results on ill-posed and ill-conditioned problems are reviewed and the formal aspects of regularization theory in the linear case are introduced. Specific topics in early vision and their regularization are then analyzed rigorously, characterizing existence, uniqueness, and stability of solutions. A fundamental difficulty that arises in almost every vision problem is scale, that is, the resolution at which to operate. Methods that have been proposed to deal with the problem include scale-space techniques that consider the behavior of the result across a continuum of scales. From the point of view of regulation theory, the concept of scale is related quite directly to the regularization parameter lambda . It suggested that methods used to obtained the optimal value of lambda may provide, either directly or after suitable modification, the optimal scale associated with the specific instance of certain problems.>
Mario Bertero, Tomaso A. Poggio, Vincent Torre
Proc. IEEE2
1987 An Optimal Scale for Edge Detection
Davi Geiger, Tomaso A. Poggio
IJCAI2
1987 Learning a Color Algorithm from Examples
Tomaso A. Poggio, Anya C. Hurlbert
NIPS1
1986 On parallel stereo
abstract
We review some of the open issues in computational stereo. In particular, we will discuss the problem of extracting better matching primitives and of dealing with occlusions. Markov Random Field models - an extension of standard regularization - suggest sophisticated stereo matching algorithms. They are, however, ill-suited to efficient, real-time applications. We will conclude reviewing a new simple but fast algorithm implemented by one of us (Drumheller, 1986) on the TMC Connection Machine (TM) computer. Some of its features are: (a) the potential for combining different primitives, including color information; (b) the use of a stronger and new formulation of the uniqueness constraint; and (c) its disparity representation that maps efficiently into the architecture of the Connection Machine computer.
Michael Drumheller, Tomaso A. Poggio
ICRA2
1986 On Edge Detection
abstract
Edge detection is the process that attempts to characterize the intensity changes in the image in terms of the physical processes that have originated them. A critical, intermediate goal of edge detection is the detection and characterization of significant intensity changes. This paper discusses this part of the edge detection problem. To characterize the types of intensity changes derivatives of different types, and possibly different scales, are needed. Thus, we consider this part of edge detection as a problem in numerical differentiation. We show that numerical differentiation of images is an ill-posed problem in the sense of Hadamard. Differentiation needs to be regularized by a regularizing filtering operation before differentiation. This shows that this part of edge detection consists of two steps, a filtering step and a differentiation step. Following this perspective, the paper discusses in detail the following theoretical aspects of edge detection. 1) The properties of different types of filters-with minimal uncertainty, with a bandpass spectrum, and with limited support-are derived. Minimal uncertainty filters optimize a tradeoff between computational efficiency and regularizing properties. 2) Relationships among several 2-D differential operators are established. In particular, we characterize the relation between the Laplacian and the second directional derivative along the gradient. Zero crossings of the Laplacian are not the only features computed in early vision. 3) Geometrical and topological properties of the zero crossings of differential operators are studied in terms of transversality and Morse theory.
Vincent Torre, Tomaso A. Poggio
IEEE Trans. Pattern Anal. Mach. Intell.2
1986 Scaling Theorems for Zero Crossings
abstract
We characterize some properties of the zero crossings of the Laplacian of signals¿in particular images¿filtered with linear filters, as a function of the scale of the filter (extending recent work by Witkin [16]). We prove that in any dimension the only filter that does not create generic zero crossings as the scale increases is the Gaussian. This result can be generalized to apply to level crossings of any linear differential operator: it applies in particular to ridges and ravines in the image intensity. In the case of the second derivative along the gradient, there is no filter that avoids creation of zero crossings, unless the filtering is performed after the derivative is applied.
Alan L. Yuille, Tomaso A. Poggio
IEEE Trans. Pattern Anal. Mach. Intell.2
1985 Early vision: From computational structure to algorithms and parallel hardware
Tomaso A. Poggio
Comput. Vis. Graph. Image Process.1
1984 Fingerprints Theorems
Alan L. Yuille, Tomaso A. Poggio
AAAI2