Aleix Martinez

dblp:m/AleixMMartinez · also Aleix M. Martinez · DBLP profile ↗
← Back
67ranked-venue papers
12as first author
9since 2021 · last 2025
0000-0002-4745-4953ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 61 · 10 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 2 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2025 Be More Specific: Evaluating Object-centric Realism in Synthetic Images
abstract
Evaluation of synthetic images is important for both model development and selection. An ideal evaluation should be specific, accurate and aligned with human perception. This paper addresses the problem of evaluating realism of objects in synthetic images. Although methods have been proposed to evaluate holistic realism, there are no methods tailored towards object-centric realism evaluation. In this work, we define a new standard for assessing object-centric realism that follows a shape-texture breakdown and proposes the first object-centric realism evaluation dataset for synthetic images. The dataset contains images generated from state-of-the-art image generative models and is richly annotated at object level across a diverse set of object categories. We then design and train the OLIP model, an architecture that considerably outperforms any existing baseline on object-centric realism evaluation.
Anqi Liang, Ciprian A. Corneanu, Qianli Feng, Giorgio Giannone, Aleix Martinez
CVPR5
2025 Structured Human Assessment of Text-to-Image Generative Models
abstract
Following the great progress in text-conditioned image generation there is a dire need for establishing clear com-parison benchmarks. Unfortunately, assessing performance of such models is highly subjective and notoriously difficult. Current automatic assessment of generated images quality and their alignment to text are approximate at best while human assessment is subjective, poorly calibrated and not very well defined. To address these concerns, we propose GenomeBench, a new framework for assessing quality of text-to-image generative models. It consists of a prompt dataset richly annotated with semantic components based on a formalized grounding of language and images. On top of it, we define a procedure to collect human assessment through a carefully guided question answering process. Fi-nally, these assessments are summarized into a novel score built around quality and alignment to text. We show the proposal achieves higher inter-annotator agreement with respect to the baseline human assessment and better cor-relation between quality and alignment compared to automatic assessment. Finally, we use this framework to dissect the performance of recent text-to-image models, providing insights on strength and weakness of each.
Ciprian A. Corneanu, Qianli Feng, Aleix Martinez
WACV3
2024 LatentPaint: Image Inpainting in Latent Space with Diffusion Models
abstract
Image inpainting using diffusion models is generally done using either preconditioned models, i.e. image conditioned models fine-tuned for the painting task, or postconditioned models, i.e. unconditioned models repurposed for the painting task at inference time. Preconditioned models are fast at inference time but extremely costly to train. Postconditioned models do not require any training but are slow during inference, requiring multiple forward and backward passes to converge to a desirable solution. Here, we derive an approach that does not require expensive training, yet is fast at inference time. To solve the costly inference computational time, we perform the forward-backward fusion step on a latent space rather than the image space. This is solved with a newly proposed propagation module in the diffusion process. Experiments on a number of domains demonstrate our approach attains or improves state-of-the-art results with the advantages of preconditioned and postconditioned models and none of their disadvantages.
Ciprian A. Corneanu, Raghudeep Gadde, Aleix Martinez
WACV3
2023 Network-Free, Unsupervised Semantic Segmentation with Synthetic Images
abstract
We derive a method that yields highly accurate semantic segmentation maps without the use of any additional neural network, layers, manually annotated training data, or supervised training. Our method is based on the observation that the correlation of a set of pixels belonging to the same semantic segment do not change when generating synthetic variants of an image using the style mixing approach in GANs. We show how we can use GAN inversion to accurately semantically segment synthetic and real photos as well as generate large training image-semantic segmentation mask pairs for downstream tasks.
Qianli Feng, Raghudeep Gadde, Wentong Liao, Eduard Ramon, Aleix Martinez
CVPR5
2022 Rayleigh EigenDirections (REDs): Nonlinear GAN Latent Space Traversals for Multidimensional Features
Guha Balakrishnan, Raghudeep Gadde, Aleix Martinez, Pietro Perona
ECCV (17)3
2021 When do GANs replicate? On the choice of dataset size
abstract
Do GANs replicate training images? Previous studies have shown that GANs do not seem to replicate training data without significant change in the training procedure. This leads to a series of research on the exact condition needed for GANs to overfit to the training data. Although a number of factors has been theoretically or empirically identified, the effect of dataset size and complexity on GANs replication is still unknown. With empirical evidence from BigGAN and StyleGAN2, on datasets CelebA, Flower and LSUN-bedroom, we show that dataset size and its complexity play an important role in GANs replication and perceptual quality of the generated images. We further quantify this relationship, discovering that replication percentage decays exponentially with respect to dataset size and complexity, with a shared decaying factor across GAN-dataset combinations. Meanwhile, the perceptual image quality follows a U-shape trend w.r.t dataset size. This finding leads to a practical tool for one-shot estimation on minimal dataset size to prevent GAN replication which can be used to guide datasets construction and selection.
Qianli Feng, Chenqi Guo, Fabian Benitez-Quiroz, Aleix Martinez
ICCV4
2021 Detail Me More: Improving GAN's photo-realism of complex scenes
abstract
Generative models can synthesize photo-realistic images of a single object. For example, for human faces, algorithms learn to model the local shape and shading of the face components, i.e., changes in the brows, eyes, nose, mouth, jaw line, etc. This is possible because all faces have two brows, two eyes, a nose and a mouth, approximately in the same location. The modeling of complex scenes is however much more challenging because the scene components and their location vary from image to image. For example, living rooms contain a varying number of products belonging to many possible categories and locations, e.g., a lamp may or may not be present in an endless number of possible locations. In the present work, we propose to add a "broker" module in Generative Adversarial Networks (GAN) to solve this problem. The broker is tasked to mediate the use of multiple discriminators in the appropriate image locales. For example, if a lamp is detected or wanted in a specific area of the scene, the broker assigns a fine-grained lamp discriminator to that image patch. This allows the generator to learn the shape and shading models of the lamp. The resulting multi-fine-grained optimization problem is able to synthesize complex scenes with almost the same level of photo-realism as single object images. We demonstrate the generability of the proposed approach on several GAN algorithms (BigGAN, ProGAN, StyleGAN, StyleGAN2), image resolutions (2562to 10242), and datasets. Our approach yields significant improvements over state-of-the-art GAN algorithms.
Raghudeep Gadde, Qianli Feng, Aleix Martinez
ICCV3
2021 Adding Knowledge to Unsupervised Algorithms for the Recognition of Intent
Stuart Synakowski, Qianli Feng, Aleix Martinez
Int. J. Comput. Vis.3
2021 Cross-Cultural and Cultural-Specific Production and Perception of Facial Expressions of Emotion in the Wild
abstract
Automatic recognition of emotion from facial expressions is an intense area of research, with a potentially long list of important application. Yet, the study of emotion requires knowing which facial expressions are used within and across cultures in the wild, not in controlled lab conditions; but such studies do not exist. Which and how many cross-cultural and cultural-specific facial expressions do people commonly use? And, what affect variables does each expression communicate to observers? If we are to design technology that understands the emotion of users, we need answers to these two fundamental questions. In this paper, we present the first large-scale study of the production and visual perception of facial expressions of emotion in the wild. We find that of the 16,384 possible facial configurations that people can theoretically produce, only 35 are successfully used to transmit emotive information across cultures, and only 8 within a smaller number of cultures. Crucially, we find that visual analysis of cross-cultural expressions yields consistent perception of emotion categories and valence, but not arousal. In contrast, visual analysis of cultural-specific expressions yields consistent perception of valence and arousal, but not of emotion categories. Additionally, we find that the number of expressions used to communicate each emotion is also different, e.g., 17 expressions transmit happiness, but only 1 is used to convey disgust.
Ramprakash Srinivasan, Aleix Martinez
IEEE Trans. Affect. Comput.2
2020 Computing the Testing Error Without a Testing Set
abstract
Deep Neural Networks (DNNs) have revolutionized computer vision. We now have DNNs that achieve top (accuracy) results in many problems, including object recognition, facial expression analysis, and semantic segmentation, to name but a few. The design of the DNNs that achieve top results is, however, non-trivial and mostly done by trail-and-error. That is, typically, researchers will derive many DNN architectures (\ie, topologies) and then test them on multiple datasets. However, there are no guarantees that the selected DNN will perform well in the real world. One can use a testing set to estimate the performance gap between the training and testing sets, but avoiding overfitting-to-the-testing-data is of concern. Using a sequestered testing data may address this problem, but this requires a constant update of the dataset, a very expensive venture. Here, we derive an algorithm to estimate the performance gap between training and testing without the need of a testing dataset. Specifically, we derive a set of persistent topology measures that identify when a DNN is learning to generalize to unseen samples. We provide extensive experimental validation on multiple networks and datasets to demonstrate the feasibility of the proposed approach.
Ciprian A. Corneanu, Sergio Escalera, Aleix Martinez
CVPR3
2020 Explainable Early Stopping for Action Unit Recognition
abstract
A common technique to avoid overfitting when training deep neural networks (DNN) is to monitor the performance in a dedicated validation data partition and to stop training as soon as it saturates. This only focuses on what the model does, while completely ignoring what happens inside it. In this work, we open the “black-box” of DNN in order to perform early stopping. We propose to use a novel theoretical framework that analyses meso-scale patterns in the topology of the functional graph of a network while it trains. Based on it, we decide when it transitions from learning towards overfitting in a more explainable way. We exemplify the benefits of this approach on a state-of-the art custom DNN that jointly learns local representations and label structure employing an ensemble of dedicated subnetworks. We show that it is practically equivalent in performance to early stopping with patience, the standard early stopping algorithm in the literature. This proves beneficial for AU recognition performance and provides new insights into how learning of AUs occurs in DNNs.
Ciprian A. Corneanu, Meysam Madadi, Sergio Escalera, Aleix Martinez
FG4
2020 GANimation: One-Shot Anatomically Consistent Facial Animation
Albert Pumarola, Antonio Agudo, Aleix Martinez, Alberto Sanfeliu, Francesc Moreno-Noguer
Int. J. Comput. Vis.3
2019 What Does It Mean to Learn in Deep Networks? And, How Does One Detect Adversarial Attacks?
abstract
The flexibility and high-accuracy of Deep Neural Networks (DNNs) has transformed computer vision. But, the fact that we do not know when a specific DNN will work and when it will fail has resulted in a lack of trust. A clear example is self-driving cars; people are uncomfortable sitting in a car driven by algorithms that may fail under some unknown, unpredictable conditions. Interpretability and explainability approaches attempt to address this by uncovering what a DNN models, i.e., what each node (cell) in the network represents and what images are most likely to activate it. This can be used to generate, for example, adversarial attacks. But these approaches do not generally allow us to determine where a DNN will succeed or fail and why. i.e., does this learned representation generalize to unseen samples? Here, we derive a novel approach to define what it means to learn in deep networks, and how to use this knowledge to detect adversarial attacks. We show how this defines the ability of a network to generalize to unseen testing samples and, most importantly, why this is the case.
Ciprian A. Corneanu, Meysam Madadi, Sergio Escalera, Aleix Martinez
CVPR4
2019 Discriminant Functional Learning of Color Features for the Recognition of Facial Action Units and Their Intensities
abstract
Color is a fundamental image feature of facial expressions. For example, when we furrow our eyebrows in anger, blood rushes in, turning some face areas red; or when one goes white in fear as a result of the drainage of blood from the face. Surprisingly, these image properties have not been exploited to recognize the facial action units (AUs) associated with these expressions. Herein, we present the first system to do recognition of AUs and their intensities using these functional color changes. These color features are shown to be robust to changes in identity, gender, race, ethnicity, and skin color. Specifically, we identify the chromaticity changes defining the transition of an AU from inactive to active and use an innovative Gabor transform-based algorithm to gain invariance to the timing of these changes. Because these image changes are given by functions rather than vectors, we use functional classifiers to identify the most discriminant color features of an AU and its intensities. We demonstrate that, using these discriminant color features, one can achieve results superior to those of the state-of-the-art. Finally, we define an algorithm that allows us to use the learned functional color representation in still images. This is done by learning the mapping between images and the identified functional color features in videos. Our algorithm works in realtime, i.e., 30 frames/second/CPU thread.
Carlos F. Benitez-Quiroz, Ramprakash Srinivasan, Aleix Martinez
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 Learning Facial Action Units From Web Images With Scalable Weakly Supervised Clustering
abstract
We present a scalable weakly supervised clustering approach to learn facial action units (AUs) from large, freely available web images. Unlike most existing methods (e.g., CNNs) that rely on fully annotated data, our method exploits web images with inaccurate annotations. Specifically, we derive a weakly-supervised spectral algorithm that learns an embedding space to couple image appearance and semantics. The algorithm has efficient gradient update, and scales up to large quantities of images with a stochastic extension. With the learned embedding space, we adopt rank-order clustering to identify groups of visually and semantically similar images, and re-annotate these groups for training AU classifiers. Evaluation on the 1 millon EmotioNet dataset demonstrates the effectiveness of our approach: (1) our learned annotations reach on average 91.3% agreement with human annotations on 7 common AUs, (2) classifiers trained with re-annotated images perform comparably to, sometimes even better than, its supervised CNN-based counterpart, and (3) our method offers intuitive outlier/noise pruning instead of forcing one annotation to every image. Code is available.
Kaili Zhao, Wen-Sheng Chu, Aleix Martinez
CVPR3
2018 GANimation: Anatomically-Aware Facial Animation from a Single Image
Albert Pumarola, Antonio Agudo, Aleix Martinez, Alberto Sanfeliu, Francesc Moreno-Noguer
ECCV (10)3
2018 Neural Model for the Visual Recognition of Animacy and Social Interaction
Mohammad Hovaidi-Ardestani, Nitin Saini, Aleix Martinez, Martin A. Giese
ICANN (3)3
2018 A Simple, Fast and Highly-Accurate Algorithm to Recover 3D Shape from 2D Landmarks on a Single Image
abstract
Three-dimensional shape reconstruction of 2D landmark points on a single image is a hallmark of human vision, but is a task that has been proven difficult for computer vision algorithms. We define a feed-forward deep neural network algorithm that can reconstruct 3D shapes from 2D landmark points almost perfectly (i.e., with extremely small reconstruction errors), even when these 2D landmarks are from a single image. Our experimental results show an improvement of up to two-fold over state-of-the-art computer vision algorithms; 3D shape reconstruction error (measured as the Procrustes distance between the reconstructed shape and the ground-truth) of human faces is , cars is .0022, human bodies is .022, and highly-deformable flags is .0004. Our algorithm was also a top performer at the 2016 3D Face Alignment in the Wild Challenge competition (done in conjunction with the European Conference on Computer Vision, ECCV) that required the reconstruction of 3D face shape from a single image. The derived algorithm can be trained in a couple hours and testing runs at more than 1,000 frames/s on an i7 desktop. We also present an innovative data augmentation approach that allows us to train the system efficiently with small number of samples. And the system is robust to noise (e.g., imprecise landmark points) and missing data (e.g., occluded or undetected landmark points).
Aleix Martinez
IEEE Trans. Pattern Anal. Mach. Intell.3
2017 Recognition of Action Units in the Wild with Deep Nets and a New Global-Local Loss
abstract
Most previous algorithms for the recognition of Action Units (AUs) were trained on a small number of sample images. This was due to the limited amount of labeled data available at the time. This meant that data-hungry deep neural networks, which have shown their potential in other computer vision problems, could not be successfully trained to detect AUs. A recent publicly available database with close to a million labeled images has made this training possible. Image and individual variability (e.g., pose, scale, illumination, ethnicity) in this set is very large. Unfortunately, the labels in this dataset are not perfect (i.e., they are noisy), making convergence of deep nets difficult. To harness the richness of this dataset while being robust to the inaccuracies of the labels, we derive a novel global-local loss. This new loss function is shown to yield fast globally meaningful convergences and locally accurate results. Comparative results with those of the EmotioNet challenge demonstrate that our newly derived loss yields superior recognition of AUs than state-of-the-art algorithms.
Carlos F. Benitez-Quiroz, Aleix Martinez
ICCV3
2016 EmotioNet: An Accurate, Real-Time Algorithm for the Automatic Annotation of a Million Facial Expressions in the Wild
abstract
Research in face perception and emotion theory requires very large annotated databases of images of facial expressions of emotion. Annotations should include Action Units (AUs) and their intensities as well as emotion category. This goal cannot be readily achieved manually. Herein, we present a novel computer vision algorithm to annotate a large database of one million images of facial expressions of emotion in the wild (i.e., face images downloaded from the Internet). First, we show that this newly proposed algorithm can recognize AUs and their intensities reliably across databases. To our knowledge, this is the first published algorithm to achieve highly-accurate results in the recognition of AUs and their intensities across multiple databases. Our algorithm also runs in real-time (>30 images/second), allowing it to work with large numbers of images and video sequences. Second, we use WordNet to download 1,000,000 images of facial expressions with associated emotion keywords from the Internet. These images are then automatically annotated with AUs, AU intensities and emotion categories by our algorithm. The result is a highly useful database that can be readily queried using semantic descriptions for applications in computer vision, affective computing, social and cognitive psychology and neuroscience, e.g., "show me all the images with happy faces" or "all images with AU 1 at intensity c".
Carlos F. Benitez-Quiroz, Ramprakash Srinivasan, Aleix Martinez
CVPR3
2016 Guest Editorial: Special Section on CVPR 2014
abstract
The papers in this special section were presented at the IEEE Computer Vision and Pattern Recognition (CVPR), June, 2014, jointly sponsored by the IEEE and the Computer Vision Foundation.
Ronen Basri, Cornelia Fermüller, Aleix Martinez, René Vidal
IEEE Trans. Pattern Anal. Mach. Intell.3
2016 Labeled Graph Kernel for Behavior Analysis
abstract
Automatic behavior analysis from video is a major topic in many areas of research, including computer vision, multimedia, robotics, biology, cognitive science, social psychology, psychiatry, and linguistics. Two major problems are of interest when analyzing behavior. First, we wish to automatically categorize observed behaviors into a discrete set of classes (i.e., classification). For example, to determine word production from video sequences in sign language. Second, we wish to understand the relevance of each behavioral feature in achieving this classification (i.e., decoding). For instance, to know which behavior variables are used to discriminate between the words apple and onion in American Sign Language (ASL). The present paper proposes to model behavior using a labeled graph, where the nodes define behavioral features and the edges are labels specifying their order (e.g., before, overlaps, start). In this approach, classification reduces to a simple labeled graph matching. Unfortunately, the complexity of labeled graph matching grows exponentially with the number of categories we wish to represent. Here, we derive a graph kernel to quickly and accurately compute this graph similarity. This approach is very general and can be plugged into any kernel-based classifier. Specifically, we derive a Labeled Graph Support Vector Machine (LGSVM) and a Labeled Graph Logistic Regressor (LGLR) that can be readily employed to discriminate between many actions (e.g., sign language concepts). The derived approach can be readily used for decoding too, yielding invaluable information for the understanding of a problem (e.g., to know how to teach a sign language). The derived algorithms allow us to achieve higher accuracy results than those of state-of-the-art algorithms in a fraction of the time. We show experimental results on a variety of problems and datasets, including multimodal data.
Aleix Martinez
IEEE Trans. Pattern Anal. Mach. Intell.2
2016 Multiple Ordinal Regression by Maximizing the Sum of Margins
abstract
Human preferences are usually measured using ordinal variables. A system whose goal is to estimate the preferences of humans and their underlying decision mechanisms requires to learn the ordering of any given sample set. We consider the solution of this ordinal regression problem using a support vector machine algorithm. Specifically, the goal is to learn a set of classifiers with common direction vectors and different biases correctly separating the ordered classes. Current algorithms are either required to solve a quadratic optimization problem, which is computationally expensive, or based on maximizing the minimum margin (i.e., a fixed-margin strategy) between a set of hyperplanes, which biases the solution to the closest margin. Another drawback of these strategies is that they are limited to order the classes using a single ranking variable (e.g., perceived length). In this paper, we define a multiple ordinal regression algorithm based on maximizing the sum of the margins between every consecutive class with respect to one or more rankings (e.g., perceived length and weight). We provide derivations of an efficient, easy-to-implement iterative solution using a sequential minimal optimization procedure. We demonstrate the accuracy of our solutions in several data sets. In addition, we provide a key application of our algorithms in estimating human subjects' ordinal classification of attribute associations to object categories. We show that these ordinal associations perform better than the binary one typically employed in the literature.
Onur C. Hamsici, Aleix Martinez
IEEE Trans. Neural Networks Learn. Syst.2
2014 Salient and non-salient fiducial detection using a probabilistic graphical model
Carlos F. Benitez-Quiroz, Samuel Rivera, Paulo F. U. Gotardo, Aleix Martinez
Pattern Recognit.4
2014 Minimizing Nearest Neighbor Classification Error for Nonparametric Dimension Reduction
abstract
In this brief, we show that minimizing nearest neighbor classification error (MNNE) is a favorable criterion for supervised linear dimension reduction (SLDR). We prove that MNNE is better than maximizing mutual information in the sense of being a proxy of the Bayes optimal criterion. Based on kernel density estimation, we derive a nonparametric algorithm for MNNE. Experiments on benchmark data sets show the superiority of MNNE over existing nonparametric SLDR methods.
Wei Bian 0003, Tianyi Zhou 0001, Aleix Martinez, George Baciu, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.3
2014 Multiobjective Optimization for Model Selection in Kernel Methods in Regression
abstract
Regression plays a major role in many scientific and engineering problems. The goal of regression is to learn the unknown underlying function from a set of sample vectors with known outcomes. In recent years, kernel methods in regression have facilitated the estimation of nonlinear functions. However, two major (interconnected) problems remain open. The first problem is given by the bias-versus-variance tradeoff. If the model used to estimate the underlying function is too flexible (i.e., high model complexity), the variance will be very large. If the model is fixed (i.e., low complexity), the bias will be large. The second problem is to define an approach for selecting the appropriate parameters of the kernel function. To address these two problems, this paper derives a new smoothing kernel criterion, which measures the roughness of the estimated function as a measure of model complexity. Then, we use multiobjective optimization to derive a criterion for selecting the parameters of that kernel. The goal of this criterion is to find a tradeoff between the bias and the variance of the learned function. That is, the goal is to increase the model fit while keeping the model complexity in check. We provide extensive experimental evaluations using a variety of problems in machine learning, pattern recognition, and computer vision. The results demonstrate that the proposed approach yields smaller estimation errors as compared with methods in the state of the art.
Di You, Carlos F. Benitez-Quiroz, Aleix Martinez
IEEE Trans. Neural Networks Learn. Syst.3
2012 Automatic selection of eye tracking variables in visual categorization for adults and infants
Samuel Rivera, Catherine A. Best, Hyungwook Yim, Aleix Martinez, Vladimir M. Sloutsky, Dirk Bernhardt-Walther
CogSci4
2012 Learning Spatially-Smooth Mappings in Non-Rigid Structure From Motion
Onur C. Hamsici, Paulo F. U. Gotardo, Aleix Martinez
ECCV (4)3
2012 A Model of the Perception of Facial Expressions of Emotion by Humans: Research Overview and Perspectives
Aleix Martinez, Shichuan Du
J. Mach. Learn. Res.1
2012 Learning deformable shape manifolds
Samuel Rivera, Aleix Martinez
Pattern Recognit.2
2011 Non-rigid structure from motion with complementary rank-3 spaces
abstract
Non-rigid structure from motion (NR-SFM) is a difficult, underconstrained problem in computer vision. This paper proposes a new algorithm that revises the standard matrix factorization approach in NR-SFM. We consider two alternative representations for the linear space spanned by a small number K of 3D basis shapes. As compared to the standard approach using general rank-3K matrix factors, we show that improved results are obtained by explicitly modeling K complementary spaces of rank-3. Our new method is positively compared to the state-of-the-art in NR-SFM, providing improved results on high-frequency deformations of both articulated and simpler deformable shapes. We also present an approach for NR-SFM with occlusion.
Paulo F. U. Gotardo, Aleix Martinez
CVPR2
2011 Kernel non-rigid structure from motion
abstract
Non-rigid structure from motion (NRSFM) is a difficult, underconstrained problem in computer vision. The standard approach in NRSFM constrains 3D shape deformation using a linear combination of K basis shapes; the solution is then obtained as the low-rank factorization of an input observation matrix. An important but overlooked problem with this approach is that non-linear deformations are often observed; these deformations lead to a weakened low-rank constraint due to the need to use additional basis shapes to linearly model points that move along curves. Here, we demonstrate how the kernel trick can be applied in standard NRSFM. As a result, we model complex, deformable 3D shapes as the outputs of a non-linear mapping whose inputs are points within a low-dimensional shape space. This approach is flexible and can use different kernels to build different non-linear models. Using the kernel trick, our model complements the low-rank constraint by capturing non-linear relationships in the shape coefficients of the linear model. The net effect can be seen as using non-linear dimensionality reduction to further compress the (shape) space of possible solutions.
Paulo F. U. Gotardo, Aleix Martinez
ICCV2
2011 Computing Smooth Time Trajectories for Camera and Deformable Shape in Structure from Motion with Occlusion
abstract
We address the classical computer vision problems of rigid and nonrigid structure from motion (SFM) with occlusion. We assume that the columns of the input observation matrix W describe smooth 2D point trajectories over time. We then derive a family of efficient methods that estimate the column space of W using compact parameterizations in the Discrete Cosine Transform (DCT) domain. Our methods tolerate high percentages of missing data and incorporate new models for the smooth time trajectories of 2D-points, affine and weak-perspective cameras, and 3D deformable shape. We solve a rigid SFM problem by estimating the smooth time trajectory of a single camera moving around the structure of interest. By considering a weak-perspective camera model from the outset, we directly compute euclidean 3D shape reconstructions without requiring postprocessing steps such as euclidean upgrade and bundle adjustment. Our results on real SFM data sets with high percentages of missing data compared positively to those in the literature. In nonrigid SFM, we propose a novel 3D shape trajectory approach that solves for the deformable structure as the smooth time trajectory of a single point in a linear shape space. A key result shows that, compared to state-of-the-art algorithms, our nonrigid SFM method can better model complex articulated deformation with higher frequency DCT components while still maintaining the low-rank factorization constraint. Finally, we also offer an approach for nonrigid SFM when W is presented with missing data.
Paulo F. U. Gotardo, Aleix Martinez
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Kernel Optimization in Discriminant Analysis
abstract
Kernel mapping is one of the most used approaches to intrinsically derive nonlinear classifiers. The idea is to use a kernel function which maps the original nonlinearly separable problem to a space of intrinsically larger dimensionality where the classes are linearly separable. A major problem in the design of kernel methods is to find the kernel parameters that make the problem linear in the mapped representation. This paper derives the first criterion that specifically aims to find a kernel representation where the Bayes classifier becomes linear. We illustrate how this result can be successfully applied in several kernel discriminant analysis algorithms. Experimental results, using a large number of databases and classifiers, demonstrate the utility of the proposed approach. The paper also shows (theoretically and experimentally) that a kernel version of Subclass Discriminant Analysis yields the highest recognition rates.
Di You, Onur C. Hamsici, Aleix Martinez
IEEE Trans. Pattern Anal. Mach. Intell.3
2010 Bayes optimal kernel discriminant analysis
abstract
Kernel methods provide an efficient mechanism to derive nonlinear algorithms. In classification problems as well as in feature extraction, kernel-based approaches map the originally nonlinearly separable data into a space of intrinsically much higher dimensionality where the data is linearly separable and can be readily classified with existing and efficient linear methods. For a given kernel function, the main challenge is to determine the parameters of the kernel which maps the original nonlinear problem to a linear one. This paper derives a Bayes optimal criterion for the selection of the kernel parameters in discriminant analysis. Our criterion selects the kernel parameters that maximize the (Bayes) classification accuracy in the kernel space. We also show how we can use the same criterion to do subclass selection in the kernel space for problems with multimodal class distributions. Extensive experimental evaluation demonstrates the superiority of the proposed criterion over the state of the art.
Di You, Aleix Martinez
CVPR2
2010 Rigid Structure from Motion from a Blind Source Separation Perspective
Jeff Fortuna, Aleix Martinez
Int. J. Comput. Vis.2
2010 Features versus Context: An Approach for Precise and Detailed Detection and Delineation of Faces and Facial Features
abstract
The appearance-based approach to face detection has seen great advances in the last several years. In this approach, we learn the image statistics describing the texture pattern (appearance) of the object class we want to detect, e.g., the face. However, this approach has had limited success in providing an accurate and detailed description of the internal facial features, i.e., eyes, brows, nose, and mouth. In general, this is due to the limited information carried by the learned statistical model. While the face template is relatively rich in texture, facial features (e.g., eyes, nose, and mouth) do not carry enough discriminative information to tell them apart from all possible background images. We resolve this problem by adding the context information of each facial feature in the design of the statistical model. In the proposed approach, the context information defines the image statistics most correlated with the surroundings of each facial component. This means that when we search for a face or facial feature, we look for those locations which most resemble the feature yet are most dissimilar to its context. This dissimilarity with the context features forces the detector to gravitate toward an accurate estimate of the position of the facial feature. Learning to discriminate between feature and context templates is difficult, however, because the context and the texture of the facial features vary widely under changing expression, pose, and illumination, and may even resemble one another. We address this problem with the use of subclass divisions. We derive two algorithms to automatically divide the training samples of each facial feature into a set of subclasses, each representing a distinct construction of the same facial component (e.g., closed versus open eyes) or its context (e.g., different hairstyles). The first algorithm is based on a discriminant analysis formulation. The second algorithm is an extension of the AdaBoost approach. We provide extensive experimental results using still images and video sequences for a total of 3,930 images. We show that the results are almost as good as those obtained with manual detection.
Liya Ding 0001, Aleix Martinez
IEEE Trans. Pattern Anal. Mach. Intell.2
2009 Support Vector Machines in face recognition with occlusions
abstract
Support vector machines (SVM) are one of the most useful techniques in classification problems. One clear example is face recognition. However, SVM cannot be applied when the feature vectors defining our samples have missing entries. This is clearly the case in face recognition when occlusions are present in the training and/or testing sets. When k features are missing in a sample vector of class 1, these define an affine subspace of k dimensions. The goal of the SVM is to maximize the margin between the vectors of class 1 and class 2 on those dimensions with no missing elements and, at the same time, maximize the margin between the vectors in class 2 and the affine subspace of class 1. This second term of the SVM criterion will minimize the overlap between the classification hyperplane and the subspace of solutions in class 1, because we do not know which values in this subspace a test vector can take. The hyperplane minimizing this overlap is obviously the one parallel to the missing dimensions. However, this condition is too restrictive, because its solution will generally contradict that obtained when maximizing the margin of the visible data. To resolve this problem, we define a criterion which minimizes the probability of overlap. The resulting optimization problem can be solved efficiently and we show how the global minimum of the error term is guaranteed under mild conditions. We provide extensive experimental results, demonstrating the superiority of the proposed approach over the state of the art.
Hongjun Jia, Aleix Martinez
CVPR2
2009 Active Appearance Models with Rotation Invariant Kernels
abstract
2D Active Appearance Models (AAM) and 3D Morphable Models (3DMM) are widely used techniques. AAM provide a fast fitting process, but may represent unwanted 3D transformations unless strictly constrained not to do so. The reverse is true for 3DMM. The two approaches also require of a pre-alignment of their 2D or 3D shapes before the modeling can be carried out which may lead to errors. Furthermore, current models are insufficient to represent nonlinear shape and texture variations. In this paper, we derive a new approach that can model nonlinear changes in examples without the need of a pre-alignment step. In addition, we show how the proposed approach carries the above mentioned advantages of AAM and 3DMM. To achieve this goal, we take advantage of the inherent properties of complex spherical distributions, which provide invariance to translation, scale and rotation. To reduce the complexity of parameter estimation we take advantage of a recent result that shows how to estimate spherical distributions using their Euclidean counterpart, e.g., the Gaussians. This leads to the definition of Rotation Invariant Kernels (RIK) for modeling nonlinear shape changes. We show the superiority of our algorithm to AAM in several face datasets. We also show how the derived algorithm can be used to model complex 3D facial expression changes observed in American Sign Language (ASL).
Onur C. Hamsici, Aleix Martinez
ICCV2
2009 Modelling and recognition of the linguistic components in American Sign Language
Liya Ding 0001, Aleix Martinez
Image Vis. Comput.2
2009 Rotation Invariant Kernels and Their Application to Shape Analysis
abstract
Shape analysis requires invariance under translation, scale, and rotation. Translation and scale invariance can be realized by normalizing shape vectors with respect to their mean and norm. This maps the shape feature vectors onto the surface of a hypersphere. After normalization, the shape vectors can be made rotational invariant by modeling the resulting data using complex scalar-rotation invariant distributions defined on the complex hypersphere, e.g., using the complex Bingham distribution. However, the use of these distributions is hampered by the difficulty in estimating their parameters and the nonlinear nature of their formulation. In the present paper, we show how a set of kernel functions that we refer to as rotation invariant kernels can be used to convert the original nonlinear problem into a linear one. As their name implies, these kernels are defined to provide the much needed rotation invariance property allowing one to bypass the difficulty of working with complex spherical distributions. The resulting approach provides an easy, fast mechanism for 2D & 3D shape analysis. Extensive validation using a variety of shape modeling and classification problems demonstrates the accuracy of this proposed approach.
Onur C. Hamsici, Aleix Martinez
IEEE Trans. Pattern Anal. Mach. Intell.2
2009 Low-Rank Matrix Fitting Based on Subspace Perturbation Analysis with Applications to Structure from Motion
abstract
The task of finding a low-rank (r) matrix that best fits an original data matrix of higher rank is a recurring problem in science and engineering. The problem becomes especially difficult when the original data matrix has some missing entries and contains an unknown additive noise term in the remaining elements. The former problem can be solved by concatenating a set of r-column matrices that share a common single r-dimensional solution space. Unfortunately, the number of possible submatrices is generally very large and, hence, the results obtained with one set of r-column matrices will generally be different from that captured by a different set. Ideally, we would like to find that solution that is least affected by noise. This requires that we determine which of the r-column matrices (i.e., which of the original feature points) are less influenced by the unknown noise term. This paper presents a criterion to successfully carry out such a selection. Our key result is to formally prove that the more distinct the r vectors of the r-column matrices are, the less they are swayed by noise. This key result is then combined with the use of a noise model to derive an upper bound for the effect that noise and occlusions have on each of the r-column matrices. It is shown how this criterion can be effectively used to recover the noise-free matrix of rank r. Finally, we derive the affine and projective structure-from-motion (SFM) algorithms using the proposed criterion. Extensive validation on synthetic and real data sets shows the superiority of the proposed approach over the state of the art.
Hongjun Jia, Aleix Martinez
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Precise detailed detection of faces and facial features
abstract
Face detection has advanced dramatically over the past three decades. Algorithms can now quite reliably detect faces in clutter in or near real time. However, much still needs to be done to provide an accurate and detailed description of external and internal features. This paper presents an approach to achieve this goal. Previous learning algorithms have had limited success on this task because the shape and texture of facial features varies widely under changing expression, pose and illumination. We address this problem with the use of subclass divisions. In this approach, we use an algorithm to automatically divide the training samples of each facial feature into a set of subclasses, each representing a distinct construction of the same facial component (e.g., close versus open eye lids). The key idea used to achieve accurate detections is to not only learn the textural information of the facial feature to be detected but that of its context (i.e., surroundings). This process permits a precise detection of key facial features. We then combine this approach with edge and color segmentation to provide an accurate and detailed detection of the shape of the major facial features (brows, eyes, nose, mouth and chin). We use this face detection algorithm to obtain precise descriptions of the facial features in video sequences of American Sign Language (ASL) sentences, where the variability in expressions can be extreme. Extensive experimental validation demonstrates our method is almost as precise as manual detection, ~ 2% error.
Liya Ding 0001, Aleix Martinez
CVPR2
2008 Face recognition with occlusions in the training and testing sets
abstract
Partial occlusions in face images pose a great problem for most face recognition algorithms. Several solutions to this problem have been proposed over the years - ranging from dividing the face image into a set of local regions to sophisticated statistical methods. In the present paper, we pose the problem as a reconstruction one. In this approach, each test image is described as a linear combination of the training samples in each class. The class samples providing the best reconstruction determine the class label. Here, ldquobest reconstructionrdquo means that reconstruction providing the smallest matching error when using an appropriate metric to compare the reconstructed and test images. A key point in our formulation is to base this reconstruction solely on the visible data in the training and testing sets. This allows to have partial occlusions in both the training and testing samples, while previous methods only dealt with occlusions in the testing set. We show extensive experimental results using a large variety of comparative studies, demonstrating the superiority of the proposed approach over the state of the art.
Hongjun Jia, Aleix Martinez
FG2
2008 Using the information embedded in the testing sample to break the limits caused by the small sample size in microarray-based classification
abstract
BACKGROUND: Microarray-based tumor classification is characterized by a very large number of features (genes) and small number of samples. In such cases, statistical techniques cannot determine which genes are correlated to each tumor type. A popular solution is the use of a subset of pre-specified genes. However, molecular variations are generally correlated to a large number of genes. A gene that is not correlated to some disease may, by combination with other genes, express itself. RESULTS: In this paper, we propose a new classiification strategy that can reduce the effect of over-fitting without the need to pre-select a small subset of genes. Our solution works by taking advantage of the information embedded in the testing samples. We note that a well-defined classification algorithm works best when the data is properly labeled. Hence, our classification algorithm will discriminate all samples best when the testing sample is assumed to belong to the correct class. We compare our solution with several well-known alternatives for tumor classification on a variety of publicly available data-sets. Our approach consistently leads to better classification results. CONCLUSION: Studies indicate that thousands of samples may be required to extract useful statistical information from microarray data. Herein, it is shown that this problem can be circumvented by using the information embedded in the testing samples.
Manli Zhu, Aleix Martinez
BMC Bioinform.2
2008 Bayes Optimality in Linear Discriminant Analysis
abstract
We present an algorithm which provides the one-dimensional subspace where the Bayes error is minimized for the C class problem with homoscedastic Gaussian distributions. Our main result shows that the set of possible one-dimensional spaces v, for which the order of the projected class means is identical, defines a convex region with associated convex Bayes error function g(v). This allows for the minimization of the error function using standard convex optimization algorithms. Our algorithm is then extended to the minimization of the Bayes error in the more general case of heteroscedastic distributions. This is done by means of an appropriate kernel mapping function. This result is further extended to obtain the d-dimensional solution for any given d, by iteratively applying our algorithm to the null space of the (d - 1)-dimensional solution. We also show how this result can be used to improve up on the outcomes provided by existing algorithms, and derive a low-computational cost, linear approximation. Extensive experimental validations are provided to demonstrate the use of these algorithms in classification, data analysis and visualization.
Onur C. Hamsici, Aleix Martinez
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Who is LB1? Discriminant analysis for the classification of specimens
Aleix Martinez, Onur C. Hamsici
Pattern Recognit.1
2008 Pruning Noisy Bases in Discriminant Analysis
abstract
The success of linear discriminant analysis (LDA) is due in part to the simplicity of its formulation, which reduces to a simultaneous diagonalization of two symmetric matrices A and B;. However, a fundamental drawback of this approach is that it cannot be efficiently applied wherever the matrix A is singular or when some of the smallest variances in A are due to noise. In this paper, we present a factorization of A(-1) and a correlation-based criterion that can be readily employed to solve these problems. We provide detailed derivations for the linear and nonlinear classification problems. The usefulness of the proposed approach is demonstrated thoroughly using a large variety of databases.
Manli Zhu, Aleix Martinez
IEEE Trans. Neural Networks2
2007 Recovering the linguistic components of the manual signs in American Sign Language
abstract
Manual signs in American sign language (ASL) are constructed using three building blocks -handshape, motion, and place of articulations. Only when these three are successfully estimated, can a sign by uniquely identified. Hence, the use of pattern recognition techniques that use only a subset of these is inappropriate. To achieve accurate classifications, the motion, the handshape and their three-dimensional position need to be recovered. In this paper, we define an algorithm to determine these three components form a single video sequence of two-dimensional pictures of a sign. We demonstrated the use of our algorithm in describing and recognizing a set of manual signs in ASL.
Liya Ding 0001, Aleix Martinez
AVSS2
2007 Sparse Kernels for Bayes Optimal Discriminant Analysis
abstract
Discriminant Analysis (DA) methods have demonstrated their utility in countless applications in computer vision and other areas of research - especially in the C class classification problem. The most popular approach is linear DA (LDA), which provides the C - 1-dimensional Bayes optimal solution, but only when all the class covariance matrices are identical. This is rarely the case in practice. To alleviate this restriction, Kernel LDA (KLDA) has been proposed. In this approach, we first (intrinsically) map the original nonlinear problem to a linear one and then use LDA to find the C - 1-dimensional Bayes optimal subspace. However, the use of KLDA is hampered by its computational cost, given by the number of training samples available and by the limitedness of LDA in providing a C - 1-dimensional solution space. In this paper, we first extend the definition of LDA to provide subspace of q < C - 1 dimensions where the Bayes error is minimized. Then, to reduce the computational burden of the derived solution, we define a sparse kernel representation, which is able to automatically select the most appropriate sample feature vectors that represent the kernel. We demonstrate the superiority of the proposed approach on several standard datasets. Comparisons are drawn with a large number of known DA algorithms.
Onur C. Hamsici, Aleix Martinez
CVPR2
2007 Spherical-Homoscedastic Shapes
abstract
Shape analysis requires invariance under translation, scale and rotation. Translation and scale invariance can be realized by normalizing shape vectors with respect to their mean and norm. This maps the shape feature vectors onto the surface of a hypersphere. After normalization, the shape vectors can be made rotational invariant by modelling the resulting data using complex scalar rotation invariant distributions defined on the complex hypersphere, e.g., using the complex Bingham distribution. However, the use of these distributions is hampered by the difficulty in estimating their parameters, which is shown to be very costly or impossible in most cases. The purpose of this paper is twofold. First, we show under which conditions the classification results obtained with complex Binghams are identical to those obtained with the easy-to-estimate complex Normal distribution. Second, we derive a kernel function which (intrinsically) maps the data into a space where the above conditions are satisfied and, hence, where the normal model can be successfully used. This results in a simple, low-cost algorithm for representing and classifying shapes. We demonstrate the use of this technique in several experimental results for object and face recognition. Comparisons to other statistical shape representation/classification approaches demonstrate the superiority of the proposed algorithms in classification accuracy and computational time.
Onur C. Hamsici, Aleix Martinez
ICCV2
2007 Spherical-Homoscedastic Distributions: The Equivalency of Spherical and Normal Distributions in Classification
Onur C. Hamsici, Aleix Martinez
J. Mach. Learn. Res.2
2006 Selecting Principal Components in a Two-Stage LDA Algorithm
abstract
Linear Discriminant Analysis (LDA) is a well-known and important tool in pattern recognition with potential applications in many areas of research. The most famous and used formulation of LDA is that given by the Fisher-Rao criterion, where the problem reduces to a simple simultaneous diagonalization of two symmetric, positive-definite matrices, A and B; i.e. B^-1 AV = VA. Here, A defines the metric to be maximized, while B defines the metric to be minimized. However, when B has near-zero eigenvalues, the Fisher-Rao criterion gets dominated by these. While this works well when such small variances describe vectors where most of the discriminant information is, the results will be incorrect when these small variances are caused by noise. Knowing which of these near-zero values are to be used and which need to be eliminated is a challenging yet fundamental task in LDA. This paper presents a criterion for the selection of those vectors of B that are best for classification. The proposed solution is based on a simple factorization of B^-1 A that permits the re-ordering of the eigenvectors of B without the need to effect the end result. This allows us to readily eliminate the noisy vectors while keeping the most discriminant ones. A theoretical basis for these results is presented along with extensive experimental results to validate the claims.
Manli Zhu, Aleix Martinez
CVPR (1)2
2006 A weighted probabilistic approach to face recognition from multiple images and video sequences
Aleix Martinez
Image Vis. Comput.2
2006 Subclass Discriminant Analysis
abstract
Over the years, many Discriminant Analysis (DA) algorithms have been proposed for the study of high-dimensional data in a large variety of problems. Each of these algorithms is tuned to a specific type of data distribution (that which best models the problem at hand). Unfortunately, in most problems the form of each class pdf is a priori unknown, and the selection of the DA algorithm that best fits our data is done over trial-and-error. Ideally, one would like to have a single formulation which can be used for most distribution types. This can be achieved by approximating the underlying distribution of each class with a mixture of Gaussians. In this approach, the major problem to be addressed is that of determining the optimal number of Gaussians per class, i.e., the number of subclasses. In this paper, two criteria able to find the most convenient division of each class into a set of subclasses are derived. Extensive experimental results are shown using five databases. Comparisons are given against Linear Discriminant Analysis (LDA), Direct LDA (DLDA), Heteroscedastic LDA (HLDA), Nonparametric DA (NDA), and Kernel-Based LDA (K-LDA). We show that our method is always the best or comparable to the best.
Manli Zhu, Aleix Martinez
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Robust motion estimation under varying illumination
Yeon-Ho Kim, Aleix Martinez, Avinash C. Kak
Image Vis. Comput.2
2005 Where Are Linear Feature Extraction Methods Applicable?
abstract
A fundamental problem in computer vision and pattern recognition is to determine where and, most importantly, why a given technique is applicable. This is not only necessary because it helps us decide which techniques to apply at each given time. Knowing why current algorithms cannot be applied facilitates the design of new algorithms robust to such problems. In this paper, we report on a theoretical study that demonstrates where and why generalized eigen-based linear equations do not work. In particular, we show that when the smallest angle between the ith eigenvector given by the metric to be maximized and the first i eigenvectors given by the metric to be minimized is close to zero, our results are not guaranteed to be correct. Several properties of such models are also presented. For illustration, we concentrate on the classical applications of classification and feature extraction. We also show how we can use our findings to design more robust algorithms. We conclude with a discussion on the broader impacts of our results.
Aleix Martinez, Manli Zhu
IEEE Trans. Pattern Anal. Mach. Intell.1
2004 A Local Approach for Robust Optical Flow Estimation under Varying Illumination
abstract
The problem of motion estimation, in general, is made difficult by large illumination variations and by motion discontinuities. In recent papers, we and others have proposed global approaches to deal with both problems simultaneously within the regularization framework. A major drawback of such global methods is that several regularization parameters responsible for the integration of the illumination and motion components need to be determined in advance. This has reduced the applicability of global methods. In this paper, a parameter-free local approach, which solves a linear regression problem using a simple parametric model, is presented. To achieve robustness for the linear regression problem, we introduce a modified version of the least median of squares algorithm. We show quantitative error comparisons between the results obtained by our local approach and those produced by several global methods. Our results show that our local method is comparable to the best results obtained by the global approaches yet does not require any manual selection of parameters. 1
Yeon-Ho Kim, Aleix Martinez, Avinash C. Kak
BMVC2
2004 On combining graph-partitioning with non-parametric clustering for image segmentation
Aleix Martinez, Pradit Mittrapiyanuruk, Avinash C. Kak
Comput. Vis. Image Underst.1
2003 Recognizing Expression Variant Faces from a Single Sample Image per Class
abstract
Although important contributions to face recognition have been reported, few focus on how to robustly recognize expression variant faces from as few as one single training sample per class. Since learning cannot generally be applied when only one sample per class is available, matching techniques (distance measures) are usually employed instead (e.g. correlations). However, distance measures generally attempt to match all features with equal importance (weighting), because not only is it difficult to know which features are more useful (for classification), but when or under which circumstances this happens. For example, when recognizing faces in the original image space (e.g. using the Euclidean distance-correlation), it is not known which pixels are more and which are less appropriate for use. We use the optical flow between the testing and sample images as a measure of how good each pixel is. Pixels that have a small flow will have high weights, pixels with a large flow will have small weights. Our experimental results show that the method proposed in this contribution outperforms the classical Euclidean distance (correlation) measure and the PCA (principal component analysis) approach.
Aleix Martinez
CVPR (1)1
2003 Special issue on face recognition
Aleix Martinez, Ming-Hsuan Yang 0001, David J. Kriegman
Comput. Vis. Image Underst.1
2002 Purdue RVL-SLLL ASL Database for Automatic Recognition of American Sign Language
abstract
This article reports on an extensive database of American Sign Language (ASL) motions, handshapes, words and sentences. Research on automatic recognition of ASL requires a suitable database for the training and the testing of algorithms. The databases that are currently available do not allow for algorithmic development that requires a step-by-step approach to ASL recognition-from the recognition of individual handshapes, to the recognition of motion primitives, and finally, to the recognition of full sentences. We have sought to remove these deficiencies in a new database-the Purdue RVL-SLLL ASL database.
Aleix Martinez, Ronnie B. Wilbur, Robin Shay, Avinash C. Kak
ICMI1
2002 Recognizing Imprecisely Localized, Partially Occluded, and Expression Variant Faces from a Single Sample per Class
abstract
The classical way of attempting to solve the face (or object) recognition problem is by using large and representative data sets. In many applications, though, only one sample per class is available to the system. In this contribution, we describe a probabilistic approach that is able to compensate for imprecisely localized, partially occluded, and expression-variant faces even when only one single training sample per class is available to the system. To solve the localization problem, we find the subspace (within the feature space, e.g., eigenspace) that represents this error for each of the training images. To resolve the occlusion problem, each face is divided into k local regions which are analyzed in isolation. In contrast with other approaches where a simple voting space is used, we present a probabilistic method that analyzes how "good" a local match is. To make the recognition system less sensitive to the differences between the facial expression displayed on the training and the testing images, we weight the results obtained on each local area on the basis of how much of this local area is affected by the expression displayed on the current test image.
Aleix Martinez
IEEE Trans. Pattern Anal. Mach. Intell.1
2001 PCA versus LDA
abstract
In the context of the appearance-based paradigm for object recognition, it is generally believed that algorithms based on LDA (linear discriminant analysis) are superior to those based on PCA (principal components analysis). In this communication, we show that this is not always the case. We present our case first by using intuitively plausible arguments and, then, by showing actual results on a face database. Our overall conclusion is that when the training data set is small, PCA can outperform LDA and, also, that PCA is less sensitive to different training data sets.
Aleix Martinez, Avinash C. Kak
IEEE Trans. Pattern Anal. Mach. Intell.1
2001 Clustering in image space for place recognition and visual annotations for human-robot interaction
abstract
The most classical way of attempting to solve the vision-guided navigation problem for autonomous robots corresponds to the use of three-dimensional (3-D) geometrical descriptions of the scene; what is known as model-based approaches. However, these approaches do not facilitate the user's task because they require that geometrically precise models of the 3-D environment be given by the user. In this paper, we propose the use of "annotations" posted on some type of blackboard or "descriptive" map to facilitate this user-robot interaction. We show that, by using this technique, user commands can be as simple as "go to label 5." To build such a mechanism, new approaches for vision-guided mobile robot navigation have to be found. We show that this can be achieved by using mixture models within an appearance-based paradigm. Mixture models are more useful in practice than other pattern recognition methods such as principal component analysis (PCA) or Fisher discriminant analysis (FDA)-also known as linear discriminant analysis (LDA), because they can represent nonlinear subspaces. However, given the fact that mixture models are usually learned using the expectation-maximization (EM) algorithm which is a gradient ascent technique, the system cannot always converge to a desired final solution, due to the local maxima problem. To resolve this, a genetic version of the EM algorithm is used. We then show the capabilities of this latest approach on a navigation task that uses the above described "annotations."
Aleix Martinez, Jordi Vitrià
IEEE Trans. Syst. Man Cybern. Part B1
2000 Recognition of Partially Occluded and/or Imprecisely Localized Faces Using a Probabilistic Approach
abstract
New face recognition approaches are needed, because although much progress has been recently achieved in the field (e.g. within the eigenspace domain), still many problems are to be robustly solved. Two of these problems are occlusions and the imprecise localization of faces (which ultimately imply a failure in identification). While little has been done to account for the first problem, almost nothing has been proposed to account for the second. This paper presents a probabilistic approach that attempts to solve both problems while using an eigenspace representation. To resolve the localization problem, we need to find the subspace (within the feature space, e.g. eigenspace) that represents this error for each of the training images. To resolve the occlusion problem, each face is divided into n local regions which are analyzed in isolation. In contrast with other previous approaches, where a simple voting space is used, we present a probabilistic method that analyzes how "good" a local match is. Our method has proven to be superior to a local voting PCA on a set of 2600 face images.
Aleix Martinez
CVPR1
2000 Learning mixture models using a genetic version of the EM algorithm
Aleix Martinez, Jordi Vitrià
Pattern Recognit. Lett.1