Àgata Lapedriza

dblp:80/2887 · DBLP profile ↗
← Back
28ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0002-5248-0443ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 8 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Automated Detection of Visual Attribute Reliance with a Self-Reflective Agent
abstract
When a vision model performs image recognition, which visual attributes drive its predictions? Detecting unintended reliance on specific visual features is critical for ensuring model robustness, preventing overfitting, and avoiding spurious correlations. We introduce an automated framework for detecting such dependencies in trained vision models. At the core of our method is a self-reflective agent that systematically generates and tests hypotheses about visual attributes that a model may rely on. This process is iterative: the agent refines its hypotheses based on experimental outcomes and uses a self-evaluation protocol to assess whether its findings accurately explain model behavior. When inconsistencies arise, the agent self-reflects over its findings and triggers a new cycle of experimentation. We evaluate our approach on a novel benchmark of 130 models designed to exhibit diverse visual attribute dependencies across 18 categories. Our results show that the agent's performance consistently improves with self-reflection, with a significant performance increase over non-reflective baselines. We further demonstrate that the agent identifies real-world visual attribute dependencies in state-of-the-art models, including CLIP's vision encoder and the YOLOv8 object detector.
Christy Li, Josep López Camuñas, Jake Thomas Touchet, Jacob Andreas, Àgata Lapedriza, Antonio Torralba 0001, Tamar Rott Shaham
NeurIPS5
2025 Analyzing the Visual Road Scene for Driver Stress Estimation
abstract
This paper studies the contribution of the visual road scene to estimate the driver-reported stress levels. Our research leverages on previous work showing that environmental factors, such as traffic congestion, weather conditions, and driving context, impact driver’s stress. Each of the models we evaluated is trained and tested with the publicly available AffectiveROAD dataset to estimate three categories of driver-reported stress level. We test three types of modelling approaches: (i) single-frame baselines (Random Forest, SVM, and Convolutional Neural Networks); (ii) Temporal Segment Networks (TSN) and two variants of it, which use learned weights (TSN-w) and LSTM (TSN-LSTM) as consensus functions; and (iii) video classification Transformers. Our experiments reveal that the TSN-w, TSN-LSTM, and Transformer models achieve statistically equivalent performances, all significantly outperforming the other models. Particularly noteworthy is TSN-w, which attains the highest performance observed with an average accuracy of 0.77. We further provide an explainability analysis using Class Activation Mapping and image semantic segmentation to identify the elements of the road scene that contribute the most to high levels of stress. Our results demonstrate that the visible road scene offers significant contextual information for estimating driver-reported stress levels, with potential implications for the design of safer urban road environments.
Cristina Bustos, Albert Solé-Ribalta, Neska El Haouij, Javier Borge-Holthoefer, Àgata Lapedriza, Rosalind W. Picard
IEEE Trans. Affect. Comput.5
2024 Analyzing Cultural Representations of Emotions in LLMs Through Mixed Emotion Survey
abstract
Large Language Models (LLMs) have gained widespread global adoption, showcasing advanced linguistic capabilities across multiple of languages. There is a growing interest in academia to use these models to simulate and study human behaviors. However, it is crucial to acknowledge that an LLM's proficiency in a specific language might not fully encapsulate the norms and values associated with its culture. Concerns have emerged regarding potential biases towards Anglo-centric cultures and values due to the predominance of Western and US-based training data. This study focuses on analyzing the cultural representations of emotions in LLMs, in the specific case of mixed-emotion situations. Our methodology is based on the studies of Miyamoto et al. (2010), which identified distinctive emotional indicators in Japanese and American human responses. We first administer their mixed emotion survey to five different LLMs and analyze their outputs. Second, we experiment with contextual variables to explore variations in responses considering both language and speaker origin. Thirdly, we expand our investigation to encompass additional East Asian and Western European origin languages to gauge their alignment with their respective cultures, anticipating a closer fit. We find that (1) models have limited alignment with the evidence in the literature; (2) written language has greater effect on LLMs' response than information on participants origin; and (3) LLMs responses were found more similar for East Asian languages than Western European languages.
Shiran Dudy, Ibrahim Said Ahmad, Ryoko Kitajima, Àgata Lapedriza
ACII4
2023 On the use of Vision-Language models for Visual Sentiment Analysis: a study on CLIP
abstract
This work presents a study on how to exploit the CLIP embedding space to perform Visual Sentiment Analysis. We experiment with two architectures built on top of the CLIP embedding space, which we denote by CLIP-E. We train the CLIP-E models with WEBEmo, the largest publicly available and manually labeled benchmark for Visual Sentiment Analysis, and perform two sets of experiments. First, we test on WEBEmo and compare the CLIP-E architectures with state-of-the-art (SOTA) models and with CLIP Zero-Shot. Second, we perform cross dataset evaluation, and test the CLIP-E architectures trained with WEBEmo on other Visual Sentiment Analysis benchmarks. Our results show that the CLIP-E approaches outperform SOTA models in WEBEmo fine grained categorization, and they also generalize better when tested on datasets that have not been seen during training. Interestingly, we observed that for the FI dataset, CLIP Zero-Shot produces better accuracies than SOTA models and CLIP-E trained on WEBEmo. These results motivate several questions that we discuss in this paper, such as how we should design new benchmarks and evaluate Visual Sentiment Analysis, and whether we should keep designing tailored Deep Learning models for Visual Sentiment Analysis or focus our efforts on better using the knowledge encoded in large vision-language models such as CLIP for this task. Our code is available at https://github.com/cristinabustos16/CLIP-E.
Cristina Bustos, Carles Civit, Brian Du, Albert Solé-Ribalta, Àgata Lapedriza
ACII5
2023 Incidents1M: A Large-Scale Dataset of Images With Natural Disasters, Damage, and Incidents
abstract
Natural disasters, such as floods, tornadoes, or wildfires, are increasingly pervasive as the Earth undergoes global warming. It is difficult to predict when and where an incident will occur, so timely emergency response is critical to saving the lives of those endangered by destructive events. Fortunately, technology can play a role in these situations. Social media posts can be used as a low-latency data source to understand the progression and aftermath of a disaster, yet parsing this data is tedious without automated methods. Prior work has mostly focused on text-based filtering, yet image and video-based filtering remains largely unexplored. In this work, we present the Incidents1M Dataset, a large-scale multi-label dataset which contains 977,088 images, with 43 incident and 49 place categories. We provide details of the dataset construction, statistics and potential biases; introduce and train a model for incident detection; and perform image-filtering experiments on millions of images on Flickr and Twitter. We also present some applications on incident analysis to encourage and enable future work in computer vision for humanitarian aid. Code, data, and models are available at http://incidentsdataset.csail.mit.edu.
Ethan Weber, Dim P. Papadopoulos, Àgata Lapedriza, Ferda Ofli, Muhammad Imran 0002, Antonio Torralba 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Deploying a robotic positive psychology coach to improve college students' psychological well-being
abstract
Despite the increase in awareness and support for mental health, college students' mental health is reported to decline every year in many countries. Several interactive technologies for mental health have been proposed and are aiming to make therapeutic service more accessible, but most of them only provide one-way passive contents for their users, such as psycho-education, health monitoring, and clinical assessment. We present a robotic coach that not only delivers interactive positive psychology interventions but also provides other useful skills to build rapport with college students. Results from our on-campus housing deployment feasibility study showed that the robotic intervention showed significant association with increases in students' psychological well-being, mood, and motivation to change. We further found that students' personality traits were associated with the intervention outcomes as well as their working alliance with the robot and their satisfaction with the interventions. Also, students' working alliance with the robot was shown to be associated with their pre-to-post change in motivation for better well-being. Analyses on students' behavioral cues showed that several verbal and nonverbal behaviors were associated with the change in self-reported intervention outcomes. The qualitative analyses on the post-study interview suggest that the robotic coach's companionship made a positive impression on students, but also revealed areas for improvement in the design of the robotic coach. Results from our feasibility study give insight into how learning users' traits and recognizing behavioral cues can help an AI agent provide personalized intervention experiences for better mental health outcomes.
Sooyeon Jeong, Laura Aymerich-Franch, Kika Arias, Sharifa Alghowinem, Àgata Lapedriza, Rosalind W. Picard, Hae Won Park 0001, Cynthia Breazeal
User Model. User Adapt. Interact.5
2022 Interpreting Face Inference Models Using Hierarchical Network Dissection
Divyang Teotia, Àgata Lapedriza, Sarah Ostadabbas
Int. J. Comput. Vis.2
2021 Predicting Driver Self-Reported Stress by Analyzing the Road Scene
abstract
Several studies have shown the relevance of biosignals in driver stress recognition. In this work, we examine something important that has been less frequently explored: We develop methods to test if the visual driving scene can be used to estimate a drivers’ subjective stress levels. For this purpose, we use the AffectiveROAD video recordings and their corresponding stress labels, a continuous human-driver-provided stress metric. We use the common class discretization for stress, dividing its continuous values into three classes: low, medium, and high. We design and evaluate three computer vision modeling approaches to classify the driver’s stress levels: (1) object presence features, where features are computed using automatic scene segmentation; (2) end-to-end image classification; and (3) end-to-end video classification. All three approaches show promising results, suggesting that it is possible to approximate the drivers’ subjective stress from the information found in the visual scene. We observe that the video classification, which processes the temporal information integrated with the visual information, obtains the highest accuracy of 0.72, compared to a random baseline accuracy of 0.33 when tested on a set of nine drivers.
Cristina Bustos, Neska El Haouij, Albert Solé-Ribalta, Javier Borge-Holthoefer, Àgata Lapedriza, Rosalind W. Picard
ACII5
2021 Recognizing Emotions evoked by Movies using Multitask Learning
abstract
Understanding the emotional impact of movies has become important for affective movie analysis, ranking, and indexing. Methods for recognizing evoked emotions are usually trained on human annotated data. Concretely, viewers watch video clips and have to manually annotate the emotions they experienced while watching the videos. Then, the common practice is to aggregate the different annotations, by computing average scores or majority voting, and train and test models on these aggregated annotations. With this procedure a single aggregated evoked emotion annotation is obtained per each video. However, emotions experienced while watching a video are subjective: different individuals might experience different emotions. In this paper, we model the emotions evoked by videos in a different manner: instead of modeling the aggregated value we jointly model the emotions experienced by each viewer and the aggregated value using a multi-task learning approach. Concretely, we propose two deep learning architectures: a Single-Task (ST) architecture and a Multi-Task (MT) architecture. Our results show that the MT approach can more accurately model each viewer and the aggregated annotation when compared to methods that are directly trained on the aggregated annotations. Furthermore, our approach outperforms the current state-of-the-art results on the COGNIMUSE benchmark.
Hassan Hayat, Carles Ventura, Àgata Lapedriza
ACII3
2021 Facial Expressions as a Vulnerability in Face Recognition
abstract
This work explores facial expression bias as a security vulnerability of face recognition systems. Despite the great performance achieved by state-of-the-art face recognition systems, the algorithms are still sensitive to a large range of covariates. We present a comprehensive analysis of how facial expression bias impacts the performance of face recognition technologies. Our study analyzes: i) facial expression biases in the most popular face recognition databases; and ii) the impact of facial expression in face recognition performances. Our experimental framework includes two face detectors, three face recognition models, and three different databases. Our results demonstrate a huge facial expression bias in the most widely used databases, as well as a related impact of face expression in the performance of state-of-the-art algorithms. This work opens the door to new research lines focused on mitigating the observed vulnerability.
Alejandro Peña, Aythami Morales, Ignacio Serna, Julian Fierrez, Àgata Lapedriza
ICIP5
2020 Detecting Natural Disasters, Damage, and Incidents in the Wild
Ethan Weber, Nuria Marzo, Dim P. Papadopoulos, Aritro Biswas, Àgata Lapedriza, Ferda Ofli, Muhammad Imran 0002, Antonio Torralba 0001
ECCV (19)5
2020 Human-centric dialog training via offline reinforcement learning
abstract
Natasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, Rosalind Picard. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Natasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson, Àgata Lapedriza, Noah Jones, Shixiang Gu, Rosalind W. Picard
EMNLP (1)5
2020 Learning Emotional-Blinded Face Representations
abstract
We propose two face representations that are blind to facial expressions associated to emotional responses. This work is in part motivated by new international regulations for personal data protection, which enforce data controllers to protect any kind of sensitive information involved in automatic processes. The advances in Affective Computing have contributed to improve human-machine interfaces but, at the same time, the capacity to monitorize emotional responses triggers potential risks for humans, both in terms of fairness and privacy. We propose two different methods to learn these expression-blinded facial features. We show that it is possible to eliminate information related to emotion recognition tasks, while the performance of subject verification, gender recognition, and ethnicity classification are just slightly affected. We also present an application to train fairer classifiers in a case study of attractiveness classification with respect to a protected facial expression attribute. The results demonstrate that it is possible to reduce emotional information in the face representation while retaining competitive performance in other face-based artificial intelligence tasks.
Alejandro Peña, Julian Fierrez, Aythami Morales, Àgata Lapedriza
ICPR4
2020 A Robotic Positive Psychology Coach to Improve College Students' Wellbeing
abstract
A significant number of college students suffer from mental health issues that impact their physical, social, and occupational outcomes. Various scalable technologies have been proposed in order to mitigate the negative impact of mental health disorders. However, the evaluation for these technologies, if done at all, often reports mixed results on improving users' mental health. We need to better understand the factors that align a user's attributes and needs with technology-based interventions for positive outcomes. In psychotherapy theory, therapeutic alliance and rapport between a therapist and a client is regarded as the basis for therapeutic success. In prior works, social robots have shown the potential to build rapport and a working alliance with users in various settings. In this work, we explore the use of a social robot coach to deliver positive psychology interventions to college students living in on-campus dormitories. We recruited 35 college students to participate in our study and deployed a social robot coach in their room. The robot delivered daily positive psychology sessions among other useful skills like delivering the weather forecast, scheduling reminders, etc. We found a statistically significant improvement in participants' psychological wellbeing, mood, and readiness to change behavior for improved wellbeing after they completed the study. Furthermore, students' personality traits were found to have a significant association with intervention efficacy. Analysis of the post-study interview revealed students' appreciation of the robot's companionship and their concerns for privacy.
Sooyeon Jeong, Sharifa Alghowinem, Laura Aymerich-Franch, Kika Arias, Àgata Lapedriza, Rosalind W. Picard, Hae Won Park 0001, Cynthia Breazeal
RO-MAN5
2020 Context Based Emotion Recognition Using EMOTIC Dataset
abstract
In our everyday lives and social interactions we often try to perceive the emotional states of people. There has been a lot of research in providing machines with a similar capacity of recognizing emotions. From a computer vision perspective, most of the previous efforts have been focusing in analyzing the facial expressions and, in some cases, also the body pose. Some of these methods work remarkably well in specific settings. However, their performance is limited in natural, unconstrained environments. Psychological studies show that the scene context, in addition to facial expression and body pose, provides important information to our perception of people's emotions. However, the processing of the context for automatic emotion recognition has not been explored in depth, partly due to the lack of proper data. In this paper we present EMOTIC, a dataset of images of people in a diverse set of natural situations, annotated with their apparent emotion. The EMOTIC dataset combines two different types of emotion representation: (1) a set of 26 discrete categories, and (2) the continuous dimensions Valence, Arousal, and Dominance. We also present a detailed statistical and algorithmic analysis of the dataset along with annotators' agreement analysis. Using the EMOTIC dataset we train different CNN models for emotion recognition, combining the information of the bounding box containing the person with the contextual information extracted from the scene. Our results show how scene context provides important information to automatically recognize emotional states and motivate further research in this direction.
Ronak Kosti, José M. Álvarez 0004, Adrià Recasens, Àgata Lapedriza
IEEE Trans. Pattern Anal. Mach. Intell.4
2019 Unintentional affective priming during labeling may bias labels
abstract
Online platforms displaying long streams of examples are often employed to gather labels from both experts and crowd workers. While previous work in crowdsourcing focused on objective tasks and estimating error parameters of annotators, collecting labels in a subjective setting (e.g. emotion recognition) is more complicated due to different interpretations of examples. These interpretations could be influenced by many factors such as annotator mood and previously seen examples. In this work, we examine two hypotheses of order-dependent biases in sequential labeling tasks: negatively auto-correlated sequential decision making and positively auto-correlated affective priming. Using controlled generation of facial expressions, we find that i) annotators achieve higher agreement when presented examples in the same sequential order, ii) the valence label of the current image positively correlates with the previous labels given. While we also observe a positive correlation between labels and the number of preceding positive and negative images seen, this correlation is highly dependent on example ordering. Our findings demonstrate that randomized examples given to annotators may produce systematic bias in labels. Future data collection should present examples in orderings which mitigate such bias.
Judy Hanwen Shen, Àgata Lapedriza, Rosalind W. Picard
ACII2
2019 Approximating Interactive Human Evaluation with Self-Play for Open-Domain Dialog Systems
abstract
Building an open-domain conversational agent is a challenging problem. Current evaluation methods, mostly post-hoc judgments of static conversation, do not capture conversation quality in a realistic interactive context. In this paper, we investigate interactive human evaluation and provide evidence for its necessity; we then introduce a novel, model-agnostic, and dataset-agnostic method to approximate it. In particular, we propose a self-play scenario where the dialog system talks to itself and we calculate a combination of proxies such as sentiment and semantic coherence on the conversation trajectory. We show that this metric is capable of capturing the human-rated quality of a dialog model better than any automated metric known to-date, achieving a significant Pearson correlation (r>.7, p<.05). To investigate the strengths of this novel metric and interactive evaluation in comparison to state-of-the-art metrics and human evaluation of static conversations, we perform extended experiments with a set of models, including several that make novel improvements to recent hierarchical dialog generation architectures through sentiment and semantic knowledge distillation on the utterance level. Finally, we open-source the interactive evaluation platform we built and the dataset we collected to allow researchers to efficiently deploy and evaluate dialog models.
Asma Ghandeharioun, Judy Hanwen Shen, Natasha Jaques, Craig Ferguson, Noah Jones, Àgata Lapedriza, Rosalind W. Picard
NeurIPS6
2018 Places: A 10 Million Image Database for Scene Recognition
abstract
The rise of multi-million-item dataset initiatives has enabled data-hungry machine learning algorithms to reach near-human semantic classification performance at tasks such as visual object and scene recognition. Here we describe the Places Database, a repository of 10 million scene photographs, labeled with scene semantic categories, comprising a large and diverse list of the types of environments encountered in the world. Using the state-of-the-art Convolutional Neural Networks (CNNs), we provide scene classification CNNs (Places-CNNs) as baselines, that significantly outperform the previous approaches. Visualization of the CNNs trained on Places shows that object detectors emerge as an intermediate representation of scene classification. With its high-coverage and high-diversity of exemplars, the Places Database along with the Places-CNNs offer a novel resource to guide future progress on scene recognition problems.
Bolei Zhou, Àgata Lapedriza, Aditya Khosla, Aude Oliva, Antonio Torralba 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2017 Emotion Recognition in Context
abstract
Understanding what a person is experiencing from her frame of reference is essential in our everyday life. For this reason, one can think that machines with this type of ability would interact better with people. However, there are no current systems capable of understanding in detail peoples emotional states. Previous research on computer vision to recognize emotions has mainly focused on analyzing the facial expression, usually classifying it into the 6 basic emotions [11]. However, the context plays an important role in emotion perception, and when the context is incorporated, we can infer more emotional states. In this paper we present the Emotions in Context Database (EMCO), a dataset of images containing people in context in non-controlled environments. In these images, people are annotated with 26 emotional categories and also with the continuous dimensions valence, arousal, and dominance [21]. With the EMCO dataset, we trained a Convolutional Neural Network model that jointly analyzes the person and the whole scene to recognize rich information about emotional states. With this, we show the importance of considering the context for recognizing peoples emotions in images, and provide a benchmark in the task of emotion recognition in visual context.
Ronak Kosti, José M. Álvarez 0004, Adrià Recasens, Àgata Lapedriza
CVPR4
2017 Winner takes all hashing for speeding up the training of neural networks in large class problems
Amir H. Bakhtiary, Àgata Lapedriza, David Masip
Pattern Recognit. Lett.2
2016 Learning Deep Features for Discriminative Localization
abstract
In this work, we revisit the global average pooling layer proposed in [13], and shed light on how it explicitly enables the convolutional neural network (CNN) to have remarkable localization ability despite being trained on imagelevel labels. While this technique was previously proposed as a means for regularizing training, we find that it actually builds a generic localizable deep representation that exposes the implicit attention of CNNs on an image. Despite the apparent simplicity of global average pooling, we are able to achieve 37.1% top-5 error for object localization on ILSVRC 2014 without training on any bounding box annotation. We demonstrate in a variety of experiments that our network is able to localize the discriminative image regions despite just being trained for solving classification task1.
Bolei Zhou, Aditya Khosla, Àgata Lapedriza, Aude Oliva, Antonio Torralba 0001
CVPR3
2015 Emotion recognition from mid-level features
David Sánchez-Mendoza, David Masip, Àgata Lapedriza
Pattern Recognit. Lett.3
2014 Learning Deep Features for Scene Recognition using Places Database
Bolei Zhou, Àgata Lapedriza, Jianxiong Xiao, Antonio Torralba 0001, Aude Oliva
NIPS2
2013 Opinion Mining on Educational Resources at the Open University of Catalonia
abstract
In order to make improvements to teaching, it is vital to know what students think of the way they are taught. With that purpose in mind, exhaustively analyzing the forums associated with the subjects taught at the Universitat Oberta de Cataluya (UOC) would be extremely helpful, as the university's students often post comments on their learning experiences in them. Exploiting the content of such forums is not a simple undertaking. The volume of data involved is very large, and performing the task manually would require a great deal of effort from lecturers. As a first step to solve this problem, we propose a tool to automatically analyze the posts in forums of communities of UOC students and teachers, with a view to systematically mining the opinions they contain. This article defines the architecture of such tool and explains how lexical-semantic and language technology resources can be used to that end. For pilot testing purposes, the tool has been used to identify students' opinions on the UOC's Business Intelligence master's degree course during the last two years. The paper discusses the results of such test. The contribution of this paper is twofold. Firstly, it demonstrates the feasibility of using natural language parsing techniques to help teachers to make decisions. Secondly, it introduces a simple tool that can be refined and adapted to a virtual environment for the purpose in question.
Isabel Guitart, Jordi Conesa, Luis Villarejo, Àgata Lapedriza, David Masip, Antoni Perez, Elena Planas
CISIS4
2013 Emotion Detection Using Hybrid Structural and Appearance Descriptors
David Sánchez-Mendoza, David Masip, Xavier Baró, Àgata Lapedriza
MDAI4
2009 Boosted Online Learning for Face Recognition
abstract
Face recognition applications commonly suffer from three main drawbacks: a reduced training set, information lying in high-dimensional subspaces, and the need to incorporate new people to recognize. In the recent literature, the extension of a face classifier in order to include new people in the model has been solved using online feature extraction techniques. The most successful approaches of those are the extensions of the principal component analysis or the linear discriminant analysis. In the current paper, a new online boosting algorithm is introduced: a face recognition method that extends a boosting-based classifier by adding new classes while avoiding the need of retraining the classifier each time a new person joins the system. The classifier is learned using the multitask learning principle where multiple verification tasks are trained together sharing the same feature space. The new classes are added taking advantage of the structure learned previously, being the addition of new classes not computationally demanding. The present proposal has been (experimentally) validated with two different facial data sets by comparing our approach with the current state-of-the-art techniques. The results show that the proposed online boosting algorithm fares better in terms of final accuracy. In addition, the global performance does not decrease drastically even when the number of classes of the base problem is multiplied by eight.
David Masip, Àgata Lapedriza, Jordi Vitrià
IEEE Trans. Syst. Man Cybern. Part B2
2008 On the use of independent tasks for face recognition
abstract
We present a method for learning discriminative linear feature extraction using independent tasks. More concretely, given a target classification task, we consider a complementary classification task that is independent of the target one. For example, in face classification field, subject recognition can be a target task while facial expression classification can be a complementary task. Then, we use labels of the complementary task in order to obtain a more robust feature extraction, being the new feature space less sensitive to the complementary classification. To learn the proposed feature extraction we use the mutual information measure between the projected data and both labels from the target and the complementary tasks. In our experiments, this framework has been applied to a face recognition problem, in order to inhibit this classification task from environmental artifacts, and to mitigate the effects of the small sample size problem. Our classification experiments show an improved feature extraction process using the proposed method.
Àgata Lapedriza, David Masip, Jordi Vitrià
CVPR1
2008 A sparse Bayesian approach for joint feature selection and classifier learning
Àgata Lapedriza, Santi Seguí, David Masip, Jordi Vitrià
Pattern Anal. Appl.1