VLDB 2026 Research / reviewers in the wild / expert
David Freire-Obregón
dblp:157/2453
· DBLP profile ↗
31ranked-venue papers
14as first author
23since 2021 · last 2026
0000-0003-2378-4277ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 12 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 15 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Emotional Modulation in Swarm Decision Dynamics
David Freire-Obregón |
ICAART (1) | 1 |
| 2026 | Multi-year long-term person re-identification using gait and HAR featuresabstract• A real-world dataset was collected from ultra-distance runners at different locations in 2020 and 2023, introducing realistic long-term Re-ID challenges like domain shift and appearance changes. • A two-stream Re-ID model combining gait and human action recognition (HAR) features through a cross-attention fusion, enriching gait-based identity cues with behavior context. • The method significantly improves over gait-only baselines, with up to 12 % mAP gain in cross-year evaluations and 11.6 % in same-year evaluations. • Cross-attention fusion allows the model to prioritize gait information while adaptively integrating activity cues from HAR, leading to faster convergence and higher Rank-1 accuracy. • Experimental results show that the fusion of motion and behavior signals outperforms traditional appearance-based Re-ID and standalone gait methods, especially in unconstrained outdoor environments. We propose a two-stream person re-identification (Re-ID) framework that integrates gait and human action recognition (HAR) through cross-attention fusion. The model processes gait sequences via a BiLSTM-based encoder to capture temporal motion dynamics. At the same time, HAR embeddings are extracted using pre-trained video backbones and distilled into compact behavioral features. These two modalities are fused using a cross-attention mechanism, enriching gait-based identity representations with context-aware activity cues. We evaluate our method on a newly curated long-term spatio-temporal dataset of ultra-distance runners captured in natural outdoor settings across multiple locations spanning three years (2020 to 2023). Experimental results demonstrate that integrating HAR significantly enhances gait-based Re-ID performance. Compared to gait-only models, our approach yields a 12 % improvement in mean Average Precision (mAP) in cross-year scenarios and up to an 11.6 % gain in same-year evaluations. The HAR-enhanced models also exhibit faster convergence and higher Rank-1 accuracy, establishing the effectiveness of multi-modal motion-based representations for long-term, real-world person Re-ID. David Freire-Obregón, Oliverio J. Santana, Javier Lorenzo-Navarro, Daniel Hernández-Sosa, Modesto Castrillón-Santana |
Pattern Recognit. | 1 |
| 2025 | An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still ImagesabstractFacial expression recognition (FER) is a key research area in computer vision and human-computer interaction. Despite recent advances, challenges persist, especially in generalizing to new scenarios. In fact, zero-shot FER significantly reduces the performance of state-of-the-art FER models. The community has recently started to explore the integration of knowledge from Large Language Models for visual tasks. In this work, we evaluate a broad collection of Visual Language Models (VLMs), avoiding the lack of task-specific knowledge by adopting a Visual Question Answering strategy. We compare the proposed pipeline with state-of-the-art FER models, both integrating and excluding VLMs, evaluating well-known FER benchmarks: AffectNet, FERPlus, and RAF-DB. The results show state-of-the-art performance for some VLMs in zero-shot FER scenarios, suggesting a research line for further exploration to improve FER generalization. José Salas-Cáceres, Modesto Castrillón-Santana, David Freire-Obregón, Oliverio J. Santana, Daniel Hernández-Sosa, Javier Lorenzo-Navarro |
VCIP | 3 |
| 2025 | Synthesizing multilevel abstraction ear sketches for enhanced biometric recognitionabstractSketch understanding poses unique challenges for general-purpose vision algorithms due to the sparse and semantically ambiguous nature of sketches. This paper introduces a novel approach to biometric recognition that leverages sketch-based representations of ears, a largely unexplored but promising area in biometric research. Specifically, we address the “ sketch-2-image ” matching problem by synthesizing ear sketches at multiple abstraction levels, achieved through a triplet-loss function adapted to integrate these levels. The abstraction level is determined by the number of strokes used, with fewer strokes reflecting higher abstraction. Our methodology combines sketch representations across abstraction levels to improve robustness and generalizability in matching. Extensive evaluations were conducted on four ear datasets (AMI, AWE, IITDII, and BIPLab) using various pre-trained neural network backbones, showing consistently superior performance over state-of-the-art methods. These results highlight the potential of ear sketch-based recognition, with cross-dataset tests confirming its adaptability to real-world conditions and suggesting applicability beyond ear biometrics. • Sketch-Based Datasets Expansion: Leveraging CLIPasso, we transformed ear images into sketches at various abstraction levels, preserving key features and introducing a novel data representation for biometric analysis. • Triplet-Loss Function Enhancement: Adapting the triplet-loss function to incorporate multiple abstraction levels significantly improves recognition performance over traditional methods. • Comparative Backbone Analysis: An exhaustive evaluation of different backbones highlights their effectiveness in sketch-based ear recognition, guiding advancements in biometric technologies. • Cross-Dataset Generalizability Tests: Training on combined datasets and testing on distinct ones validate our approach’s robustness and effectiveness against unseen data distributions. David Freire-Obregón, João C. Neves 0001, Ziga Emersic, Blaz Meden, Modesto Castrillón-Santana, Hugo Proença 0001 |
Image Vis. Comput. | 1 |
| 2025 | Exploring biometric domain adaptation in human action recognition models for unconstrained environmentsabstractAbstract In conventional machine learning (ML), a fundamental assumption is that the training and test sets share identical feature distributions, a reasonable premise drawn from the same dataset. However, real-world scenarios often defy this assumption, as data may originate from diverse sources, causing disparities between training and test data distributions. This leads to a domain shift, where variations emerge between the source and target domains. This study delves into human action recognition (HAR) models within an unconstrained, real-world setting, scrutinizing the impact of input data variations related to contextual information and video encoding. The objective is to highlight the intricacies of model performance and interpretability in this context. Additionally, the study explores the domain adaptability of HAR models, specifically focusing on their potential for re-identifying individuals within uncontrolled environments. The experiments involve seven pre-trained backbone models and introduce a novel analytical approach by linking domain-related (HAR) and domain-unrelated (re-identification (re-ID)) tasks. Two key analyses addressing contextual information and encoding strategies reveal that maintaining the same encoding approach during training results in high task correlation while incorporating richer contextual information enhances performance. A notable outcome of this study is the comprehensive evaluation of a novel transformer-based architecture driven by a HAR backbone, which achieves a robust re-ID performance superior to state-of-the-art (SOTA). However, it faces challenges when other encoding schemes are applied, highlighting the role of the HAR classifier in performance variations. David Freire-Obregón, Paola Barra, Modesto Castrillón-Santana, Maria De Marsico |
Multim. Tools Appl. | 1 |
| 2025 | Multimodal emotion recognition based on a fusion of audiovisual information with temporal dynamicsabstractAbstract In the Human-Machine Interactions (HMI) landscape, understanding user emotions is pivotal for elevating user experiences. This paper explores Facial Expression Recognition (FER) within HMI, employing a distinctive multimodal approach that integrates visual and auditory information. Recognizing the dynamic nature of HMI, where situations evolve, this study emphasizes continuous emotion analysis. This work assesses various fusion strategies that involve the addition to the main network of different architectures, such as autoencoders (AE) or an Embracement module, to combine the information of multiple biometric cues. In addition to the multimodal approach, this paper introduces a new architecture that prioritizes temporal dynamics by incorporating Long Short-Term Memory (LSTM) networks. The final proposal, which integrates different multimodal approaches with the temporal focus capabilities of the LSTM architecture, was tested across three public datasets: RAVDESS, SAVEE, and CREMA-D. It showcased state-of-the-art accuracy of 88.11%, 86.75%, and 80.27%, respectively, and outperformed other existing approaches. José Salas-Cáceres, Javier Lorenzo-Navarro, David Freire-Obregón, Modesto Castrillón-Santana |
Multim. Tools Appl. | 3 |
| 2024 | Towards Bi-Hemispheric Emotion Mapping Through EEG: A Dual-Stream Neural Network ApproachabstractEmotion classification through EEG signals plays a significant role in psychology, neuroscience, and human-computer interaction. This paper addresses the challenge of mapping human emotions using EEG data in the Mapping Human Emotions through EEG Signals FG24 competition. Subjects mimic the facial expressions of an avatar, displaying fear, joy, anger, sadness, disgust, and surprise in a VR setting. EEG data is captured using a multi-channel sensor system to discern brain activity patterns. We propose a novel two-stream neural network employing a Bi-Hemispheric approach for emotion inference, surpassing baseline methods and enhancing emotion recognition accuracy. Additionally, we conduct a temporal analysis revealing that specific signal intervals at the beginning and end of the emotion stimulus sequence contribute significantly to improve accuracy. Leveraging insights gained from this temporal analysis, our approach offers enhanced performance in capturing subtle variations in the states of emotions. Code is available at https://github.com/davidfreire/FG24-EmoNeuroDB/ David Freire-Obregón, Daniel Hernández-Sosa, Oliverio J. Santana, Javier Lorenzo-Navarro, Modesto Castrillón-Santana |
FG | 1 |
| 2024 | An Evaluation of General-Purpose Optical Character Recognizers and Digit Detectors for Race Bib Number Recognition
Modesto Castrillón-Santana, David Freire-Obregón, Daniel Hernández-Sosa, Oliverio J. Santana, Francisco Ortega-Zamorano, José Isern González, Javier Lorenzo-Navarro |
ICPRAM | 2 |
| 2024 | Classifying Soccer Ball-on-Goal Position Through Kicker Shooting Action
Javier Torón-Artiles, Daniel Hernández-Sosa, Oliverio J. Santana, Javier Lorenzo-Navarro, David Freire-Obregón |
ICPRAM | 5 |
| 2024 | Heterogeneous Transfer Learning in Sports: Human Action Recognition for Gender and Outcome Prediction
Javier Torón-Artiles, Daniel Hernández-Sosa, Oliverio J. Santana, Javier Lorenzo-Navarro, David Freire-Obregón |
ICPRAM | 5 |
| 2024 | Applying deep learning image enhancement methods to improve person re-identificationabstractPerson re-identification has gained significant attention in recent years due to its numerous practical applications in video surveillance. However, while artificial intelligence and deep learning methods have enabled substantial progress in particular aspects of this domain, putting together those individual advances to generate practical systems remains a computer vision challenge. Existing methods are typically designed assuming the target person’s images are captured under uniform, stable conditions with similar lighting levels, but this assumption may not hold in real-world scenarios, such as outdoor monitoring over 24 h, as image quality can vary considerably throughout day and night. In this paper, we propose a framework that incorporates image enhancement techniques to improve the performance of a person re-identification model. The proposed approach achieves a significant improvement in a demanding re-identification dataset, raising the mAP from 9.0% using a zero-shot baseline to 65.8% through the combined use of low-light image enhancement methods and noise reduction. Oliverio J. Santana, Javier Lorenzo-Navarro, David Freire-Obregón, Daniel Hernández-Sosa, Modesto Castrillón-Santana |
Neurocomputing | 3 |
| 2024 | Introduction to the special issue on "Computer vision solutions for part-based image analysis and classification (CV_PARTIAL)"
Fabio Narducci, Piercalo Dondi, David Freire-Obregón, Florin Pop |
Pattern Recognit. Lett. | 3 |
| 2023 | Evaluation of a Visual Question Answering Architecture for Pedestrian Attribute Recognition
Modesto Castrillón-Santana, Elena Sánchez-Nielsen, David Freire-Obregón, Oliverio J. Santana, Daniel Hernández-Sosa, Javier Lorenzo-Navarro |
CAIP (1) | 3 |
| 2023 | Novelty Detection in Human-Machine Interaction Through a Multimodal Approach
José Salas-Cáceres, Javier Lorenzo-Navarro, David Freire-Obregón, Modesto Castrillón-Santana |
CIARP | 3 |
| 2023 | A Large-Scale Re-identification Analysis in Sporting Scenarios: the Betrayal of Reaching a Critical PointabstractRe-identifying participants in ultra-distance running competitions can be daunting due to the extensive distances and constantly changing terrain. To overcome these challenges, computer vision techniques have been developed to analyze runners’ faces, numbers on their bibs, and clothing. However, our study presents a novel gait-based approach for runners’ re-identification (re-ID) by leveraging various pre-trained human action recognition (HAR) models and loss functions. Our results show that this approach provides promising results for re-identifying runners in ultra-distance competitions. Furthermore, we investigate the significance of distinct human body movements when athletes are approaching their endurance limits and their potential impact on re-ID accuracy. Our study examines how the recognition of a runner’s gait is affected by a competition’s critical point (CP), defined as a moment of severe fatigue and the point where the finish line comes into view, just a few kilometers away from this location. We aim to determine how this CP can improve the accuracy of athlete re-ID. Our experimental results demonstrate that gait recognition can be significantly enhanced (up to a 9% increase in mAP) as athletes approach this point. This highlights the potential of utilizing gait recognition in real-world scenarios, such as ultra-distance competitions or long-duration surveillance tasks. David Freire-Obregón, Javier Lorenzo-Navarro, Oliverio J. Santana, Daniel Hernández-Sosa, Modesto Castrillón-Santana |
IJCB | 1 |
| 2023 | Deep Learning for Diagonal Earlobe Crease DetectionabstractAn article published on Medical News Today in June 2022 presented a \nfundamental question in its title: Can an earlobe crease predict heart attacks? \nThe author explained that end arteries supply the heart and ears. In other \nwords, if they lose blood supply, no other arteries can take over, resulting in \ntissue damage. Consequently, some earlobes have a diagonal crease, line, or \ndeep fold that resembles a wrinkle. In this paper, we take a step toward \ndetecting this specific marker, commonly known as DELC or Frank's Sign. For \nthis reason, we have made the first DELC dataset available to the public. In \naddition, we have investigated the performance of numerous cutting-edge \nbackbones on annotated photos. Experimentally, we demonstrate that it is \npossible to solve this challenge by combining pre-trained encoders with a \ncustomized classifier to achieve 97.7% accuracy. Moreover, we have analyzed the \nbackbone trade-off between performance and size, estimating MobileNet as the \nmost promising encoder. Sara L. Almonacid-Uribe, Oliverio J. Santana, Daniel Hernández-Sosa, David Freire-Obregón |
ICPRAM | 4 |
| 2023 | Evaluating the Impact of Low-Light Image Enhancement Methods on Runner Re-Identification in the WildabstractPerson re-identification (ReID) is a trending topic in computer vision. Significant developments have been achieved, but most rely on datasets with subjects captured statically within a short period of time in rather good lighting conditions. In the wild scenarios, such as long-distance races that involve widely varying lighting conditions, from full daylight to night, present a considerable challenge. This issue cannot be addressed by increasing the exposure time on the capture device, as the runners' motion will lead to blurred images, hampering any ReID attempts. In this paper, we survey some low-light image enhancement methods. Our results show that including an image processing step in a ReID pipeline before extracting the distinctive body appearance features from the subjects can provide significant performance improvements. Oliverio J. Santana, Javier Lorenzo-Navarro, David Freire-Obregón, Daniel Hernández-Sosa, Modesto Castrillón-Santana |
ICPRAM | 3 |
| 2023 | Facial expression analysis in a wild sporting environmentabstractThe scientific community and mass media have already reported the use of nonverbal behavior analysis in sports for athletes' performance. Their conclusions stated that certain emotional expressions are linked to athlete's performance, or even that psychological strategies serve to improve endurance performance. This paper examines the portrayal of well-known emotions and their relationship to the participants of an ultra-distance race in a high-stake environment. For this purpose, we analyzed almost 600 runners captured when they passed through a set of locations placed along the race track. We have observed a correlation between the runners' facial expressions and their performance along the track. Moreover, we have analyzed Action Unit activations and aligned our findings with the state-of-the-art psychological baseline. Oliverio J. Santana, David Freire-Obregón, Daniel Hernández-Sosa, Javier Lorenzo-Navarro, Elena Sánchez-Nielsen, Modesto Castrillón-Santana |
Multim. Tools Appl. | 2 |
| 2023 | Zero-shot ear cross-dataset transfer for person recognition on mobile devicesabstractSmartphones contain personal and private data to be protected, such as everyday communications or bank accounts. Several biometric techniques have been developed to unlock smartphones, among which ear biometrics represents a natural and promising opportunity even though the ear can be used in other biometric and multi-biometric applications. A problem in generalizing research results to real-world applications is that the available ear datasets present different characteristics and some bias. This paper stems from a study about the effect of mixing multiple datasets during the training of an ear recognition system. The main contribution is the evaluation of a robust pipeline that learns to combine data from different sources and highlights the importance of pre-training encoders on auxiliary tasks. The reported experiments exploit eight diverse training datasets to demonstrate the generalization capabilities of the proposed approach. Performance evaluation includes testing with collections not seen during training and assessing zero-shot cross-dataset transfer. The results confirm that mixing different sources provides an insightful perspective on the datasets and competitive results with some existing benchmarks. David Freire-Obregón, Maria De Marsico, Paola Barra, Javier Lorenzo-Navarro, Modesto Castrillón-Santana |
Pattern Recognit. Lett. | 1 |
| 2022 | Towards cumulative race time regression in sports: I3D ConvNet transfer learning in ultra-distance running eventsabstractPredicting an athlete’s performance based on short footage is highly challenging. Performance prediction requires high domain knowledge and enough evidence to infer an appropriate quality assessment. Sports pundits can often infer this kind of information in real-time. In this paper, we propose regressing an ultra-distance runner cumulative race time (CRT), i.e., the time the runner has been in action since the race start, by using only a few seconds of footage as input. We modified the I3D ConvNet backbone slightly and trained a newly added regressor for that purpose. We use appropriate pre-processing of the visual input to enable transfer learning from a specific runner. We show that the resulting neural network can provide a remarkable performance for short input footage: 18 minutes and a half mean absolute error in estimating the CRT for runners who have been in action from 8 to 20 hours. Our methodology has several favorable properties: it does not require a human expert to provide any insight, it can be used at any moment during the race by just observing a runner, and it can inform the race staff about a runner at any given time. David Freire-Obregón, Javier Lorenzo-Navarro, Oliverio J. Santana, Daniel Hernández-Sosa, Modesto Castrillón-Santana |
ICPR | 1 |
| 2022 | Boosting Re-identification in the Ultra-running Scenario
Miguel Angel Medina, Javier Lorenzo-Navarro, David Freire-Obregón, Oliverio J. Santana, Daniel Hernández-Sosa, Modesto Castrillón-Santana |
ICPRAM | 3 |
| 2022 | Inflated 3D ConvNet context analysis for violence detectionabstractAbstract According to the Wall Street Journal, one billion surveillance cameras will be deployed around the world by 2021. This amount of information can be hardly managed by humans. Using a Inflated 3D ConvNet as backbone, this paper introduces a novel automatic violence detection approach that outperforms state-of-the-art existing proposals. Most of those proposals consider a pre-processing step to only focus on some regions of interest in the scene, i.e., those actually containing a human subject. In this regard, this paper also reports the results of an extensive analysis on whether and how the context can affect or not the adopted classifier performance. The experiments show that context-free footage yields substantial deterioration of the classifier performance (2% to 5%) on publicly available datasets. However, they also demonstrate that performance stabilizes in context-free settings, no matter the level of context restriction applied. Finally, a cross-dataset experiment investigates the generalizability of results obtained in a single-collection experiment (same dataset used for training and testing) to cross-collection settings (different datasets used for training and testing). David Freire-Obregón, Paola Barra, Modesto Castrillón-Santana, Maria De Marsico |
Mach. Vis. Appl. | 1 |
| 2021 | Improving user verification in human-robot interaction from audio or image inputs through sample quality assessmentabstractIn this paper, we tackle the task of improving biometric verification in the context of Human-Robot Interaction (HRI). A robot that wants to identify a specific person to provide a service can do so by either image verification or, if light conditions are not favourable, through voice verification. In our approach, we will take advantage of the possibility a robot has of recovering further data until it is sure of the identity of the person. The key contribution is that we select from both image and audio signals the parts that are of higher confidence. For images we use a system that looks at the face of each person and selects frames in which the confidence is high while keeping those frames separate in time to avoid using very similar facial appearance . For audio our approach tries to find the parts of the signal that contain a person talking, avoiding those in which noise is present by segmenting the signal. Once the parts of interest are found, each input is described with an independent deep learning architecture that obtains a descriptor for each kind of input (face/voice). We also present in this paper fusion methods that improve performance by combining the features from both face and voice, results to validate this are shown for each independent input and for the fusion methods. David Freire-Obregón, Kevin Rosales-Santana, Pedro A. Marín-Reyes, Adrián Peñate Sánchez, Javier Lorenzo-Navarro, Modesto Castrillón-Santana |
Pattern Recognit. Lett. | 1 |
| 2020 | TGCRBNW: A Dataset for Runner Bib Number Detection (and Recognition) in the WildabstractRacing bib number (RBN) detection and recognition is a specific problem related to text recognition in natural scenes. In this paper, we present a novel dataset created after registering participants in a real ultrarunning competition which comprises a wide range of acquisition conditions in five different recording points, including nightlight and daylight. The dataset contains more than 3K samples of over 400 different individuals. The aim is to provide an “in the wild” benchmark for both RBN detection and recognition problems. To illustrate the present difficulties, the dataset is evaluated for RBN detection using different Faster R-CNN specific detection models, filtering its output with heuristics based on body detection to improve the overall detection performance. Initial results are promising, but there is still significant room for improvement. And detection is just the first step to accomplish “in the wild” RBN recognition. Pablo Hernández-Carrascosa, Adrián Peñate Sánchez, Javier Lorenzo-Navarro, David Freire-Obregón, Modesto Castrillón-Santana |
ICPR | 4 |
| 2020 | An attention recurrent model for human cooperation detection
David Freire-Obregón, Modesto Castrillón-Santana, Paola Barra, Carmen Bisogni, Michele Nappi |
Comput. Vis. Image Underst. | 1 |
| 2020 | TGC20ReId: A dataset for sport event re-identification in the wild
Adrián Peñate Sánchez, David Freire-Obregón, Adrián Lorenzo-Melián, Javier Lorenzo-Navarro, Modesto Castrillón-Santana |
Pattern Recognit. Lett. | 2 |
| 2019 | Deep learning for source camera identification on mobile devices
David Freire-Obregón, Fabio Narducci, Silvio Barra, Modesto Castrillón-Santana |
Pattern Recognit. Lett. | 1 |
| 2018 | Evaluation of local descriptors and CNNs for non-adult detection in visual contentabstractThe recent evolution of storage devices, digital embedded cameras and the Internet have collaterally allowed sexual predators to take advantage of these technological breakthroughs to gather illegal media, which is exhibited uncensored through Peer-to-Peer file sharing networks. In this paper, we are particularly concerned about the increasing availability of Child Abuse Material. Therefore, we have explored alternatives to detect non-adults in visual content. Initially, different age estimations and underage detection techniques are reviewed by analyzing existing datasets. Finally, several local descriptors and Convolutional Neural Networks for underage detection are evaluated. The experimental results obtained for a large dataset that combines collections such as FG-Net, Adience, GenderChildren, The Image of Groups and Boys2Men evidence the complementary information contained in both local descriptors and neural networks , as their fusion boosts the accuracy of non-adult detection to over 93%. Modesto Castrillón-Santana, Javier Lorenzo-Navarro, Carlos Manuel Travieso-González, David Freire-Obregón, Jesús B. Alonso |
Pattern Recognit. Lett. | 4 |
| 2015 | An Evolutive Approach for Smile Recognition in Video SequencesabstractFacial expression recognition is one of the most challenging research areas in the image recognition field and has been actively studied since the 70's. For instance, smile recognition has been studied due to the fact that it is considered an important facial expression in human communication, it is therefore likely useful for human–machine interaction. Moreover, if a smile can be detected and also its intensity estimated, it will raise the possibility of new applications in the future. We are talking about quantifying the emotion at low computation cost and high accuracy. For this aim, we have used a new support vector machine (SVM)-based approach that integrates a weighted combination of local binary patterns (LBPs)-and principal component analysis (PCA)-based approaches. Furthermore, we construct this smile detector considering the evolution of the emotion along its natural life cycle. As a consequence, we achieved both low computation cost and high performance with video sequences. David Freire-Obregón, Modesto Castrillón-Santana |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2014 | Automatic clothes segmentation for soft biometricsabstractDuring the last decade, researchers have verified that clothing can provide information for gender recognition. However, before extracting features, it is necessary to segment the clothing region. We introduce a new clothes segmentation method based on the application of the GrabCut technique over a trixel mesh, obtaining very promising results for a close to real time system. Finally, the clothing features are combined with facial and head context information to outperform previous results in gender recognition with a public database. David Freire-Obregón, Modesto Castrillón-Santana, Enrique Ramón-Balmaseda, Javier Lorenzo-Navarro |
ICIP | 1 |
| 2010 | Learning to recognize gender using experienceabstractAutomatic facial analysis abilities are commonly integrated in a system by a previous off-line learning stage. In this paper we argue that a facial analysis system would improve its facial analysis capabilities based on its own experience similarly to the way a biological system, i.e. the human system, does throughout the years. The approach described, focused on gender classification, updates its knowledge according to the classification results. The presented gender experiments suggest that this approach is promising, even when just a short simulation of what for humans would take years of acquisition experience was performed. Modesto Castrillón-Santana, Javier Lorenzo-Navarro, David Freire-Obregón, Oscar Déniz-Suárez |
ICIP | 3 |