VLDB 2026 Research / reviewers in the wild / expert
Modesto Castrillón-Santana
dblp:s/ModestoCastrillonSantana · also Modesto Castrillón, Modesto Fernando Castrillón-Santana
· DBLP profile ↗
57ranked-venue papers
16as first author
22since 2021 · last 2026
0000-0002-8673-2725ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 10 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 11 first-author · 12 since 2021Security and privacy · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An annotation assistant for monitoring the electrical grid using aerial imagesabstractMonitoring the electrical grid is essential to ensure reliable service and prevent accidents. This supervision is performed by aerial vehicles for image collection; later, these collected images are processed and analyzed by expert annotators. Due to the high costs of manually handling such as large datasets, we present a novel hybrid methodology that leverages deep learning to reduce and optimize annotation workload. The approach uses annotator-provided labels to train a neural network that makes annotation suggestions and gradually reduces the manual workload. Our work is closely related to active learning, but with a key difference: all data must be labeled and verified to guarantee correctness. Therefore, our methodology focuses on reducing the annotation time rather maximizing model performance. Our hybrid method assists annotators by suggesting annotations on high-confidence images that only need verification instead of being created from scratch. Using the proposed approach, annotators can complete their task at least 2.67x faster than with the previous fully manual labeling procedure. Cristina Benlliure-Jimenez, Adrián Peñate Sánchez, Javier Lorenzo-Navarro, Modesto Castrillón-Santana, Francisco-Mario Hernández Tejera |
Knowl. Based Syst. | 4 |
| 2026 | Multi-year long-term person re-identification using gait and HAR featuresabstract• A real-world dataset was collected from ultra-distance runners at different locations in 2020 and 2023, introducing realistic long-term Re-ID challenges like domain shift and appearance changes. • A two-stream Re-ID model combining gait and human action recognition (HAR) features through a cross-attention fusion, enriching gait-based identity cues with behavior context. • The method significantly improves over gait-only baselines, with up to 12 % mAP gain in cross-year evaluations and 11.6 % in same-year evaluations. • Cross-attention fusion allows the model to prioritize gait information while adaptively integrating activity cues from HAR, leading to faster convergence and higher Rank-1 accuracy. • Experimental results show that the fusion of motion and behavior signals outperforms traditional appearance-based Re-ID and standalone gait methods, especially in unconstrained outdoor environments. We propose a two-stream person re-identification (Re-ID) framework that integrates gait and human action recognition (HAR) through cross-attention fusion. The model processes gait sequences via a BiLSTM-based encoder to capture temporal motion dynamics. At the same time, HAR embeddings are extracted using pre-trained video backbones and distilled into compact behavioral features. These two modalities are fused using a cross-attention mechanism, enriching gait-based identity representations with context-aware activity cues. We evaluate our method on a newly curated long-term spatio-temporal dataset of ultra-distance runners captured in natural outdoor settings across multiple locations spanning three years (2020 to 2023). Experimental results demonstrate that integrating HAR significantly enhances gait-based Re-ID performance. Compared to gait-only models, our approach yields a 12 % improvement in mean Average Precision (mAP) in cross-year scenarios and up to an 11.6 % gain in same-year evaluations. The HAR-enhanced models also exhibit faster convergence and higher Rank-1 accuracy, establishing the effectiveness of multi-modal motion-based representations for long-term, real-world person Re-ID. David Freire-Obregón, Oliverio J. Santana, Javier Lorenzo-Navarro, Daniel Hernández-Sosa, Modesto Castrillón-Santana |
Pattern Recognit. | 5 |
| 2025 | An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still ImagesabstractFacial expression recognition (FER) is a key research area in computer vision and human-computer interaction. Despite recent advances, challenges persist, especially in generalizing to new scenarios. In fact, zero-shot FER significantly reduces the performance of state-of-the-art FER models. The community has recently started to explore the integration of knowledge from Large Language Models for visual tasks. In this work, we evaluate a broad collection of Visual Language Models (VLMs), avoiding the lack of task-specific knowledge by adopting a Visual Question Answering strategy. We compare the proposed pipeline with state-of-the-art FER models, both integrating and excluding VLMs, evaluating well-known FER benchmarks: AffectNet, FERPlus, and RAF-DB. The results show state-of-the-art performance for some VLMs in zero-shot FER scenarios, suggesting a research line for further exploration to improve FER generalization. José Salas-Cáceres, Modesto Castrillón-Santana, David Freire-Obregón, Oliverio J. Santana, Daniel Hernández-Sosa, Javier Lorenzo-Navarro |
VCIP | 2 |
| 2025 | Synthesizing multilevel abstraction ear sketches for enhanced biometric recognitionabstractSketch understanding poses unique challenges for general-purpose vision algorithms due to the sparse and semantically ambiguous nature of sketches. This paper introduces a novel approach to biometric recognition that leverages sketch-based representations of ears, a largely unexplored but promising area in biometric research. Specifically, we address the “ sketch-2-image ” matching problem by synthesizing ear sketches at multiple abstraction levels, achieved through a triplet-loss function adapted to integrate these levels. The abstraction level is determined by the number of strokes used, with fewer strokes reflecting higher abstraction. Our methodology combines sketch representations across abstraction levels to improve robustness and generalizability in matching. Extensive evaluations were conducted on four ear datasets (AMI, AWE, IITDII, and BIPLab) using various pre-trained neural network backbones, showing consistently superior performance over state-of-the-art methods. These results highlight the potential of ear sketch-based recognition, with cross-dataset tests confirming its adaptability to real-world conditions and suggesting applicability beyond ear biometrics. • Sketch-Based Datasets Expansion: Leveraging CLIPasso, we transformed ear images into sketches at various abstraction levels, preserving key features and introducing a novel data representation for biometric analysis. • Triplet-Loss Function Enhancement: Adapting the triplet-loss function to incorporate multiple abstraction levels significantly improves recognition performance over traditional methods. • Comparative Backbone Analysis: An exhaustive evaluation of different backbones highlights their effectiveness in sketch-based ear recognition, guiding advancements in biometric technologies. • Cross-Dataset Generalizability Tests: Training on combined datasets and testing on distinct ones validate our approach’s robustness and effectiveness against unseen data distributions. David Freire-Obregón, João C. Neves 0001, Ziga Emersic, Blaz Meden, Modesto Castrillón-Santana, Hugo Proença 0001 |
Image Vis. Comput. | 5 |
| 2025 | Exploring biometric domain adaptation in human action recognition models for unconstrained environmentsabstractAbstract In conventional machine learning (ML), a fundamental assumption is that the training and test sets share identical feature distributions, a reasonable premise drawn from the same dataset. However, real-world scenarios often defy this assumption, as data may originate from diverse sources, causing disparities between training and test data distributions. This leads to a domain shift, where variations emerge between the source and target domains. This study delves into human action recognition (HAR) models within an unconstrained, real-world setting, scrutinizing the impact of input data variations related to contextual information and video encoding. The objective is to highlight the intricacies of model performance and interpretability in this context. Additionally, the study explores the domain adaptability of HAR models, specifically focusing on their potential for re-identifying individuals within uncontrolled environments. The experiments involve seven pre-trained backbone models and introduce a novel analytical approach by linking domain-related (HAR) and domain-unrelated (re-identification (re-ID)) tasks. Two key analyses addressing contextual information and encoding strategies reveal that maintaining the same encoding approach during training results in high task correlation while incorporating richer contextual information enhances performance. A notable outcome of this study is the comprehensive evaluation of a novel transformer-based architecture driven by a HAR backbone, which achieves a robust re-ID performance superior to state-of-the-art (SOTA). However, it faces challenges when other encoding schemes are applied, highlighting the role of the HAR classifier in performance variations. David Freire-Obregón, Paola Barra, Modesto Castrillón-Santana, Maria De Marsico |
Multim. Tools Appl. | 3 |
| 2025 | Multimodal emotion recognition based on a fusion of audiovisual information with temporal dynamicsabstractAbstract In the Human-Machine Interactions (HMI) landscape, understanding user emotions is pivotal for elevating user experiences. This paper explores Facial Expression Recognition (FER) within HMI, employing a distinctive multimodal approach that integrates visual and auditory information. Recognizing the dynamic nature of HMI, where situations evolve, this study emphasizes continuous emotion analysis. This work assesses various fusion strategies that involve the addition to the main network of different architectures, such as autoencoders (AE) or an Embracement module, to combine the information of multiple biometric cues. In addition to the multimodal approach, this paper introduces a new architecture that prioritizes temporal dynamics by incorporating Long Short-Term Memory (LSTM) networks. The final proposal, which integrates different multimodal approaches with the temporal focus capabilities of the LSTM architecture, was tested across three public datasets: RAVDESS, SAVEE, and CREMA-D. It showcased state-of-the-art accuracy of 88.11%, 86.75%, and 80.27%, respectively, and outperformed other existing approaches. José Salas-Cáceres, Javier Lorenzo-Navarro, David Freire-Obregón, Modesto Castrillón-Santana |
Multim. Tools Appl. | 4 |
| 2025 | Guest editorial: special issue on pedestrian attribute recognition and person re-identification
Antonio Greco 0001, Modesto Castrillón-Santana, Bruno Vento |
Pattern Anal. Appl. | 2 |
| 2024 | Towards Bi-Hemispheric Emotion Mapping Through EEG: A Dual-Stream Neural Network ApproachabstractEmotion classification through EEG signals plays a significant role in psychology, neuroscience, and human-computer interaction. This paper addresses the challenge of mapping human emotions using EEG data in the Mapping Human Emotions through EEG Signals FG24 competition. Subjects mimic the facial expressions of an avatar, displaying fear, joy, anger, sadness, disgust, and surprise in a VR setting. EEG data is captured using a multi-channel sensor system to discern brain activity patterns. We propose a novel two-stream neural network employing a Bi-Hemispheric approach for emotion inference, surpassing baseline methods and enhancing emotion recognition accuracy. Additionally, we conduct a temporal analysis revealing that specific signal intervals at the beginning and end of the emotion stimulus sequence contribute significantly to improve accuracy. Leveraging insights gained from this temporal analysis, our approach offers enhanced performance in capturing subtle variations in the states of emotions. Code is available at https://github.com/davidfreire/FG24-EmoNeuroDB/ David Freire-Obregón, Daniel Hernández-Sosa, Oliverio J. Santana, Javier Lorenzo-Navarro, Modesto Castrillón-Santana |
FG | 5 |
| 2024 | An Evaluation of General-Purpose Optical Character Recognizers and Digit Detectors for Race Bib Number Recognition
Modesto Castrillón-Santana, David Freire-Obregón, Daniel Hernández-Sosa, Oliverio J. Santana, Francisco Ortega-Zamorano, José Isern González, Javier Lorenzo-Navarro |
ICPRAM | 1 |
| 2024 | Applying deep learning image enhancement methods to improve person re-identificationabstractPerson re-identification has gained significant attention in recent years due to its numerous practical applications in video surveillance. However, while artificial intelligence and deep learning methods have enabled substantial progress in particular aspects of this domain, putting together those individual advances to generate practical systems remains a computer vision challenge. Existing methods are typically designed assuming the target person’s images are captured under uniform, stable conditions with similar lighting levels, but this assumption may not hold in real-world scenarios, such as outdoor monitoring over 24 h, as image quality can vary considerably throughout day and night. In this paper, we propose a framework that incorporates image enhancement techniques to improve the performance of a person re-identification model. The proposed approach achieves a significant improvement in a demanding re-identification dataset, raising the mAP from 9.0% using a zero-shot baseline to 65.8% through the combined use of low-light image enhancement methods and noise reduction. Oliverio J. Santana, Javier Lorenzo-Navarro, David Freire-Obregón, Daniel Hernández-Sosa, Modesto Castrillón-Santana |
Neurocomputing | 5 |
| 2023 | Evaluation of a Visual Question Answering Architecture for Pedestrian Attribute Recognition
Modesto Castrillón-Santana, Elena Sánchez-Nielsen, David Freire-Obregón, Oliverio J. Santana, Daniel Hernández-Sosa, Javier Lorenzo-Navarro |
CAIP (1) | 1 |
| 2023 | Novelty Detection in Human-Machine Interaction Through a Multimodal Approach
José Salas-Cáceres, Javier Lorenzo-Navarro, David Freire-Obregón, Modesto Castrillón-Santana |
CIARP | 4 |
| 2023 | A Large-Scale Re-identification Analysis in Sporting Scenarios: the Betrayal of Reaching a Critical PointabstractRe-identifying participants in ultra-distance running competitions can be daunting due to the extensive distances and constantly changing terrain. To overcome these challenges, computer vision techniques have been developed to analyze runners’ faces, numbers on their bibs, and clothing. However, our study presents a novel gait-based approach for runners’ re-identification (re-ID) by leveraging various pre-trained human action recognition (HAR) models and loss functions. Our results show that this approach provides promising results for re-identifying runners in ultra-distance competitions. Furthermore, we investigate the significance of distinct human body movements when athletes are approaching their endurance limits and their potential impact on re-ID accuracy. Our study examines how the recognition of a runner’s gait is affected by a competition’s critical point (CP), defined as a moment of severe fatigue and the point where the finish line comes into view, just a few kilometers away from this location. We aim to determine how this CP can improve the accuracy of athlete re-ID. Our experimental results demonstrate that gait recognition can be significantly enhanced (up to a 9% increase in mAP) as athletes approach this point. This highlights the potential of utilizing gait recognition in real-world scenarios, such as ultra-distance competitions or long-duration surveillance tasks. David Freire-Obregón, Javier Lorenzo-Navarro, Oliverio J. Santana, Daniel Hernández-Sosa, Modesto Castrillón-Santana |
IJCB | 5 |
| 2023 | Evaluating the Impact of Low-Light Image Enhancement Methods on Runner Re-Identification in the WildabstractPerson re-identification (ReID) is a trending topic in computer vision. Significant developments have been achieved, but most rely on datasets with subjects captured statically within a short period of time in rather good lighting conditions. In the wild scenarios, such as long-distance races that involve widely varying lighting conditions, from full daylight to night, present a considerable challenge. This issue cannot be addressed by increasing the exposure time on the capture device, as the runners' motion will lead to blurred images, hampering any ReID attempts. In this paper, we survey some low-light image enhancement methods. Our results show that including an image processing step in a ReID pipeline before extracting the distinctive body appearance features from the subjects can provide significant performance improvements. Oliverio J. Santana, Javier Lorenzo-Navarro, David Freire-Obregón, Daniel Hernández-Sosa, Modesto Castrillón-Santana |
ICPRAM | 5 |
| 2023 | Refactoring and performance analysis of the main CNN architectures: using false negative rate minimization to solve the clinical images melanoma detection problemabstractBACKGROUND: Melanoma is one of the deadliest tumors in the world. Early detection is critical for first-line therapy in this tumor pathology and it remains challenging due to the need for histological analysis to ensure correctness in diagnosis. Therefore, multiple computer-aided diagnosis (CAD) systems working on melanoma images were proposed to mitigate the need of a biopsy. However, although the high global accuracy is declared in literature results, the CAD systems for the health fields must focus on the lowest false negative rate (FNR) possible to qualify as a diagnosis support system. The final goal must be to avoid classification type 2 errors to prevent life-threatening situations. Another goal could be to create an easy-to-use system for both physicians and patients. RESULTS: To achieve the minimization of type 2 error, we performed a wide exploratory analysis of the principal convolutional neural network (CNN) architectures published for the multiple image classification problem; we adapted these networks to the melanoma clinical image binary classification problem (MCIBCP). We collected and analyzed performance data to identify the best CNN architecture, in terms of FNR, usable for solving the MCIBCP problem. Then, to provide a starting point for an easy-to-use CAD system, we used a clinical image dataset (MED-NODE) because clinical images are easier to access: they can be taken by a smartphone or other hand-size devices. Despite the lower resolution than dermoscopic images, the results in the literature would suggest that it would be possible to achieve high classification performance by using clinical images. In this work, we used MED-NODE, which consists of 170 clinical images (70 images of melanoma and 100 images of naevi). We optimized the following CNNs for the MCIBCP problem: Alexnet, DenseNet, GoogleNet Inception V3, GoogleNet, MobileNet, ShuffleNet, SqueezeNet, and VGG16. CONCLUSIONS: The results suggest that a CNN built on the VGG or AlexNet structure can ensure the lowest FNR (0.07) and (0.13), respectively. In both cases, discrete global performance is ensured: 73% (accuracy), 82% (sensitivity) and 59% (specificity) for VGG; 89% (accuracy), 87% (sensitivity) and 90% (specificity) for AlexNet. Luigi Di Biasi, Fabiola De Marco, Alessia Auriemma Citarella, Modesto Castrillón-Santana, Paola Barra, Genny Tortora |
BMC Bioinform. | 4 |
| 2023 | Facial expression analysis in a wild sporting environmentabstractThe scientific community and mass media have already reported the use of nonverbal behavior analysis in sports for athletes' performance. Their conclusions stated that certain emotional expressions are linked to athlete's performance, or even that psychological strategies serve to improve endurance performance. This paper examines the portrayal of well-known emotions and their relationship to the participants of an ultra-distance race in a high-stake environment. For this purpose, we analyzed almost 600 runners captured when they passed through a set of locations placed along the race track. We have observed a correlation between the runners' facial expressions and their performance along the track. Moreover, we have analyzed Action Unit activations and aligned our findings with the state-of-the-art psychological baseline. Oliverio J. Santana, David Freire-Obregón, Daniel Hernández-Sosa, Javier Lorenzo-Navarro, Elena Sánchez-Nielsen, Modesto Castrillón-Santana |
Multim. Tools Appl. | 6 |
| 2023 | Zero-shot ear cross-dataset transfer for person recognition on mobile devicesabstractSmartphones contain personal and private data to be protected, such as everyday communications or bank accounts. Several biometric techniques have been developed to unlock smartphones, among which ear biometrics represents a natural and promising opportunity even though the ear can be used in other biometric and multi-biometric applications. A problem in generalizing research results to real-world applications is that the available ear datasets present different characteristics and some bias. This paper stems from a study about the effect of mixing multiple datasets during the training of an ear recognition system. The main contribution is the evaluation of a robust pipeline that learns to combine data from different sources and highlights the importance of pre-training encoders on auxiliary tasks. The reported experiments exploit eight diverse training datasets to demonstrate the generalization capabilities of the proposed approach. Performance evaluation includes testing with collections not seen during training and assessing zero-shot cross-dataset transfer. The results confirm that mixing different sources provides an insightful perspective on the datasets and competitive results with some existing benchmarks. David Freire-Obregón, Maria De Marsico, Paola Barra, Javier Lorenzo-Navarro, Modesto Castrillón-Santana |
Pattern Recognit. Lett. | 5 |
| 2022 | Towards cumulative race time regression in sports: I3D ConvNet transfer learning in ultra-distance running eventsabstractPredicting an athlete’s performance based on short footage is highly challenging. Performance prediction requires high domain knowledge and enough evidence to infer an appropriate quality assessment. Sports pundits can often infer this kind of information in real-time. In this paper, we propose regressing an ultra-distance runner cumulative race time (CRT), i.e., the time the runner has been in action since the race start, by using only a few seconds of footage as input. We modified the I3D ConvNet backbone slightly and trained a newly added regressor for that purpose. We use appropriate pre-processing of the visual input to enable transfer learning from a specific runner. We show that the resulting neural network can provide a remarkable performance for short input footage: 18 minutes and a half mean absolute error in estimating the CRT for runners who have been in action from 8 to 20 hours. Our methodology has several favorable properties: it does not require a human expert to provide any insight, it can be used at any moment during the race by just observing a runner, and it can inform the race staff about a runner at any given time. David Freire-Obregón, Javier Lorenzo-Navarro, Oliverio J. Santana, Daniel Hernández-Sosa, Modesto Castrillón-Santana |
ICPR | 5 |
| 2022 | Boosting Re-identification in the Ultra-running Scenario
Miguel Angel Medina, Javier Lorenzo-Navarro, David Freire-Obregón, Oliverio J. Santana, Daniel Hernández-Sosa, Modesto Castrillón-Santana |
ICPRAM | 6 |
| 2022 | Inflated 3D ConvNet context analysis for violence detectionabstractAbstract According to the Wall Street Journal, one billion surveillance cameras will be deployed around the world by 2021. This amount of information can be hardly managed by humans. Using a Inflated 3D ConvNet as backbone, this paper introduces a novel automatic violence detection approach that outperforms state-of-the-art existing proposals. Most of those proposals consider a pre-processing step to only focus on some regions of interest in the scene, i.e., those actually containing a human subject. In this regard, this paper also reports the results of an extensive analysis on whether and how the context can affect or not the adopted classifier performance. The experiments show that context-free footage yields substantial deterioration of the classifier performance (2% to 5%) on publicly available datasets. However, they also demonstrate that performance stabilizes in context-free settings, no matter the level of context restriction applied. Finally, a cross-dataset experiment investigates the generalizability of results obtained in a single-collection experiment (same dataset used for training and testing) to cross-collection settings (different datasets used for training and testing). David Freire-Obregón, Paola Barra, Modesto Castrillón-Santana, Maria De Marsico |
Mach. Vis. Appl. | 3 |
| 2022 | Editorial for the special issue on implicit biometric authentication and monitoring through Internet of Biometric Things (I-BIO)
Stefano Ricciardi, Modesto Castrillón-Santana |
Pattern Recognit. Lett. | 2 |
| 2021 | Improving user verification in human-robot interaction from audio or image inputs through sample quality assessmentabstractIn this paper, we tackle the task of improving biometric verification in the context of Human-Robot Interaction (HRI). A robot that wants to identify a specific person to provide a service can do so by either image verification or, if light conditions are not favourable, through voice verification. In our approach, we will take advantage of the possibility a robot has of recovering further data until it is sure of the identity of the person. The key contribution is that we select from both image and audio signals the parts that are of higher confidence. For images we use a system that looks at the face of each person and selects frames in which the confidence is high while keeping those frames separate in time to avoid using very similar facial appearance . For audio our approach tries to find the parts of the signal that contain a person talking, avoiding those in which noise is present by segmenting the signal. Once the parts of interest are found, each input is described with an independent deep learning architecture that obtains a descriptor for each kind of input (face/voice). We also present in this paper fusion methods that improve performance by combining the features from both face and voice, results to validate this are shown for each independent input and for the fusion methods. David Freire-Obregón, Kevin Rosales-Santana, Pedro A. Marín-Reyes, Adrián Peñate Sánchez, Javier Lorenzo-Navarro, Modesto Castrillón-Santana |
Pattern Recognit. Lett. | 6 |
| 2020 | TGCRBNW: A Dataset for Runner Bib Number Detection (and Recognition) in the WildabstractRacing bib number (RBN) detection and recognition is a specific problem related to text recognition in natural scenes. In this paper, we present a novel dataset created after registering participants in a real ultrarunning competition which comprises a wide range of acquisition conditions in five different recording points, including nightlight and daylight. The dataset contains more than 3K samples of over 400 different individuals. The aim is to provide an “in the wild” benchmark for both RBN detection and recognition problems. To illustrate the present difficulties, the dataset is evaluated for RBN detection using different Faster R-CNN specific detection models, filtering its output with heuristics based on body detection to improve the overall detection performance. Initial results are promising, but there is still significant room for improvement. And detection is just the first step to accomplish “in the wild” RBN recognition. Pablo Hernández-Carrascosa, Adrián Peñate Sánchez, Javier Lorenzo-Navarro, David Freire-Obregón, Modesto Castrillón-Santana |
ICPR | 5 |
| 2020 | An attention recurrent model for human cooperation detection
David Freire-Obregón, Modesto Castrillón-Santana, Paola Barra, Carmen Bisogni, Michele Nappi |
Comput. Vis. Image Underst. | 2 |
| 2020 | TGC20ReId: A dataset for sport event re-identification in the wild
Adrián Peñate Sánchez, David Freire-Obregón, Adrián Lorenzo-Melián, Javier Lorenzo-Navarro, Modesto Castrillón-Santana |
Pattern Recognit. Lett. | 5 |
| 2019 | AveRobot: An Audio-visual Dataset for People Re-identification and Verification in Human-Robot InteractionabstractIntelligent technologies have pervaded our daily life, making it easier for people to complete their activities. One emerging application is involving the use of robots for assisting people in various tasks (e.g., visiting a museum). In this context, it is crucial to enable robots to correctly identify people. Existing robots often use facial information to establish the identity of a person of interest. But, the face alone may not offer enough relevant information due to variations in pose, illumination, resolution and recording distance. Other biometric modalities like the voice can improve the recognition performance in these conditions. However, the existing datasets in robotic scenarios usually do not include the audio cue and tend to suffer from one or more limitations: most of them are acquired under controlled conditions, limited in number of identities or samples per user, collected by the same recording device, and/or not freely available. In this paper, we propose AveRobot, an audio-visual dataset of 111 participants vocalizing short sentences under robot assistance scenarios. The collection took place into a three-floor building through eight different cameras with built-in microphones. The performance for face and voice re-identification and verification was evaluated on this dataset with deep learning baselines, and compared against audio-visual datasets from diverse scenarios. The results showed that AveRobot is a challenging dataset for people re-identification and verification. Mirko Marras, Pedro A. Marín-Reyes, Javier Lorenzo-Navarro, Modesto Castrillón-Santana, Gianni Fenu |
ICPRAM | 4 |
| 2019 | Deep learning for source camera identification on mobile devices
David Freire-Obregón, Fabio Narducci, Silvio Barra, Modesto Castrillón-Santana |
Pattern Recognit. Lett. | 4 |
| 2018 | Automatic Counting and Classification of Microplastic ParticlesabstractMicroplastic particles have become an important ecological problem due to the huge amount of plastics debris that ends up in the sea. An additional impact is the ingestion of microplastics by marine species, and thus microplastics enter into the food chain with unpredictable effects on humans. In addition to the exploration of their presence in fishes, researchers are studying the presence of microplastics in coastal areas. The workload is therefore time consuming, due to the need to carry out regular campaigns to quantify their presence in the samples. So, in this work a method for automatic counting and classifying microplastic particles is presented. To the best of our knowledge, this is the first proposal to address this challenging problem. The method makes use of Computer Vision techniques for analyzing the acquired images of the samples; and Machine Learning techniques to develop accurate classifiers of the different types of microplastic particles that are considered. The obtained results show that making use of color based and shape based features along with a Random Forest classifier, an accuracy of 96.6% is achieved recognizing four types of particles: pellets, fragments, tar and line. Javier Lorenzo-Navarro, Modesto Castrillón-Santana, May Gómez, Alicia Herrera, Pedro A. Marín-Reyes |
ICPRAM | 2 |
| 2018 | Evaluation of local descriptors and CNNs for non-adult detection in visual contentabstractThe recent evolution of storage devices, digital embedded cameras and the Internet have collaterally allowed sexual predators to take advantage of these technological breakthroughs to gather illegal media, which is exhibited uncensored through Peer-to-Peer file sharing networks. In this paper, we are particularly concerned about the increasing availability of Child Abuse Material. Therefore, we have explored alternatives to detect non-adults in visual content. Initially, different age estimations and underage detection techniques are reviewed by analyzing existing datasets. Finally, several local descriptors and Convolutional Neural Networks for underage detection are evaluated. The experimental results obtained for a large dataset that combines collections such as FG-Net, Adience, GenderChildren, The Image of Groups and Boys2Men evidence the complementary information contained in both local descriptors and neural networks , as their fusion boosts the accuracy of non-adult detection to over 93%. Modesto Castrillón-Santana, Javier Lorenzo-Navarro, Carlos Manuel Travieso-González, David Freire-Obregón, Jesús B. Alonso |
Pattern Recognit. Lett. | 1 |
| 2017 | MEG: Texture operators for multi-expert gender classification
Modesto Castrillón-Santana, Maria De Marsico, Michele Nappi, Daniel Riccio |
Comput. Vis. Image Underst. | 1 |
| 2017 | Descriptors and regions of interest fusion for in- and cross-database gender classification in the wild
Modesto Castrillón-Santana, Javier Lorenzo-Navarro, Enrique Ramón-Balmaseda |
Image Vis. Comput. | 1 |
| 2017 | A multimedia system to produce and deliver video fragments on demand on parliamentary websites
Elena Sánchez-Nielsen, Francisco Chávez-Gutiérrez, Javier Lorenzo-Navarro, Modesto Castrillón-Santana |
Multim. Tools Appl. | 4 |
| 2017 | Multi-scale score level fusion of local descriptors for gender classification in the wild
Modesto Castrillón-Santana, Javier Lorenzo-Navarro, Enrique Ramón-Balmaseda |
Multim. Tools Appl. | 1 |
| 2017 | Periocular and iris local descriptors for identity verification in mobile applications
Naiara Aginako, Modesto Castrillón-Santana, Javier Lorenzo-Navarro, José María Martínez-Otzeta, Basilio Sierra |
Pattern Recognit. Lett. | 2 |
| 2016 | Local descriptors fusion for mobile iris verificationabstractThis paper summarizes the proposal submitted by the joint team conformed by researchers from UPV and ULPGC to the Mobile Iris CHallenge Evaluation II. The approach makes use of a state-of-the-art iris segmentation technique, to later extract features making use of local descriptors. Those suitable to the problem are selected after evaluating a collection of 15 local descriptors, covering a range of different grid configuration setups. A Machine Learning approach is used, learning a supervised classifier to deal with the descriptors data. A classifier is obtained for each descriptor, and the best ones are combined in a multi-classifier system. The final step fuses the classifier outputs obtained for 5 different local descriptors, to compute the dissimilarity measure for a pair of iris images. Naiara Aginako, José María Martínez-Otzeta, Basilio Sierra, Modesto Castrillón-Santana, Javier Lorenzo-Navarro |
ICPR | 4 |
| 2016 | Mobile Iris CHallenge Evaluation II: Results from the ICPR competitionabstractThe growing interest for mobile biometrics stems from the increasing need to secure personal data and services, which are often stored or accessed from there. Modern user mobile devices, with acquisition and computation resources to support related operations, are nowadays widely available. This makes this research topic very attracting and promising. Iris recognition plays a major role in this scenario. However, mobile biometrics still suffer from some hindering factors. The resolution of captured images and the computational power are not comparable to desktop systems yet. Furthermore, the acquisition setting is generally uncontrolled, with users who are not that expert to autonomously generate biometric samples of sufficient quality. Mobile Iris CHallenge Evaluation aims at providing a testbed to assess the progress of mobile iris recognition, and to evaluate the extent of its present limitations. This paper presents the results of the competition launched at the 2016 edition of the International Conference on Pattern Recognition (ICPR). Modesto Castrillón-Santana, Maria De Marsico, Michele Nappi, Fabio Narducci, Hugo Proença 0001 |
ICPR | 1 |
| 2016 | On using periocular biometric for gender classification in the wild
Modesto Castrillón-Santana, Javier Lorenzo-Navarro, Enrique Ramón-Balmaseda |
Pattern Recognit. Lett. | 1 |
| 2015 | An Evolutive Approach for Smile Recognition in Video SequencesabstractFacial expression recognition is one of the most challenging research areas in the image recognition field and has been actively studied since the 70's. For instance, smile recognition has been studied due to the fact that it is considered an important facial expression in human communication, it is therefore likely useful for human–machine interaction. Moreover, if a smile can be detected and also its intensity estimated, it will raise the possibility of new applications in the future. We are talking about quantifying the emotion at low computation cost and high accuracy. For this aim, we have used a new support vector machine (SVM)-based approach that integrates a weighted combination of local binary patterns (LBPs)-and principal component analysis (PCA)-based approaches. Furthermore, we construct this smile detector considering the evolution of the emotion along its natural life cycle. As a consequence, we achieved both low computation cost and high performance with video sequences. David Freire-Obregón, Modesto Castrillón-Santana |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2014 | Kinship verification in the wild: The first kinship verification competitionabstractKinship verification from facial images in wild conditions is a relatively new and challenging problem in face analysis. Several datasets and algorithms have been proposed in recent years. However, most existing datasets are of small sizes and one standard evaluation protocol is still lack so that it is difficult to compare the performance of different kinship verification methods. In this paper, we present the Kinship Verification in the Wild Competition: the first kinship verification competition which is held in conjunction with the International Joint Conference on Biometrics 2014, Clearwater, Florida, USA. The key goal of this competition is to compare the performance of different methods on a new-collected dataset with the same evaluation protocol and develop the first standardized benchmark for kinship verification in the wild. Jiwen Lu, Junlin Hu 0001, Xiuzhuang Zhou, Jie Zhou 0001, Modesto Castrillón-Santana, Javier Lorenzo-Navarro, Lu Kou, Andrea Bottino, Tiago F. Vieira |
IJCB | 5 |
| 2014 | Automatic clothes segmentation for soft biometricsabstractDuring the last decade, researchers have verified that clothing can provide information for gender recognition. However, before extracting features, it is necessary to segment the clothing region. We introduce a new clothes segmentation method based on the application of the GrabCut technique over a trixel mesh, obtaining very promising results for a close to real time system. Finally, the clothing features are combined with facial and head context information to outperform previous results in gender recognition with a public database. David Freire-Obregón, Modesto Castrillón-Santana, Enrique Ramón-Balmaseda, Javier Lorenzo-Navarro |
ICIP | 2 |
| 2014 | People Semantic Description and Re-identification from Point Cloud GeometryabstractThe automatic extraction of biometric descriptors of anonymous people is a challenging scenario in camera networks. This task is typically accomplished making use of visual information. Calibrated RGBD sensors make possible the extraction of point cloud information. We present a novel approach for people semantic description and re-identification using the individual point cloud information. The proposal combines the use of simple geometric features with point cloud features based on surface normals. To test the system validity, we have collected a new and challenging dataset using a RGBD sensor in a top view configuration containing up to 63 identities captured in different sessions in different days within a two weeks period. The results achieved outperform the previous literature based exclusively on geometric features for re-identification, providing additionally very promising results in people description related to gender and hair style. Modesto Castrillón-Santana, Javier Lorenzo-Navarro, Daniel Hernández-Sosa |
ICPR | 1 |
| 2013 | Improving Gender Classification Accuracy in the Wild
Modesto Castrillón-Santana, Javier Lorenzo-Navarro, Enrique Ramón-Balmaseda |
CIARP (2) | 1 |
| 2012 | Gender Classification in Large Databases
Enrique Ramón-Balmaseda, Javier Lorenzo-Navarro, Modesto Castrillón-Santana |
CIARP | 3 |
| 2012 | Combining Face and Facial Feature Detectors for Face Detection Performance Improvement
Modesto Castrillón-Santana, Daniel Hernández-Sosa, Javier Lorenzo-Navarro |
CIARP | 1 |
| 2011 | Experiments in Short-term Wind Power Prediction using Variable Selection
Javier Lorenzo-Navarro, Juan Méndez, Daniel Hernández-Sosa, Modesto Castrillón-Santana |
ICAART (1) | 4 |
| 2011 | Competition on counter measures to 2-D facial spoofing attacksabstractSpoofing identities using photographs is one of the most common techniques to attack 2-D face recognition systems. There seems to exist no comparative studies of different techniques using the same protocols and data. The motivation behind this competition is to compare the performance of different state-of-the-art algorithms on the same database using a unique evaluation method. Six different teams from universities around the world have participated in the contest. Use of one or multiple techniques from motion, texture analysis and liveness detection appears to be the common trend in this competition. Most of the algorithms are able to clearly separate spoof attempts from real accesses. The results suggest the investigation of more complex attacks. Murali Mohan Chakka, André Anjos, Sébastien Marcel, Roberto Tronci, Daniele Muntoni, Gianluca Fadda, Maurizio Pili, Nicola Sirena, Gabriele Murgia, Marco Ristori, Fabio Roli, Dong Yi, Zhen Lei 0001, Stan Z. Li, William Robson Schwartz, Anderson Rocha 0001, Hélio Pedrini, Javier Lorenzo-Navarro, Modesto Castrillón-Santana, Jukka Komulainen, Abdenour Hadid, Matti Pietikäinen |
IJCB | 21 |
| 2011 | Face and eye detection on hard datasetsabstractFace and eye detection algorithms are deployed in a wide variety of applications. Unfortunately, there has been no quantitative comparison of how these detectors perform under difficult circumstances. We created a dataset of low light and long distance images which possess some of the problems encountered by face and eye detectors solving real world problems. The dataset we created is composed of reimaged images (photohead) and semi-synthetic heads imaged under varying conditions of low light, atmospheric blur, and distances of 3m, 50m, 80m, and 200m. This paper analyzes the detection and localization performance of the participating face and eye algorithms compared with the Viola Jones detector and four leading commercial face detectors. Performance is characterized under the different conditions and parameterized by per-image brightness and contrast. In localization accuracy for eyes, the groups/companies focusing on long-range face detection outperform leading commercial applications. Jonathan Parris, Kimberly Wilber, Brian Heflin, Ham M. Rara, Ahmed El-Barkouky, Aly A. Farag, Javier R. Movellan, Modesto Castrillón-Santana, Javier Lorenzo-Navarro, Mohammad Nayeem Teli, Sébastien Marcel, Cosmin Atanasoaei, Terrance E. Boult |
IJCB | 9 |
| 2011 | A comparison of face and facial feature detectors based on the Viola-Jones general object detection framework
Modesto Castrillón-Santana, Oscar Déniz-Suárez, Daniel Hernández-Sosa, Javier Lorenzo-Navarro |
Mach. Vis. Appl. | 1 |
| 2010 | Learning to recognize gender using experienceabstractAutomatic facial analysis abilities are commonly integrated in a system by a previous off-line learning stage. In this paper we argue that a facial analysis system would improve its facial analysis capabilities based on its own experience similarly to the way a biological system, i.e. the human system, does throughout the years. The approach described, focused on gender classification, updates its knowledge according to the classification results. The presented gender experiments suggest that this approach is promising, even when just a short simulation of what for humans would take years of acquisition experience was performed. Modesto Castrillón-Santana, Javier Lorenzo-Navarro, David Freire-Obregón, Oscar Déniz-Suárez |
ICIP | 1 |
| 2010 | Computer vision based eyewear selectorabstractThe widespread availability of portable computing power and inexpensive digital cameras are opening up new possibilities for retailers in some markets. One example is in optical shops, where a number of systems exist that facilitate eyeglasses selection. These systems are now more necessary as the market is saturated with an increasingly complex array of lenses, frames, coatings, tints, photochromic and polarizing treatments, etc. Research challenges encompass Computer Vision, Multimedia and Human-Computer Interaction. Cost factors are also of importance for widespread product acceptance. This paper describes a low-cost system that allows the user to visualize different glasses models in live video. The user can also move the glasses to adjust its position on the face. The system, which runs at 9.5 frames/s on general-purpose hardware, has a homeostatic module that keeps image parameters controlled. This is achieved by using a camera with motorized zoom, iris, white balance, etc. This feature can be specially useful in environments with changing illumination and shadows, like in an optical shop. The system also includes a face and eye detection module and a glasses management module. Oscar Déniz-Suárez, Modesto Castrillón-Santana, Javier Lorenzo-Navarro, Luis Antón-Canalís, Mario Hernández-Tejera, Gloria Bueno García |
J. Zhejiang Univ. Sci. C | 2 |
| 2008 | Automatic Initialization for Facial Analysis in Interactive Robotics
Ahmad Rabie, Christian Lang 0002, Marc Hanheide, Modesto Castrillón-Santana, Gerhard Sagerer |
ICVS | 4 |
| 2007 | An Analysis of Automatic Gender Classification
Modesto Castrillón-Santana, Quoc C. Vuong |
CIARP | 1 |
| 2007 | An engineering approach to sociable robotsabstractRobotics researchers and cognitive scientists are becoming more and more interested in so-called sociable robots. These machines normally have expressive power (facial features, voice, …) as well as abilities for locating, paying attention to, and addressing people. The design objective is to make robots which are able to sustain natural interactions with people. This capacity falls within the range classed as social intelligence in humans. This position paper argues that the reproduction of social intelligence, as opposed to other types of human ability, may lead to fragile performance, in the sense that tested cases may produce rather different performances to future (untested) cases and situations. This limitation stems from the fact that our social abilities, which appear early in life, are mainly unconscious in origin. This is in contrast with other human abilities that we carry out using conscious effort, and for which we can easily conceive algorithms and representations. This novel perspective is deemed useful for defining the obstacles and limitations of a field that is generating increasing interest. Taking into account the mentioned issues, a development approach suited to the problem is proposed. The use of this approach is demonstrated in the development of CASIMIRO, a robotic head with basic interaction abilities. Oscar Déniz-Suárez, Mario Hernández-Tejera, Javier Lorenzo-Navarro, Modesto Castrillón-Santana |
J. Exp. Theor. Artif. Intell. | 4 |
| 2007 | ENCARA2: Real-time detection of multiple faces at different resolutions in video streams
Modesto Castrillón-Santana, Oscar Déniz-Suárez, Cayetano Guerra, Mario Hernández-Tejera |
J. Vis. Commun. Image Represent. | 1 |
| 2003 | ENCARA: real-time detection of frontal facesabstractThis paper describes a real-time approach for face detection and selection of frontal views, for further processing. Typically, face detection papers provide results for a set of single images but the problem of face detection in video streams rarely is tackled. Instead of performing an exhaustive search for every video stream frame a set of opportunistic ideas applied in a cascade fashion and based on temporal and spatial coherence provide promising results in real-time. Modesto Castrillón-Santana, Mario Hernández-Tejera, Jorge Cabrera-Gámez 0001 |
ICIP (3) | 1 |
| 2003 | Face recognition using independent component analysis and support vector machines
Oscar Déniz-Suárez, Modesto Castrillón-Santana, Mario Hernández-Tejera |
Pattern Recognit. Lett. | 2 |
| 1999 | DESEO: An Active Vision System for Detection, Tracking and Recognition
Mario Hernández-Tejera, Jorge Cabrera-Gámez 0001, Antonio Carlos Domínguez-Brito, Modesto Castrillón-Santana, Cayetano Guerra, Daniel Hernández-Sosa, José Isern González |
ICVS | 4 |