Sergio Escalera

dblp:77/5527 · also Sergio Escalera Guerrero · DBLP profile ↗
← Back
214ranked-venue papers
29as first author
81since 2021 · last 2026
0000-0003-0617-8873ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 164 · 23 first-author · 53 since 2021Graphics, computer vision, multimedia, augmented reality and games · 99 · 9 first-author · 41 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 4D Point Cloud Segmentation via Active Test-Time Adaptation
abstract
4D point cloud segmentation is crucial for autonomous driving with continuous LiDAR streams. While test-time adaptation (TTA) is the standard approach for handling dynamic environments, current methods suffer from catastrophic error accumulation due to over-reliance on pseudo-labels. Active learning could provide reliable annotations for critical samples, but combining it with TTA faces severe challenges: realtime processing requirements and expensive 3D labeling costs. In this paper, we propose ATTA-4DSeg, the first framework to achieve efficient active test-time adaptation for 4D point cloud segmentation under extreme budget constraints. Our key insight is a self-reinforcing loop: oracle annotations refine adaptation prototypes, which then guide the selection of subsequent high-value samples from regions with severe distribution shifts, maximizing each annotation’s impact. Specifically, we propose three key innovations: (1) dual-prototype comparison that precisely localizes distribution shift boundaries to narrow annotation scope, (2) Class-Inverse Budget Allocation (CIBA) ensuring balanced adaptation across all categories, coupled with hybrid uncertainty scoring combining voxel-level geometry and point-wise variance for optimal sample selection, and (3) a refinement strategy leveraging sparse oracle annotations to improve predictions on unlabeled points, maximizing annotation utility. Extensive experiments show ATTA-4DSeg improves mIoU by 18.87%, 19.92%, and 3.6% on three domain adaptation benchmarks using only 1% annotation budget. Our method operates 2.28× faster than state-of-the-art methods. Remarkably, our approach reaches 90% of fully-supervised performance using only 5% annotation budget.
Mingrong Gong, Chaoqi Chen, Luyao Tang, Sergio Escalera
AAAI5
2026 What Matters in Virtual Try-Off? Dual-UNet Diffusion Models for Garment Reconstruction
Loc-Phat Truong, Meysam Madadi, Sergio Escalera
ICPR (4)3
2026 DIGEST: Dynamic Graph Refinement with Dual Contrastive Semantic Transfer for Multimodal Recommendation
abstract
Multimodal recommendation benefits from leveraging rich content signals such as images and texts to alleviate interaction sparsity, yet existing graph-based approaches are still hindered by (i) noisy user—item edges that are treated as static during training and (ii) inconsistent representation spaces across interaction-driven and modality-induced graph views. To address these issues, we propose DIGEST, a multi-graph framework that propagates trainable ID embeddings on a denoised user—item graph and a fused modality-induced item—item graph, and interleaves message passing with dynamic graph refinement that iteratively reweights existing edges to suppress noisy connections. To enable reliable semantic transfer across views, DIGEST further introduces a dual contrastive alignment that (i) aligns the collaborative and semantic item views and (ii) constrains the semantic graph representations to projected multimodal features, together with a lightweight dimension decorrelation regularizer and adaptive gated fusion to reduce redundancy and stabilize multi-view learning. Extensive experiments on three Amazon benchmark datasets demonstrate that DIGEST consistently outperforms state-of-the-art multimodal recommenders, achieving up to 8.43% relative improvement on NDCG@20 and 7.66% on Recall@20 over the strongest baselines.
Xiangyu Sai, Meysam Madadi, Sergio Escalera, Yong Xu 0007
SIGIR3
2026 UniAttack: Unified Physical-Digital Face Attack Detection
Shunxin Chen, Ajian Liu 0001, Haocheng Yuan, Junze Zheng, Dingheng Zeng, Jiankang Deng, Sergio Escalera, Xiaoming Liu 0002, Jun Wan 0001, Zhen Lei 0001
Int. J. Comput. Vis.9
2026 Guest Editorial: Special Issue for the British Machine Vision Conference (BMVC), 2024 (Glasgow, Scotland, UK)
Carlos Francisco Moreno-García, Gerardo Aragon-Camarasa, Edmond S. L. Ho, Paul Henderson, Nicolas Pugeault, Jungong Han, Sergio Escalera
Int. J. Comput. Vis.7
2026 Personalized continuous sign language production via a motion-aware federated diffusion model
Razieh Rastgoo, Kourosh Kiani, Sergio Escalera
Neurocomputing3
2026 DGPDL: Domain-Guided Prompt Distribution Learning for Generalizable Face Anti-Spoofing
abstract
The overfitting of domain signals results in poor domain generalization of face anti-spoofing. The current methods usually improve the diversity of source domains to alleviate this overfitting. However, this benefit is minimal, as even the most diverse domain signals will also be absent in the target domain. In this work, we propose a Domain-Guided Prompt Distribution Learning (DGPDL) built on Vision-Language Models like CLIP, which explores a unified representation of domain signals as a prompt across the source and target domain to alleviate the understanding bias caused by domain gaps. Specifically, we first define a learnable Domain-Specific Distribution (DSD) that covers as many domain elements as possible, such as image quality, color tone, camera settings, etc., which establish connections between different domains and linearly combinable prompt in any domain; Then, based on the style statistics of the given sample, we construct its optimal Domain-Specific Prompts (DSPs) from the defined DSD through the designed Prompt Assemble Attention (PAA) with the similarity matching; Finally, the assembled DSPs will act as carrier or agent to perform on both the vision and language branches, synergistically improving the model's recognition of domain signals. By using the prompt to represent domain signals uniformly, if the model can be robust to DSPs in the source domain, it should be applicable to target domain, as they share the same DSD. By representing domain signals as prompts rather than instantiation features, DGPDL effectively reduces the reliance on specific domain appearances. This design enables the model to dynamically adapt to unseen target domains without the need for retraining. Extensive experiments show that the DGPDL is effective and outperforms the state-of-the-art methods on several cross-domain benchmarks.
Ajian Liu 0001, Xun Lin, Ruicong Zhi, Yanyan Liang 0001, Xinshan Zhu, Zhanchuan Cai, Jun Wan 0001, Sergio Escalera, Zhen Lei 0001
IEEE Trans. Pattern Anal. Mach. Intell.8
2025 Mixture-of-Attack-Experts with Class Regularization for Unified Physical-Digital Face Attack Detection
abstract
Unified detection of digital and physical attacks in facial recognition systems has become a focal point of research in recent years. However, current multi-modal methods typically ignore the intra-class and inter-class variability across different types of attacks, leading to degraded performance. To address this limitation, we propose MoAE-CR, a framework that effectively leverages class-aware information for improved attack detection. Our improvements manifest at two levels, i.e., the feature and loss level. At the feature level, we propose Mixture-of-Attack-Experts (MoAEs) to capture more subtle differences among various types of fake faces. At the loss level, we introduce Class Regularization (CR) through the Disentanglement Module (DM) and the Cluster Distillation Module (CDM). The DM enhances class separability by increasing the distance between the centers of live and fake face classes. However, center-to-center constraints alone are insufficient to ensure distinctive representations for individual features. Thus, we propose the CDM to further cluster features around their class centers while maintaining separation from other classes. Moreover, specific attacks that significantly deviate from common attack patterns are often overlooked. To address this issue, our distance calculation prioritizes more distant features. Extensive experiments on two unified physical-digital attack datasets demonstrate the state-of-the-art performance of the proposed method.
Shunxin Chen, Ajian Liu 0001, Junze Zheng, Jun Wan 0001, Kailai Peng, Sergio Escalera, Zhen Lei 0001
AAAI6
2025 From Sparse Signal to Smooth Motion: Real-Time Motion Generation with Rolling Prediction Models
abstract
In extended reality (XR), generating full-body motion of the users is important to understand their actions, drive their virtual avatars for social interaction, and convey a realistic sense of presence. While prior works focused on spatially sparse and always-on input signals from motion controllers, many XR applications opt for vision-based hand tracking for reduced user friction and better immersion. Compared to controllers, hand tracking signals are less accurate and can even be missing for an extended period of time. To handle such unreliable inputs, we present Rolling Prediction Model (RPM), an online and real-time approach that generates smooth full-body motion from temporally and spatially sparse input signals. Our model generates 1) accurate motion that matches the inputs (i.e., tracking mode) and 2) plausible motion when inputs are missing (i.e., synthesis mode). More importantly, RPM generates seamless transitions from tracking to synthesis, and vice versa. To demonstrate the practical importance of handling noisy and missing inputs, we present GORP, the first dataset of realistic sparse inputs from a commercial virtual reality (VR) headset with paired high quality body motion ground truth. GORP provides >14 hours of VR gameplay data from 28 people using motion controllers (spatially sparse) and hand tracking (spatially and temporally sparse). We benchmark RPM against the state of the art on both synthetic data and GORP to highlight how we can bridge the gap for real-world applications with a realistic dataset and by handling unreliable input signals. Our code, pretrained models, and GORP dataset are available in the project webpage.
Germán Barquero, Nadine Bertsch, Manojkumar Marramreddy, Carlos Chacón, Filippo Arcadu, Ferran Rigual, Nicky He, Cristina Palmero, Sergio Escalera, Yuting Ye, Robin Kips
CVPR9
2025 L-SWAG: Layer-Sample Wise Activation with Gradients Information for Zero-Shot NAS on Vision Transformers
abstract
Training-free Neural Architecture Search (NAS) efficiently identifies high-performing neural networks using zero-cost (ZC) proxies. Unlike multi-shot and one-shot NAS approaches, ZC-NAS is both (i) time-efficient, eliminating the need for model training, and (ii) interpretable, with proxy designs often theoretically grounded. Despite rapid developments in the field, current SOTA ZC proxies are typically constrained to well-established convolutional search spaces. With the rise of Large Language Models shaping the future of deep learning, this work extends ZC proxy applicability to Vision Transformers (ViTs). We present a new benchmark using the Autoformer search space evaluated on 6 distinct tasks and propose Layer-Sample Wise Activation with Gradients information (L-SWAG), a novel, generalizable metric that characterizes both convolutional and transformer architectures across 14 tasks. Additionally, previous works highlighted how different proxies contain complementary information, motivating the need for a ML model to identify useful combinations. To further enhance ZC-NAS, we therefore introduce LIBRA-NAS (Low Information gain and Bias Re-Alignment), a method that strategically combines proxies to best represent a specific benchmark. Integrated into the NAS search, LIBRA-NAS outperforms evolution and gradient-based NAS techniques by identifying an architecture with a 17.0% test error on ImageNet1k in just 0.1 GPU days.
Sofia Casarin, Sergio Escalera, Oswald Lanz
CVPR2
2025 MixerMDM: Learnable Composition of Human Motion Diffusion Models
abstract
Generating human motion guided by conditions such as textual descriptions is challenging due to the need for datasets with pairs of high-quality motion and their corresponding conditions. The difficulty increases when aiming for finer control in the generation. To that end, prior works have proposed to combine several motion diffusion models pre-trained on datasets with different types of conditions, thus allowing control with multiple conditions. However, the proposed merging strategies overlook that the optimal way to combine the generation processes might depend on the particularities of each pre-trained generative model and also the specific textual descriptions. In this context, we introduce MixerMDM, the first learnable model composition technique for combining pre-trained text-conditioned human motion diffusion models. Unlike previous approaches, MixerMDM provides a dynamic mixing strategy that is trained in an adversarial fashion to learn to combine the denoising process of each model depending on the set of conditions driving the generation. By using MixerMDM to combine single- and multi-person motion diffusion models, we achieve fine-grained control on the dynamics of every person individually, and also on the overall interaction. Furthermore, we propose a new evaluation technique that, for the first time in this task, measures the interaction and individual quality by computing the alignment between the mixed generated motions and their conditions as well as the capabilities of MixerMDM to adapt the mixing throughout the denoising process depending on the motions to mix.
Pablo Ruiz-Ponce, Germán Barquero, Cristina Palmero, Sergio Escalera, José García Rodríguez 0001
CVPR4
2025 Sparse-Dense Side-Tuner for Efficient Video Temporal Grounding
abstract
Video Temporal Grounding (VTG) involves Moment Retrieval (MR) and Highlight Detection (HD) based on textual queries. For this, most methods rely solely on final-layer features of frozen large pre-trained backbones, limiting their adaptability to new domains. While full fine-tuning is often impractical, parameter-efficient fine-tuning -- and particularly side-tuning (ST) -- has emerged as an effective alternative. However, prior ST approaches this problem from a frame-level refinement perspective, overlooking the inherent sparse nature of MR. To address this, we propose the Sparse-Dense Side-Tuner (SDST), the first anchor-free ST architecture for VTG. We also introduce the Reference-based Deformable Self-Attention, a novel mechanism that enhances the context modeling of the deformable attention -- a key limitation of existing anchor-free methods. Additionally, we present the first effective integration of InternVideo2 backbone into an ST framework, showing its profound implications in performance. Overall, our method significantly improves existing ST methods, achieving highly competitive or SOTA results on QVHighlights, TACoS, and Charades-STA, while reducing up to a 73% the parameter count w.r.t. the existing SOTA methods. The code is publicly accessible at https://github.com/davidpujol/SDST.
David Pujol-Perich, Sergio Escalera, Albert Clapés
ICCV2
2025 MANTRA: The Manifold Triangulations Assemblage
abstract
The rising interest in leveraging higher-order interactions present in complex systems has led to a surge in more expressive models exploiting higher-order structures in the data, especially in topological deep learning (TDL), which designs neural networks on higher-order domains such as simplicial complexes. However, progress in this field is hindered by the scarcity of datasets for benchmarking these architectures. To address this gap, we introduce MANTRA, the first large-scale, diverse, and intrinsically higher-order dataset for benchmarking higher-order models, comprising over 43,000 and 250,000 triangulations of surfaces and three-dimensional manifolds, respectively. With MANTRA, we assess several graph- and simplicial complex-based models on three topological classification tasks. We demonstrate that while simplicial complex-based neural networks generally outperform their graph-based counterparts in capturing simple topological invariants, they also struggle, suggesting a rethink of TDL. Thus, MANTRA serves as a benchmark for assessing and advancing topological methods, paving the way towards more effective higher-order models.
Rubén Ballester, Ernst Röell, Daniel Bin Schmid, Mathieu Alain, Sergio Escalera, Carles Casacuberta, Bastian Rieck
ICLR5
2025 REACT 2025: the Third Multiple Appropriate Facial Reaction Generation Challenge
abstract
In dyadic interactions, a broad spectrum of human facial reactions might be appropriate for responding to each human speaker behaviour. Following the successful organisation of the REACT 2023 and REACT 2024 challenges, we are proposing the REACT 2025 challenge encouraging the development and benchmarking of Machine Learning (ML) models that can be used to generate multiple appropriate, diverse, realistic and synchronised human-style facial reactions expressed by human listeners in response to an input stimulus (i.e., audio-visual behaviours expressed by their corresponding speakers). As a key of the challenge, we provide challenge participants with the first natural and large-scale multi-modal Multiple Appropriate Facial Reaction Generation (MAFRG) dataset (called MARS) recording 136 human-human dyadic interactions containing a total of 2856 interaction sessions covering five different topics. In addition, this paper also presents the challenge guidelines and the performance of our baselines on the two proposed sub-challenges: Offline MAFRG and Online MAFRG, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2025
Siyang Song, Micol Spitale, Xiangyu Kong 0001, Hengde Zhu, Cristina Palmero, Germán Barquero, Sergio Escalera, Michel F. Valstar, Mohamed Daoudi, Tobias Baur 0001, Fabien Ringeval, Andrew Howes 0001, Elisabeth André, Hatice Gunes
ACM Multimedia8
2025 SADA: Semantic Adversarial Unsupervised Domain Adaptation for Temporal Action Localization
abstract
Temporal Action Localization (TAL) is a complex task that poses relevant challenges, particularly when attempting to generalize on new-unseen-domains in realworld applications. These scenarios, despite realistic, are often neglected in the literature, exposing these solutions to important performance degradation. In this work, we tackle this issue by introducing, for the first time, an approach for Unsupervised Domain Adaptation (UDA) in sparse TAL, which we refer to as Semantic Adversarial unsupervised Domain Adaptation (SADA). Our contributions are threefold: (1) we pioneer the development of a domain adaptation model that operates on realistic sparse action detection benchmarks; (2) we tackle the limitations of global-distribution alignment techniques by introducing a novel adversarial loss that is sensitive to local class distributions, ensuring finer-grained adaptation; and (3) we present a novel set of benchmarks based on EpicKitchens100 and CharadesEgo, that evaluate multiple domain shifts in a comprehensive manner. Our experiments indicate that SADA improves the adaptation across domains when compared to fully supervised state-of-the-art and alternative UDA methods, attaining a performance boost of up to 6.14% mAP. The code is publicly available at https://github.com/davidpujol/SADA.
David Pujol-Perich, Albert Clapés, Sergio Escalera
WACV3
2025 Housed pig identification and tracking for precision livestock farming
Albert Compte, Yudong Yan, Xavier Cortés, Sergio Escalera, Júlio C. S. Jacques Júnior
Expert Syst. Appl.4
2025 Guest Editorial: Special Issue on Biometrics Security and Privacy
Jun Wan 0001, Arun Ross, Sergio Escalera
Int. J. Comput. Vis.3
2025 A deep generative Skeleton-based dynamic hand gesture production model
Razieh Rastgoo, Kourosh Kiani, Sergio Escalera
Multim. Tools Appl.3
2025 Exploring Emotion Expression Recognition in Older Adults Interacting With a Virtual Coach
abstract
The EMPATHIC project aimed to design an emotionally expressive virtual coach capable of engaging healthy seniors to improve well-being and promote independent aging. In particular, the system's human sensing capabilities allow for the perception of emotional states to provide a personalized experience. This paper outlines the development of the emotion expression recognition module of the virtual coach, encompassing data collection, annotation design, and a first methodological approach, all tailored to the project requirements. With the latter, we investigate the role of various modalities, individually and combined, for discrete emotion expression recognition in this context: speech from audio, and facial expressions, gaze, and head dynamics from video. The collected corpus includes users from Spain, France, and Norway, and was annotated separately for the audio and video channels with distinct emotional labels, allowing for a performance comparison across cultures and label types. Results confirm the informative power of the modalities studied for the emotional categories considered, with multimodal methods generally outperforming others (around 68% accuracy with audio labels and 72-74% with video labels). The findings are expected to contribute to the limited literature on emotion recognition applied to older adults in conversational human-machine interaction, and guide the development of future systems.
Cristina Palmero, Mikel de Velasco-Vázquez, Mohamed Amine Hmani, Aymen Mtibaa, Leila Ben Letaifa, Pau Buch-Cardona, Raquel Justo, Terry Amorese, Eduardo Gonzalez-Fraile, Begoña Fernández-Ruanova, Jofre Tenorio-Laranga, Anna Torp Johansen, Micaela Rodrigues da Silva, L. J. Martinussen, Maria Stylianou Korsnes, Gennaro Cordasco, Anna Esposito, Mounim A. El-Yacoubi, Dijana Petrovska-Delacrétaz, M. Inés Torres, Sergio Escalera
IEEE Trans. Affect. Comput.21
2025 Introduction to the Special Issue on Text-Multimedia Retrieval: Retrieving Multimedia Data by Means of Natural Language
Alex Falcon, Giuseppe Serra 0001, Sergio Escalera, Michael Wray
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Seamless Human Motion Composition with Blended Positional Encodings
abstract
Conditional human motion generation is an important topic with many applications in virtual reality, gaming, and robotics. While prior works have focused on generating motion guided by text, music, or scenes, these typically result in isolated motions confined to short durations. Instead, we address the generation of long, continuous sequences guided by a series of varying textual descriptions. In this context, we introduce FlowMDM, the first diffusion-based model that generates seamless Human Motion Compositions (HMC) without any postprocessing or redundant denoising steps. For this, we introduce the Blended Positional Encodings, a technique that leverages both absolute and relative positional encodings in the denoising chain. More specifically, global motion coherence is recovered at the absolute stage, whereas smooth and realistic transitions are built at the relative stage. As a result, we achieve state-of-the-art results in terms of accuracy, realism, and smoothness on the Babel and HumanML3D datasets. FlowMDM excels when trained with only a single description per motion sequence thanks to its Pose-Centric Cross-ATtention, which makes it robust against varying text descriptions at inference time. Finally, to address the limitations of existing HMC metrics, we propose two new metrics: the Peak Jerk and the Area Under the Jerk, to detect abrupt transitions.
Germán Barquero, Sergio Escalera, Cristina Palmero
CVPR2
2024 Your Image Is My Video: Reshaping the Receptive Field via Image-to-Video Differentiable AutoAugmentation and Fusion
abstract
The landscape of deep learning research is moving towards innovative strategies to harness the true potential of data. Traditionally, emphasis has been on scaling model architectures, resulting in large and complex neural networks, which can be difficult to train with limited computational resources. However, independently of the model size, data quality (i.e. amount and variability) is still a major factor that affects model generalization. In this work, we propose a novel technique to exploit available data through the use of automatic data augmentation for the tasks of image classification and semantic segmentation. We introduce the first Differentiable Augmentation Search method (DAS) to generate variations of images that can be processed as videos. Compared to previous approaches, DAS is extremely fast and flexible, allowing the search on very large search spaces in less than a GPU day. Our intuition is that the increased receptive field in the temporal dimension provided by DAS could lead to benefits also to the spatial receptive field. More specifically, we leverage DAS to guide the reshaping of the spatial receptive field by selecting task-dependant transformations. As a result, compared to standard augmentation alternatives, we improve in terms of accuracy on ImageNet, Cifar10, Cijar100, Tiny-ImageNet, Pascal-VOC-2012 and CityScapes datasets when plugging-in our DAS over different light-weight video backbones.
Sofia Casarin, Cynthia Ifeyinwa Ugwu, Sergio Escalera, Oswald Lanz
CVPR3
2024 A Noisy Elephant in the Room: Is Your out-of-Distribution Detector Robust to Label Noise?
abstract
The ability to detect unfamiliar or unexpected images is essential for safe deployment of computer vision systems. In the context of classification, the task of detecting images outside of a model's training domain is known as out-of-distribution (OOD) detection. While there has been a growing research interest in developing post-hoc OOD detection methods, there has been comparably little discussion around how these methods perform when the underlying classifier is not trained on a clean, carefully curated dataset. In this work, we take a closer look at 20 state-of-the-art OOD detection methods in the (more realistic) scenario where the labels used to train the underlying classifier are unreliable (e.g. crowd-sourced or web-scraped labels). Extensive experiments across different datasets, noise types & levels, architectures and checkpointing strategies provide insights into the effect of class label noise on OOD detection, and show that poor separation between incorrectly classified ID samples vs. OOD samples is an overlooked yet important limitation of existing methods. Code: https://github.com/glhr/ood-labelnoise
Galadrielle Humblot-Renaux, Sergio Escalera, Thomas B. Moeslund
CVPR2
2024 CFPL-FAS: Class Free Prompt Learning for Generalizable Face Anti-Spoofing
abstract
Domain generalization (DG) based Face Anti-Spoofing (FAS) aims to improve the model's performance on unseen domains. Existing methods either rely on domain labels to align domain-invariant feature spaces, or disentangle generalizable features from the whole sample, which inevitably lead to the distortion of semantic feature structures and achieve limited generalization. In this work, we make use of large-scale VLMs like CLIP and leverage the textual feature to dynamically adjust the classifier's weights for exploring generalizable visual features. Specifically, we propose a novel Class Free Prompt Learning (CFPL) paradigm for DG FAS, which utilizes two lightweight transformers, namely Content Q-Former (CQF) and Style Q-Former (SQF), to learn the different semantic prompts conditioned on content and style features by using a set of learnable query vectors, respectively. Thus, the generalizable prompt can be learned by two improvements: (1) A Prompt-Text Matched (PTM) supervision is introduced to ensure CQF learns visual representation that is most informative of the content description. (2) A Diversified Style Prompt (DSP) technology is proposed to diversify the learning of style prompts by mixing feature statistics between instance-specific styles. Finally, the learned text features modulate visual features to generalization through the designed Prompt Modulation (PM). Extensive experiments show that the CFPL is effective and outperforms the state-of-the-art methods on several cross-domain datasets.
Ajian Liu 0001, Jianwen Gan, Jun Wan 0001, Yanyan Liang 0001, Jiankang Deng, Sergio Escalera, Zhen Lei 0001
CVPR7
2024 Agglomerative Token Clustering
Joakim Bruslund Haurum, Sergio Escalera, Graham W. Taylor, Thomas B. Moeslund
ECCV (57)2
2024 Brain Responses to Emotional Avatars Challenge: Dataset and Results
abstract
Signals from electroencephalography (EEG) are considered to be very useful in identifying affective states. This study introduces a novel methodology for eliciting affective responses using virtual reality, featuring an immersive virtual environment in which subjects encounter an avatar displaying the six cardinal emotions: fear, joy, anger, sadness, disgust, and surprise. Subjects are asked to mimic the avatar's expressions while their cerebral activity is recorded using a multi-channel EEG apparatus with advanced dry sensor technology to reduce interference from facial kinetics and musculature. The study's cohort consisted of 40 individuals, each contributing 720 data points per emotional stimulus. A key aspect of this research is creating and releasing a novel dataset (EmoNeuroDB) re-sulting from this endeavor, which served as the foundation for a competition at the FG 2024. This paper describes the winner's methodological framework, emphasizing the dataset's importance as a primary contribution to the field.
Agnieszka Dubiel, Dorota Kaminska, Grzegorz Zwolinski, Akbar Anbar Jafari, Prasoon Kumar Vinodkumar, Egils Avots, Júlio C. S. Jacques Júnior, Sergio Escalera, Gholamreza Anbarjafari
FG8
2024 REACT 2024: the Second Multiple Appropriate Facial Reaction Generation Challenge
abstract
In dyadic interactions, humans communicate their intentions and state of mind using verbal and non-verbal cues, where multiple different facial reactions might be appropriate in response to a specific speaker behaviour. Then, how to develop a machine learning (ML) model that can automatically generate multiple appropriate, diverse, realistic and synchronised human facial reactions from an previously unseen speaker behaviour is a challenging task. Following the successful organisation of the first REACT challenge (REACT 2023), this edition of the challenge (REACT 2024) employs a subset used by the previous challenge, which contains segmented 30-secs dyadic interaction clips originally recorded as part of the NOXI and RECOLA datasets, encouraging participants to develop and benchmark Machine Learning (ML) models that can generate multiple appropriate facial reactions (including facial image sequences and their attributes) given an input conversational partner's stimulus under various dyadic video conference scenarios. This paper presents: (i) the guidelines of the REACT 2024 challenge; (ii) the dataset utilized in the challenge; and (iii) the performance of the baseline systems on the two proposed sub-challenges: Offline Multiple Appropriate Facial Reaction Generation and Online Multiple Appropriate Facial Reaction Generation, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2024.
Siyang Song, Micol Spitale, Cristina Palmero, Germán Barquero, Hengde Zhu, Sergio Escalera, Michel F. Valstar, Tobias Baur 0001, Fabien Ringeval, Elisabeth André, Hatice Gunes
FG7
2024 DualH: A Dual Hierarchical Model for Temporal Action Localization
abstract
Temporal action localization aims to detect action boundaries and classify action labels in untrimmed videos. Recent efforts have focused on utilizing Transformers to encode extracted features into a bottom-up pyramid feature map and localizing actions from all levels of the pyramid while only considering features from those specific levels. A limitation of this bottom-up encoding is that the lower-level features lack broader contexts, while the upper-level features lose local boundary information. Consequently, the performance of the model may be hindered. In this work, we propose a dual hierarchical model to mitigate this issue. The first hierarchy operates on the full temporal sequence to encode features at multiple scales. These features are fused to ensure all temporal locations consider both local boundary information and broader contexts. Next, the fused feature is downsampled to a pyramid representation for localizing actions at multiple resolutions. Experimental results on THUMOS14, ActivityNet-1.3, and EPIC-KITCHENS-100 demonstrate that our dual hierarchical design improves the performance with respect to the conventional bottom-up pyramid Transformer-based models.
Zejian Zhang, Cristina Palmero, Sergio Escalera
FG3
2024 Unified Physical-Digital Face Attack Detection
Ajian Liu 0001, Haocheng Yuan, Junze Zheng, Dingheng Zeng, Jiankang Deng, Sergio Escalera, Xiaoming Liu 0002, Jun Wan 0001, Zhen Lei 0001
IJCAI8
2024 FM-CLIP: Flexible Modal CLIP for Face Anti-Spoofing
abstract
In this work, borrowing a solution from the large-scale vision-language models (VLMs) instead of directly removing modality-specific signals from visual features, we propose a novel Flexible Modal CLIP (FM-CLIP) for flexible modal FAS, that can utilize text features to dynamically adjust visual features to be modality independent. In the visual branch, considering the huge visual differences of the same attack in different modalities, which makes it difficult for classifiers to flexibly identify subtle spoofing clues in different test modalities, we propose Cross-Modal Spoofing Enhancer (CMS-Enhancer). It includes a Frequency Extractor (FE) and Cross-Modal Interactor (CMI), aiming to map different modal attacks in a shared frequency space to reduce interference from modality-specific signals and enhance spoofing clues by leveraging cross-modal learning from the shared frequency space. In the text branch, we introduce a Language-Guided Patch Alignment (LGPA) based on prompt learning, which further guides the image encoder to focus on patch-level spoofing representations through dynamic weighting by text features. Thus, our FM-CLIP can flexibly test different modal samples by identifying and enhancing modality-agnostic spoofing cues. Finally, extensive experiments show that FM-CLIP is effective and outperforms state-of-the-art methods on multiple multi-modal datasets.
Ajian Liu 0001, Hui Ma 0018, Junze Zheng, Haocheng Yuan, Xiaoyuan Yu, Yanyan Liang 0001, Sergio Escalera, Jun Wan 0001, Zhen Lei 0001
ACM Multimedia7
2024 A Generative Multi-Resolution Pyramid and Normal-Conditioning 3D Cloth Draping
abstract
RGB cloth generation has been deeply studied in the related literature, however, 3D garment generation remains an open problem. In this paper, we build a conditional variational autoencoder for 3D garment generation and draping. We propose a pyramid network to add garment details progressively in a canonical space, i.e. unposing and unshaping the garments w.r.t. the body. We study conditioning the network on surface normal UV maps, as an intermediate representation, which is an easier problem to optimize than 3D coordinates. Our results on two public datasets, CLOTH3D and CAPE, show that our model is robust, controllable in terms of detail generation by the use of multi-resolution pyramids, and achieves state-of-the-art results that can highly generalize to unseen garments, poses, and shapes even when training with small amounts of data. The code can be found at: https://github.com/HunorLaczko/pyramid-drape
Hunor Laczkó, Meysam Madadi, Sergio Escalera, Jordi Gonzàlez 0001
WACV3
2024 Word separation in continuous sign language using isolated signs and post-processing
Razieh Rastgoo, Kourosh Kiani, Sergio Escalera
Expert Syst. Appl.3
2024 A survey on recent advances in Sign Language Production
Razieh Rastgoo, Kourosh Kiani, Sergio Escalera, Vassilis Athitsos, Mohammad Sabokrou
Expert Syst. Appl.3
2024 Multi-modal zero-shot dynamic hand gesture recognition
Razieh Rastgoo, Kourosh Kiani, Sergio Escalera, Mohammad Sabokrou
Expert Syst. Appl.3
2024 Predicting the generalization gap in neural networks using topological data analysis
abstract
Understanding how neural networks generalize on unseen data is crucial for designing more robust and reliable models. In this paper, we study the generalization gap of neural networks using methods from topological data analysis. For this purpose, we compute homological persistence diagrams of weighted graphs constructed from neuron activation correlations after a training phase, aiming to capture patterns that are linked to the generalization capacity of the network. We compare the usefulness of different numerical summaries from persistence diagrams and show that a combination of some of them can accurately predict and partially explain the generalization gap without the need of a test set. Evaluation on two computer vision recognition tasks (CIFAR10 and SVHN) shows competitive generalization gap prediction when compared against state-of-the-art methods.
Rubén Ballester, Xavier Arnal Clemente, Carles Casacuberta, Meysam Madadi, Ciprian A. Corneanu, Sergio Escalera
Neurocomputing6
2024 TopoX: A Suite of Python Packages for Machine Learning on Topological Domains
abstract
We introduce TopoX, a Python software suite that provides reliable and user-friendly building blocks for computing and machine learning on topological domains that extend graphs: hypergraphs, simplicial, cellular, path and combinatorial complexes. TopoX consists of three packages: TopoNetX facilitates constructing and computing on these domains, including working with nodes, edges and higher-order cells; TopoEmbedX provides methods to embed topological domains into vector spaces, akin to popular graph-based embedding algorithms such as node2vec; TopoModelX is built on top of PyTorch and offers a comprehensive toolbox of higher-order message passing functions for neural networks on topological domains. The extensively documented and unit-tested source code of TopoX is available under MIT license at https://pyt-team.github.io.
Mustafa Hajij, Mathilde Papillon, Florian Frantzen, Jens Agerberg, Ibrahem AlJabea, Rubén Ballester, Claudio Battiloro, Guillermo Bernárdez, Tolga Birdal, Aiden Brent, Sang (Peter) Chin, Sergio Escalera, Simone Fiorellino, Odin Hoff Gardaa, Gurusankar Gopalakrishnan, Devendra Govil, Josef Hoppe, Maneel Reddy Karri, Jude Khouja, Manuel Lecha, Neal Livesay, Jan Meißner, Alexander Nikitin 0002, Theodore Papamarkou, Jaro Prílepok, Karthikeyan Natesan Ramamurthy, Paul Rosen 0001, Aldo Guzmán-Sáenz, Alessandro Salatiello, Shreyas N. Samaga, Simone Scardapane, Michael T. Schaub, Luca Scofano, Indro Spinelli, Lev Telyatnikov, Quang Truong, Robin Walters 0001, Maosheng Yang, Olga Zaghen, Ghada Zamzmi, Ali Zia, Nina Miolane
J. Mach. Learn. Res.12
2024 A transformer model for boundary detection in continuous sign language
Razieh Rastgoo, Kourosh Kiani, Sergio Escalera
Multim. Tools Appl.3
2024 Surveillance Face Anti-Spoofing
abstract
Face Anti-spoofing (FAS) is essential to secure face recognition systems from various physical attacks. However, recent research generally focuses on short-distance applications (i.e., phone unlocking) while lacking consideration of long-distance scenes (i.e., surveillance security checks). In order to promote relevant research and fill this gap in the community, we collect a large-scale Su rveillance Hi gh-Fi delity Mask (SuHiFiMask) dataset captured under 40 surveillance scenes, which has 101 subjects from different age groups with$232~3\text{D}$attacks (high-fidelity masks),$200~2\text{D}$attacks (posters, portraits, and screens), and 2 adversarial attacks. In this scene, low image resolution and noise interference are new challenges faced in surveillance FAS. Together with the SuHiFiMask dataset, we propose a Contrastive Quality-Invariance Learning (CQIL) network to alleviate the performance degradation caused by image quality from three aspects: 1) An Image Quality Variable module (IQV) is introduced to recover image information associated with discrimination by combining the super-resolution network. 2) Using generated sample pairs to simulate quality variance distributions to help contrastive learning strategies obtain robust feature representation under quality variation. 3) A Separate Quality Network (SQN) is designed to learn discriminative features independent of image quality. Finally, a large number of experiments verify the quality of the SuHiFiMask dataset and the superiority of the proposed CQIL.
Ajian Liu 0001, Jun Wan 0001, Sergio Escalera, Stan Z. Li, Zhen Lei 0001
IEEE Trans. Inf. Forensics Secur.4
2023 Blowing in the Wind: CycleNet for Human Cinemagraphs from Still Images
abstract
Cinemagraphs are short looping videos created by adding subtle motions to a static image. This kind of media is popular and engaging. However, automatic generation of cinemagraphs is an underexplored area and current solutions require tedious low-level manual authoring by artists. In this paper, we present an automatic method that allows generating human cinemagraphs from single RGB images. We investigate the problem in the context of dressed humans under the wind. At the core of our method is a novel cyclic neural network that produces looping cinemagraphs for the target loop duration. To circumvent the problem of collecting real data, we demonstrate that it is possible, by working in the image normal space, to learn garment motion dynamics on synthetic data and generalize to real data. We evaluate our method on both synthetic and real data and demonstrate that it is possible to create compelling and plausible cinemagraphs from single RGB images.
Hugo Bertiche, Niloy J. Mitra, Kuldeep Kulkarni, Chun-Hao Paul Huang, Tuanfeng Y. Wang, Meysam Madadi, Sergio Escalera, Duygu Ceylan
CVPR7
2023 Multi-Rate Sensor Fusion for Unconstrained Near-Eye Gaze Estimation
abstract
The power requirements of video-oculography systems can be prohibitive for high-speed operation on portable devices. Recently, low-power alternatives such as photosensors have been evaluated, providing gaze estimates at high frequency with a trade-off in accuracy and robustness. Potentially, an approach combining slow/high-fidelity and fast/low-fidelity sensors should be able to exploit their complementarity to track fast eye motion accurately and robustly. To foster research on this topic, we introduce OpenSFEDS, a near-eye gaze estimation dataset containing approximately 2M synthetic camera-photosensor image pairs sampled at 500 Hz under varied appearance and camera position. We also formulate the task of sensor fusion for gaze estimation, proposing a deep learning framework consisting in appearance-based encoding and temporal eye-state dynamics. We evaluate several single- and multi-rate fusion baselines on OpenSFEDS, achieving 8.7% error decrease when tracking fast eye movements with a multi-rate approach vs. a gaze forecasting approach operating with a low-speed sensor alone.
Cristina Palmero, Oleg V. Komogortsev, Sergio Escalera, Sachin S. Talathi
ETRA3
2023 BeLFusion: Latent Diffusion for Behavior-Driven Human Motion Prediction
abstract
Stochastic human motion prediction (HMP) has generally been tackled with generative adversarial networks and variational autoencoders. Most prior works aim at predicting highly diverse motion in terms of the skeleton joints' dispersion. This has led to methods predicting fast and divergent movements, which are often unrealistic and incoherent with past motion. Such methods also neglect scenarios where anticipating diverse short-range behaviors with subtle joint displacements is important. To address these issues, we present BeLFusion, a model that, for the first time, leverages latent diffusion models in HMP to sample from a behavioral latent space where behavior is disentangled from pose and motion. Thanks to our behavior coupler, which is able to transfer sampled behavior to ongoing motion, BeLFusion’s predictions display a variety of behaviors that are significantly more realistic, and coherent with past motion than the state of the art. To support it, we introduce two metrics, the Area of the Cumulative Motion Distribution, and the Average Pairwise Distance Error, which are correlated to realism according to a qualitative study (126 participants). Finally, we prove BeLFusion’s generalization power in a new cross-dataset scenario for stochastic HMP.
Germán Barquero, Sergio Escalera, Cristina Palmero
ICCV2
2023 Gloss-free Sign Language Translation: Improving from Visual-Language Pretraining
abstract
Sign Language Translation (SLT) is a challenging task due to its cross-domain nature, involving the translation of visual-gestural language to text. Many previous methods employ an intermediate representation, i.e., gloss sequences, to facilitate SLT, thus transforming it into a two-stage task of sign language recognition (SLR) followed by sign language translation (SLT). However, the scarcity of gloss-annotated sign language data, combined with the information bottleneck in the mid-level gloss representation, has hindered the further development of the SLT task. To address this challenge, we propose a novel Gloss-Free SLT based on Visual-Language Pretraining (GFSLT-VLP), which improves SLT by inheriting language-oriented prior knowledge from pre-trained models, without any gloss annotation assistance. Our approach involves two stages: (i) integrating Contrastive Language-Image Pre-training (CLIP) with masked self-supervised learning to create pre-tasks that bridge the semantic gap between visual and textual representations and restore masked sentences, and (ii) constructing an end-to-end architecture with an encoder-decoder-like structure that inherits the parameters of the pre-trained Visual Encoder and Text Decoder from the first stage. The seamless combination of these novel designs forms a robust sign language representation and significantly improves gloss-free sign language translation. In particular, we have achieved unprecedented improvements in terms of BLEU-4 score on the PHOENIX14T dataset (≥+5) and the CSL-Daily dataset (≥+3) compared to state-of-the-art gloss-free SLT methods. Furthermore, our approach also achieves competitive results on the PHOENIX14T dataset when compared with most of the gloss-based methods1.
Benjia Zhou, Albert Clapés, Jun Wan 0001, Yanyan Liang 0001, Sergio Escalera, Zhen Lei 0001
ICCV6
2023 REACT2023: The First Multiple Appropriate Facial Reaction Generation Challenge
abstract
The Multiple Appropriate Facial Reaction Generation Challenge (REACT2023) is the first competition event focused on evaluating multimedia processing and machine learning techniques for generating human-appropriate facial reactions in various dyadic interaction scenarios, with all participants competing strictly under the same conditions. The goal of the challenge is to provide the first benchmark test set for multi-modal information processing and to foster collaboration among the audio, visual, and audio-visual behaviour analysis and behaviour generation (a.k.a generative AI) communities, to compare the relative merits of the approaches to automatic appropriate facial reaction generation under different spontaneous dyadic interaction conditions. This paper presents: (i) the novelties, contributions and guidelines of the REACT2023 challenge; (ii) the dataset utilized in the challenge; and (iii) the performance of the baseline systems on the two proposed sub-challenges: Offline Multiple Appropriate Facial Reaction Generation and Online Multiple Appropriate Facial Reaction Generation, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2023.
Siyang Song, Micol Spitale, Germán Barquero, Cristina Palmero, Sergio Escalera, Michel F. Valstar, Tobias Baur 0001, Fabien Ringeval, Elisabeth André, Hatice Gunes
ACM Multimedia6
2023 CodaLab Competitions: An Open Source Platform to Organize Scientific Challenges
abstract
CodaLab Competitions is an open source web platform designed to help data scientists and research teams to crowd-source the resolution of machine learning problems through the organization of competitions, also called challenges or contests. CodaLab Competitions provides useful features such as multiple phases, results and code submissions, multi-score leaderboards, and jobs running inside Docker containers. The platform is very flexible and can handle large scale experiments, by allowing organizers to upload large datasets and provide their own CPU or GPU compute workers.
Adrien Pavão, Isabelle Guyon, Anne-Catherine Letournel, Dinh-Tuan Tran, Xavier Baró, Hugo Jair Escalante, Sergio Escalera, Tyler Thomas, Zhen Xu 0007
J. Mach. Learn. Res.7
2023 MyoPS: A benchmark of myocardial pathology segmentation combining three-sequence cardiac magnetic resonance images
Lei Li 0020, Fuping Wu, Xinzhe Luo, Carlos Martín-Isla, Shuwei Zhai, Zhen Zhang 0057, Markus J. Ankenbrand, Haochuan Jiang, Linhong Wang, Tewodros Weldebirhan Arega, Elif Altunok, Jun Ma 0016, Xiaoping Yang 0001, Élodie Puybareau, Ilkay Öksüz, Stéphanie Bricq, Weisheng Li 0001, Kumaradevan Punithakumar, Sotirios A. Tsaftaris, Laura Maria Schreiber, Guocai Liu, Yong Xia 0001, Guotai Wang, Sergio Escalera, Xiahai Zhuang
Medical Image Anal.31
2023 CrossMoDA 2021 challenge: Benchmark of cross-modality domain adaptation techniques for vestibular schwannoma and cochlea segmentation
abstract
Domain Adaptation (DA) has recently been of strong interest in the medical imaging community. While a large variety of DA techniques have been proposed for image segmentation, most of these techniques have been validated either on private datasets or on small publicly available datasets. Moreover, these datasets mostly addressed single-class problems. To tackle these limitations, the Cross-Modality Domain Adaptation (crossMoDA) challenge was organised in conjunction with the 24th International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI 2021). CrossMoDA is the first large and multi-class benchmark for unsupervised cross-modality Domain Adaptation. The goal of the challenge is to segment two key brain structures involved in the follow-up and treatment planning of vestibular schwannoma (VS): the VS and the cochleas. Currently, the diagnosis and surveillance in patients with VS are commonly performed using contrast-enhanced T1 (ceT1) MR imaging. However, there is growing interest in using non-contrast imaging sequences such as high-resolution T2 (hrT2) imaging. For this reason, we established an unsupervised cross-modality segmentation benchmark. The training dataset provides annotated ceT1 scans (N=105) and unpaired non-annotated hrT2 scans (N=105). The aim was to automatically perform unilateral VS and bilateral cochlea segmentation on hrT2 scans as provided in the testing set (N=137). This problem is particularly challenging given the large intensity distribution gap across the modalities and the small volume of the structures. A total of 55 teams from 16 countries submitted predictions to the validation leaderboard. Among them, 16 teams from 9 different countries submitted their algorithm for the evaluation phase. The level of performance reached by the top-performing teams is strikingly high (best median Dice score — VS: 88.4%; Cochleas: 85.7%) and close to full supervision (median Dice score — VS: 92.5%; Cochleas: 87.7%). All top-performing methods made use of an image-to-image translation approach to transform the source-domain images into pseudo-target-domain images. A segmentation network was then trained using these generated images and the manual annotations provided for the source image.
Reuben Dorent, Aaron Kujawa, Marina Ivory, Spyridon Bakas, Nicola Rieke, Samuel Joutard, Ben Glocker, Manuel Jorge Cardoso, Marc Modat, Kayhan Batmanghelich, Arseniy Belkov, Maria G. Baldeon Calisto, Jae Won Choi, Benoit M. Dawant, Hexin Dong, Sergio Escalera, Yubo Fan, Lasse Hansen, Mattias P. Heinrich, Smriti Joshi, Victoriya Kashtanova, Hyeongyu Kim, Satoshi Kondo, Christian N. Kruse, Susana K. Lai-Yuen, Hao Li 0108, Buntheng Ly, Ipek Oguz, Hyungseob Shin, Boris Shirokikh, Zixian Su, Guotai Wang, Jianghao Wu 0001, Yanwu Xu 0001, Li Zhang 0047, Sébastien Ourselin, Jonathan Shapey, Tom Vercauteren
Medical Image Anal.16
2023 A deep co-attentive hand-based video question answering framework using multi-view skeleton
Razieh Rastgoo, Kourosh Kiani, Sergio Escalera
Multim. Tools Appl.3
2023 ZS-GR: zero-shot gesture recognition from RGB-D videos
Razieh Rastgoo, Kourosh Kiani, Sergio Escalera
Multim. Tools Appl.3
2023 Video Transformers: A Survey
abstract
Transformer models have shown great success handling long-range interactions, making them a promising tool for modeling video. However, they lack inductive biases and scale quadratically with input length. These limitations are further exacerbated when dealing with the high dimensionality introduced by the temporal dimension. While there are surveys analyzing the advances of Transformers for vision, none focus on an in-depth analysis of video-specific designs. In this survey, we analyze the main contributions and trends of works leveraging Transformers to model video. Specifically, we delve into how videos are handled at the input level first. Then, we study the architectural changes made to deal with video more efficiently, reduce redundancy, re-introduce useful inductive biases, and capture long-term temporal dynamics. In addition, we provide an overview of different training regimes and explore effective self-supervised learning strategies for video. Finally, we conduct a performance comparison on the most common benchmark for Video Transformers (i.e., action classification), finding them to outperform 3D ConvNets even with less computational complexity.
Javier Selva, Anders Skaarup Johansen, Sergio Escalera, Kamal Nasrollahi, Thomas B. Moeslund, Albert Clapés
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Learning to Recognize Actions on Objects in Egocentric Video With Attention Dictionaries
abstract
We present EgoACO, a deep neural architecture for video action recognition that learns to pool action-context-object descriptors from frame level features by leveraging the verb-noun structure of action labels in egocentric video datasets. The core component is class activation pooling (CAP), a differentiable pooling layer that combines ideas from bilinear pooling for fine-grained recognition and from feature learning for discriminative localization. CAP uses self-attention with a dictionary of learnable weights to pool from the most relevant feature regions. Through CAP, EgoACO learns to decode object and scene context descriptors from video frame features. For temporal modeling we design a recurrent version of class activation pooling termed Long Short-Term Attention (LSTA). LSTA extends convolutional gated LSTM with built-in spatial attention and a re-designed output gate. Action, object and context descriptors are fused by a multi-head prediction that accounts for the inter-dependencies between noun-verb-action structured labels in egocentric video datasets. EgoACO features built-in visual explanations, helping learning and interpretation of discriminative information in video. Results on the two largest egocentric action recognition datasets currently available, EPIC-KITCHENS and EGTEA Gaze+, show that by decoding action-context-object descriptors, the model achieves state-of-the-art recognition performance.
Swathikiran Sudhakaran, Sergio Escalera, Oswald Lanz
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Gate-Shift-Fuse for Video Action Recognition
abstract
Convolutional Neural Networks are the de facto models for image recognition. However 3D CNNs, the straight forward extension of 2D CNNs for video recognition, have not achieved the same success on standard action recognition benchmarks. One of the main reasons for this reduced performance of 3D CNNs is the increased computational complexity requiring large scale annotated datasets to train them in scale. 3D kernel factorization approaches have been proposed to reduce the complexity of 3D CNNs. Existing kernel factorization approaches follow hand-designed and hard-wired techniques. In this paper we propose Gate-Shift-Fuse (GSF), a novel spatio-temporal feature extraction module which controls interactions in spatio-temporal decomposition and learns to adaptively route features through time and combine them in a data dependent manner. GSF leverages grouped spatial gating to decompose input tensor and channel weighting to fuse the decomposed tensors. GSF can be inserted into existing 2D CNNs to convert them into an efficient and high performing spatio-temporal feature extractor, with negligible parameter and compute overhead. We perform an extensive analysis of GSF using two popular 2D CNN families and achieve state-of-the-art or competitive performance on five standard action recognition benchmarks.
Swathikiran Sudhakaran, Sergio Escalera, Oswald Lanz
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Deep Learning Segmentation of the Right Ventricle in Cardiac MRI: The M&Ms Challenge
abstract
In recent years, several deep learning models have been proposed to accurately quantify and diagnose cardiac pathologies. These automated tools heavily rely on the accurate segmentation of cardiac structures in MRI images. However, segmentation of the right ventricle is challenging due to its highly complex shape and ill-defined borders. Hence, there is a need for new methods to handle such structure's geometrical and textural complexities, notably in the presence of pathologies such as Dilated Right Ventricle, Tricuspid Regurgitation, Arrhythmogenesis, Tetralogy of Fallot, and Inter-atrial Communication. The last MICCAI challenge on right ventricle segmentation was held in 2012 and included only 48 cases from a single clinical center. As part of the 12th Workshop on Statistical Atlases and Computational Models of the Heart (STACOM 2021), the M&Ms-2 challenge was organized to promote the interest of the research community around right ventricle segmentation in multi-disease, multi-view, and multi-center cardiac MRI. Three hundred sixty CMR cases, including short-axis and long-axis 4-chamber views, were collected from three Spanish hospitals using nine different scanners from three different vendors, and included a diverse set of right and left ventricle pathologies. The solutions provided by the participants show that nnU-Net achieved the best results overall. However, multi-view approaches were able to capture additional information, highlighting the need to integrate multiple cardiac diseases, views, scanners, and acquisition protocols to produce reliable automatic cardiac segmentation algorithms.
Carlos Martín-Isla, Víctor M. Campello, Cristian Izquierdo, Kaisar Kushibar, Carla Sendra-Balcells, Polyxeni Gkontra, Alireza Sojoudi, Mitchell J. Fulton, Tewodros Weldebirhan Arega, Kumaradevan Punithakumar, Lei Li 0020, Xiaowu Sun, Yasmina Alkhalil, Di Liu 0003, Sana Jabbar, Sandro F. Queiros, Francesco Galati, Moona Mazher, Zheyao Gao, Marcel Beetz, Lennart Tautz, Christoforos Galazis, Marta Varela, Markus Hüllebrand, Vicente Grau, Xiahai Zhuang, Domenec Puig, Maria A. Zuluaga, Hassan Mohy-ud-Din, Dimitris N. Metaxas, Marcel Breeuwer, Rob J. van der Geest, Michelle Noga, Stéphanie Bricq, Mark Rentschler, Andrea Guala 0002, Steffen E. Petersen, Sergio Escalera, Jose Rodriguez-Palomares, Karim Lekadir
IEEE J. Biomed. Health Informatics38
2022 Towards Self-Supervised Gaze Estimation
Arya Farkhondeh, Cristina Palmero, Simone Scardapane, Sergio Escalera
BMVC4
2022 Relevance-based Margin for Contrastively-trained Video Retrieval Models
abstract
Video retrieval using natural language queries has attracted increasing interest due to its relevance in real-world applications, from intelligent access in private media galleries to web-scale video search. Learning the cross-similarity of video and text in a joint embedding space is the dominant approach. To do so, a contrastive loss is usually employed because it organizes the embedding space by putting similar items close and dissimilar items far. This framework leads to competitive recall rates, as they solely focus on the rank of the groundtruth items. Yet, assessing the quality of the ranking list is of utmost importance when considering intelligent retrieval systems, since multiple items may share similar semantics, hence a high relevance. Moreover, the aforementioned framework uses a fixed margin to separate similar and dissimilar items, treating all non-groundtruth items as equally irrelevant. In this paper we propose to use a variable margin: we argue that varying the margin used during training based on how much relevant an item is to a given query, i.e. a relevance-based margin, easily improves the quality of the ranking lists measured through nDCG and mAP. We demonstrate the advantages of our technique using different models on EPIC-Kitchens-100 and YouCook2. We show that even if we carefully tuned the fixed margin, our technique (which does not have the margin as a hyper-parameter) would still achieve better performance. Finally, extensive ablation studies and qualitative analysis support the robustness of our approach. Code will be released at \urlhttps://github.com/aranciokov/RelevanceMargin-ICMR22.
Alex Falcon, Swathikiran Sudhakaran, Giuseppe Serra 0001, Sergio Escalera, Oswald Lanz
ICMR4
2022 Meta-Album: Multi-domain Meta-Dataset for Few-Shot Image Classification
abstract
We introduce Meta-Album, an image classification meta-dataset designed to facilitate few-shot learning, transfer learning, meta-learning, among other tasks. It includes 40 open datasets, each having at least 20 classes with 40 examples per class, with verified licences. They stem from diverse domains, such as ecology (fauna and flora), manufacturing (textures, vehicles), human actions, and optical character recognition, featuring various image scales (microscopic, human scales, remote sensing). All datasets are preprocessed, annotated, and formatted uniformly, and come in 3 versions (Micro $\subset$ Mini $\subset$ Extended) to match users’ computational resources. We showcase the utility of the first 30 datasets on few-shot learning problems. The other 10 will be released shortly after. Meta-Album is already more diverse and larger (in number of datasets) than similar efforts, and we are committed to keep enlarging it via a series of competitions. As competitions terminate, their test data are released, thus creating a rolling benchmark, available through OpenML.org. Our website https://meta-album.github.io/ contains the source code of challenge winning methods, baseline methods, data loaders, and instructions for contributing either new datasets or algorithms to our expandable meta-dataset.
Dustin Carrión-Ojeda, Sergio Escalera, Isabelle Guyon, Mike Huisman, Felix Mohr, Jan N. van Rijn, Haozhe Sun, Joaquin Vanschoren, Phan Anh Vu
NeurIPS3
2022 Multi-Task Classification of Sewer Pipe Defects and Properties using a Cross-Task Graph Neural Network Decoder
abstract
The sewerage infrastructure is one of the most important and expensive infrastructures in modern society. In order to efficiently manage the sewerage infrastructure, automated sewer inspection has to be utilized. However, while sewer defect classification has been investigated for decades, little attention has been given to classifying sewer pipe properties such as water level, pipe material, and pipe shape, which are needed to evaluate the level of sewer pipe deterioration.In this work we classify sewer pipe defects and properties concurrently and present a novel decoder-focused multi-task classification architecture Cross-Task Graph Neural Network (CT-GNN), which refines the disjointed per-task predictions using cross-task information. The CT-GNN architecture extends the traditional disjointed task-heads decoder, by utilizing a cross-task graph and unique class node embeddings. The cross-task graph can either be determined a priori based on the conditional probability between the task classes or determined dynamically using self-attention. CT-GNN can be added to any backbone and trained end-to-end at a small increase in the parameter count. We achieve state-of-the-art performance on all four classification tasks in the Sewer-ML dataset, improving defect classification and water level classification by 5.3 and 8.0 percentage points, respectively. We also outperform the single task methods as well as other multi-task classification approaches while introducing 50 times fewer parameters than previous model-focused approaches. The code and models are available at the project page http://vap.aau.dk/ctgnn.
Joakim Bruslund Haurum, Meysam Madadi, Sergio Escalera, Thomas B. Moeslund
WACV3
2022 End-to-end global to local convolutional neural network learning for hand pose recovery in depth data
abstract
Abstract Despite recent advances in 3‐D pose estimation of human hands, thanks to the advent of convolutional neural networks (CNNs) and depth cameras, this task is still far from being solved in uncontrolled setups. This is mainly due to the highly non‐linear dynamics of fingers and self‐occlusions, which make hand model training a challenging task. In this study, a novel hierarchical tree‐like structured CNN is exploited, in which branches are trained to become specialised in predefined subsets of hand joints called local poses. Further, local pose features, extracted from hierarchical CNN branches, are fused to learn higher order dependencies among joints in the final pose by end‐to‐end training. Lastly, the loss function used is also defined to incorporate appearance and physical constraints about doable hand motions and deformations. Finally, a non‐rigid data augmentation approach is introduced to increase the amount of training depth data. Experimental results suggest that feeding a tree‐shaped CNN, specialised in local poses, into a fusion network for modelling joints' correlations and dependencies, helps to increase the precision of final estimations, showing competitive results on NYU, MSRA, Hands17 and SyntheticHand datasets.
Meysam Madadi, Sergio Escalera, Xavier Baró, Jordi Gonzàlez 0001
IET Comput. Vis.2
2022 Guest Editorial: Special issue on computer vision and machine learning for healthcare applications
Cristina Palmero, M. Inés Torres, Anna Esposito, Sergio Escalera
Pattern Anal. Appl.4
2022 Modeling, Recognizing, and Explaining Apparent Personality From Videos
abstract
Explainability and interpretability are two critical aspects of decision support systems. Despite their importance, it is only recently that researchers are starting to explore these aspects. This paper provides an introduction to explainability and interpretability in the context of apparent personality recognition. To the best of our knowledge, this is the first effort in this direction. We describe a challenge we organized on explainability in first impressions analysis from video. We analyze in detail the newly introduced data set, evaluation protocol, proposed solutions and summarize the results of the challenge. We investigate the issue of bias in detail. Finally, derived from our study, we outline research opportunities that we foresee will be relevant in this area in the near future.
Hugo Jair Escalante, Heysem Kaya, Albert Ali Salah, Sergio Escalera, Yagmur Güçlütürk, Umut Güçlü, Xavier Baró, Isabelle Guyon, Júlio C. S. Jacques Júnior, Meysam Madadi, Stéphane Ayache, Evelyne Viegas, Furkan Gürpinar, Achmadnoer Sukma Wicaksana, Cynthia C. S. Liem, Marcel van Gerven, Rob van Lier
IEEE Trans. Affect. Comput.4
2022 First Impressions: A Survey on Vision-Based Apparent Personality Trait Analysis
abstract
Personality analysis has been widely studied in psychology, neuropsychology, and signal processing fields, among others. From the past few years, it also became an attractive research area in visual computing. From the computational point of view, by far speech and text have been the most considered cues of information for analyzing personality. However, recently there has been an increasing interest from the computer vision community in analyzing personality from visual data. Recent computer vision approaches are able to accurately analyze human faces, body postures and behaviors, and use these information to inferapparentpersonality traits. Because of the overwhelming research interest in this topic, and of the potential impact that this sort of methods could have in society, we present in this paper an up-to-date review of existing vision-based approaches for apparent personality trait recognition. We describe seminal and cutting edge works on the subject, discussing and comparing their distinctive features and limitations. Future venues of research in the field are identified and discussed. Furthermore, aspects on the subjectivity in data labeling/evaluation, as well as current datasets and challenges organized to push the research on the field are reviewed.
Júlio C. S. Jacques Júnior, Yagmur Güçlütürk, Marc Pérez 0001, Umut Güçlü, Carlos Andújar, Xavier Baró, Hugo Jair Escalante, Isabelle Guyon, Marcel van Gerven, Rob van Lier, Sergio Escalera
IEEE Trans. Affect. Comput.11
2022 ChaLearn Looking at People: IsoGD and ConGD Large-Scale RGB-D Gesture Recognition
abstract
The ChaLearn large-scale gesture recognition challenge has run twice in two workshops in conjunction with the International Conference on Pattern Recognition (ICPR) 2016 and International Conference on Computer Vision (ICCV) 2017, attracting more than 200 teams around the world. This challenge has two tracks, focusing on isolated and continuous gesture recognition, respectively. It describes the creation of both benchmark datasets and analyzes the advances in large-scale gesture recognition based on these two datasets. In this article, we discuss the challenges of collecting large-scale ground-truth annotations of gesture recognition and provide a detailed analysis of the current methods for large-scale isolated and continuous gesture recognition. In addition to the recognition rate and mean Jaccard index (MJI) as evaluation metrics used in previous challenges, we introduce the corrected segmentation rate (CSR) metric to evaluate the performance of temporal segmentation for continuous gesture recognition. Furthermore, we propose a bidirectional long short-term memory (Bi-LSTM) method, determining video division points based on skeleton points. Experiments show that the proposed Bi-LSTM outperforms state-of-the-art methods with an absolute improvement of 8.1% (from 0.8917 to 0.9639) of CSR.
Jun Wan 0001, Chi Lin 0002, Longyin Wen, Yunan Li 0001, Qiguang Miao, Sergio Escalera, Gholamreza Anbarjafari, Isabelle Guyon, Guodong Guo, Stan Z. Li
IEEE Trans. Cybern.6
2022 Contrastive Context-Aware Learning for 3D High-Fidelity Mask Face Presentation Attack Detection
abstract
Face presentation attack detection (PAD) is essential to secure face recognition systems primarily from high-fidelity mask attacks. Most existing 3D mask PAD benchmarks suffer from several drawbacks: 1) a limited number of mask identities, types of sensors, and a total number of videos; 2) low-fidelity quality of facial masks. Basic deep models and remote photoplethysmography (rPPG) methods achieved acceptable performance on these benchmarks but still far from the needs of practical scenarios. To bridge the gap to real-world applications, we introduce a large-scale High-Fidelity Mask dataset, namely HiFiMask. Specifically, a total amount of 54,600 videos are recorded from 75 subjects with 225 realistic masks by 7 new kinds of sensors. Along with the dataset, we propose a novel Contrastive Context-aware Learning (CCL) framework. CCL is a new training methodology for supervised PAD tasks, which is able to learn by leveraging rich contexts accurately (e.g., subjects, mask material and lighting) among pairs of live faces and high-fidelity mask attacks. Extensive experimental evaluations on HiFiMask and three additional 3D mask datasets demonstrate the effectiveness of our method. The codes and dataset will be released soon.
Ajian Liu 0001, Zitong Yu, Jun Wan 0001, Anyang Su, Zichang Tan, Sergio Escalera, Junliang Xing, Yanyan Liang 0001, Guodong Guo, Zhen Lei 0001, Stan Z. Li
IEEE Trans. Inf. Forensics Secur.8
2022 Neural Cloth Simulation
abstract
We present a general framework for the garment animation problem through unsupervised deep learning inspired in physically based simulation. Existing trends in the literature already explore this possibility. Nonetheless, these approaches do not handle cloth dynamics. Here, we propose the first methodology able to learn realistic cloth dynamics unsupervisedly, and henceforth, a general formulation for neural cloth simulation. The key to achieve this is to adapt an existing optimization scheme for motion from simulation based methodologies to deep learning. Then, analyzing the nature of the problem, we devise an architecture able to automatically disentangle static and dynamic cloth subspaces by design. We will show how this improves model performance. Additionally, this opens the possibility of a novel motion augmentation technique that greatly improves generalization. Finally, we show it also allows to control the level of motion in the predictions. This is a useful, never seen before, tool for artists. We provide of detailed analysis of the problem to establish the bases of neural cloth simulation and guide future research into the specifics of this domain.
Hugo Bertiche, Meysam Madadi, Sergio Escalera
ACM Trans. Graph.3
2021 Deep Parametric Surfaces for 3D Outfit Reconstruction from Single View Image
abstract
We present a methodology to retrieve analytical surfaces parametrized as a neural network. Previous works on 3D reconstruction yield point clouds, voxelized objects or meshes. Instead, our approach yields 2-manifolds in the euclidean space through deep learning. To this end, we implement a novel formulation for fully connected layers as parametrized manifolds that allows continuous predictions with differential geometry. Based on this property we propose a novel smoothness loss. Results on CLOTH3D++ dataset show the possibility to infer different topologies and the benefits of the smoothness term based on differential geometry.
Hugo Bertiche, Meysam Madadi, Sergio Escalera
FG3
2021 UV-based reconstruction of 3D garments from a single RGB image
abstract
Garments are highly detailed and dynamic objects made up of particles that interact with each other and with other objects, making the task of 2D to 3D garment reconstruction extremely challenging. Therefore, having a lightweight 3D representation capable of modelling fine details is of great importance. This work presents a deep learning framework based on Generative Adversarial Networks (GANs) to reconstruct 3D garment models from a single RGB image. It has the peculiarity of using UV maps to represent 3D data, a lightweight representation capable of dealing with high-resolution details and wrinkles. With this model and kind of 3D representation, we achieve state-of-the-art results on the CLOTH3D++ dataset, generating good quality and realistic garment reconstructions regardless of the garment topology and shape, human pose, occlusions and lightning.
Albert Rial-Farràs, Meysam Madadi, Sergio Escalera
FG3
2021 DeePSD: Automatic Deep Skinning And Pose Space Deformation For 3D Garment Animation
abstract
We present a novel solution to the garment animation problem through deep learning. Our contribution allows animating any template outfit with arbitrary topology and geometric complexity. Recent works develop models for garment edition, resizing and animation at the same time by leveraging the support body model (encoding garments as body homotopies). This leads to complex engineering solutions that suffer from scalability, applicability and compatibility. By limiting our scope to garment animation only, we are able to propose a simple model that can animate any outfit, independently of its topology, vertex order or connectivity. Our proposed architecture maps outfits to animated 3D models into the standard format for 3D animation (blend weights and blend shapes matrices), automatically providing of compatibility with any graphics engine. We also propose a methodology to complement supervised learning with an unsupervised physically based learning that implicitly solves collisions and enhances cloth quality.
Hugo Bertiche, Meysam Madadi, Emilio Tylson, Sergio Escalera
ICCV4
2021 The EMPATHIC Virtual Coach: a demo
abstract
The main objective of the EMPATHIC project has been the design and development of a virtual coach to engage the healthy-senior user and to enhance well-being through awareness of personal status. The EMPATHIC approach addresses this objective through multimodal interactions supported by the GROW coaching model. The paper summarizes the main components of the EMPATHIC Virtual Coach (EMPATHIC-VC) and introduces a demonstration of the coaching sessions in selected scenarios.
Javier Mikel Olaso, Alain Vázquez, Leila Ben Letaifa, Mikel de Velasco-Vázquez, Aymen Mtibaa, Mohamed Amine Hmani, Dijana Petrovska-Delacrétaz, Gérard Chollet, César Montenegro, Asier López-Zorrilla, Raquel Justo, Roberto Santana 0001, Jofre Tenorio-Laranga, Eduardo Gonzalez-Fraile, Begoña Fernández-Ruanova, Gennaro Cordasco, Anna Esposito, Kristin Beck Gjellesvik, Anna Torp Johansen, Maria Stylianou Korsnes, Colin Pickard, Cornelius Glackin, Gary Cahalane, Pau Buch-Cardona, Cristina Palmero, Sergio Escalera, Olga Gordeeva, Olivier Deroo, Anaïs Fernández, Daria Kyslitska, José Antonio Lozano 0001, M. Inés Torres, Stephan Schlögl
ICMI26
2021 Mobile eHealth Platform for Home Monitoring of Bipolar Disorder
Joan Codina, Sergio Escalera, Joan Escudero, Coen Antens, Pau Buch-Cardona, Mireia Farrús
MMM (2)2
2021 CASIA-SURF CeFA: A Benchmark for Multi-modal Cross-ethnicity Face Anti-spoofing
abstract
The issue of ethnic bias has proven to affect the performance of face recognition in previous works, while it still remains to be vacant in face anti-spoofing. Therefore, in order to study the ethnic bias for face anti-spoofing, we introduce the largest CASIA-SURF Cross-ethnicity Face Anti-spoofing (CeFA) dataset, covering 3 ethnicities, 3 modalities, 1,607 subjects, and 2D plus 3D attack types. Five protocols are introduced to measure the affect under varied evaluation conditions, such as cross-ethnicity, unknown spoofs or both of them. As our knowledge, CASIA-SURF CeFA is the first dataset including explicit ethnic labels in current released datasets. Then, we propose a novel multi-modal fusion method as a strong baseline to alleviate the ethnic bias, which employs a partially shared fusion strategy to learn complementary information from multiple modalities. Extensive experiments have been conducted on the proposed dataset to verify its significance and generalization capability for other existing datasets, i.e., CASIA-SURF, OULU-NPU and SiW datasets. The dataset is available at https://sites.google.com/qq.com/face-anti-spoofing/welcome/challengecvpr2020?authuser=0.
Ajian Liu 0001, Zichang Tan, Jun Wan 0001, Sergio Escalera, Guodong Guo, Stan Z. Li
WACV4
2021 DECONbench: a benchmarking platform dedicated to deconvolution methods for tumor heterogeneity quantification
abstract
BACKGROUND: Quantification of tumor heterogeneity is essential to better understand cancer progression and to adapt therapeutic treatments to patient specificities. Bioinformatic tools to assess the different cell populations from single-omic datasets as bulk transcriptome or methylome samples have been recently developed, including reference-based and reference-free methods. Improved methods using multi-omic datasets are yet to be developed in the future and the community would need systematic tools to perform a comparative evaluation of these algorithms on controlled data. RESULTS: We present DECONbench, a standardized unbiased benchmarking resource, applied to the evaluation of computational methods quantifying cell-type heterogeneity in cancer. DECONbench includes gold standard simulated benchmark datasets, consisting of transcriptome and methylome profiles mimicking pancreatic adenocarcinoma molecular heterogeneity, and a set of baseline deconvolution methods (reference-free algorithms inferring cell-type proportions). DECONbench performs a systematic performance evaluation of each new methodological contribution and provides the possibility to publicly share source code and scoring. CONCLUSION: DECONbench allows continuous submission of new methods in a user-friendly fashion, each novel contribution being automatically compared to the reference baseline methods, which enables crowdsourced benchmarking. DECONbench is designed to serve as a reference platform for the benchmarking of deconvolution methods in the evaluation of cancer heterogeneity. We believe it will contribute to leverage the benchmarking practices in the biomedical and life science communities. DECONbench is hosted on the open source Codalab competition platform. It is freely available at: https://competitions.codalab.org/competitions/27453 .
Clémentine Decamps, Alexis Arnaud, Florent Petitprez, Mira Ayadi, Aurélia Baurès, Lucile Armenoult, Sergio Escalera, Isabelle Guyon, Rémy Nicolle, Richard Tomasini, Aurélien de Reyniès, Jérôme Cros, Yuna Blum, Magali Richard
BMC Bioinform.7
2021 Introduction to the special issue of the ECML PKDD 2021 journal track
Annalisa Appice, Sergio Escalera, José A. Gámez 0001, Heike Trautmann
Data Min. Knowl. Discov.2
2021 Sign Language Recognition: A Deep Survey
Razieh Rastgoo, Kourosh Kiani, Sergio Escalera
Expert Syst. Appl.3
2021 Deep Unsupervised 3D Human Body Reconstruction from a Sparse set of Landmarks
Meysam Madadi, Hugo Bertiche, Sergio Escalera
Int. J. Comput. Vis.3
2021 Introduction to the special issue of the ECML PKDD 2021 journal track
Annalisa Appice, Sergio Escalera, José A. Gámez 0001, Heike Trautmann
Mach. Learn.2
2021 Hand pose aware multimodal isolated sign language recognition
Razieh Rastgoo, Kourosh Kiani, Sergio Escalera
Multim. Tools Appl.3
2021 Winning Solutions and Post-Challenge Analyses of the ChaLearn AutoDL Challenge 2019
abstract
This paper reports the results and post-challenge analyses of ChaLearn's AutoDL challenge series, which helped sorting out a profusion of AutoML solutions for Deep Learning (DL) that had been introduced in a variety of settings, but lacked fair comparisons. All input data modalities (time series, images, videos, text, tabular) were formatted as tensors and all tasks were multi-label classification problems. Code submissions were executed on hidden tasks, with limited time and computational resources, pushing solutions that get results quickly. In this setting, DL methods dominated, though popular Neural Architecture Search (NAS) was impractical. Solutions relied on fine-tuned pre-trained networks, with architectures matching data modality. Post-challenge tests did not reveal improvements beyond the imposed time limit. While no component is particularly original or novel, a high level modular organization emerged featuring a "meta-learner", "data ingestor", "model selector", "model/learner", and "evaluator". This modularity enabled ablation studies, which revealed the importance of (off-platform) meta-learning, ensembling, and efficient data management. Experiments on heterogeneous module combinations further confirm the (local) optimality of the winning solutions. Our challenge legacy includes an ever-lasting benchmark (http://autodl.chalearn.org), the open-sourced code of the winners, and a free "AutoDL self-service."
Zhengying Liu, Adrien Pavão, Zhen Xu 0007, Sergio Escalera, Fabio Ferreira, Isabelle Guyon, Sirui Hong, Frank Hutter, Rongrong Ji, Júlio C. S. Jacques Júnior, Marius Lindauer, Meysam Madadi, Thomas Nierhoff, Kangning Niu, Chunguang Pan, Danny Stoll, Sébastien Treguer, Peng Wang 0095, Chenglin Wu 0001, Youcheng Xiong, Arber Zela, Yang Zhang 0079
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Automatic Recognition of Facial Displays of Unfelt Emotions
abstract
Humans modify their facial expressions in order to communicate their internal states and sometimes to mislead observers regarding their true emotional states. Evidence in experimental psychology shows that discriminative facial responses are short and subtle. This suggests that such behavior would be easier to distinguish when captured in high resolution at an increased frame rate. We are proposing SASE-FE, the first dataset of facial expressions that are either congruent or incongruent with underlying emotion states. We show that overall the problem of recognizing whether facial movements are expressions of authentic emotions or not can be successfully addressed by learning spatio-temporal representations of the data. For this purpose, we propose a method that aggregates features along fiducial trajectories in a deeply learnt space. Performance of the proposed model shows that on average, it is easier to distinguish among genuine facial expressions of emotion than among unfelt facial expressions of emotion and that certain emotion pairs such as contempt and disgust are more difficult to distinguish than the rest. Furthermore, the proposed methodology improves state of the art results on CK+ and OULU-CASIA datasets for video emotion recognition, and achieves competitive results when classifying facial action units on BP4D datase.
Kaustubh Kulkarni, Ciprian A. Corneanu, Ikechukwu Ofodile, Sergio Escalera, Xavier Baró, Sylwia Julia Hyniewska, Juri Allik, Gholamreza Anbarjafari
IEEE Trans. Affect. Comput.4
2021 Survey on Emotional Body Gesture Recognition
abstract
Automatic emotion recognition has become a trending research topic in the past decade. While works based on facial expressions or speech abound, recognizing affect from body gestures remains a less explored topic. We present a new comprehensive survey hoping to boost research in the field. We first introduce emotional body gestures as a component of what is commonly known as ”body language” and comment general aspects as gender differences and culture dependence. We then define a complete framework for automatic emotional body gesture recognition. We introduce person detection and comment static and dynamic body pose estimation methods both in RGB and 3D. We then comment the recent literature related to representation learning and emotion recognition from images of emotionally expressive gestures. We also discuss multi-modal approaches that combine speech or face with body gestures for improved emotion recognition. While pre-processing methodologies (e.g., human detection and pose estimation) are nowadays mature technologies fully developed for robust large scale analysis, we show that for emotion recognition the quantity of labelled data is scarce. There is no agreement on clearly defined output spaces and the representations are shallow and largely based on naive geometrical representations.
Fatemeh Noroozi, Ciprian A. Corneanu, Dorota Kaminska, Tomasz Sapinski, Sergio Escalera, Gholamreza Anbarjafari
IEEE Trans. Affect. Comput.5
2021 On the Effect of Observed Subject Biases in Apparent Personality Analysis From Audio-Visual Signals
abstract
Personality perception is implicitly biased due to many subjective factors, such as cultural, social, contextual, gender, and appearance. Approaches developed for automatic personality perception are not expected to predict the real personality of the target but the personality external observers attributed to it. Hence, they have to deal with human bias, inherently transferred to the training data. However, bias analysis in personality computing is an almost unexplored area. In this article, we study different possible sources of bias affecting personality perception, including emotions from facial expressions, attractiveness, age, gender, and ethnicity, as well as their influence on prediction ability for apparent personality estimation. To this end, we propose a multimodal deep neural network that combines raw audio and visual information alongside predictions of attribute-specific models to regress apparent personality. We also analyze spatio-temporal aggregation schemes and the effect of different time intervals on first impressions. We base our study on the ChaLearn first impressions dataset, consisting of one-person conversational videos. Our model shows state-of-the-art results regressing apparent personality based on the Big-Five model. Furthermore, given the interpretability nature of our network design, we provide an incremental analysis on the impact of each possible source of bias on final network predictions.
Ricardo Darío Pérez Principi, Cristina Palmero, Júlio C. S. Jacques Júnior, Sergio Escalera
IEEE Trans. Affect. Comput.4
2021 Multi-Centre, Multi-Vendor and Multi-Disease Cardiac Segmentation: The M&Ms Challenge
abstract
The emergence of deep learning has considerably advanced the state-of-the-art in cardiac magnetic resonance (CMR) segmentation. Many techniques have been proposed over the last few years, bringing the accuracy of automated segmentation close to human performance. However, these models have been all too often trained and validated using cardiac imaging samples from single clinical centres or homogeneous imaging protocols. This has prevented the development and validation of models that are generalizable across different clinical centres, imaging conditions or scanner vendors. To promote further research and scientific benchmarking in the field of generalizable deep learning for cardiac segmentation, this paper presents the results of the Multi-Centre, Multi-Vendor and Multi-Disease Cardiac Segmentation (M&Ms) Challenge, which was recently organized as part of the MICCAI 2020 Conference. A total of 14 teams submitted different solutions to the problem, combining various baseline models, data augmentation strategies, and domain adaptation techniques. The obtained results indicate the importance of intensity-driven data augmentation, as well as the need for further research to improve generalizability towards unseen scanner vendors or new imaging protocols. Furthermore, we present a new resource of 375 heterogeneous CMR datasets acquired by using four different scanner vendors in six hospitals and three different countries (Spain, Canada and Germany), which we provide as open-access for the community to enable future research in the field.
Víctor M. Campello, Polyxeni Gkontra, Cristian Izquierdo, Carlos Martín-Isla, Alireza Sojoudi, Peter M. Full, Klaus H. Maier-Hein, Yao Zhang 0010, Zhiqiang He 0002, Jun Ma 0016, Mario Parreño, Alberto Albiol, Fanwei Kong, Shawn C. Shadden, Jorge Corral Acero, Vaanathi Sundaresan, Mina Saber, Mustafa A. Alattar, Hongwei Li 0004, Bjoern Menze, Firas Khader, Christoph Haarburger, Cian M. Scannell, Mitko Veta, Adam Carscadden, Kumaradevan Punithakumar, Xiao Liu 0037, Sotirios A. Tsaftaris, Xiaoqiong Huang, Xin Yang 0009, Lei Li 0020, Xiahai Zhuang, David Viladés, Martín Luís Descalzo, Andrea Guala 0002, Lucia La Mura, Matthias G. W. Friedrich, Ria Garg, Julie Lebel, Filipe Henriques, Mahir Karakas, Ersin Çavus, Steffen E. Petersen, Sergio Escalera, Santi Seguí, Jose Rodriguez-Palomares, Karim Lekadir
IEEE Trans. Medical Imaging44
2021 PBNS: physically based neural simulation for unsupervised garment pose space deformation
abstract
We present a methodology to automatically obtain Pose Space Deformation (PSD) basis for rigged garments through deep learning. Classical approaches rely on Physically Based Simulations (PBS) to animate clothes. These are general solutions that, given a sufficiently fine-grained discretization of space and time, can achieve highly realistic results. However, they are computationally expensive and any scene modification prompts the need of re-simulation. Linear Blend Skinning (LBS) with PSD offers a lightweight alternative to PBS, though, it needs huge volumes of data to learn proper PSD. We propose using deep learning, formulated as an implicit PBS, to un-supervisedly learn realistic cloth Pose Space Deformations in a constrained scenario: dressed humans. Furthermore, we show it is possible to train these models in an amount of time comparable to a PBS of a few sequences. To the best of our knowledge, we are the first to propose a neural simulator for cloth. While deep-based approaches in the domain are becoming a trend, these are data-hungry models. Moreover, authors often propose complex formulations to better learn wrinkles from PBS data. Supervised learning leads to physically inconsistent predictions that require collision solving to be used. Also, dependency on PBS data limits the scalability of these solutions, while their formulation hinders its applicability and compatibility. By proposing an unsupervised methodology to learn PSD for LBS models (3D animation standard), we overcome both of these drawbacks. Results obtained show cloth-consistency in the animated garments and meaningful pose-dependant folds and wrinkles. Our solution is extremely efficient, handles multiple layers of cloth, allows unsupervised outfit resizing and can be easily applied to any custom 3D avatar.
Hugo Bertiche, Meysam Madadi, Sergio Escalera
ACM Trans. Graph.3
2020 Computing the Testing Error Without a Testing Set
abstract
Deep Neural Networks (DNNs) have revolutionized computer vision. We now have DNNs that achieve top (accuracy) results in many problems, including object recognition, facial expression analysis, and semantic segmentation, to name but a few. The design of the DNNs that achieve top results is, however, non-trivial and mostly done by trail-and-error. That is, typically, researchers will derive many DNN architectures (\ie, topologies) and then test them on multiple datasets. However, there are no guarantees that the selected DNN will perform well in the real world. One can use a testing set to estimate the performance gap between the training and testing sets, but avoiding overfitting-to-the-testing-data is of concern. Using a sequestered testing data may address this problem, but this requires a constant update of the dataset, a very expensive venture. Here, we derive an algorithm to estimate the performance gap between training and testing without the need of a testing dataset. Specifically, we derive a set of persistent topology measures that identify when a DNN is learning to generalize to unseen samples. We provide extensive experimental validation on multiple networks and datasets to demonstrate the feasibility of the proposed approach.
Ciprian A. Corneanu, Sergio Escalera, Aleix Martinez
CVPR2
2020 Gate-Shift Networks for Video Action Recognition
abstract
Deep 3D CNNs for video action recognition are designed to learn powerful representations in the joint spatio-temporal feature space. In practice however, because of the large number of parameters and computations involved, they may under-perform in the lack of sufficiently large datasets for training them at scale. In this paper we introduce spatial gating in spatial-temporal decomposition of 3D kernels. We implement this concept with Gate-Shift Module (GSM). GSM is lightweight and turns a 2D-CNN into a highly efficient spatio-temporal feature extractor. With GSM plugged in, a 2D-CNN learns to adaptively route features through time and combine them, at almost no additional parameters and computational overhead. We perform an extensive evaluation of the proposed module to study its effectiveness in video action recognition, achieving state-of-the-art results on Something Something-V1 and Diving48 datasets, and obtaining competitive results on EPIC-Kitchens with far less model complexity.
Swathikiran Sudhakaran, Sergio Escalera, Oswald Lanz
CVPR2
2020 CLOTH3D: Clothed 3D Humans
Hugo Bertiche, Meysam Madadi, Sergio Escalera
ECCV (20)3
2020 ChaLearn LAP 2020 Challenge on Identity-preserved Human Detection: Dataset and Results
abstract
This paper summarizes the ChaLearn Looking at People 2020 Challenge on Identity-preserved Human Detection (IPHD). For the purpose, we released a large novel dataset containing more than 112K pairs of spatiotemporally aligned depth and thermal frames (and 175K instances of humans) sampled from 780 sequences. The sequences contain hundreds of non-identifiable people appearing in a mix of in-the-wild and scripted scenarios recorded in public and private places. The competition was divided into three tracks depending on the modalities exploited for the detection: (1) depth, (2) thermal, and (3) depth-thermal fusion. Color was also captured but only used to facilitate the groundtruth annotation. Still the temporal synchronization of three sensory devices is challenging, so bad temporal matches across modalities can occur. Hence, the labels provided should considered “weak”, although test frames were carefully selected to minimize this effect and ensure the fairest comparison of the participants' results. Despite this added difficulty, the results got by the participants demonstrate current fully-supervised methods can deal with that and achieve outstanding detection performance when measured in terms of [email protected].
Albert Clapés, Júlio C. S. Jacques Júnior, Carla Morral, Sergio Escalera
FG4
2020 Explainable Early Stopping for Action Unit Recognition
abstract
A common technique to avoid overfitting when training deep neural networks (DNN) is to monitor the performance in a dedicated validation data partition and to stop training as soon as it saturates. This only focuses on what the model does, while completely ignoring what happens inside it. In this work, we open the “black-box” of DNN in order to perform early stopping. We propose to use a novel theoretical framework that analyses meso-scale patterns in the topology of the functional graph of a network while it trains. Based on it, we decide when it transitions from learning towards overfitting in a more explainable way. We exemplify the benefits of this approach on a state-of-the art custom DNN that jointly learns local representations and label structure employing an ensemble of dedicated subnetworks. We show that it is practically equivalent in performance to early stopping with patience, the standard early stopping algorithm in the literature. This proves beneficial for AU recognition performance and provides new insights into how learning of AUs occurs in DNNs.
Ciprian A. Corneanu, Meysam Madadi, Sergio Escalera, Aleix Martinez
FG3
2020 Seniors' ability to decode differently aged facial emotional expressions
abstract
The present investigation aims at assessing elders' ability to decode facial emotional expressions conveyed by differently aged people in order to confirm (or disconfirm) the appropriateness of the “own age bias” theory, as well as investigate effects of different ages and different emotional categories. The study, involves 44 healthy elders (23 females), aged 65+ (mean age=75.09; SD=±7.9) which were requested to label 76 pictures depicting elders, middle-aged and young women and men displaying the six facial emotional expressions of disgust, anger, fear, sadness, happiness and neutrality. Results show a complex pattern of influences that calls for more deep investigations on the features to be accounted by providing socially and emotionally believable interfaces of effective and efficient algorithms to detect and decode their users' emotional facial expressions.
Anna Esposito, Terry Amorese, Mauro N. Maldonato, Alessandro Vinciarelli, M. Inés Torres, Sergio Escalera, Gennaro Cordasco
FG6
2020 Impairments in decoding facial and vocal emotional expressions in high functioning autistic adults and adolescents
abstract
The present investigation shows that gender of stimuli, age, and emotional categories affects the ability of adults and adolescent with Autistic Spectrum Conditions (ASC) to decode facial and vocal emotional expressions. A total of 60 subjects participated to the research: 15 ASC and 15 control adolescents aged between 10-14 years; and 15 ASC and 15 control young adults aged between 20-24 years. Their tasks consisted in decoding: a) 24 adults and 24 children contemporary facial emotional expressions of happiness, sadness, anger, fear, surprise, and disgust; and b) 20 adult's vocal emotional expressions of the same abovementioned emotions (except disgust). Significant differences were observed between ASC and typically developed peers. The data suggest that gender, type (voices or faces) of stimuli, and participants' age affect the emotion recognition process making difficult the definition of a common and shared pattern of emotional expression's recognition compliance among autistic and control groups. These results suggest that efficient and effective e-health technologies need to be able to learn and adapt to user individual traits and subjective needs to offer personalized assistance and support.
Anna Esposito, Italia Cirillo, Antonietta Maria Esposito, Leopoldina Fortunati, Gian Luca Foresti, Sergio Escalera, Nikolaos G. Bourbakis
FG6
2020 Generative Video Face Reenactment by AUs and Gaze Regularization
abstract
In this work, we propose an encoder-decoder-like architecture to perform face reenactment in image sequences. Our goal is to transfer the training subject identity to a given test subject. We regularize face reenactment by facial action unit intensity and 3D gaze vector regression. This way, we enforce the network to transfer subtle facial expressions and eye dynamics, providing a more lifelike result. The proposed encoder-decoder receives as input the previous sequence frame stacked to the current frame image of facial landmarks. Thus, the generated frames benefit from appearance and geometry, while keeping temporal coherence for the generated sequence. At test stage, a new target subject with the facial performance of the source subject and the appearance of the training subject is reenacted. Principal component analysis is applied to project the test subject geometry to the closest training subject geometry before reenactment. Evaluation of our proposal shows faster convergence, and more accurate and realistic results in comparison to other architectures without action units and gaze regularization.
Josep Famadas, Meysam Madadi, Cristina Palmero, Sergio Escalera
FG4
2020 Message from the General and Program Chairs FG 2020
abstract
Welcome to the 15th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2020). FG is the premier international conference on vision-based automatic face and body behavior analysis and applications. Since its first meeting in Zurich in 1994, the conference has been held fourteen times throughout the world. This is the 15th conference.
Juan P. Wachs, Sergio Escalera, Jeffrey F. Cohn, Albert Ali Salah, Arun Ross
FG2
2020 Hand sign language recognition using multi-view hand skeleton
Razieh Rastgoo, Kourosh Kiani, Sergio Escalera
Expert Syst. Appl.3
2020 CR-Net: A Deep Classification-Regression Network for Multimodal Apparent Personality Analysis
Yunan Li 0001, Jun Wan 0001, Qiguang Miao, Sergio Escalera, Huijuan Fang, Huizhou Chen, Xiangda Qi, Guodong Guo
Int. J. Comput. Vis.4
2020 Video-based isolated hand sign language recognition using a deep cascaded model
Razieh Rastgoo, Kourosh Kiani, Sergio Escalera
Multim. Tools Appl.3
2020 Guest Editorial: Image and Video Inpainting and Denoising
abstract
The papers in this special issue comprise all aspects of computer vision and pattern recognition devoted to image and video inpainting, including related tasks like denoising, debluring, sampling, super-resolutkon enhancement, restoration, hallucination, etc. The special issue was associated to the 2018 Chalearn Looking at People Satellite ECCV Workshop1 and the 2018 ChaLearn Challenges on Image and Video Inpainting.
Sergio Escalera, Hugo Jair Escalante, Xavier Baró, Isabelle Guyon, Meysam Madadi, Jun Wan 0001, Stéphane Ayache, Yagmur Güçlütürk, Umut Güçlü
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 SMPLR: Deep learning based SMPL reverse for 3D human pose and shape recovery
Meysam Madadi, Hugo Bertiche, Sergio Escalera
Pattern Recognit.3
2020 Towards automated computer vision: analysis of the AutoCV challenges 2019
Zhengying Liu, Zhen Xu 0007, Sergio Escalera, Isabelle Guyon, Júlio C. S. Jacques Júnior, Meysam Madadi, Adrien Pavão, Sébastien Treguer, Wei-Wei Tu
Pattern Recognit. Lett.3
2019 What Does It Mean to Learn in Deep Networks? And, How Does One Detect Adversarial Attacks?
abstract
The flexibility and high-accuracy of Deep Neural Networks (DNNs) has transformed computer vision. But, the fact that we do not know when a specific DNN will work and when it will fail has resulted in a lack of trust. A clear example is self-driving cars; people are uncomfortable sitting in a car driven by algorithms that may fail under some unknown, unpredictable conditions. Interpretability and explainability approaches attempt to address this by uncovering what a DNN models, i.e., what each node (cell) in the network represents and what images are most likely to activate it. This can be used to generate, for example, adversarial attacks. But these approaches do not generally allow us to determine where a DNN will succeed or fail and why. i.e., does this learned representation generalize to unseen samples? Here, we derive a novel approach to define what it means to learn in deep networks, and how to use this knowledge to detect adversarial attacks. We show how this defines the ability of a network to generalize to unseen testing samples and, most importantly, why this is the case.
Ciprian A. Corneanu, Meysam Madadi, Sergio Escalera, Aleix Martinez
CVPR3
2019 LSTA: Long Short-Term Attention for Egocentric Action Recognition
abstract
Egocentric activity recognition is one of the most challenging tasks in video analysis. It requires a fine-grained discrimination of small objects and their manipulation. While some methods base on strong supervision and attention mechanisms, they are either annotation consuming or do not take spatio-temporal patterns into account. In this paper we propose LSTA as a mechanism to focus on features from spatial relevant parts while attention is being tracked smoothly across the video sequence. We demonstrate the effectiveness of LSTA on egocentric activity recognition with an end-to-end trainable two-stream architecture, achieving state-of-the-art performance on four standard benchmarks.
Swathikiran Sudhakaran, Sergio Escalera, Oswald Lanz
CVPR2
2019 A Dataset and Benchmark for Large-Scale Multi-Modal Face Anti-Spoofing
abstract
Face anti-spoofing is essential to prevent face recognition systems from a security breach. Much of the progresses have been made by the availability of face anti-spoofing benchmark datasets in recent years. However, existing face anti-spoofing benchmarks have limited number of subjects (≤170) and modalities (≤2), which hinder the further development of the academic community. To facilitate face anti-spoofing research, we introduce a large-scale multi-modal dataset, namely CASIA-SURF, which is the largest publicly available dataset for face anti-spoofing in terms of both subjects and visual modalities. Specifically, it consists of 1,000 subjects with 21,000 videos and each sample has 3 modalities (i.e., RGB, Depth and IR). We also provide a measurement set, evaluation protocol and training/validation/testing subsets, developing a new benchmark for face anti-spoofing. Moreover, we present a new multi-modal fusion method as baseline, which performs feature re-weighting to select the more informative channel features while suppressing the less useful ones for each modal. Extensive experiments have been conducted on the proposed dataset to verify its significance and generalization capability. The dataset is available at https://sites.google.com/qq.com/chalearnfacespoofingattackdete/.
Xiaobo Wang 0001, Ajian Liu 0001, Jun Wan 0001, Sergio Escalera, Hailin Shi, Stan Z. Li
CVPR6
2019 Multi-task human analysis in still images: 2D/3D pose, depth map, and multi-part segmentation
abstract
While many individual tasks in the domain of human analysis have recently received an accuracy boost from deep learning approaches, multi-task learning has mostly been ignored due to a lack of data. New synthetic datasets are being released, filling this gap with synthetic generated data. In this work, we analyze four related human analysis tasks in still images in a multi-task scenario by leveraging such datasets. Specifically, we study the correlation of 2D/3D pose estimation, body part segmentation and full-body depth estimation. These tasks are learned via the well-known Stacked Hourglass module such that each of the task-specific streams shares information with the others. The main goal is to analyze how training together these four related tasks can benefit each individual task for a better generalization. Results on the newly released SURREAL dataset show that all four tasks benefit from the multi-task approach, but with different combinations of tasks: while combining all four tasks improves 2D pose estimation the most, 2D pose improves neither 3D pose nor full-body depth estimation. On the other hand 2D parts segmentation can benefit from 2D pose but not from 3D pose. In all cases, as expected, the maximum improvement is achieved on those human body parts that show more variability in terms of spatial distribution, appearance and shape, e.g. wrists and ankles.
Daniel Sánchez 0005, Marc Oliu, Meysam Madadi, Xavier Baró, Sergio Escalera
FG5
2019 On the effect of age perception biases for real age regression
abstract
Automatic age estimation from facial images represents an important task in computer vision. This paper analyses the effect of gender, age, ethnic, makeup and expression attributes of faces as sources of bias to improve deep apparent age prediction. Following recent works where it is shown that apparent age labels benefit real age estimation, rather than direct real to real age regression, our main contribution is the integration, in an end-to-end architecture, of face attributes for apparent age prediction with an additional loss for real age regression. Experimental results on the APPA-REAL dataset indicate the proposed network successfully take advantage of the adopted attributes to improve both apparent and real age estimation. Our model outperformed a state-of-the-art architecture proposed to separately address apparent and real age regression. Finally, we present preliminary results and discussion of a proof of concept application using the proposed model to regress the apparent age of an individual based on the gender of an external observer.
Júlio C. S. Jacques Júnior, Cagri Ozcinar, Marina Marjanovic, Xavier Baró, Gholamreza Anbarjafari, Sergio Escalera
FG6
2019 From 2D to 3D geodesic-based garment matching
Egils Avots, Meysam Madadi, Sergio Escalera, Jordi Gonzàlez 0001, Xavier Baró, Paul Pällin, Gholamreza Anbarjafari
Multim. Tools Appl.3
2019 A novel deep network architecture for reconstructing RGB facial images from thermal for face recognition
Andre Litvin, Kamal Nasrollahi, Sergio Escalera, Cagri Ozcinar, Thomas B. Moeslund, Gholamreza Anbarjafari
Multim. Tools Appl.3
2019 Guest editorial: special issue on human abnormal behavioural analysis
Gholamreza Anbarjafari, Sergio Escalera, Kamal Nasrollahi, Hugo Jair Escalante, Xavier Baró, Jun Wan 0001, Thomas B. Moeslund
Mach. Vis. Appl.2
2019 Audio-Visual Emotion Recognition in Video Clips
abstract
This paper presents a multimodal emotion recognition system, which is based on the analysis of audio and visual cues. From the audio channel, Mel-Frequency Cepstral Coefficients, Filter Bank Energies and prosodic features are extracted. For the visual part, two strategies are considered. First, facial landmarks' geometric relations, i.e., distances and angles, are computed. Second, we summarize each emotional video into a reduced set of key-frames, which are taught to visually discriminate between the emotions. In order to do so, a convolutional neural network is applied to key-frames summarizing videos. Finally, confidence outputs of all the classifiers from all the modalities are used to define a new feature space to be learned for final emotion label prediction, in a late fusion/stacking fashion. The experiments conducted on the SAVEE, eNTERFACE'05, and RML databases show significant performance improvements by our proposed system in comparison to current alternatives, defining the current state-of-the-art in all three databases.
Fatemeh Noroozi, Marina Marjanovic, Angelina Njegus, Sergio Escalera, Gholamreza Anbarjafari
IEEE Trans. Affect. Comput.4
2019 Dynamic 3D Hand Gesture Recognition by Learning Weighted Depth Motion Maps
abstract
Hand gesture recognition (HGR) from sequences of depth maps is a challenging computer vision task because of the low inter-class and high intra-class variability, different execution rates of each gesture, and the high articulated nature of the human hand. In this paper, a multilevel temporal sampling (MTS) method is first proposed that is based on the motion energy of keyframes of depth sequences. As a result, long, middle, and short sequences are generated that contain the relevant gesture information. The MTS results in increasing the intra-class similarity while raising the inter-class dissimilarities. The weighted depth motion map (WDMM) is then proposed to extract the spatiotemporal information from generated summarized sequences by an accumulated weighted absolute difference of consecutive frames. The histogram of gradient and local binary pattern are exploited to extract features from WDMM. The obtained results define the current state-of-the-art on three public benchmark datasets of: MSR Gesture 3D, SKIG, and MSR Action 3D, for 3D HGR. We also achieve competitive results on NTU action dataset.
Reza Azad, Maryam Asadi-Aghbolaghi, Shohreh Kasaei, Sergio Escalera
IEEE Trans. Circuits Syst. Video Technol.4
2018 Recurrent CNN for 3D Gaze Estimation using Appearance and Shape Cues
Cristina Palmero, Javier Selva, Mohammad Ali Bagheri, Sergio Escalera
BMVC4
2018 Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future Goals
abstract
In this paper, we strive to answer two questions: What is the current state of 3D hand pose estimation from depth images? And, what are the next challenges that need to be tackled? Following the successful Hands In the Million Challenge (HIM2017), we investigate the top 10 state-of-the-art methods on three tasks: single frame 3D pose estimation, 3D hand tracking, and hand pose estimation during object interaction. We analyze the performance of different CNN structures with regard to hand shape, joint visibility, view point and articulation distributions. Our findings include: (1) isolated 3D hand pose estimation achieves low mean errors (10 mm) in the view point range of [70, 120] degrees, but it is far from being solved for extreme view points; (2) 3D volumetric representations outperform 2D CNNs, better capturing the spatial structure of the depth data; (3) Discriminative methods still generalize poorly to unseen hand shapes; (4) While joint occlusions pose a challenge for most methods, explicit modeling of structure constraints can significantly narrow the gap between errors on visible and occluded joints.
Shanxin Yuan, Guillermo Garcia-Hernando, Björn Stenger, Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee, Pavlo Molchanov 0001, Jan Kautz, Sina Honari, Liuhao Ge, Junsong Yuan 0001, Xinghao Chen 0001, Guijin Wang, Fan Yang 0032, Kai Akiyama, Yang Wu 0001, Qingfu Wan, Meysam Madadi, Sergio Escalera, Shile Li, Dongheui Lee, Iasonas Oikonomidis, Antonis A. Argyros, Tae-Kyun Kim 0001
CVPR19
2018 Deep Structure Inference Network for Facial Action Unit Recognition
Ciprian A. Corneanu, Meysam Madadi, Sergio Escalera
ECCV (12)3
2018 Folded Recurrent Neural Networks for Future Video Prediction
Marc Oliu, Javier Selva, Sergio Escalera
ECCV (14)3
2018 Changes in Facial Expression as Biometric: A Database and Benchmarks of Identification
abstract
Facial dynamics can be considered as unique signatures for discrimination between people. These have started to become important topic since many devices have the possibility of unlocking using face recognition or verification. In this work, we evaluate the efficacy of the transition frames of video in emotion as compared to the peak emotion frames for identification. For experiments with transition frames we extract features from each frame of the video from a fine-tuned VGG-Face Convolutional Neural Network (CNN) and geometric features from facial landmark points. To model the temporal context of the transition frames we train a Long-Short Term Memory (LSTM) on the geometric and the CNN features. Furthermore, we employ two fusion strategies: first, an early fusion, in which the geometric and the CNN features are stacked and fed to the LSTM. Second, a late fusion, in which the prediction of the LSTMs, trained independently on the two features, are stacked and used with a Support Vector Machine (SVM). Experimental results show that the late fusion strategy gives the best results and the transition frames give better identification results as compared to the peak emotion frames.
Rain Eric Haamer, Kaustubh Kulkarni, Nasrin Imanpour, Mohammad A. Haque, Egils Avots, Michelle Breisch, Kamal Nasrollahi, Sergio Escalera, Cagri Ozcinar, Xavier Baró, Ahmad Reza Naghsh-Nilchi, Thomas B. Moeslund, Gholamreza Anbarjafari
FG8
2018 Deep Multimodal Pain Recognition: A Database and Comparison of Spatio-Temporal Visual Modalities
abstract
Pain is a symptom of many disorders associated with actual or potential tissue damage in human body. Managing pain is not only a duty but also highly cost prone. The most primitive state of pain management is the assessment of pain. Traditionally it was accomplished by self-report or visual inspection by experts. However, automatic pain assessment systems from facial videos are also rapidly evolving due to the need of managing pain in a robust and cost effective way. Among different challenges of automatic pain assessment from facial video data two issues are increasingly prevalent: first, exploiting both spatial and temporal information of the face to assess pain level, and second, incorporating multiple visual modalities to capture complementary face information related to pain. Most works in the literature focus on merely exploiting spatial information on chromatic (RGB) video data on shallow learning scenarios. However, employing deep learning techniques for spatio-temporal analysis considering Depth (D) and Thermal (T) along with RGB has high potential in this area. In this paper, we present the first state-of-the-art publicly available database, 'Multimodal Intensity Pain (MIntPAIN)' database, for RGBDT pain level recognition in sequences. We provide a first baseline results including 5 pain levels recognition by analyzing independent visual modalities and their fusion with CNN and LSTM models. From the experimental evaluation we observe that fusion of modalities helps to enhance recognition performance of pain levels in comparison to isolated ones. In particular, the combination of RGB, D, and T in an early fusion fashion achieved the best recognition rate.
Mohammad A. Haque, Rubén Ballester, Fatemeh Noroozi, Kaustubh Kulkarni, Christian B. Laursen, Ramin Irani, Marco Bellantonio, Sergio Escalera, Gholamreza Anbarjafari, Kamal Nasrollahi, Ole Kæseler Andersen, Erika G. Spaich, Thomas B. Moeslund
FG8
2018 RGB-D-based human motion recognition with deep learning: A survey
Pichao Wang, Wanqing Li 0001, Philip Ogunbona, Jun Wan 0001, Sergio Escalera
Comput. Vis. Image Underst.5
2018 Expert systems: Special issue on "Machine Learning Methods Neural Networks applied to Vision and Robotics (MLMVR)"
abstract
The International Joint Conference on Neural Networks (IJCNN) was held in Anchorage (Alaska) in May 2017. This top conference in the field of neural networks included many tracks and special sessions. In particular, a special session on Machine Learning Methods Neural Networks applied to Vision and Robotics (MLMVR) was organized by the authors receiving a large volume of excellent contributions. Only a small set of outstanding papers presented at this special session were invited to submit extended versions of their work. After a rigorous revision process, four of these papers were accepted.
José García Rodríguez 0001, Sergio Escalera, Alexandra Psarrou, Isabelle Guyon, Andrew Lewis 0004, Jürgen Leitner
Expert Syst. J. Knowl. Eng.2
2018 Recurrent neural networks for remote sensing image classification
abstract
Automatically classifying an image has been a central problem in computer vision for decades. A plethora of models has been proposed, from handcrafted feature solutions to more sophisticated approaches such as deep learning. The authors address the problem of remote sensing image classification, which is an important problem to many real world applications. They introduce a novel deep recurrent architecture that incorporates high‐level feature descriptors to tackle this challenging problem. Their solution is based on the general encoder–decoder framework. To the best of the authors’ knowledge, this is the first study to use a recurrent network structure on this task. The experimental results show that the proposed framework outperforms the previous works in the three datasets widely used in the literature. They have achieved a state‐of‐the‐art accuracy rate of 97.29% on the UC Merced dataset.
Mohamed Ilyes Lakhal, Hakan Çevikalp, Sergio Escalera, Ferda Ofli
IET Comput. Vis.3
2018 Back-dropout transfer learning for action recognition
abstract
Transfer learning aims at adapting a model learned from source dataset to target dataset. It is a beneficial approach especially when annotating on the target dataset is expensive or infeasible. Transfer learning has demonstrated its powerful learning capabilities in various vision tasks. Despite transfer learning being a promising approach, it is still an open question how to adapt the model learned from the source dataset to the target dataset. One big challenge is to prevent the impact of category bias on classification performance. Dataset bias exists when two images from the same category, but from different datasets, are not classified as the same. To address this problem, a transfer learning algorithm has been proposed, called negative back‐dropout transfer learning (NB‐TL), which utilizes images that have been misclassified and further performs back‐dropout strategy on them to penalize errors. Experimental results demonstrate the effectiveness of the proposed algorithm. In particular, the authors evaluate the performance of the proposed NB‐TL algorithm on UCF 101 action recognition dataset, achieving 88.9% recognition rate.
Huamin Ren, Nattiya Kanhabua, Andreas Møgelmose, Weifeng Liu 0002, Kaustubh Kulkarni, Sergio Escalera, Xavier Baró, Thomas B. Moeslund
IET Comput. Vis.6
2018 Looking at People Special Issue
Sergio Escalera, Jordi Gonzàlez 0001, Hugo Jair Escalante, Xavier Baró, Isabelle Guyon
Int. J. Comput. Vis.1
2018 Exploiting feature representations through similarity learning, post-ranking and ranking aggregation for person re-identification
Júlio C. S. Jacques Júnior, Xavier Baró, Sergio Escalera
Image Vis. Comput.3
2018 Top-down model fitting for hand pose recovery in sequences of depth images
Meysam Madadi, Sergio Escalera, Alex Carruesco, Carlos Andújar, Xavier Baró, Jordi Gonzàlez 0001
Image Vis. Comput.2
2018 Beyond one-hot encoding: Lower dimensional target embedding
Pau Rodríguez, Miguel Ángel Bautista 0001, Jordi Gonzàlez 0001, Sergio Escalera
Image Vis. Comput.4
2018 Action detection fusing multiple Kinects and a WIMU: an application to in-home assistive technology for the elderly
Albert Clapés, Àlex Pardo, Oriol Pujol, Sergio Escalera
Mach. Vis. Appl.4
2018 Error-Correcting Factorization
abstract
Error Correcting Output Codes (ECOC) is a successful technique in multi-class classification, which is a core problem in Pattern Recognition and Machine Learning. A major advantage of ECOC over other methods is that the multi-class problem is decoupled into a set of binary problems that are solved independently. However, literature defines a general error-correcting capability for ECOCs without analyzing how it distributes among classes, hindering a deeper analysis of pair-wise error-correction. To address these limitations this paper proposes an Error-Correcting Factorization (ECF) method. Our contribution is three fold: (I) We propose a novel representation of the error-correction capability, called the design matrix, that enables us to build an ECOC on the basis of allocating correction to pairs of classes. (II) We derive the optimal code length of an ECOC using rank properties of the design matrix. (III) ECF is formulated as a discrete optimization problem, and a relaxed solution is found using an efficient constrained block coordinate descent approach. (IV) Enabled by the flexibility introduced with the design matrix we propose to allocate the error-correction on classes that are prone to confusion. Experimental results in several databases show that when allocating the error-correction to confusable classes ECF outperforms state-of-the-art approaches.
Miguel Ángel Bautista 0001, Oriol Pujol, Fernando De la Torre, Sergio Escalera
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 Guest Editorial: The Computational Face
abstract
The papers in this special section examine the concept of automated face analysis (AFA). AFA has received special attention from the computer vision and pattern recognition communities. Research progress often gives the impression that problems such as face recognition and face detection are solved, at least for some scenarios. Several aspects of face analysis remain open problems, including the implementation of large scale face recognition/detection methods for in the wild images, emotion recognition, micro-expression analysis, and others. The community keeps making rapid progress on these topics, with continual improvement of current methods and creation of new ones that push the state-of-the-art. Applications are countless, including security and video surveillance, human computer/robot interaction, communication, entertainment, and commerce, while having an important social impact in assistive technologies for education and health. The importance of face analysis, together with the vast amount of work on the subject and the latest developments in the field, motivated us to organize a special section on this theme. The scope of the compilation comprises all aspects of face analysis from a computer vision perspective. Including, but not limited to: recognition, detection, alignment, reconstruction of faces, pose estimation of faces, gaze analysis, age, emotion, gender, and facial attributes estimation, and applications among others.
Sergio Escalera, Xavier Baró, Isabelle Guyon, Hugo Jair Escalante, Georgios Tzimiropoulos, Michel F. Valstar, Maja Pantic, Jeffrey F. Cohn, Takeo Kanade
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Articulated motion and deformable objects
Jun Wan 0001, Sergio Escalera, Francisco José Perales López, Josef Kittler
Pattern Recognit.2
2018 Guest Editorial: Apparent Personality Analysis
abstract
The papers in this special section focus on personality analysis. Automatic analysis of videos to characterize human behavior has become an active area of research with applications in affective computing, human-machine interfaces, gaming, security, marketing, and health, just to mention a few. Research advances in multimedia information processing, computer vision and pattern recognition have lead to established methodologies that are able to successfully recognize consciously executed actions, or intended movements (e.g., gestures, actions, interactions with objects and other people). However, recently there has been much progress in terms of computational approaches to characterize sub-conscious behaviors, which may be revealing aptitudes or competence, hidden intentions, and personality traits. Much remains to be done still, but it is essential nowadays to have a compilation of cutting edge work in this direction to identify potential opportunities and challenges involved. In this line, we edited a special issue on automatic methods for apparent personality analysis. Personality refers to individual differences in characteristic patterns of thinking, feeling and behaving.
Sergio Escalera, Xavier Baró, Isabelle Guyon, Hugo Jair Escalante
IEEE Trans. Affect. Comput.1
2018 Multimodal First Impression Analysis with Deep Residual Networks
abstract
People form first impressions about the personalities of unfamiliar individuals even after very brief interactions with them. In this study we present and evaluate several models that mimic this automatic social behavior. Specifically, we present several models trained on a large dataset of short YouTube video blog posts for predicting apparent Big Five personality traits of people and whether they seem suitable to be recommended to a job interview. Along with presenting our audiovisual approach and results that won the third place in the ChaLearn First Impressions Challenge, we investigate modeling in different modalities including audio only, visual only, language only, audiovisual, and combination of audiovisual and language. Our results demonstrate that the best performance could be obtained using a fusion of all data modalities. Finally, in order to promote explainability in machine learning and to provide an example for the upcoming ChaLearn challenges, we present a simple approach for explaining the predictions for job interview recommendations.
Yagmur Güçlütürk, Umut Güçlü, Xavier Baró, Hugo Jair Escalante, Isabelle Guyon, Sergio Escalera, Marcel van Gerven, Rob van Lier
IEEE Trans. Affect. Comput.6
2017 Apparent and Real Age Estimation in Still Images with Deep Residual Regressors on Appa-Real Database
abstract
After decades of research, the real (biological) age estimation from a single face image reached maturity thanks to the availability of large public face databases and impressive accuracies achieved by recently proposed methods. The estimation of “apparent age” is a related task concerning the age perceived by human observers. Significant advances have been also made in this new research direction with the recent Looking At People challenges. In this paper we make several contributions to age estimation research. (i) We introduce APPA-REAL, a large face image database with both real and apparent age annotations. (ii)We study the relationship between real and apparent age. (iii) We develop a residual age regression method to further improve the performance. (iv) We show that real age estimation can be successfully tackled as an apparent age estimation followed by an apparent to real age residual regression. (v) We graphically reveal the facial regions on which the CNN focuses in order to perform apparent and real age estimation tasks.
Eirikur Agustsson, Radu Timofte, Sergio Escalera, Xavier Baró, Isabelle Guyon, Rasmus Rothe
FG3
2017 A Survey on Deep Learning Based Approaches for Action and Gesture Recognition in Image Sequences
abstract
The interest in action and gesture recognition has grown considerably in the last years. In this paper, we present a survey on current deep learning methodologies for action and gesture recognition in image sequences. We introduce a taxonomy that summarizes important aspects of deep learning for approaching both tasks. We review the details of the proposed architectures, fusion strategies, main datasets, and competitions. We summarize and discuss the main works proposed so far with particular interest on how they treat the temporal dimension of data, discussing their main features and identify opportunities and challenges for future research.
Maryam Asadi-Aghbolaghi, Albert Clapés, Marco Bellantonio, Hugo Jair Escalante, Víctor Ponce-López, Xavier Baró, Isabelle Guyon, Shohreh Kasaei, Sergio Escalera
FG9
2017 Exploiting feature Representations Through Similarity Learning and Ranking Aggregation for Person Re-identification
abstract
Person re-identification has received special attentionby the human analysis community in the last few years.To address the challenges in this field, many researchers haveproposed different strategies, which basically exploit eithercross-view invariant features or cross-view robust metrics. Inthis work we propose to combine different feature representationsthrough ranking aggregation. Spatial information, whichpotentially benefits the person matching, is represented usinga 2D body model, from which color and texture informationare extracted and combined. We also consider contextualinformation (background and foreground data), automaticallyextracted via Deep Decompositional Network, and the usage ofConvolutional Neural Network (CNN) features. To describe thematching between images we use the polynomial feature map,also taking into account local and global information. Finally,the Stuart ranking aggregation method is employed to combinecomplementary ranking lists obtained from different featurerepresentations. Experimental results demonstrated that weimprove the state-of-the-art on VIPeR and PRID450s datasets,achieving 58.77% and 71.56% on top-1 rank recognitionrate, respectively, as well as obtaining competitive results onCUHK01 dataset.
Júlio C. S. Jacques Júnior, Xavier Baró, Sergio Escalera
FG3
2017 Dominant and Complementary Multi-Emotional Facial Expression Recognition Using C-Support Vector Classification
abstract
We are proposing a new facial expression recognition model which introduces 30+ detailed facial expressions recognisable by any artificial intelligence interacting with a human. Throughout this research, we introduce two categories for the emotions, namely, dominant emotions and complementary emotions. In this research paper the complementary emotion is recognised by using the eye region if the dominant emotion is angry, fearful or sad, and if the dominant emotion is disgust or happiness the complementary emotion is mainly conveyed by the mouth. In order to verify the tagged dominant and complementary emotions, randomly chosen people voted for the recognised multi-emotional facial expressions. The average results of voting are showing that 73.88% of the voters agree on the correctness of the recognised multi-emotional facial expressions.
Christer Loob, Pejman Rasti, Iiris Lüsi, Júlio C. S. Jacques Júnior, Xavier Baró, Sergio Escalera, Tomasz Sapinski, Dorota Kaminska, Gholamreza Anbarjafari
FG6
2017 Joint Challenge on Dominant and Complementary Emotion Recognition Using Micro Emotion Features and Head-Pose Estimation: Databases
abstract
In this work two databases for the Joint Challenge on Dominant and Complementary Emotion Recognition Using Micro Emotion Features and Head-Pose Estimation1 are introduced. Head pose estimation paired with and detailed emotion recognition have become very important in relation to human-computer interaction. The 3D head pose database, SASE, is a 3D database acquired with Microsoft Kinect 2 camera, including RGB and depth information of different head poses which is composed by a total of 30000 frames with annotated markers, including 32 male and 18 female subjects. For the dominant and complementary emotion database, iCVMEFED, includes 31250 images with different emotions of 115 subjects whose gender distribution is almost uniform. For each subject there are 5 samples. The emotions are composed by 7 basic emotions plus neutral, being defined as complementary and dominant pairs. The emotion associated to the images were labeled with the support of psychologists.
Iiris Lüsi, Júlio C. S. Jacques Júnior, Jelena Gorbova, Xavier Baró, Sergio Escalera, Hasan Demirel, Juri Allik, Cagri Ozcinar, Gholamreza Anbarjafari
FG5
2017 Occlusion Aware Hand Pose Recovery from Sequences of Depth Images
abstract
State-of-the-art approaches on hand pose estimation from depth images have reported promising results under quite controlled considerations. In this paper we propose a two-step pipeline for recovering the hand pose from a sequence of depth images. The pipeline has been designed to deal with images taken from any viewpoint and exhibiting a high degree of finger occlusion. In a first step we initialize the hand pose using a part-based model, fitting a set of hand components in the depth images. In a second step we consider temporal data and estimate the parameters of a trained bilinear model consisting of shape and trajectory bases. Results on a synthetic, highly-occluded dataset demonstrate that the proposed method outperforms most recent pose recovering approaches, including those based on CNNs.
Meysam Madadi, Sergio Escalera, Alex Carruesco, Carlos Andújar, Xavier Baró, Jordi Gonzàlez 0001
FG2
2017 Wordfence: Text detection in natural images with border awareness
abstract
In recent years, text recognition has achieved remarkable success in recognizing scanned document text. However, word recognition in natural images is still an open problem, which generally requires time consuming post-processing steps. We present a novel architecture for individual word detection in scene images based on semantic segmentation. Our contributions are twofold: the concept of WordFence, which detects border areas surrounding each individual word and a novel pixelwise weighted softmax loss function which penalizes background and emphasizes small text regions. WordFence ensures that each word is detected individually, and the new loss function provides a strong training signal to both text and word border localization. The proposed technique avoids intensive post-processing, producing an end-to-end word detection system. We achieve superior localization recall on common benchmark datasets - 92% recall on ICDAR11 and ICDAR13 and 63% recall on SVT. Furthermore, our end-to-end word recognition system achieves state-of-the-art 86% F-Score on ICDAR13.
Andrei Polzounov, Artsiom Ablavatski, Sergio Escalera, Shijian Lu, Jianfei Cai 0001
ICIP3
2017 Design of an explainable machine learning challenge for video interviews
abstract
This paper reviews and discusses research advances on “explainable machine learning” in computer vision. We focus on a particular area of the “Looking at People” (LAP) thematic domain: first impressions and personality analysis. Our aim is to make the computational intelligence and computer vision communities aware of the importance of developing explanatory mechanisms for computer-assisted decision making applications, such as automating recruitment. Judgments based on personality traits are being made routinely by human resource departments to evaluate the candidates' capacity of social insertion and their potential of career growth. However, inferring personality traits and, in general, the process by which we humans form a first impression of people, is highly subjective and may be biased. Previous studies have demonstrated that learning machines can learn to mimic human decisions. In this paper, we go one step further and formulate the problem of explaining the decisions of the models as a means of identifying what visual aspects are important, understanding how they relate to decisions suggested, and possibly gaining insight into undesirable negative biases. We design a new challenge on explainability of learning machines for first impressions analysis. We describe the setting, scenario, evaluation metrics and preliminary outcomes of the competition. To the best of our knowledge this is the first effort in terms of challenges for explainability in computer vision. In addition our challenge design comprises several other quantitative and qualitative elements of novelty, including a “coopetition” setting, which combines competition and collaboration.
Hugo Jair Escalante, Isabelle Guyon, Sergio Escalera, Júlio C. S. Jacques Júnior, Meysam Madadi, Xavier Baró, Stéphane Ayache, Evelyne Viegas, Yagmur Güçlütürk, Umut Güçlü, Marcel van Gerven, Rob van Lier
IJCNN3
2017 ChaLearn looking at people: A review of events and resources
abstract
This paper reviews the historic of ChaLearn Looking at People (LAP) events. We started in 2011 (with the release of the first Kinect device) to run challenges related to human action/activity and gesture recognition. Since then we have regularly organized events in a series of competitions covering all aspects of visual analysis of humans. So far we have organized more than 10 international challenges and events in this field. This paper reviews associated events, and introduces the ChaLearn LAP platform where public resources (including code, data and preprints of papers) related to the organized events are available. We also provide a discussion on our main findings and perspectives of ChaLearn LAP activities.
Sergio Escalera, Xavier Baró, Hugo Jair Escalante, Isabelle Guyon
IJCNN1
2017 Locality regularized group sparse coding for action recognition
Mohammad Ali Bagheri, Qigang Gao, Sergio Escalera, Thomas B. Moeslund, Huamin Ren, Elham Etemad
Comput. Vis. Image Underst.3
2017 Automatic Sleep System Recommendation by Multi-modal RBG-Depth-Pressure Anthropometric Analysis
Cristina Palmero, Jordi Esquirol, Vanessa Bayo, Miquel Àngel Cos, Pouya Ahmadmonfared, Joan Salabert, Sergio Escalera
Int. J. Comput. Vis.8
2017 Subspace Procrustes Analysis
Xavier Perez-Sala, Fernando De la Torre, Laura Igual, Sergio Escalera, Cecilio Angulo
Int. J. Comput. Vis.4
2017 Evolving weighting schemes for the Bag of Visual Words
Hugo Jair Escalante, Víctor Ponce-López, Sergio Escalera, Xavier Baró, Alicia Morales-Reyes, José Martínez-Carranza
Neural Comput. Appl.3
2017 Editorial: special issue on computational intelligence for vision and robotics
José García Rodríguez 0001, Isabelle Guyon, Sergio Escalera, Alexandra Psarrou, Andrew Lewis 0004, Miguel Cazorla
Neural Comput. Appl.3
2016 Continuous Supervised Descent Method for Facial Landmark Localisation
Marc Oliu, Ciprian A. Corneanu, László A. Jeni, Jeffrey F. Cohn, Takeo Kanade, Sergio Escalera
ACCV (2)6
2016 ChaLearn Joint Contest on Multimedia Challenges Beyond Visual Analysis: An overview
abstract
This paper provides an overview of the Joint Contest on Multimedia Challenges Beyond Visual Analysis. We organized an academic competition that focused on four problems that require effective processing of multimodal information in order to be solved. Two tracks were devoted to gesture spotting and recognition from RGB-D video, two fundamental problems for human computer interaction. Another track was devoted to a second round of the first impressions challenge of which the goal was to develop methods to recognize personality traits from short video clips. For this second round we adopted a novel collaborative-competitive (i.e., coopetition) setting. The fourth track was dedicated to the problem of video recommendation for improving user experience. The challenge was open for about 45 days, and received outstanding participation: almost 200 participants registered to the contest, and 20 teams sent predictions in the final stage. The main goals of the challenge were fulfilled: the state of the art was advanced considerably in the four tracks, with novel solutions to the proposed problems (mostly relying on deep learning). However, further research is still required. The data of the four tracks will be available to allow researchers to keep making progress in the four tracks.
Hugo Jair Escalante, Víctor Ponce-López, Jun Wan 0001, Michael Riegler 0001, Albert Clapés, Sergio Escalera, Isabelle Guyon, Xavier Baró, Pål Halvorsen, Henning Müller, Martha A. Larson
ICPR7
2016 Fusion of classifier predictions for audio-visual emotion recognition
abstract
In this paper is presented a novel multimodal emotion recognition system which is based on the analysis of audio and visual cues. MFCC-based features are extracted from the audio channel and facial landmark geometric relations are computed from visual data. Both sets of features are learnt separately using state-of-the-art classifiers. In addition, we summarise each emotion video into a reduced set of key-frames, which are learnt in order to visually discriminate emotions by means of a Convolutional Neural Network. Finally, confidence outputs of all classifiers from all modalities are used to define a new feature space to be learnt for final emotion prediction, in a late fusion/stacking fashion. The conducted experiments on eNTERFACE'05 database show significant performance improvements of our proposed system in comparison to state-of-the-art approaches.
Fatemeh Noroozi, Marina Marjanovic, Angelina Njegus, Sergio Escalera, Gholamreza Anbarjafari
ICPR4
2016 Support vector machines with time series distance kernels for action classification
abstract
Despite the outperformance of Support Vector Machine (SVM) on many practical classification problems, the algorithm is not directly applicable to multi-dimensional trajectories having different lengths. In this paper, a new class of SVM that is applicable to trajectory classification, such as action recognition, is developed by incorporating two efficient time-series distances measures into the kernel function. Dynamic Time Warping and Longest Common Subsequence distance measures along with their derivatives are employed as the SVM kernel. In addition, the pairwise proximity learning strategy is utilized in order to make use of non-positive semi-definite kernels in the SVM formulation. The proposed method is employed for a challenging classification problem: action recognition by depth cameras using only skeleton data; and evaluated on three benchmark action datasets. Experimental results demonstrate the outperformance of our methodology compared to the state-of-the-art on the considered datasets.
Mohammad Ali Bagheri, Qigang Gao, Sergio Escalera
WACV3
2016 Automatic garment retexturing based on infrared information
Egils Avots, Morteza Daneshmand, Andres Traumann, Sergio Escalera, Gholamreza Anbarjafari
Comput. Graph.4
2016 A real-time Human-Robot Interaction system based on gestures for assistive scenarios
Gerard Canal, Sergio Escalera, Cecilio Angulo
Comput. Vis. Image Underst.2
2016 Poselet-Based Contextual Rescoring for Human Pose Estimation via Pictorial Structures
Antonio Hernández-Vela, Stan Sclaroff, Sergio Escalera
Int. J. Comput. Vis.3
2016 Multi-modal RGB-Depth-Thermal Human Body Segmentation
Cristina Palmero, Albert Clapés, Chris Bahnsen, Andreas Møgelmose, Thomas B. Moeslund, Sergio Escalera
Int. J. Comput. Vis.6
2016 Challenges in multimodal gesture recognition
abstract
This paper surveys the state of the art on multimodal gesture recognition and introduces the JMLR special topic on gesture recognition 2011-2015. We began right at the start of the \kinect revolution when inexpensive infrared cameras providing image depth recordings became available. We published papers using this technology and other more conventional methods, including regular video cameras, to record data, thus providing a good overview of uses of machine learning and computer vision using multimodal data in this area of application. Notably, we organized a series of challenges and made available several datasets we recorded for that purpose, including tens of thousands of videos, which are available to conduct further research. We also overview recent state of the art works on gesture recognition based on a proposed taxonomy for gesture recognition, discussing challenges and future lines of research.
Sergio Escalera, Vassilis Athitsos, Isabelle Guyon
J. Mach. Learn. Res.1
2016 Robust non-blind color video watermarking using QR decomposition and entropy analysis
Pejman Rasti, Salma Samiei, Mary Agoyi, Sergio Escalera, Gholamreza Anbarjafari
J. Vis. Commun. Image Represent.4
2016 Survey on RGB, 3D, Thermal, and Multimodal Approaches for Facial Expression Recognition: History, Trends, and Affect-Related Applications
abstract
Facial expressions are an important way through which humans interact socially. Building a system capable of automatically recognizing facial expressions from images and video has been an intense field of study in recent years. Interpreting such expressions remains challenging and much research is needed about the way they relate to human affect. This paper presents a general overview of automatic RGB, 3D, thermal and multimodal facial expression analysis. We define a new taxonomy for the field, encompassing all steps from face detection to facial expression recognition, and describe and classify the state of the art methods accordingly. We also present the important datasets and the bench-marking of most influential methods. We conclude with a general discussion about trends, important questions and future lines of research.
Ciprian A. Corneanu, Marc Oliu, Jeffrey F. Cohn, Sergio Escalera
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 Guest Editors' Introduction to the Special Issue on Multimodal Human Pose Recovery and Behavior Analysis
abstract
The sixteen papers in this special section focus on human pose recovery and behavior analysis (HuPBA). This is one of the most challenging topics in computer vision, pattern analysis, and machine learning. It is of critical importance for application areas that include gaming, computer interaction, human robot interaction, security, commerce, assistive technologies and rehabilitation, sports, sign language recognition, and driver assistance technology, to mention just a few. In essence, HuPBA requires dealing with the articulated nature of the human body, changes in appearance due to clothing, and the inherent problems of clutter scenes, such as background artifacts, occlusions, and illumination changes. These papers represent the most recent research in this field, including new methods considering still images, image sequences, depth data, stereo vision, 3D vision, audio, and IMUs, among others.
Sergio Escalera, Jordi Gonzàlez 0001, Xavier Baró, Jamie Shotton
IEEE Trans. Pattern Anal. Mach. Intell.1
2016 Segmentation of RGB-D indoor scenes by stacking random forests and conditional random fields
Mikkel Thøgersen, Sergio Escalera, Jordi Gonzàlez 0001, Thomas B. Moeslund
Pattern Recognit. Lett.2
2016 A Gesture Recognition System for Detecting Behavioral Patterns of ADHD
abstract
We present an application of gesture recognition using an extension of dynamic time warping (DTW) to recognize behavioral patterns of attention deficit hyperactivity disorder (ADHD). We propose an extension of DTW using one-class classifiers in order to be able to encode the variability of a gesture category, and thus, perform an alignment between a gesture sample and a gesture class. We model the set of gesture samples of a certain gesture category using either Gaussian mixture models or an approximation of convex hulls. Thus, we add a theoretical contribution to classical warping path in DTW by including local modeling of intraclass gesture variability. This methodology is applied in a clinical context, detecting a group of ADHD behavioral patterns defined by experts in psychology/psychiatry, to provide support to clinicians in the diagnose procedure. The proposed methodology is tested on a novel multimodal dataset (RGB plus depth) of ADHD children recordings with behavioral patterns. We obtain satisfying results when compared to standard state-of-the-art approaches in the DTW context.
Miguel Ángel Bautista 0001, Antonio Hernández-Vela, Sergio Escalera, Laura Igual, Oriol Pujol, Josep Moya, Verónica Violant Holz, María Teresa Anguera
IEEE Trans. Cybern.3
2015 Gesture and Action Recognition by Evolved Dynamic Subgestures
abstract
This paper introduces a framework for gesture and action recognition based on the \nevolution of temporal gesture primitives, or subgestures. Our work is inspired on the \nprinciple of producing genetic variations within a population of gesture subsequences, \nwith the goal of obtaining a set of gesture units that enhance the generalization capability of standard gesture recognition approaches. In our context, gesture primitives are evolved over time using dynamic programming and generative models in order to recognize complex actions. In few generations, the proposed subgesture-based representation \nof actions and gestures outperforms the state of the art results on the MSRDaily3D and \nMSRAction3D datasets.
Víctor Ponce-López, Hugo Jair Escalante, Sergio Escalera, Xavier Baró
BMVC3
2015 Unsupervised Behavior-Specific Dictionary Learning for Abnormal Event Detection
abstract
Abnormal event detection has been a challenge due to the lack of complete normal information in the training data and the volatility of the definitions of both normality and abnormality. Recent research applying sparse representation has shown its effectiveness in the expression of normal patterns. Despite progress in this area, the relationship of atoms within the dictionary is commonly neglected, thereafter anomalies which are detected based on reconstruction error could brings high false alarm - noise or infrequent normal visual features could be wrongly detected as anomalies, especially when the training data is only a small proportion of the surveillance data. Therefore, we propose behavior-specific dictionaries (BSD) through unsupervised learning, pursuing atoms from the same type of behavior to represent one behavior dictionary. To further improve the dictionary by introducing information from potential infrequent normal patterns, we refine the dictionary by searching ‘missed atoms’ that have compact coefficients. Experimental results show that our BSD algorithm outperforms state-of-the-art dictionaries in abnormal event detection on the public UCSD dataset. Moreover, BSD has less false alarms compared to state-of-the-art dictionaries especially when the training set is small, which is demonstrated on Anomaly Stairs dataset.
Huamin Ren, Weifeng Liu 0002, Søren I. Olsen, Sergio Escalera, Thomas B. Moeslund
BMVC4
2015 Gesture based human multi-robot interaction
abstract
The emergence of robot applications for non-technical users implies designing new ways of interaction between robotic platforms and users. The main goal of this work is the development of a gestural interface to interact with robots in a similar way as humans do, allowing the user to provide information of the task with non-verbal communication. The gesture recognition application has been implemented using the Microsoft's Kinect™v2 sensor. Hence, a real-time algorithm based on skeletal features is described to deal with both, static gestures and dynamic ones, being the latter recognized using a weighted Dynamic Time Warping method. The gesture recognition application has been implemented in a multi-robot case. A NAO humanoid robot is in charge of interacting with the users and respond to the visual signals they produce. Moreover, a wheeled Wifibot robot carries both the sensor and the NAO robot, easing navigation when necessary. A broad set of user tests have been carried out demonstrating that the system is, indeed, a natural approach to human robot interaction, with a fast response and easy to use, showing high gesture recognition rates.
Gerard Canal, Cecilio Angulo, Sergio Escalera
IJCNN3
2015 Improving bag of visual words representations with genetic programming
abstract
The bag of visual words is a well established representation in diverse computer vision problems. Taking inspiration from the fields of text mining and retrieval, this representation has proved to be very effective in a large number of domains. In most cases, a standard term-frequency weighting scheme is considered for representing images and videos in computer vision. This is somewhat surprising, as there are many alternative ways of generating bag of words representations within the text processing community. This paper explores the use of alternative weighting schemes for landmark tasks in computer vision: image categorization and gesture recognition. We study the suitability of using well-known supervised and unsupervised weighting schemes for such tasks. More importantly, we devise a genetic program that learns new ways of representing images and videos under the bag of visual words representation. The proposed method learns to combine term-weighting primitives trying to maximize the classification performance. Experimental results are reported in standard image and video data sets showing the effectiveness of the proposed evolutionary algorithm.
Hugo Jair Escalante, José Martínez-Carranza, Sergio Escalera, Víctor Ponce-López, Xavier Baró
IJCNN3
2015 ChaLearn looking at people 2015 new competitions: Age estimation and cultural event recognition
abstract
Following previous series on Looking at People (LAP) challenges [1], [2], [3], in 2015 ChaLearn runs two new competitions within the field of Looking at People: age and cultural event recognition in still images. We propose the first crowd-sourcing application to collect and label data about apparent age of people instead of the real age. In terms of cultural event recognition, tens of categories have to be recognized. This involves scene understanding and human analysis. This paper summarizes both challenges and data, providing some initial baselines. The results of the first round of the competition were presented at ChaLearn LAP 2015 IJCNN special session on computer vision and robotics http://www.dtic.ua.es/~jgarcia/IJCNN2015. Details of the ChaLearn LAP competitions can be found at http://gesture.chalearn.org/.
Sergio Escalera, Jordi Gonzàlez 0001, Xavier Baró, Pablo Pardo, Junior Fabian, Marc Oliu, Hugo Jair Escalante, Ivan Huerta Casado, Isabelle Guyon
IJCNN1
2015 Design of the 2015 ChaLearn AutoML challenge
abstract
ChaLearn is organizing the Automatic Machine Learning (AutoML) contest for IJCNN 2015, which challenges participants to solve classification and regression problems without any human intervention. Participants' code is automatically run on the contest servers to train and test learning machines. However, there is no obligation to submit code; half of the prizes can be won by submitting prediction results only. Datasets of progressively increasing difficulty are introduced throughout the six rounds of the challenge. (Participants can enter the competition in any round.) The rounds alternate phases in which learners are tested on datasets participants have not seen, and phases in which participants have limited time to tweak their algorithms on those datasets to improve performance. This challenge will push the state of the art in fully automatic machine learning on a wide range of real-world problems. The platform will remain available beyond the termination of the challenge.
Isabelle Guyon, Kristin P. Bennett, Gavin C. Cawley, Hugo Jair Escalante, Sergio Escalera, Tin Kam Ho, Núria Macià, Bisakha Ray, Mehreen Saeed, Alexander R. Statnikov, Evelyne Viegas
IJCNN5
2015 Spatial codification of label predictions in multi-scale stacked sequential learning: a case study on multi-class medical volume segmentation
abstract
In this study, the authors propose the spatial codification of label predictions within the multi‐scale stacked sequential learning (MSSL) framework, a successful learning scheme to deal with non‐independent identically distributed data entries. After providing a motivation for this objective, they describe its theoretical framework based on the introduction of the blurred shape model as a smart descriptor to codify the spatial distribution of the predicted labels and define the new extended feature set for the second stacked classifier. They then particularise this scheme to be applied in volume segmentation applications. Finally, they test the implementation of the proposed framework in two medical volume segmentation datasets, obtaining significant performance improvements (with a 95% of confidence) in comparison to standard Adaboost classifier and classical MSSL approaches.
Frederic Sampedro, Sergio Escalera
IET Comput. Vis.2
2015 HuPBA8k+: Dataset and ECOC-Graph-Cut based segmentation of human limbs
Daniel Sánchez 0005, Miguel Ángel Bautista 0001, Sergio Escalera
Neurocomputing3
2015 Combining local and global learners in the pairwise multiclass classification
Mohammad Ali Bagheri, Qigang Gao, Sergio Escalera
Pattern Anal. Appl.3
2015 Generalized multi-scale stacked sequential learning for multi-class classification
Eloi Puertas, Sergio Escalera, Oriol Pujol
Pattern Anal. Appl.2
2015 Multi-part body segmentation based on depth maps for soft biometry analysis
Meysam Madadi, Sergio Escalera, Jordi Gonzàlez 0001, F. Xavier Roca, Felipe Lumbreras
Pattern Recognit. Lett.2
2015 Non-verbal communication analysis in Victim-Offender Mediations
Víctor Ponce-López, Sergio Escalera, Marc Pérez 0001, Oriol Janés, Xavier Baró
Pattern Recognit. Lett.2
2014 Contextual Rescoring for Human Pose Estimation
Antonio Hernández-Vela, Sergio Escalera, Stan Sclaroff
BMVC2
2014 Generic Subclass Ensemble: A Novel Approach to Ensemble Classification
abstract
Multiple classifier systems, also known as classifier ensembles, have received great attention in recent years because of their improved classification accuracy in different applications. In this paper, we propose a new general approach to ensemble classification, named generic subclass ensemble, in which each base classifier is trained with data belonging to a subset of classes, and thus discriminates among a subset of target categories. The ensemble classifiers are then fused using a combination rule. The proposed approach differs from existing methods that manipulate the target attribute, since in our approach individual classification problems are not restricted to two-class problems. We perform a series of experiments to evaluate the efficiency of the generic subclass approach on a set of benchmark datasets. Experimental results with multilayer perceptrons show that the proposed approach presents a viable alternative to the most commonly used ensemble classification approaches.
Mohammad Ali Bagheri, Qigang Gao, Sergio Escalera
ICPR3
2014 A Framework of Multi-classifier Fusion for Human Action Recognition
abstract
The performance of different action-recognition methods using skeleton joint locations have been recently studied by several computer vision researchers. However, the potential improvement in classification through classifier fusion by ensemble-based methods has remained unattended. In this work, we evaluate the performance of an ensemble of five action learning techniques, each performing the recognition task from a different perspective. The underlying rationale of the fusion approach is that different learners employ varying structures of input descriptors/features to be trained. These varying structures cannot be attached and used by a single learner. In addition, combining the outputs of several learners can reduce the risk of an unfortunate selection of a poorly performing learner. This leads to having a more robust and general-applicable framework. Also, we propose two simple, yet effective, action description techniques. In order to improve the recognition performance, a powerful combination strategy is utilized based on the Dempster-Shafer theory, which can effectively make use of diversity of base learners trained on different sources of information. The recognition results of the individual classifiers are compared with those obtained from fusing the classifiers' output, showing advanced performance of the proposed methodology.
Mohammad Ali Bagheri, Gang Hu 0011, Qigang Gao, Sergio Escalera
ICPR4
2014 On the design of an ECOC-Compliant Genetic Algorithm
Miguel Ángel Bautista 0001, Sergio Escalera, Xavier Baró, Oriol Pujol
Pattern Recognit.2
2014 Continuous Generalized Procrustes analysis
Laura Igual, Xavier Perez-Sala, Sergio Escalera, Cecilio Angulo, Fernando De la Torre
Pattern Recognit.3
2014 Probability-based Dynamic Time Warping and Bag-of-Visual-and-Depth-Words for Human Gesture Recognition in RGB-D
Antonio Hernández-Vela, Miguel Ángel Bautista 0001, Xavier Perez-Sala, Víctor Ponce-López, Sergio Escalera, Xavier Baró, Oriol Pujol, Cecilio Angulo
Pattern Recognit. Lett.5
2014 Iterative multi-class multi-scale stacked sequential learning: Definition and application to medical volume segmentation
Frederic Sampedro, Sergio Escalera, Anna Puig
Pattern Recognit. Lett.2
2014 Spherical Blurred Shape Model for 3-D Object and Pose Recognition: Quantitative Analysis and HCI Applications in Smart Environments
abstract
The use of depth maps is of increasing interest after the advent of cheap multisensor devices based on structured light, such as Kinect. In this context, there is a strong need of powerful 3-D shape descriptors able to generate rich object representations. Although several 3-D descriptors have been already proposed in the literature, the research of discriminative and computationally efficient descriptors is still an open issue. In this paper, we propose a novel point cloud descriptor called spherical blurred shape model (SBSM) that successfully encodes the structure density and local variabilities of an object based on shape voxel distances and a neighborhood propagation strategy. The proposed SBSM is proven to be rotation and scale invariant, robust to noise and occlusions, highly discriminative for multiple categories of complex objects like the human hand, and computationally efficient since the SBSM complexity is linear to the number of object voxels. Experimental evaluation in public depth multiclass object data, 3-D facial expressions data, and a novel hand poses data sets show significant performance improvements in relation to state-of-the-art approaches. Moreover, the effectiveness of the proposal is also proved for object spotting in 3-D scenes and for real-time automatic hand pose recognition in human computer interaction scenarios.
Oscar Lopes, Miguel Reyes, Sergio Escalera, Jordi Gonzàlez 0001
IEEE Trans. Cybern.3
2013 Multi-modal descriptors for multi-class hand pose recognition in human computer interaction systems
abstract
Hand pose recognition in advanced Human Computer Interaction systems (HCI) is becoming more feasible thanks to the use of affordable multi-modal RGB-Depth cameras. Depth data generated by these sensors is a very valuable input information, although the representation of 3D descriptors is still a critical step to obtain robust object representations. This paper presents an overview of different multi-modal descriptors, and provides a comparative study of two feature descriptors called Multi-modal Hand Shape (MHS) and Fourier-based Hand Shape (FHS), which compute local and global 2D-3D hand shape statistics to robustly describe hand poses. A new dataset of 38K hand poses has been created for real-time hand pose and gesture recognition, corresponding to five hand shape categories recorded from eight users. Experimental results show good performance of the fused MHS and FHS descriptors, improving recognition accuracy while assuring real-time computation in HCI scenarios.
Jordi Abella, Raúl Alcaide, Anna Sabaté, Joan Mas Romeu, Sergio Escalera, Jordi Gonzàlez 0001, Coen Antens
ICMI5
2013 ChaLearn multi-modal gesture recognition 2013: grand challenge and workshop summary
abstract
We organized a Grand Challenge and Workshop on Multi-Modal Gesture Recognition.
Sergio Escalera, Jordi Gonzàlez 0001, Xavier Baró, Miguel Reyes, Isabelle Guyon, Vassilis Athitsos, Hugo Jair Escalante, Leonid Sigal, Antonis A. Argyros, Cristian Sminchisescu, Richard Bowden, Stan Sclaroff
ICMI1
2013 Multi-modal gesture recognition challenge 2013: dataset and results
abstract
The recognition of continuous natural gestures is a complex and challenging problem due to the multi-modal nature of involved visual cues (e.g. fingers and lips movements, subtle facial expressions, body pose, etc.), as well as technical limitations such as spatial and temporal resolution and unreliable depth cues. In order to promote the research advance on this field, we organized a challenge on multi-modal gesture recognition. We made available a large video database of 13,858 gestures from a lexicon of 20 Italian gesture categories recorded with a Kinect™ camera, providing the audio, skeletal model, user mask, RGB and depth images. The focus of the challenge was on user independent multiple gesture learning. There are no resting positions and the gestures are performed in continuous sequences lasting 1-2 minutes, containing between 8 and 20 gesture instances in each sequence. As a result, the dataset contains around 1.720.800 frames. In addition to the 20 main gesture categories, "distracter" gestures are included, meaning that additional audio and gestures out of the vocabulary are included. The final evaluation of the challenge was defined in terms of the Levenshtein edit distance, where the goal was to indicate the real order of gestures within the sequence. 54 international teams participated in the challenge, and outstanding results were obtained by the first ranked participants.
Sergio Escalera, Jordi Gonzàlez 0001, Xavier Baró, Miguel Reyes, Oscar Lopes, Isabelle Guyon, Vassilis Athitsos, Hugo Jair Escalante
ICMI1
2013 Multi-modal social signal analysis for predicting agreement in conversation settings
abstract
In this paper we present a non-invasive ambient intelligence framework for the analysis of non-verbal communication applied to conversational settings. In particular, we apply feature extraction techniques to multi-modal audio-RGB-depth data. We compute a set of behavioral indicators that define communicative cues coming from the fields of psychology and observational methodology. We test our methodology over data captured in victim-offender mediation scenarios. Using different state-of-the-art classification approaches, our system achieve upon 75% of recognition predicting agreement among the parts involved in the conversations, using as ground truth the experts opinions.
Víctor Ponce-López, Sergio Escalera, Xavier Baró
ICMI2
2013 A Framework towards the Unification of Ensemble Classification Methods
abstract
Multiple classifier systems, also known as classifier ensembles, have received great attention in recent years because of the improved classification accuracy in different applications. A large variety of ensemble methods have been proposed in order to exploit strengths of individual classifiers. In this paper, we present a unifying framework for multiple classifier systems, which unites most classification methods by an ensemble of classifiers. Specifically, we link two research lines in machine learning: multiclass classification based on the class binarization techniques and the strategies of ensemble classification. With the proposed framework, the various ensemble classification strategies will be broadly categorized into four main approaches. Then, we provide a brief survey of ensemble methods based on these main approaches as well as principle techniques proposed to combine them.
Mohammad Ali Bagheri, Qigang Gao, Sergio Escalera
ICMLA (2)3
2013 A genetic-based subspace analysis method for improving Error-Correcting Output Coding
Mohammad Ali Bagheri, Qigang Gao, Sergio Escalera
Pattern Recognit.3
2013 Multi-modal user identification and object recognition surveillance system
Albert Clapés, Miguel Reyes, Sergio Escalera
Pattern Recognit. Lett.3
2012 Graph cuts optimization for multi-limb human segmentation in depth maps
abstract
We present a generic framework for object segmentation using depth maps based on Random Forest and Graph-cuts theory, and apply it to the segmentation of human limbs in depth maps. First, from a set of random depth features, Random Forest is used to infer a set of label probabilities for each data sample. This vector of probabilities is used as unary term in α-β swap Graph-cuts algorithm. Moreover, depth of spatio-temporal neighboring data points are used as boundary potentials. Results on a new multi-label human depth data set show high performance in terms of segmentation overlapping of the novel methodology compared to classical approaches.
Antonio Hernández-Vela, Nadezhda Zlateva, Alexander Marinov, Miguel Reyes, Petia Radeva, Dimo Dimov 0001, Sergio Escalera
CVPR7
2012 Rough Set Subspace Error-Correcting Output Codes
abstract
Among the proposed methods to deal with multi-class classification problems, the Error-Correcting Output Codes (ECOC) represents a powerful framework. The key factor in designing any ECOC matrix is the independency of the binary classifiers, without which the ECOC method would be ineffective. This paper proposes an efficient new approach to the ECOC framework in order to improve independency among classifiers. The underlying rationale for our work is that we design three-dimensional codematrix, where the third dimension is the feature space of the problem domain. Using rough set-based feature selection, a new algorithm, named "Rough Set Subspace ECOC (RSS-ECOC)" is proposed. We introduce the Quick Multiple Reduct algorithm in order to generate a set of reducts for a binary problem, where each reduct is used to train a dichotomizer. In addition to creating more independent classifiers, ECOC matrices with longer codes can be built. The numerical experiments in this study compare the classification accuracy of the proposed RSS-ECOC with classical ECOC, one-versus-one, and one-versus-all methods on 24 UCI datasets. The results show that the proposed technique increases the classification accuracy in comparison with the state of the art coding methods.
Mohammad Ali Bagheri, Qigang Gao, Sergio Escalera
ICDM3
2012 BoVDW: Bag-of-Visual-and-Depth-Words for gesture recognition
Antonio Hernández-Vela, Miguel Ángel Bautista 0001, Xavier Perez-Sala, Víctor Ponce-López, Xavier Baró, Oriol Pujol, Cecilio Angulo, Sergio Escalera
ICPR8
2012 Minimal design of error-correcting output codes
Miguel Ángel Bautista 0001, Sergio Escalera, Xavier Baró, Petia Radeva, Jordi Vitrià, Oriol Pujol
Pattern Recognit. Lett.2
2012 Accurate Coronary Centerline Extraction, Caliber Estimation, and Catheter Detection in Angiographies
abstract
Segmentation of coronary arteries in X-Ray angiography is a fundamental tool to evaluate arterial diseases and choose proper coronary treatment. The accurate segmentation of coronary arteries has become an important topic for the registration of different modalities which allows physicians rapid access to different medical imaging information from Computed Tomography (CT) scans or Magnetic Resonance Imaging (MRI). In this paper, we propose an accurate fully automatic algorithm based on Graph-cuts for vessel centerline extraction, caliber estimation, and catheter detection. Vesselness, geodesic paths, and a new multi-scale edgeness map are combined to customize the Graph-cuts approach to the segmentation of tubular structures, by means of a global optimization of the Graph-cuts energy function. Moreover, a novel supervised learning methodology that integrates local and contextual information is proposed for automatic catheter detection. We evaluate the method performance on three datasets coming from different imaging systems. The method performs as good as the expert observer w.r.t. centerline detection and caliber estimation. Moreover, the method discriminates between arteries and catheter with an accuracy of 96.5%, sensitivity of 72%, and precision of 97.4%.
Antonio Hernández-Vela, Carlo Gatta, Sergio Escalera, Laura Igual, Victoria Martin-Yuste, Manel Sabate, Petia Radeva
IEEE Trans. Inf. Technol. Biomed.3
2012 Increasing Retrieval Quality in Conversational Recommenders
abstract
A major task of research in conversational recommender systems is personalization. Critiquing is a common and powerful form of feedback, where a user can express her feature preferences by applying a series of directional critiques over the recommendations instead of providing specific preference values. Incremental Critiquing (IC) is a conversational recommender system that uses critiquing as a feedback to efficiently personalize products. The expectation is that in each cycle the system retrieves the products that best satisfy the user's soft product preferences from a minimal information input. In this paper, we present a novel technique that increases retrieval quality based on a combination of compatibility and similarity scores. Under the hypothesis that a user learns during the recommendation process, we propose two novel exponential Reinforcement Learning (RL) approaches for compatibility that take into account both the instant at which the user makes a critique and the number of satisfied critiques. Moreover, we consider that the impact of features on the similarity differs according to the preferences manifested by the user. We propose a Global Weighting (GW) approach that uses a common weight for nearest cases in order to focus on groups of relevant products. We show that our methodology significantly improves recommendation efficiency in four data sets of different sizes in terms of session length in comparison with state-of-the-art approaches. Moreover, our recommender shows higher robustness against noisy user data when compared to classical approaches.
Maria Salamó, Sergio Escalera
IEEE Trans. Knowl. Data Eng.2
2011 Human Behavior Analysis from Video Data Using Bag-of-Gestures
Víctor Ponce-López, Mario Gorga, Xavier Baró, Sergio Escalera
IJCAI4
2011 Accurate and Robust Fully-Automatic QCA: Method and Numerical Validation
Antonio Hernández-Vela, Carlo Gatta, Sergio Escalera, Laura Igual, Victoria Martin-Yuste, Petia Radeva
MICCAI (3)3
2011 Intelligent GPGPU Classification in Volume Visualization: A framework based on Error-Correcting Output Codes
abstract
Abstract In volume visualization, the definition of the regions of interest is inherently an iterative trial‐and‐error process finding out the best parameters to classify and render the final image. Generally, the user requires a lot of expertise to analyze and edit these parameters through multi‐dimensional transfer functions. In this paper, we present a framework of intelligent methods to label on‐demand multiple regions of interest. These methods can be split into a two‐level GPU‐based labelling algorithm that computes in time of rendering a set of labelled structures using the Machine Learning Error‐Correcting Output Codes (ECOC) framework. In a pre‐processing step, ECOC trains a set of Adaboost binary classifiers from a reduced pre‐labelled data set. Then, at the testing stage, each classifier is independently applied on the features of a set of unlabelled samples and combined to perform multi‐class labelling. We also propose an alternative representation of these classifiers that allows to highly parallelize the testing stage. To exploit that parallelism we implemented the testing stage in GPU‐OpenCL. The empirical results on different data sets for several volume structures shows high computational performance and classification accuracy.
Sergio Escalera, Anna Puig, Oscar Amoros, Maria Salamó
Comput. Graph. Forum1
2011 Online error correcting output codes
Sergio Escalera, David Masip, Eloi Puertas, Petia Radeva, Oriol Pujol
Pattern Recognit. Lett.1
2011 Circular Blurred Shape Model for Multiclass Symbol Recognition
abstract
In this paper, we propose a circular blurred shape model descriptor to deal with the problem of symbol detection and classification as a particular case of object recognition. The feature extraction is performed by capturing the spatial arrangement of significant object characteristics in a correlogram structure. The shape information from objects is shared among correlogram regions, where a prior blurring degree defines the level of distortion allowed in the symbol, making the descriptor tolerant to irregular deformations. Moreover, the descriptor is rotation invariant by definition. We validate the effectiveness of the proposed descriptor in both the multiclass symbol recognition and symbol detection domains. In order to perform the symbol detection, the descriptors are learned using a cascade of classifiers. In the case of multiclass categorization, the new feature space is learned using a set of binary classifiers which are embedded in an error-correcting output code design. The results over four symbol data sets show the significant improvements of the proposed descriptor compared to the state-of-the-art descriptors. In particular, the results are even more significant in those cases where the symbols suffer from elastic deformations.
Sergio Escalera, Alicia Fornés, Oriol Pujol, Josep Lladós 0001, Petia Radeva
IEEE Trans. Syst. Man Cybern. Part B1
2010 Adding Classes Online in Error Correcting Output Codes Framework
abstract
This article proposes a general extension of the Error Correcting Output Codes (ECOC) framework to the online learning scenario. As a result, the final classifier handles the addition of new classes independently of the base classifier used. Validation on UCI database and two real machine vision applications show that the online problem-dependent ECOC proposal provides a feasible and robust way for handling new classes using any base classifier.
Sergio Escalera, David Masip, Eloi Puertas, Petia Radeva, Oriol Pujol
ICPR1
2010 Symbol Classification Using Dynamic Aligned Shape Descriptor
abstract
Shape representation is a difficult task because of several symbol distortions, such as occlusions, elastic deformations, gaps or noise. In this paper, we propose a new descriptor and distance computation for coping with the problem of symbol recognition in the domain of Graphical Document Image Analysis. The proposed D-Shape descriptor encodes the arrangement information of object parts in a circular structure, allowing different levels of distortion. The classification is performed using a cyclic Dynamic Time Warping based method, allowing distortions and rotation. The methodology has been validated on different data sets, showing very high recognition rates.
Alicia Fornés, Sergio Escalera, Josep Lladós 0001, Ernest Valveny
ICPR2
2010 Error-Correcting Ouput Codes Library
Sergio Escalera, Oriol Pujol, Petia Radeva
J. Mach. Learn. Res.1
2010 Traffic sign recognition system with beta -correction
Sergio Escalera, Oriol Pujol, Petia Radeva
Mach. Vis. Appl.1
2010 On the Decoding Process in Ternary Error-Correcting Output Codes
abstract
A common way to model multiclass classification problems is to design a set of binary classifiers and to combine them. Error-Correcting Output Codes (ECOC) represent a successful framework to deal with these type of problems. Recent works in the ECOC framework showed significant performance improvements by means of new problem-dependent designs based on the ternary ECOC framework. The ternary framework contains a larger set of binary problems because of the use of a "do not care" symbol that allows us to ignore some classes by a given classifier. However, there are no proper studies that analyze the effect of the new symbol at the decoding step. In this paper, we present a taxonomy that embeds all binary and ternary ECOC decoding strategies into four groups. We show that the zero symbol introduces two kinds of biases that require redefinition of the decoding design. A new type of decoding measure is proposed, and two novel decoding strategies are defined. We evaluate the state-of-the-art coding and decoding strategies over a set of UCI Machine Learning Repository data sets and into a real traffic sign categorization problem. The experimental results show that, following the new decoding strategies, the performance of the ECOC design is significantly improved.
Sergio Escalera, Oriol Pujol, Petia Radeva
IEEE Trans. Pattern Anal. Mach. Intell.1
2010 Re-coding ECOCs without re-training
Sergio Escalera, Oriol Pujol, Petia Radeva
Pattern Recognit. Lett.1
2009 Contextual-Guided Bag-of-Visual-Words Model for Multi-class Object Categorization
Mehdi Mirza-Mohammadi, Sergio Escalera, Petia Radeva
CAIP2
2009 Quality Enhancement Based on Reinforcement Learning and Feature Weighting for a Critiquing-Based Recommender
Maria Salamó, Sergio Escalera, Petia Radeva
ICCBR2
2009 Circular Blurred Shape Model for symbol spotting in documents
abstract
Symbol spotting problem requires feature extraction strategies able to generalize from training samples and to localize the target object while discarding most part of the image. In the case of document analysis, symbol spotting techniques have to deal with a high variability of symbols' appearance. In this paper, we propose the Circular Blurred Shape Model descriptor. Feature extraction is performed capturing the spatial arrangement of significant object characteristics in a correlogram structure. Shape information from objects is shared among correlogram regions, being tolerant to the irregular deformations. Descriptors are learnt using a cascade of classifiers and Abadoost as the base classifier. Finally, symbol spotting is performed by means of a windowing strategy using the learnt cascade over plan and old musical score documents. Spotting and multi-class categorization results show better performance comparing with the state-of-the-art descriptors.
Sergio Escalera, Alicia Fornés, Oriol Pujol, Alberto Escudero, Petia Radeva
ICIP1
2009 Visual content layer for scalable object recognition in urban image databases
abstract
Rich online map interaction represents a useful tool to get multimedia information related to physical places. With this type of systems, users can automatically compute the optimal route for a trip or to look for entertainment places or hotels near their actual position. Standard maps are defined as a fusion of layers, where each one contains specific data such height, streets, or a particular business location. In this paper we propose the construction of a visual content layer which describes the visual appearance of geographic locations in a city. We captured, by means of a mobile mapping system, a huge set of georeferenced images (> 500 K) which cover the whole city of Barcelona. For each image, hundreds of region descriptions are computed off-line and described as a hash code. This allows an efficient and scalable way of accessing maps by visual content.
Xavier Baró, Sergio Escalera, Petia Radeva, Jordi Vitrià
ICME2
2009 Blurred Shape Model for binary and grey-level symbol recognition
Sergio Escalera, Alicia Fornés, Oriol Pujol, Petia Radeva, Gemma Sánchez, Josep Lladós 0001
Pattern Recognit. Lett.1
2009 Separability of ternary codes for sparse designs of error-correcting output codes
Sergio Escalera, Oriol Pujol, Petia Radeva
Pattern Recognit. Lett.1
2009 Traffic Sign Recognition Using Evolutionary Adaboost Detection and Forest-ECOC Classification
abstract
The high variability of sign appearance in uncontrolled environments has made the detection and classification of road signs a challenging problem in computer vision. In this paper, we introduce a novel approach for the detection and classification of traffic signs. Detection is based on a boosted detectors cascade, trained with a novel evolutionary version of Adaboost, which allows the use of large feature spaces. Classification is defined as a multiclass categorization problem. A battery of classifiers is trained to split classes in an Error-Correcting Output Code (ECOC) framework. We propose an ECOC design through a forest of optimal tree structures that are embedded in the ECOC matrix. The novel system offers high performance and better accuracy than the state-of-the-art strategies and is potentially better in terms of noise, affine deformation, partial occlusions, and reduced illumination.
Xavier Baró, Sergio Escalera, Jordi Vitrià, Oriol Pujol, Petia Radeva
IEEE Trans. Intell. Transp. Syst.2
2008 Separability of ternary Error-Correcting Output Codes
abstract
Error correcting output codes (ECOC) represent a successful framework to deal with multi-class categorization problems based on combining binary classifiers. In this paper, we present a new formulation of the ternary ECOC distance and the error-correcting capabilities in the ternary ECOC framework. Based on the new measure, we stress on how to design coding matrices preventing codification ambiguity and propose a new sparse random coding matrix with ternary distance maximization. The results on the UCI Repository and in a real speed traffic categorization problem show that when the coding design satisfies the new ternary measures, significant performance improvement is obtained independently of the decoding strategy applied.
Sergio Escalera, Oriol Pujol, Petia Radeva
ICPR1
2008 Error-Correcting output coding for chagasic patients characterization
abstract
The Chagas¿ disease is endemic in all Latin America, affecting millions of people in the continent. In order to diagnose and treat the Chagas¿ disease, it is important to detect and measure the coronary damage of the patient. In this paper, we analyze and categorize patients into different groups based on the coronary damage produced by the disease. Based on the features of the heart cycle extracted using high resolution ECG, a multi-class scheme of error-correcting output codes (ECOC) is formulated and successfully applied. The results show that the proposed scheme obtains significant performance improvements compared to previous works and state-of-the-art ECOC designs.
Sergio Escalera, Oriol Pujol, Petia Radeva
ICPR1
2008 Sub-class Error-Correcting Output Codes
Sergio Escalera, Oriol Pujol, Petia Radeva
ICVS1
2008 Subclass Problem-Dependent Design for Error-Correcting Output Codes
abstract
A common way to model multi-class classification problems is by means of Error-Correcting Output Codes (ECOC). Given a multi-class problem, the ECOC technique designs a code word for each class, where each position of the code identifies the membership of the class for a given binary problem. A classification decision is obtained by assigning the label of the class with the closest code. One of the main requirements of the ECOC design is that the base classifier is capable of splitting each sub-group of classes from each binary problem. However, we can not guarantee that a linear classifier model convex regions. Furthermore, non-linear classifiers also fail to manage some type of surfaces. In this paper, we present a novel strategy to model multi-class classification problems using sub-class information in the ECOC framework. Complex problems are solved by splitting the original set of classes into sub-classes, and embedding the binary problems in a problem-dependent ECOC design. Experimental results show that the proposed splitting procedure yields a better performance when the class overlap or the distribution of the training objects conceil the decision boundaries for the base classifier. The results are even more significant when one has a sufficiently large training size.
Sergio Escalera, David M. J. Tax, Oriol Pujol, Petia Radeva, Robert P. W. Duin
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 An incremental node embedding technique for error correcting output codes
Oriol Pujol, Sergio Escalera, Petia Radeva
Pattern Recognit.2
2007 Multi-class Binary Object Categorization Using Blurred Shape Models
Sergio Escalera, Alicia Fornés, Oriol Pujol, Josep Lladós 0001, Petia Radeva
CIARP1
2007 Complex Salient Regions for Computer Vision Problems
abstract
The good of interest point detectors is to find, in an unsupervised way, keypoints easy to extract and at the same time robust to image transformations. We present a novel set of saliency feathers based on image singularities that takes into account the region content in terms of intensity and local structure. The region complexity is estimated by means of the entropy of the grey-level information; shape information is obtained by measuring the entropy of significant orientations. The regions are located in their representative scale and categorized by their complexity level. Thus, the regions are highly discriminable and less sensitive to confusion and false alalarm than the traditional approaches. We compare the novel complex salient regions with the state-of-the-art keypoint detectors. The presented interest points show robustness to a wide set of image transformations and high repeatability, as well as allows matching from different camera points of view. Beside. We show the temporal robustness of the novel salient regions in real video sequences, being potentially useful for matching, image retrieval, and object categorization problems.
Sergio Escalera, Petia Radeva, Oriol Pujol
CVPR1
2007 Boosted Landmarks of Contextual Descriptors and Forest-ECOC: A novel framework to detect and classify objects in cluttered scenes
Sergio Escalera, Oriol Pujol, Petia Radeva
Pattern Recognit. Lett.1
2006 Decoding of Ternary Error Correcting Output Codes
Sergio Escalera, Oriol Pujol, Petia Radeva
CIARP1