EDBT 2026 Demo / reviewers in the wild / expert
Miguel Bordallo López
dblp:117/8168
· DBLP profile ↗
28ranked-venue papers
3as first author
22since 2021 · last 2026
0000-0002-5707-9085ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Computer networks · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multimodal Sensing-Enabled Digital Twin for 6G RAN in Indoor Environments
Shakthi Gimhana, Taufiq Ahmed, Niklas Vaara, Praneeth Susarla, Dileepa Marasinghe, Vlad-Costin Andrei, Miguel Bordallo López, Antti Pauanne, R. M. A. P. Rajatheva, Ari Pouttu |
INFOCOM | 7 |
| 2026 | A comprehensive survey on contactless vital sign monitoring using vision-based, radio-based, and fusion approaches
Zichen Li, Xiaoting Wu, Constantino Álvarez Casado, Ville Lindholm, Kristina Mikkonen, Zhaoqiang Xia, Xiaoyi Feng, Miguel Bordallo López |
Neurocomputing | 8 |
| 2026 | Alzheimer's disease classification based on multimodal consistent distribution and trusted fusion
Xiaoyan Kui, Yulan Dai, Beiji Zou 0001, Chengzhang Zhu, Yang Li 0111, Zexin Ji, Liming Chen 0002, Miguel Bordallo López |
Neural Networks | 8 |
| 2025 | Exploring Facial Kinship Verification through Contactless Heart Activity AnalysisabstractFacial Kinship Verification (FKV) aims at automatically determining whether two subjects have a kinship relation based on human faces. It has potential applications in finding missing children and social media analysis. Traditional FKV faces challenges as it is vulnerable to spoof attacks and raises privacy issues. In this paper, we explore for the first time the FKV by analyzing cardiac activity through physiological signals, with a specific focus on remote Photoplethysmography (rPPG). rPPG signals are extracted from facial videos, resulting in a one-dimensional signal that measures the changes in visible light reflection emitted to and detected from the skin caused by the heartbeat. Specifically, in this paper, we employed a straightforward one-dimensional Convolutional Neural Network (1DCNN) with a 1DCNN-Attention module and kinship contrastive loss to learn the kin similarity from rPPGs. The network takes multiple rPPG signals extracted from various facial Regions of Interest (ROIs) as inputs. Additionally, the 1DCNN attention module is designed to learn and capture the discriminative kin features from feature embeddings. Finally, we demonstrate the feasibility of rPPG to detect kinship with the experiment evaluation on the UvANEMO Smile Database from different kin relations. Xiaoting Wu, Xiaoyi Feng, Constantino Álvarez Casado, Miguel Bordallo López |
ICASSP | 5 |
| 2025 | Exponentially Weighted Instance-Aware Repeat Factor Sampling for Long-Tailed Object Detection Model Training in Unmanned Aerial Vehicles Surveillance ScenariosabstractObject detection models often struggle with class imbalance, where rare categories appear significantly less frequently than common ones. Existing sampling-based rebalancing strategies, such as Repeat Factor Sampling (RFS) and Instance-Aware Repeat Factor Sampling (IRFS), mitigate this issue by adjusting sample frequencies based on image and instance counts. However, these methods are based on linear adjustments, which limit their effectiveness in long-tailed distributions. This work introduces Exponentially Weighted Instance-Aware Repeat Factor Sampling (E-IRFS), an extension of IRFS that applies exponential scaling to better differentiate between rare and frequent classes. E-IRFS adjusts sampling probabilities using an exponential function applied to the geometric mean of image and instance frequencies, ensuring a more adaptive rebalancing strategy. We evaluate E-IRFS on a dataset derived from the Fireman-UAV-RGBT Dataset and four additional public datasets, using YOLOv11 object detection models to identify fire, smoke, people and lakes in emergency scenarios. The results show that E-IRFS improves detection performance by 22% over the baseline and outperforms RFS and IRFS, particularly for rare categories. The analysis also highlights that E-IRFS has a stronger effect on lightweight models with limited capacity, as these models rely more on data sampling strategies to address class imbalance. The findings demonstrate that E-IRFS improves rare object detection in resource-constrained environments, making it a suitable solution for real-time applications such as UAV-based emergency monitoring. The code is available at: https://github.com/futurians/E-IRFS. Taufiq Ahmed, Abhishek Kumar 0011, Constantino Álvarez Casado, Anlan Zhang, Tuomo Hänninen, Lauri Lovén, Miguel Bordallo López, Sasu Tarkoma |
IROS | 7 |
| 2025 | APML: Adaptive Probabilistic Matching Loss for Robust 3D Point Cloud ReconstructionabstractTraining deep learning models for point cloud prediction tasks such as shape completion and generation depends critically on loss functions that measure discrepancies between predicted and ground-truth point sets. Commonly used functions such as Chamfer Distance (CD), HyperCD, InfoCD and Density-aware CD rely on nearest-neighbor assignments, which often induce many-to-one correspondences, leading to point congestion in dense regions and poor coverage in sparse regions. These losses also involve non-differentiable operations due to index selection, which may affect gradient-based optimization. Earth Mover Distance (EMD) enforces one-to-one correspondences and captures structural similarity more effectively, but its cubic computational complexity limits its practical use. We propose the Adaptive Probabilistic Matching Loss (APML), a fully differentiable approximation of one-to-one matching that leverages Sinkhorn iterations on a temperature-scaled similarity matrix derived from pairwise distances. We analytically compute the temperature to guarantee a minimum assignment probability, eliminating manual tuning. APML achieves near-quadratic runtime, comparable to Chamfer-based losses, and avoids non-differentiable operations. When integrated into state-of-the-art architectures (PoinTr, PCN, FoldingNet) on ShapeNet benchmarks and on a spatio‑temporal Transformer (CSI2PC) that generates 3‑D human point clouds from WiFi‑CSI measurements, APM loss yields faster convergence, superior spatial distribution, especially in low-density regions, and improved or on-par quantitative performance without additional hyperparameter search. The code is available at: https://github.com/apm-loss/apml. Sasan Sharifipour, Constantino Álvarez Casado, Mohammad Sabokrou, Miguel Bordallo López |
NeurIPS | 4 |
| 2025 | A Survey on Sensor-Based Techniques for Continuous Stress Monitoring in Knowledge Work EnvironmentsabstractProlonged work stress has an extensive negative impact on modern society. Recently, it has become an increasing issue, specifically in cognitively demanding knowledge-intensive professions. To address the global necessity of timely detection and reduction of work stress, sensor-based automated methods for measuring stress are emerging. Physiological and behavioral sensor data enable the potential for continuous stress detection, but challenges still exist concerning the effort required from the user and the sufficiency of available information, especially for models that want to adapt to personal traits and stress perceptions. This survey paper focuses on sensor-based stress recognition enabling continuous unobtrusive stress monitoring in the knowledge work environment, with a user acceptance and load suitable for sustainable long-term adoption. We provide an overview of the theoretical background of work stress and review the recent developments of sensor-based stress assessment, emphasizing real-world studies using physiological, behavioral, and environmental data. In addition, we discuss the applicability and challenges of different monitoring methods, including user acceptance. The presented survey provides insights into automating the assessment of work stress and related factors to advance the development of personalized well-being solutions based on pervasive data. Johanna Kallio, Elena Vildjiounaite, Jaakko Tervonen, Miguel Bordallo López |
ACM Trans. Comput. Heal. | 4 |
| 2024 | Strong Multimodal Representation Learner through Cross-domain Distillation for Alzheimer's Disease ClassificationabstractVision-language foundational models have achieved commendable results on related tasks. However, their application to medical tasks is still limited due to issues arising from data biases. Currently, leveraging existing foundational models to improve medical tasks remains a challenge. To this end, this paper proposes a strong multimodal representation learning method based on cross-domain distillation handling structural Magnetic Resonance Imaging (sMRI), Positron Emission Computed Tomograph (PET) images, and mini-mental state examination (MMSE) score for Alzheimer’s disease (AD) classification. Specifically, we establish a text-to-image cross-domain distillation learning framework, enabling a text encoder pre-trained on general visual recognition tasks to guide the training of sMRI and PET image feature extractors. Simultaneously, positional encoding is used to extract the magnitude features of MMSE scores. Based on the multimodal representations extracted from sMRI, PET images, and MMSE scores, we perform a self-attention operation equipped with a gating mechanism for multimodal feature fusion. This mechanism controls the contribution of each modality representation to the classification decision, dynamically strengthening or weakening specific modality representations and helping construct stronger fused features for AD classification. Our method undergoes 5-fold cross-validation on the widely used ADNI dataset, and comparative experimental results demonstrate that our method achieves advanced performance in two AD-related binary classification tasks. Yulan Dai, Beiji Zou 0001, Xiaoyan Kui, Qinsong Li, Wei Zhao 0040, Jun Liu 0075, Miguel Bordallo López |
BIBM | 7 |
| 2024 | FireMan-UAV-RGBT: A Novel UAV-Based RGB-Thermal Video Dataset for the Detection of Wildfires in the Finnish ForestsabstractWildfire detection in the densely forested and remote regions of Finland presents substantial challenges. This paper introduces a new publicly available dataset, FireMan-UAV-RGBT, comprising UAV-captured RGB and thermal video data to advance wildfire detection methodologies. The dataset includes high-resolution images of boreal forests that have been carefully annotated both manually and using a semi-automatic method that leverages thermal information for improved RGB image segmentation. The utility of the dataset is assessed by applying established deep learning models (ResN et50 and YOLOv8), and comparing their performance in unimodal and multimodal detection approaches. The performance is evaluated using both intra-set validation on the novel dataset and inter-set evaluation through cross-validation with the Flame-1 and Flame-2 datasets, demonstrating the usability of our dataset in wildfire detection scenarios. The FireMan-UAV-RGBT dataset represents a step forward in wildfire management, offering a resource that may contribute to cost-effective and environmentally sensitive solutions in remote sensing and emergency response strategies. S. D. M. W. Kularatne, Constantino Álvarez Casado, Janne Rajala, Tuomo Hänninen, Miguel Bordallo López, Le Ngu Nguyen |
ETFA | 5 |
| 2024 | Estimating Exercise-Induced Fatigue from Thermal Facial ImagesabstractExercise-induced fatigue resulting from physical activity can be an early indicator of overtraining, illness, or other health issues. In this paper, we present an automated method for estimating exercise-induced fatigue levels through the use of thermal imaging and facial analysis techniques utilizing deep learning models. Leveraging a novel dataset comprising over 400,000 thermal facial images of rested and fatigued users, our results suggest that exercise-induced fatigue levels could be predicted with only one static thermal frame with an average error smaller than 15%. The results emphasize the viability of using thermal imaging in conjunction with deep learning for reliable exercise-induced fatigue estimation. Manuel Lage Cañellas, Constantino Álvarez Casado, Le Nguyen, Miguel Bordallo López |
ICASSP | 4 |
| 2024 | Polynomial Solvers for mmWave Radio BeamformingabstractMillimeter (mmWave) beamforming is an integral component of fifth-generation (5G) and beyond radio commu-nications. 5G beamforming involves the initial beam selection procedure using a codebook with multiple radio beam directions. Conventional codebook-based alignment schemes involve exhaustive sweeping over the predefined beam directions, the number of which increases significantly with large numbers of antennas resulting in undesirable latency and communications signal overhead. In this paper, we propose a novel algebraic-based codebook using Gröbner basis polynomial solvers to reduce the signal overhead during beam alignment. We also analyze the complexity-performance tradeoff between the proposed algebraic-based codebook and the exhaustive-based beam alignment across different monomial thresholds, multiple antenna configurations and radio contextual location information. Our results show that the proposed approach reduces the beam-search overhead at an average complexity reduction ratio of 73.95% with a performance tradeoff error of 32.25%. Praneeth Susarla, Snehal Bhayani, S. S. Krishna Chaitanya Bulusu, Miguel Bordallo López, Janne Heikkilä, Markku Juntti, Olli Silvén |
ICC | 4 |
| 2024 | Facial expression analysis using Decomposed Multiscale Spatiotemporal Networks
Wheidima C. Melo, Eric Granger, Miguel Bordallo López |
Expert Syst. Appl. | 3 |
| 2024 | Audio-Visual Kinship Verification: A New Dataset and a Unified Adaptive Adversarial Multimodal Learning ApproachabstractFacial kinship verification refers to automatically determining whether two people have a kin relation from their faces. It has become a popular research topic due to potential practical applications. Over the past decade, many efforts have been devoted to improving the verification performance from human faces only while lacking other biometric information, for example, speaking voice. In this article, to interpret and benefit from multiple modalities, we propose for the first time to combine human faces and voices to verify kinship, which we refer it as the audio-visual kinship verification study. We first establish a comprehensive audio-visual kinship dataset that consists of familial talking facial videos under various scenarios, called TALKIN-Family. Based on the dataset, we present the extensive evaluation of kinship verification from faces and voices. In particular, we propose a deep-learning-based fusion method, called unified adaptive adversarial multimodal learning (UAAML). It consists of the adversarial network and the attention module on the basis of unified multimodal features. Experiments show that audio (voice) information is complementary to facial features and useful for the kinship verification problem. Furthermore, the proposed fusion method outperforms baseline methods. In addition, we also evaluate the human verification ability on a subset of TALKIN-Family. It indicates that humans have higher accuracy when they have access to both faces and voices. The machine-learning methods could effectively and efficiently outperform the human ability. Finally, we include the future work and research opportunities with the TALKIN-Family dataset. Xiaoting Wu, Xueyi Zhang 0001, Xiaoyi Feng, Miguel Bordallo López, Li Liu 0002 |
IEEE Trans. Cybern. | 4 |
| 2023 | Non-Contact Heart Rate Measurement from Deteriorated VideosabstractRemote photoplethysmography (rPPG) offers a state-of-the-art, non-contact methodology for estimating human pulse by analyzing facial videos. Despite its potential, rPPG methods can be susceptible to various artifacts, such as noise, occlusions, and other obstructions caused by sunglasses, masks, or even involuntary face touching. In this study, we apply image processing transformations to intentionally degrade video quality, mimicking these challenging conditions, and subsequently evaluate the performance of both non-learning and learning-based rPPG methods on the deteriorated data. Our results reveal a significant decrease in accuracy in the presence of these artifacts, prompting us to propose the application of restoration techniques, such as denoising and inpainting, to improve heart-rate estimation outcomes. By addressing these challenging conditions and occlusion artifacts, our approach aims to make rPPG methods more robust and adaptable to real-world situations. To assess the effectiveness of our proposed methods, we undertake comprehensive experiments on three publicly available datasets, encompassing a wide range of scenarios and artifact types. Our findings underscore the potential to construct a robust rPPG system by employing an optimal combination of restoration algorithms and rPPG techniques. Moreover, our study contributes to the advancement of privacy-conscious rPPG methodologies, thereby bolstering the overall utility and impact of this innovative technology in the field of remote heart-rate estimation under realistic and diverse conditions. Nhi Nguyen, Le Ngu Nguyen, Constantino Álvarez Casado, Olli Silvén, Miguel Bordallo López |
ETFA | 5 |
| 2023 | Semantic Slicing across the Distributed Intelligent 6G Wireless NetworksabstractIn the age of the Internet of Things (IoT) and the expanding computing continuum, it’s crucial to manage and share resources at the edges of networks. This position paper presents a new concept known as ’semantic slicing’. This approach harnesses the power of artificial intelligence (AI), wireless networks, edge computing, and sensing technologies to enable novel applications, optimize resource allocation, and streamline data processing and decision-making across complex systems spanning the computing continuum. Semantic slicing applies a deep understanding of the data and specific application requirements to intelligently allocate resources and distribute processing tasks in the computing continuum. This strategy allows for the creation of systems that are not only more efficient and responsive, but also better equipped to adapt to a variety of applications and services. Lauri Lovén, Hafiz Faheem Shahid, Le Ngu Nguyen, Erkki Harjula, Olli Silvén, Susanna Pirttikangas, Miguel Bordallo López |
SECON | 7 |
| 2023 | Depression Recognition Using Remote Photoplethysmography From Facial VideosabstractDepression is a mental illness that may be harmful to an individual's health. The detection of mental health disorders in the early stages and a precise diagnosis are critical to avoid social, physiological, or psychological side effects. This work analyzes physiological signals to observe if different depressive states have a noticeable impact on the blood volume pulse (BVP) and the heart rate variability (HRV) response. Although typically, HRV features are calculated from biosignals obtained with contact-based sensors such as wearables, we propose instead a novel scheme that directly extracts them from facial videos, just based on visual information, removing the need for any contact-based device. Our solution is based on a pipeline that is able to extract complete remote photoplethysmography signals (rPPG) in a fully unsupervised manner. We use these rPPG signals to calculate over 60 statistical, geometrical, and physiological features that are further used to train several machine learning regressors to recognize different levels of depression. Experiments on two benchmark datasets indicate that this approach offers comparable results to other audiovisual modalities based on voice or facial expression, potentially complementing them. In addition, the results achieved for the proposed method show promising and solid performance that outperforms hand-engineered methods and is comparable to deep learning-based approaches. Constantino Álvarez Casado, Manuel Lage Cañellas, Miguel Bordallo López |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | MDN: A Deep Maximization-Differentiation Network for Spatio-Temporal Depression DetectionabstractDeep learning (DL) models have been successfully applied in video-based affective computing, allowing, for instance, to recognize emotions and mood, or to estimate the intensity of pain or stress of individuals based on their facial expressions. Despite the recent advances with state-of-the-art DL models for spatio-temporal recognition of facial expressions associated with depressive behaviour, some key challenges remain in the cost-effective application of 3D-CNNs: (1) 3D convolutions usually employ structures with fixed temporal depth that decreases the potential to extract discriminative representations due to the usually small difference of spatio-temporal variations along different depression levels; and (2) the computational complexity of these models with consequent susceptibility to overfitting. To address these challenges, we propose a novel DL architecture called the Maximization and Differentiation Network (MDN) in order to effectively represent facial expression variations that are relevant for depression assessment. The MDN, operating without 3D convolutions, explores multiscale temporal information using a maximization block that captures smooth facial variations and a difference block that encodes sudden facial variations. Extensive experiments using our proposed MDN with models with 100 and 152 layers result in improved performance while reducing the number of parameters by more than$3\times$when compared with 3D ResNet models. Our model also outperforms other 3D models and achieves state-of-the-art results for depression detection. Code available at:https://github.com/wheidima/MDN. Wheidima C. Melo, Eric Granger, Miguel Bordallo López |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | Face2PPG: An Unsupervised Pipeline for Blood Volume Pulse Extraction From FacesabstractPhotoplethysmography (PPG) signals have become a key technology in many fields, such as medicine, well-being, or sports. Our work proposes a set of pipelines to extract remote PPG signals (rPPG) from the face robustly, reliably, and configurably. We identify and evaluate the possible choices in the critical steps of unsupervised rPPG methodologies. We assess a state-of-the-art processing pipeline in six different datasets, incorporating important corrections in the methodology that ensure reproducible and fair comparisons. In addition, we extend the pipeline by proposing three novel ideas; 1) a new method to stabilize the detected face based on a rigid mesh normalization; 2) a new method to dynamically select the different regions in the face that provide the best raw signals, and 3) a new RGB to rPPG transformation method, called Orthogonal Matrix Image Transformation (OMIT) based on QR decomposition, that increases robustness against compression artifacts. We show that all three changes introduce noticeable improvements in retrieving rPPG signals from faces, obtaining state-of-the-art results compared with unsupervised, non-learning-based methodologies and, in some databases, very close to supervised, learning-based methods. We perform a comparative study to quantify the contribution of each proposed idea. In addition, we depict a series of observations that could help in future implementations. Constantino Álvarez Casado, Miguel Bordallo López |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Video2IMU: Realistic IMU features and signals from videosabstractHuman Activity Recognition (HAR) from wearable sensor data identifies movements or activities in unconstrained environments. HAR is a challenging problem as it presents great variability across subjects. Obtaining large amounts of labelled data is not straightforward, since wearable sensor signals are not easy to label upon simple human inspection. In our work, we propose the use of neural networks for the generation of realistic signals and features using human activity monocular videos. We show how these generated features and signals can be utilized, instead of their real counterparts, to train HAR models that can recognize activities using signals obtained with wearable sensors. To prove the validity of our methods, we perform experiments on an activity recognition dataset created for the improvement of industrial work safety. We show that our model is able to realistically generate virtual sensor signals and features usable to train a HAR classifier with comparable performance as the one trained using real sensor data. Our results enable the use of available, labeled video data for training HAR models to classify signals from wearable sensors. Arttu Lämsä, Jaakko Tervonen, Jussi Liikka, Constantino Álvarez Casado, Miguel Bordallo López |
BSN | 5 |
| 2022 | Identification, Activity, and Biometric Classification using Radar-based SensingabstractWe explore the possibility of leveraging radar-based sensing systems to analyze vital signs for classification, user identification, and regression tasks. Specifically, we extract time-domain and frequency-domain features from distance, respiration, and pulse signals obtained by filtering radio-frequency signals. Our Random Forest classification models are trained on these features to recognize scenarios in which the radar data were collected, categorize individuals into age groups, and classify human activities. For classification, we achieved up to 94.7% of accuracy when distinguishing apnea and normal breathing in the lying position. We then show the feasibility of identifying individuals in a small group using vital signs, which can support model fine-tuning with data acquired from new users. Furthermore, we used a Random Forest regression model to estimate the Body Mass Index, height, and weight of subjects. These classification, identification, and regression models benefit smart systems that can simultaneously identify users, recognize their behaviours, and extract their vital signs from radar sensors. Le Ngu Nguyen, Constantino Álvarez Casado, Olli Silvén, Miguel Bordallo López |
ETFA | 4 |
| 2022 | Facial Kinship Verification: A Comprehensive Review and OutlookabstractThe goal of Facial Kinship Verification (FKV) is to automatically determine whether two individuals have a kin relationship or not from their given facial images or videos. It is an emerging and challenging problem that has attracted increasing attention due to its practical applications. Over the past decade, significant progress has been achieved in this new field. Handcrafted features and deep learning techniques have been widely studied in FKV. The goal of this paper is to conduct a comprehensive review of the problem of FKV. We cover different aspects of the research, including problem definition, challenges, applications, benchmark datasets, a taxonomy of existing methods, and state-of-the-art performance. In retrospect of what has been achieved so far, we identify gaps in current research and discuss potential future research directions. Xiaoting Wu, Xiaoyi Feng, Xiaochun Cao, Xin Xu 0001, Dewen Hu, Miguel Bordallo López, Li Liu 0002 |
Int. J. Comput. Vis. | 6 |
| 2022 | Crowdsourcing sensitive data using public displays - opportunities, challenges, and considerationsabstractAbstract Interactive public displays are versatile two-way interfaces between the digital world and passersby. They can convey information and harvest purposeful data from their users. Surprisingly little work has exploited public displays for collecting tagged data that might be useful beyond a single application. In this work, we set to fill this gap and present two studies: (1) a field study where we investigated collecting biometrically tagged video-selfies using public kiosk-sized screens, and (2) an online narrative transportation study that further elicited rich qualitative insights on key emerging aspects from the first study. In the first study, a 61-day deployment resulted in 199 video-selfies with consent to leverage the videos in any non-profit research. The field study indicates that people are willing to donate even highly sensitive data about themselves in public. The subsequent online narrative transportation study provides a deeper understanding of a variety of issues arising from the first study that can be leveraged in the future design of such systems. The two studies combined in this article pave the way forward towards a vision where volunteers can, should they so choose, ethically and serendipitously help unleash advances in data-driven areas such as computer vision and machine learning in health care. Andy Alorwu, Niels van Berkel, Jorge Gonçalves 0001, Jonas Oppenlaender, Miguel Bordallo López, Mahalakshmy Seetharaman, Simo Hosio |
Pers. Ubiquitous Comput. | 5 |
| 2020 | Encoding Temporal Information For Automatic Depression Recognition From Facial AnalysisabstractDepression is a mental illness that may be harmful to an individual's health. Using deep learning models to recognize the facial expressions of individuals captured in videos has shown promising results for automatic depression detection. Typically, depression levels are recognized using 2D-Convolutional Neural Networks (CNNs) that are trained to extract static features from video frames, which impairs the capture of dynamic spatio-temporal relations. As an alternative, 3D-CNNs may be employed to extract spatiotemporal features from short video clips, although the risk of overfitting increases due to the limited availability of labeled depression video data. To address these issues, we propose a novel temporal pooling method to capture and encode the spatio-temporal dynamic of video clips into an image map. This approach allows fine-tuning a pre-trained 2D CNN to model facial variations, and thereby improving the training process and model accuracy. Our proposed method is based on two-stream model that performs late fusion of appearance and dynamic information. Extensive experiments on two benchmark AVEC datasets indicate that the proposed method is efficient and outperforms the state-of-the-art schemes. Wheidima C. Melo, Eric Granger, Miguel Bordallo López |
ICASSP | 3 |
| 2020 | Self-supervised pain intensity estimation from facial videos via statistical spatiotemporal distillation
Mohammad Tavakolian, Miguel Bordallo López, Li Liu 0002 |
Pattern Recognit. Lett. | 2 |
| 2018 | Kinship verification from facial images and videos: human versus machine
Miguel Bordallo López, Abdenour Hadid, Elhocine Boutellaa, Jorge Gonçalves 0001, Vassilis Kostakos, Simo Hosio |
Mach. Vis. Appl. | 1 |
| 2018 | A Survey on Computer Vision for Assistive Medical Diagnosis From FacesabstractAutomatic medical diagnosis is an emerging center of interest in computer vision as it provides unobtrusive objective information on a patient's condition. The face, as a mirror of health status, can reveal symptomatic indications of specific diseases. Thus, the detection of facial abnormalities or atypical features is at upmost importance when it comes to medical diagnostics. This survey aims to give an overview of the recent developments in medical diagnostics from facial images based on computer vision methods. Various approaches have been considered to assess facial symptoms and to eventually provide further help to the practitioners. However, the developed tools are still seldom used in clinical practice, since their reliability is still a concern due to the lack of clinical validation of the methodologies and their inadequate applicability. Nonetheless, efforts are being made to provide robust solutions suitable for healthcare environments, by dealing with practical issues such as real-time assessment or patients positioning. This survey provides an updated collection of the most relevant and innovative solutions in facial images analysis. The findings show that with the help of computer vision methods, over 30 medical conditions can be preliminarily diagnosed from the automatic detection of some of their symptoms. Furthermore, future perspectives, such as the need for interdisciplinary collaboration and collecting publicly available databases, are highlighted. Jérôme Thevenot, Miguel Bordallo López, Abdenour Hadid |
IEEE J. Biomed. Health Informatics | 2 |
| 2016 | Comments on the "Kinship Face in the Wild" Data SetsabstractThe Kinship Face in the Wild data sets, recently published in TPAMI, are currently used as a benchmark for the evaluation of kinship verification algorithms. We recommend that these data sets are no longer used in kinship verification research unless there is a compelling reason that takes into account the nature of the images. We note that most of the image kinship pairs are cropped from the same photographs. Exploiting this cropping information, competitive but biased performance can be obtained using a simple scoring approach, taking only into account the nature of the image pairs rather than any features about kin information. To illustrate our motives, we provide classification results utilizing a simple scoring method based on the image similarity of both images of a kinship pair. Using simply the distance of the chrominance averages of the images in the Lab color space without any training or using any specific kin features, we achieve performance comparable to state-of-the-art methods. We provide the source code to prove the validity of our claims and ensure the repeatability of our experiments. Miguel Bordallo López, Elhocine Boutellaa, Abdenour Hadid |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Interactive multi-frame reconstruction for mobile devices
Miguel Bordallo López, Jari Hannuksela, Olli Silvén, Markku Vehviläinen |
Multim. Tools Appl. | 1 |