Federico Sukno

dblp:49/4830 · also Federico M. Sukno · DBLP profile ↗
← Back
42ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-2029-1576ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5Security and privacy · 2 · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Unlocking 3D baby face photogrammetry: Multi-view BabyMorph reconstruction from uncalibrated photographs
abstract
• First multi-view infant 3D face reconstruction algorithm from triplets images (frontal, left and right) • Transformer-based multi-view 2D-3D reconstruction network • Geometric deep learning autoencoder replaces PCA-based models for richer, nonlinear facial shape representation • Cost-effective and simple approach, suitable for limited-resource settings Craniofacial anomalies are important diagnostic markers in early life. Recent studies emphasize the value of 3D imaging for extracting robust facial features that are potential indicators of disease. However, widespread availability of 3D scanning devices in hospitals remains a challenge and this type of technology may not be available in limited-resource settings. For this reason, we present a new approach to generate precise baby 3D face reconstructions from multiple uncalibrated 2D photographs acquired with a smartphone camera. The novel multi-view transformer network presented takes as input three uncalibrated photographs of the baby in frontal, left, and right pose. It then maps these images to a previously learned latent space that captures the baby’s 3D facial morphology, using a 2D vision transformer encoder. Subsequently, the estimated 3D geometry is recovered by decoding the latent vector using a 3D graph convolutional network decoder. Our network demonstrates a normalized mean error of 3.29% and a root mean square error of 2.62 mm between reconstructed and true 3D faces in the baby test dataset. These outcomes are comparable with those reported by both single-view and multi-view 2D-3D reconstruction state-of-the-art errors in adult models. To conclude, the presented Multi-view BabyMorph generates highly accurate 3D baby facial reconstructions from uncalibrated photographs, simplifying the process of 3D photogrammetry. This innovation expands access to advanced baby diagnosis tools, particularly in resource-limited settings.
Antònia Alomar, Gemma Piella, Esperanza Mantilla-Rivas, Austin Tapp, Antonio R. Porras, Ricardo Rubio, Silvia Maya-Enero, Federico Sukno, Marius George Linguraru
Expert Syst. Appl.8
2026 Deep pulse-signal magnification for remote heart rate estimation in compressed videos
abstract
• A deep learning framework to reduce video compression effects on rPPG signals. • Selectively enhances rPPG-relevant info, unlike standard video quality methods. • Combines an rPPG Estimator and a magnifier for pulse signals in compressed video. • Validated on four public datasets, showing strong robustness and generalization. Recent advancements in data-driven approaches for remote photoplethysmography (rPPG) have significantly improved the accuracy of remote heart rate estimation. However, the performance of such approaches worsens considerably under video compression, which is nevertheless necessary to store and transmit video data efficiently. In this paper, we investigate the impact of video compression on the recovery of physiological signals from camera-based recordings. To mitigate the negative effects of compression on rPPG estimation, we propose a novel two-stage training strategy. This approach incorporates a pulse-signal magnification transformation, which adapts compressed video data into an uncompressed domain, where the rPPG signal is amplified. We validate the effectiveness of our model through comprehensive evaluations on two publicly available datasets, UCLA-rPPG and UBFC-rPPG, assessing both intra- and cross-database performance across various compression rates. Additionally, we assess the robustness of our approach on two additional highly compressed and widely-used datasets, MAHNOB-HCI and COHFACE, which reveal outstanding heart rate estimation results.
Joaquim Comas, Adria Ruiz, Federico Sukno
Expert Syst. Appl.3
2026 Deep learnable spectral decomposition of 3D baby faces
abstract
In this paper, we introduce a novel, deep 3D morphable model for meshes with common triangulation. Specifically, we apply it to reconstruct baby faces. The proposed algorithm is simple, adaptable, and specifically targeted to perform well on small datasets. We combine Graph-Laplacian based spectral decomposition with a learnable, transformer-like component. The decomposition matrices are applied as skip-connections, providing our architecture with a prior that encodes both local and global information of the underlying mesh structure. The learnable component does not make any domain-specific assumptions and can override the prior, if necessary. This flexibility also allows our model to perform well on larger datasets. We further modify the decomposition matrices to create deeper versions of this architecture and introduce a data augmentation strategy: flipping and rotations are applied to the deviations from the mean, rather than directly to the samples. In our experiments, we compare the reconstruction error of the proposed architecture against the state of the art, examine the effect of data augmentation across a small baby face dataset and a larger adult dataset and inspect our model’s capabilities to generate new samples from the encoded distribution. We show that our method outperforms current baby face models, as well as state of the art 3D morphable models, especially on the raw data. Additionally, we demonstrate that the proposed data augmentation substantially improves existing models.
Michael Zappe, Antònia Alomar, Marius George Linguraru, Gemma Piella, Federico Sukno
Pattern Recognit.5
2025 NLML-HPE: Head Pose Estimation with Limited Data via Manifold Learning
abstract
Head pose estimation (HPE) plays a critical role in various computer vision applications such as human-computer interaction and facial recognition. In this paper, we propose a novel deep learning approach for head pose estimation with limited training data via non-linear manifold learning called NLML-HPE. This method is based on the combination of tensor decomposition (i.e., Tucker decomposition) and feed forward neural networks. Unlike traditional classification-based approaches, our method formulates head pose estimation as a regression problem, mapping input landmarks into a continuous representation of pose angles. To this end, our method uses tensor decomposition to split each Euler angle (yaw, pitch, roll) to separate subspaces and models each dimension of the underlying manifold as a cosine curve. We address two key challenges: 1. Almost all HPE datasets suffer from incorrect and inaccurate pose annotations. Hence, we generated a precise and consistent 2D head pose dataset for our training set by rotating 3D head models for a fixed set of poses and rendering the corresponding 2D images. 2. We achieved real-time performance with limited training data as our method accurately captures the nature of rotation of an object from facial landmarks. Once the underlying manifold for rotation around each axis is learned, the model is very fast in predicting unseen data. Our training and testing code is available online along with our trained models: https: //github.com/MahdiGhafoorian/NLML_HPE.
Mahdi Ghafourian, Federico Sukno
IJCB2
2024 PhysFlow: Skin tone transfer for remote heart rate estimation through conditional normalizing flows
Joaquim Comas, Antònia Alomar, Adria Ruiz, Federico Sukno
BMVC4
2024 Deep adaptative spectral zoom for improved remote heart rate estimation
abstract
Recent advances in remote heart rate measure-ment, motivated by data-driven approaches, have notably enhanced accuracy. However, these improvements primarily focus on recovering the rPPG signal, overlooking the implicit challenges of estimating the heart rate (HR) from the derived signal. While many methods employ the Fast Fourier Transform (FFT) for HR estimation, the performance of the FFT is inher-ently affected by a limited frequency resolution. In contrast, the Chirp-Z Transform (CZT), a generalization form of FFT, can refine the spectrum to the narrow-band range of interest for heart rate, providing improved frequential resolution and, consequently, more accurate estimation. This paper presents the advantages of employing the CZT for remote HR estimation and introduces a novel data-driven adaptive CZT estimator. The objective of our proposed model is to tailor the CZT to match the characteristics of each specific dataset sensor, facilitating a more optimal and accurate estimation of HR from the rPPG signal without compromising generalization across diverse datasets. This is achieved through a Sparse Matrix Optimization (SMO). We validate the effectiveness of our model through exhaustive evaluations on three publicly available datasets -UCLA-rPPG, PURE, and UBFC-rPPG-employing both intra- and cross-database performance metrics. The results reveal outstanding heart rate estimation capabilities, establishing the proposed approach as a robust and versatile estimator for any rPPG method.
Joaquim Comas, Adria Ruiz, Federico Sukno
FG3
2023 BabyNet: Reconstructing 3D faces of babies from uncalibrated photographs
abstract
We present a 3D face reconstruction system that aims at recovering the 3D facial geometry of babies from uncalibrated photographs, BabyNet. Since the 3D facial geometry of babies differs substantially from that of adults, baby-specific facial reconstruction systems are needed. BabyNet consists of two stages: 1) a 3D graph convolutional autoencoder learns a latent space of the baby 3D facial shape; and 2) a 2D encoder that maps photographs to the 3D latent space based on representative features extracted using transfer learning. In this way, using the pre-trained 3D decoder, we can recover a 3D face from 2D images. We evaluate BabyNet and show that 1) methods based on adult datasets cannot model the 3D facial geometry of babies, which proves the need for a baby-specific method, and 2) BabyNet outperforms classical model-fitting methods even when a baby-specific 3D morphable model, such as BabyFM, is used.
Araceli Morales, Antònia Alomar, Antonio R. Porras, Marius George Linguraru, Gemma Piella, Federico Sukno
Pattern Recognit.6
2023 Audio-Visual Gated-Sequenced Neural Networks for Affect Recognition
abstract
The interest in automatic emotion recognition and the larger field of Affective Computing has recently gained momentum. The current emergence of large, video-based affect datasets offering rich multi-modal inputs facilitates the development of deep learning-based models for automatic affect analysis that currently holds the state of the art. However, recent approaches to process these modalities cannot fully exploit them due to the use of oversimplified fusion schemes. Furthermore, the efficient use of temporal information inherent to these huge data are also largely unexplored hindering their potential progress. In this work, we propose a multi-modal, sequence-based neural network with gating mechanisms for Valence and Arousal based affect recognition. Our model consists of three major networks: Firstly, a latent-feature generator that extracts compact representations from both modalities that have been artificially degraded to add robustness. Secondly, a multi-task discriminator that estimates both input identity and a first step emotion quadrant estimation. Thirdly, a sequence-based predictor with attention and gating mechanisms that effectively merges both modalities and uses this information through sequence modelling. In our experiments on the SEMAINE and SEWA affect datasets, we observe the impact of both proposed methods with progressive increase in accuracy. We further show in our ablation studies how the internal attention weight and gating coefficient impact our models’ estimates quality. Finally, we demonstrate state of the art accuracy through comparisons with current alternatives on both datasets.
Decky Aspandi, Federico Sukno, Björn W. Schuller, Xavier Binefa
IEEE Trans. Affect. Comput.2
2022 End-to-End Lip-Reading Without Large-Scale Data
abstract
The development of Automatic Lip-Reading (ALR) systems for continuous speech recognition has so far limited their applicability to English since this is the only language with large-scale datasets sufficient to train end-to-end ALR systems. In this work, we show that it is possible to train competitive end-to-end ALR systems in alternative languages with challenging small-scale data as long as the appropriate restrictions are made to the learning process of the visual front-end objective. To this end, we hypothesize that the visual front-end should be trained in a self-supervised setting, allowing it to target its ownvisual units. We specifically definevisual unitsas a collection of visually similar images constrained by linguistics and provide an algorithmic implementation to automatically generate them. We show thatvisual unitscan be used to add an intermediate classification task between the visual and temporal modules that facilitates meaningful learning of visual features and, as a consequence, reduces the amount of data required to train an end-to-end ALR system. Additionally, we present a data augmentation strategy for enriching the temporal context. We synthesize realistic video sequences by appropriately combining characters-like sub-sequences from existing videos. We test the proposed ALR system on i) the VLRF dataset, a small-scale database that is one of the largest in Spanish, and achieve 44.77$\%$CER and 72.90% WER, which are competitive with the state-of-the-art and significant for this volume of training material; ii) the TCD-TIMIT dataset, a comparable medium-scale database in English, where we achieve 36.58% CER and 56.29% WER, which are also state-of-the-art results on speaker-dependent experiments.
Adriana Fernandez-Lopez, Federico Sukno
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Composite recurrent network with internal denoising for facial alignment in still and video images in the wild
abstract
Facial alignment is an essential task for many higher level facial analysis applications, such as animation, human activity recognition and human - computer interaction. Although the recent availability of big datasets and powerful deep-learning approaches have enabled major improvements on the state of the art accuracy, the performance of current approaches can severely deteriorate when dealing with images in highly unconstrained conditions, which limits the real-life applicability of such models. In this paper, we propose a composite recurrent tracker with internal denoising that jointly address both single image facial alignment and deformable facial tracking in the wild. Specifically, we incorporate multilayer LSTMs to model temporal dependencies with variable length and introduce an internal denoiser which selectively enhances the input images to improve the robustness of our overall model. We achieve this by combining 4 different sub-networks that specialize in each of the key tasks that are required, namely face detection, bounding-box tracking, facial region validation and facial alignment with internal denoising. These blocks are endowed with novel algorithms resulting in a facial tracker that is both accurate, robust to in-the-wild settings and resilient against drifting. We demonstrate this by testing our model on 300-W and Menpo datasets for single image facial alignment, and 300-VW dataset for deformable facial tracking. Comparison against 20 other state of the art methods demonstrates the excellent performance of the proposed approach.
Decky Aspandi, Oriol Martínez, Federico Sukno, Xavier Binefa
Image Vis. Comput.3
2020 Spectral Correspondence Framework for Building a 3D Baby Face Model
abstract
Early detection of facial dysmorphology - variations of the normal facial geometry - is essential for the timely detection of genetic conditions, which has a significant impact in the reduction of the mortality and morbidity associated with them. A model encoding the normal variability in the healthy population can serve as a reference to quantify the often subtle facial abnormalities that are present in young patients with such conditions. In this paper, we present the first facial model constructed exclusively from newborn data, the Baby Face Model (BabyFM). Our model is built from 3D scans with an innovative pipeline based on least squared conformal maps (LSCM). LSCM are piece-wise linear mappings that project the training faces to a common 2D space minimising the conformal distortion. This process allows improving the correspondences between 3D faces, which is particularly important for the identification of subtle dysmorphology. We evaluate the ability of our BabyFM to recover the babys facial morphology from a set of 2D images by comparing it to state-of-the-art facial models. We also compare it to models built following an analogous pipeline to the one proposed in this paper but using nonrigid iterative closest point (NICP) to establish dense correspondences between the training faces. The results show that our model reconstructs the facial morphology of babies with significantly smaller errors than the state-of-the-art models (p = 10-4) and the “NICP models” (p <; 0.01).
Araceli Morales, Antonio R. Porras, Liyun Tu, Marius George Linguraru, Gemma Piella, Federico Sukno
FG6
2020 Cogans For Unsupervised Visual Speech Adaptation To New Speakers
abstract
Audio-Visual Speech Recognition (AVSR) faces the difficult task of exploiting acoustic and visual cues simultaneously. Augmenting speech with the visual channel creates its own challenges, e.g. every person has unique mouth movements, making the generalization of visual models very difficult. This factor motivates our focus on the generalization of speaker-independent (SI) AVSR systems especially in noisy environments by exploiting the visual domain. Specifically, we are the first to explore the visual adaptation of an SI-AVSR system to an unknown and unlabelled speaker. We adapt an AVSR system trained in a source domain to decode samples in a target domain without the need for labels in the target domain. For the domain adaptation of the unknown speaker, we use Coupled Generative Adversarial Networks to automatically learn a joint distribution of multi-domain images. We evaluate our character-based AVSR system on the TCD-TIMIT dataset and obtain up to a 10% average improvement with respect to its AVSR system equivalent.
Adriana Fernandez-Lopez, Ali Karaali, Naomi Harte, Federico Sukno
ICASSP4
2019 Fully End-to-End Composite Recurrent Convolution Network for Deformable Facial Tracking In The Wild
abstract
Human facial tracking is an important task in computer vision, which has recently lost pace compared to other facial analysis tasks. The majority of current available tracker possess two major limitations: their little use of temporal information and the widespread use of handcrafted features, without taking full advantage of the large annotated datasets that have recently become available. In this paper we present a fully end-to-end facial tracking model based on current state of the art deep model architectures that can be effectively trained from the available annotated facial landmark datasets. We build our model from the recently introduced general object tracker Re3, which allows modeling the short and long temporal dependency between frames by means of its internal Long Short Term Memory (LSTM) layers. Facial tracking experiments on the challenging 300-VW dataset show that our model can produce state of the art accuracy and far lower failure rates than competing approaches. We specifically compare the performance of our approach modified to work in tracking-by-detection mode and showed that, as such, it can produce results that are comparable to state of the art trackers. However, upon activation of our tracking mechanism, the results improve significantly, confirming the advantage of taking into account temporal dependencies.
Decky Aspandi, Oriol Martínez, Federico Sukno, Xavier Binefa
FG3
2019 Tensor Decomposition and Non-linear Manifold Modeling for 3D Head Pose Estimation
Dmytro Derkach, Adria Ruiz, Federico Sukno
Int. J. Comput. Vis.3
2018 3D Head Pose Estimation Using Tensor Decomposition and Non-linear Manifold Modeling
abstract
Head pose estimation is a challenging computer vision problem with important applications in different scenarios such as human-computer interaction or face recognition. In this paper, we present an algorithm for 3D head pose estimation using only depth information from Kinect sensors. A key feature of the proposed approach is that it allows modeling the underlying 3D manifold that results from the combination of pitch, yaw and roll variations. To do so, we use tensor decomposition to generate separate subspaces for each variation factor and show that each of them has a clear structure that can be modeled with cosine functions from a unique shared parameter per angle. Such representation provides a deep understanding of data behavior and angle estimations can be performed by optimizing combination of these cosine functions. We evaluate our approach on two publicly available databases, and achieve top state-of-the-art performance.
Dmytro Derkach, Adria Ruiz, Federico Sukno
3DV3
2018 A Quantitative Comparison of Methods for 3D Face Reconstruction from 2D Images
abstract
In the past years, many studies have highlighted the relation between deviations from normal facial morphology (dysmorphology) and some genetic and mental disorders. Recent advances in methods for reconstructing the 3D geometry of the face from 2D images opens new possibilities for dysmorphology research without the need for specialized 3D imaging equipment. However, it is unclear whether these methods could reconstruct the facial geometry with the required accuracy. In this paper we present a comparative study of some of the most relevant approaches for 3D face reconstruction from 2D images, including photometric-stereo, deep learning and 3D Morphable Model fitting. We address the comparison in qualitatively and quantitatively terms using a public database consisting of 2D images and 3D scans from 100 people. Interestingly, we find that some methods produce quite noisy reconstructions that do not seem realistic, whereas others look more natural. However, the latter do not seem to adequately capture the geometric variability that exists between different subjects and produce reconstructions that look always very similar across individuals, thus questioning their fidelity.
Araceli Morales, Gemma Piella, Oriol Martínez, Federico Sukno
FG4
2018 Improving the Quality of Video-to-Language Models by Optimizing Annotation of the Training Material
Laura Pérez-Mayos, Federico Sukno, Leo Wanner
MMM (1)2
2018 Automatic local shape spectrum analysis for 3D facial expression recognition
Dmytro Derkach, Federico Sukno
Image Vis. Comput.2
2018 Survey on automatic lip-reading in the era of deep learning
Adriana Fernandez-Lopez, Federico Sukno
Image Vis. Comput.2
2018 Weighted regularized statistical shape space projection for breast 3D model reconstruction
Guillermo Ruiz, Eduard Ramon, Jaime García 0001, Federico Sukno, Miguel Ángel González Ballester
Medical Image Anal.4
2017 Head Pose Estimation Based on 3-D Facial Landmarks Localization and Regression
abstract
In this paper we present a system that is able to estimate head pose using only depth information from consumer RGB-D cameras such as Kinect 2. In contrast to most approaches addressing this problem, we do not rely on tracking and produce pose estimation in terms of pitch, yaw and roll angles using single depth frames as input. Our system combines three different methods for pose estimation: two of them are based on state-of-the-art landmark detection and the third one is a dictionarybased approach that is able to work in especially challenging scans where landmarks or mesh correspondences are too difficult to obtain. We evaluated our system on the SASE database, which consists of ~30K frames from 50 subjects. We obtained average pose estimation errors between 5 and 8 degrees per angle, achieving the best performance in the FG2017 Head Pose Estimation Challenge. Full code of the developed system is available on-line.
Dmytro Derkach, Adria Ruiz, Federico Sukno
FG3
2017 Local Shape Spectrum Analysis for 3D Facial Expression Recognition
abstract
We investigate the problem of facial expression recognition using 3D data. Building from one of the most successful frameworks for facial analysis using exclusively 3D geometry, we extend the analysis from a curve-based representation into a spectral representation, which allows a complete description of the underlying surface that can be further tuned to the desired level of detail. Spectral representations are based on the decomposition of the geometry in its spatial frequency components, much like a Fourier transform, which are related to intrinsic characteristics of the surface. In this work, we propose the use of Graph Laplacian Features (GLF), which results from the projection of local surface patches into a common basis obtained from the Graph Laplacian eigenspace. We test the proposed approach in the BU-3DFE database in terms of expressions and Action Units recognition. Our results confirm that the proposed GLF produces consistently higher recognition rates than the curves-based approach, thanks to a more complete description of the surface, while requiring a lower computational complexity. We also show that the GLF outperform the most popular alternative approach for spectral representation, ShapeDNA, which is based on the Laplace Beltrami Operator and cannot provide a stable basis that guarantee that the extracted signatures for the different patches are directly comparable.
Dmytro Derkach, Federico Sukno
FG2
2017 Towards Estimating the Upper Bound of Visual-Speech Recognition: The Visual Lip-Reading Feasibility Database
abstract
Speech is the most used communication method between humans and it involves the perception of auditory and visual channels. Automatic speech recognition focuses on interpreting the audio signals, although the video can provide information that is complementary to the audio. Exploiting the visual information, however, has proven challenging. On one hand, researchers have reported that the mapping between phonemes and visemes (visual units) is one-to-many because there are phonemes which are visually similar and indistinguishable between them. On the other hand, it is known that some people are very good lip-readers (e.g: deaf people). We study the limit of visual only speech recognition in controlled conditions. With this goal, we designed a new database in which the speakers are aware of being read and aim to facilitate lip-reading. In the literature, there are discrepancies on whether hearing-impaired people are better lip-readers than normal-hearing people. Then, we analyze if there are differences between the lip-reading abilities of 9 hearing-impaired and 15 normal-hearing people. Finally, human abilities are compared with the performance of a visual automatic speech recognition system. In our tests, hearing-impaired participants outperformed the normal-hearing participants but without reaching statistical significance. Human observers were able to decode 44% of the spoken message. In contrast, the visual only automatic system achieved 20% of word recognition rate. However, if we repeat the comparison in terms of phonemes both obtained very similar recognition rates, just above 50%. This suggests that the gap between human lip-reading and automatic speech-reading might be more related to the use of context than to the ability to interpret mouth appearance.
Adriana Fernandez-Lopez, Oriol Martínez, Federico Sukno
FG3
2017 Fusion of Valence and Arousal Annotations through Dynamic Subjective Ordinal Modelling
abstract
An essential issue when training and validating computer vision systems for affect analysis is how to obtain reliable ground-truth labels from a pool of subjective annotations. In this paper, we address this problem when labels are given in an ordinal scale and annotated items are structured as temporal sequences. This problem is of special importance in affective computing, where collected data is typically formed by videos of human interactions annotated according to the Valence and Arousal (V-A) dimensions. Moreover, recent works have shown that inter-observer agreement of V-A annotations can be considerably improved if these are given in a discrete ordinal scale. In this context, we propose a novel framework which explicitly introduces ordinal constraints to model the subjective perception of annotators. We also incorporate dynamic information to take into account temporal correlations between ground-truth labels. In our experiments over synthetic and real data with V-A annotations, we show that the proposed method outperforms alternative approaches which do not take into account either the ordinal structure of labels or their temporal correlation.
Adria Ruiz, Oriol Martínez, Xavier Binefa, Federico Sukno
FG4
2016 Weighted regularized ASM for face alignment
abstract
Active Shape Models are a powerful and well known method to perform face alignment. In some applications it is common to have shape information available beforehand, such as previously detected landmarks. Introducing this prior knowledge to the statistical model may result of great advantage but it is challenging to maintain this priors unchanged once the statistical model constraints are applied. We propose a new weighted-regularized projection into the parameter space which allows us to obtain shapes that at the same time fulfill the imposed shape constraints and are plausible according to the statistical model. The performed experiments show how using this projection better performance than competing state of the art methods is achieved.
Guillermo Ruiz, Eduard Ramon, Jaime García 0001, Miguel Ángel González Ballester, Federico Sukno
ICIP5
2015 3-D Facial Landmark Localization With Asymmetry Patterns and Shape Regression from Incomplete Local Features
abstract
We present a method for the automatic localization of facial landmarks that integrates nonrigid deformation with the ability to handle missing points. The algorithm generates sets of candidate locations from feature detectors and performs combinatorial search constrained by a flexible shape model. A key assumption of our approach is that for some landmarks there might not be an accurate candidate in the input set. This is tackled by detecting partial subsets of landmarks and inferring those that are missing, so that the probability of the flexible model is maximized. The ability of the model to work with incomplete information makes it possible to limit the number of candidates that need to be retained, drastically reducing the number of combinations to be tested with respect to the alternative of trying to always detect the complete set of landmarks. We demonstrate the accuracy of the proposed method in the face recognition grand challenge database, where we obtain average errors of approximately 3.5 mm when targeting 14 prominent facial landmarks. For the majority of these our method produces the most accurate results reported to date in this database. Handling of occlusions and surfaces with missing parts is demonstrated with tests on the Bosphorus database, where we achieve an overall error of 4.81 and 4.25 mm for data with and without occlusions, respectively. To investigate potential limits in the accuracy that could be reached, we also report experiments on a database of 144 facial scans acquired in the context of clinical research, with manual annotations performed by experts, where we obtain an overall error of 2.3 mm, with averages per landmark below 3.4 mm for all 14 targeted points and within 2 mm for half of them. The coordinates of automatically located landmarks are made available on-line.
Federico Sukno, John Waddington, Paul F. Whelan
IEEE Trans. Cybern.1
2013 A High-Resolution Atlas and Statistical Model of the Human Heart From Multislice CT
abstract
Atlases and statistical models play important roles in the personalization and simulation of cardiac physiology. For the study of the heart, however, the construction of comprehensive atlases and spatio-temporal models is faced with a number of challenges, in particular the need to handle large and highly variable image datasets, the multi-region nature of the heart, and the presence of complex as well as small cardiovascular structures. In this paper, we present a detailed atlas and spatio-temporal statistical model of the human heart based on a large population of 3D+time multi-slice computed tomography sequences, and the framework for its construction. It uses spatial normalization based on nonrigid image registration to synthesize a population mean image and establish the spatial relationships between the mean and the subjects in the population. Temporal image registration is then applied to resolve each subject-specific cardiac motion and the resulting transformations are used to warp a surface mesh representation of the atlas to fit the images of the remaining cardiac phases in each subject. Subsequently, we demonstrate the construction of a spatio-temporal statistical model of shape such that the inter-subject and dynamic sources of variation are suitably separated. The framework is applied to a 3D+time data set of 138 subjects. The data is drawn from a variety of pathologies, which benefits its generalization to new subjects and physiological studies. The obtained level of detail and the extendability of the atlas present an advantage over most cardiac models published previously.
Corné Hoogendoorn, Nicolas Duchateau, Damian Sánchez-Quintana, Tristan Whitmarsh, Federico Sukno, Mathieu De Craene, Karim Lekadir, Alejandro F. Frangi
IEEE Trans. Medical Imaging5
2012 Suicide attempters classification: Toward predictive models of suicidal behavior
David Delgado-Gómez, Hilario Blasco-Fontecilla, Federico Sukno, Maria Socorro Ramos-Plasencia, Enrique Baca-García
Neurocomputing3
2012 An Experimental Evaluation of Three Classifiers for Use in Self-Updating Face Recognition Systems
abstract
Previous studies have shown that the accuracy of Face Recognition Systems (FRSs) decreases with the time elapsed between enrollment and testing. The main reason for the decrease is the changes in appearance of the user due to factors such as ageing, beard growth, sun-tan etc. Self-update procedure, where the system learns the biometric characteristics of the user every time he/she interacts with it, can be used to automatically update the system. However, a commonly acknowledged problem is the corruption of biometric traits due to misclassification. In this article, we test FRS, based on three classification algorithms, on two challenging databases, GEFA and YT, with 14 279 and 31 951 images, respectively. Our results suggest that complex, state-of-the-art classifiers that make use of user-specific models, need not be the best choice for use in self updating systems. In other words, tolerance to corrupted training data decreases as the complexity of the classifier increases.
Sri-Kaushik Pavani, Federico Sukno, David Delgado-Gómez, Constantine Butakoff, Xavier Planes, Alejandro F. Frangi
IEEE Trans. Inf. Forensics Secur.2
2010 Automatic Cardiac MRI Segmentation Using a Biventricular Deformable Medial Model
Alejandro F. Frangi, Hongzhi Wang 0002, Federico Sukno, Catalina Tobon-Gomez, Paul A. Yushkevich
MICCAI (1)4
2010 Individual identification using personality traits
David Delgado-Gómez, Federico Sukno, David Aguado, Carlos Santa Cruz, Antonio Artés-Rodríguez
J. Netw. Comput. Appl.2
2010 The Multiscenario Multienvironment BioSecure Multimodal Database (BMDB)
abstract
A new multimodal biometric database designed and acquired within the framework of the European BioSecure Network of Excellence is presented. It is comprised of more than 600 individuals acquired simultaneously in three scenarios: 1) over the Internet, 2) in an office environment with desktop PC, and 3) in indoor/outdoor environments with mobile portable hardware. The three scenarios include a common part of audio/video data. Also, signature and fingerprint data have been acquired both with desktop PC and mobile portable hardware. Additionally, hand and iris data were acquired in the second scenario using desktop PC. Acquisition has been conducted by 11 European institutions. Additional features of the BioSecure Multimodal Database (BMDB) are: two acquisition sessions, several sensors in certain modalities, balanced gender and age distributions, multimodal realistic scenarios with simple and quick tasks per modality, cross-European diversity, availability of demographic data, and compatibility with other multimodal databases. The novel acquisition conditions of the BMDB allow us to perform new challenging research and evaluation of either monomodal or multimodal biometric systems, as in the recent BioSecure Multimodal Evaluation campaign. A description of this campaign including baseline results of individual modalities from the new database is also given. The database is expected to be available for research purposes through the BioSecure Association during 2008.
Javier Ortega-Garcia, Julian Fierrez, Fernando Alonso-Fernandez, Javier Galbally, Manuel R. Freire, Joaquín González-Rodríguez, Carmen García-Mateo, José Luis Alba-Castro, Elisardo González-Agulla, Enrique Otero Muras, Sonia Garcia-Salicetti, Lorène Allano, Van-Bao Ly, Bernadette Dorizzi, Josef Kittler, Thirimachos Bourlai, Norman Poh, Farzin Deravi, Ming W. R. Ng, Michael C. Fairhurst, Jean Hennebert, Andreas Humm, Massimo Tistarelli, Linda Brodo, Jonas Richiardi, Andrzej Drygajlo, Harald Ganster, Federico Sukno, Sri-Kaushik Pavani, Alejandro F. Frangi, Lale Akarun, Arman Savran
IEEE Trans. Pattern Anal. Mach. Intell.28
2010 Projective active shape models for pose-variant image analysis of quasi-planar objects: Application to facial analysis
Federico Sukno, Josechu J. Guerrero, Alejandro F. Frangi
Pattern Recognit.1
2009 Automatic Assessment of Eye Blinking Patterns through Statistical Shape Models
Federico Sukno, Sri-Kaushik Pavani, Constantine Butakoff, Alejandro F. Frangi
ICVS1
2009 Bilinear Models for Spatio-Temporal Point Distribution Analysis
Corné Hoogendoorn, Federico Sukno, Sebastián Ordas, Alejandro F. Frangi
Int. J. Comput. Vis.2
2009 Similarity-based Fisherfaces
David Delgado-Gómez, Jens Fagertun, Bjarne K. Ersbøll, Federico Sukno, Alejandro F. Frangi
Pattern Recognit. Lett.4
2008 Cardiac Medial Modeling and Time-Course Heart Wall Thickness Analysis
Brian B. Avants, Alejandro F. Frangi, Federico Sukno, James C. Gee, Paul A. Yushkevich
MICCAI (2)4
2008 Reliability Estimation for Statistical Shape Models
abstract
One of the drawbacks of statistical shape models is their occasional failure to converge. Although visually this fact is usually easy to recognize, there is no automatic way to detect it. In this paper, we introduce a generic reliability measure for statistical shape models. It is based on a probabilistic framework and uses information extracted by the model itself during the matching process. The proposed method was validated with two variants of Active Shape Models in the context facial image analysis. Experimental results on more than 3700 facial images showed a high degree of correlation between the segmentation accuracy and the estimated reliability metric.
Federico Sukno, Alejandro F. Frangi
IEEE Trans. Image Process.1
2008 Automatic Construction of 3D-ASM Intensity Models by Simulating Image Acquisition: Application to Myocardial Gated SPECT Studies
abstract
Active shape models bear a great promise for model-based medical image analysis. Their practical use, though, is undermined due to the need to train such models on large image databases. Automatic building of point distribution models (PDMs) has been successfully addressed and a number of autolandmarking techniques are currently available. However, the need for strategies to automatically build intensity models around each landmark has been largely overlooked in the literature. This work demonstrates the potential of creating intensity models automatically by simulating image generation. We show that it is possible to reuse a 3D PDM built from computed tomography (CT) to segment gated single photon emission computed tomography (gSPECT) studies. Training is performed on a realistic virtual population where image acquisition and formation have been modeled using the SIMIND Monte Carlo simulator and ASPIRE image reconstruction software, respectively. The dataset comprised 208 digital phantoms (4D-NCAT) and 20 clinical studies. The evaluation is accomplished by comparing point-to-surface and volume errors against a proper gold standard. Results show that gSPECT studies can be successfully segmented by models trained under this scheme with subvoxel accuracy. The accuracy in estimated LV function parameters, such as end diastolic volume, end systolic volume, and ejection fraction, ranged from 90.0% to 94.5% for the virtual population and from 87.0% to 89.5% for the clinical population.
Catalina Tobon-Gomez, Constantine Butakoff, S. Aguade, Federico Sukno, G. Moragas, Alejandro F. Frangi
IEEE Trans. Medical Imaging4
2007 Bilinear Models for Spatio-Temporal Point Distribution Analysis: Application to Extrapolation of Whole Heart Cardiac Dynamics
abstract
In this work we introduce the usage of bilinear models as a means of factorising the shape variation induced by subject variability and the contraction of the human heart. We show that it is feasible to reconstruct the shape of the heart at a certain point in the cardiac cycle if we are given a small number of shapes representing the same heart at different points in the same cycle, using the bilinear model. Depending on pathology and the ratios between healthy and pathological hearts in the training set, RMS reconstruction errors measured between 1.39 and 16.58 millimetres, with a median of 6.79 and 90th percentile of 9.95 millimetres.
Corné Hoogendoorn, Federico Sukno, Sebastián Ordas, Alejandro F. Frangi
ICCV2
2007 Active Shape Models with Invariant Optimal Features: Application to Facial Analysis
abstract
This work is framed in the field of statistical face analysis. In particular, the problem of accurate segmentation of prominent features of the face in frontal view images is addressed. We propose a method that generalizes linear Active Shape Models (ASMs), which have already been used for this task. The technique is built upon the development of a nonlinear intensity model, incorporating a reduced set of differential invariant features as local image descriptors. These features are invariant to rigid transformations, and a subset of them is chosen by Sequential Feature Selection for each landmark and resolution level. The new approach overcomes the unimodality and Gaussianity assumptions of classical ASMs regarding the distribution of the intensity values across the training set. Our methodology has demonstrated a significant improvement in segmentation precision as compared to the linear ASM and Optimal Features ASM (a nonlinear extension of the pioneer algorithm) in the tests performed on AR, XM2VTS, and EQUINOX databases.
Federico Sukno, Sebastián Ordas, Constantine Butakoff, Santiago Cruz, Alejandro F. Frangi
IEEE Trans. Pattern Anal. Mach. Intell.1
2004 AV@CAR: A Spanish Multichannel Multimodal Corpus for In-Vehicle Automatic Audio-Visual Speech Recognition
Alfonso Ortega Giménez, Federico Sukno, Eduardo Lleida, Alejandro F. Frangi, Antonio Miguel, Luis Buera, Ernesto Zacur
LREC2