EDBT 2026 Demo / reviewers in the wild / expert
Claudio Ferrari
dblp:155/3208
· DBLP profile ↗
40ranked-venue papers
15as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 12 first-author · 19 since 2021Artificial intelligence and machine learning · 25 · 9 first-author · 16 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Polyglot: Multilingual Style Preserving Speech-Driven Facial Animation
Federico Nocentini, Kwanggyoon Seo, Qingju Liu, Claudio Ferrari, Stefano Berretti, David Ferman, Hyeongwoo Kim, Pablo Garrido 0001, Akin Caliskan |
FG | 4 |
| 2026 | Spatio-temporal transformers for action unit classification with event cameras
Luca Cultrera, Federico Becattini, Lorenzo Berlincioni, Claudio Ferrari, Alberto Del Bimbo |
Comput. Vis. Image Underst. | 4 |
| 2026 | A video dataset for Multi-Emotion and Interpersonal relation analysis
Hajer Guerdelli, Claudio Ferrari, Stefano Berretti, Walid Barhoumi, Alberto Del Bimbo |
Comput. Vis. Image Underst. | 2 |
| 2026 | Beyond Fixed Topologies: Unregistered Training and Comprehensive Evaluation Metrics for 3D Talking Heads
Federico Nocentini, Thomas Besnier, Claudio Ferrari, Sylvain Arguillère, Mohamed Daoudi, Stefano Berretti |
Int. J. Comput. Vis. | 3 |
| 2025 | 3D Face Reconstruction Error Decomposed: A Modular Benchmark for Fair and Fast Method EvaluationabstractComputing the standard benchmark metric for 3D face reconstruction, namely geometric error, requires a number of steps, such as mesh cropping, rigid alignment, or point correspondence. Current benchmark tools are monolithic (they implement a specific combination of these steps), even though there is no consensus on the best way to measure error. We present a toolkit for a Modularized 3D Face reconstruction Benchmark (M3DFB), where the fundamental components of error computation are segregated and interchangeable, allowing one to quantify the effect of each. Furthermore, we propose a new component, namely correction, and present a computationally efficient approach that penalizes for mesh topology inconsistency. Using this toolkit, we test 16 error estimators with 10 reconstruction methods on two real and two synthetic datasets. Critically, the widely used ICP-based estimator provides the worst benchmarking performance, as it significantly alters the true ranking of the top- 5 reconstruction methods. Notably, the correlation of ICP with the true error can be as low as 0.41. Moreover, non-rigid alignment leads to significant improvement (correlation larger than 0.90), highlighting the importance of annotating 3D landmarks on datasets. Finally, the proposed correction scheme, together with non-rigid warping, leads to an accuracy on a par with the best non-rigid ICP-based estimators, but runs an order of magnitude faster. Our open-source codebase is designed for researchers to easily compare alternatives for each component, thus helping accelerating progress in benchmarking for 3D face reconstruction and, furthermore, supporting the improvement of learned reconstruction methods, which depend on accurate error estimation for effective training. Evangelos Sariyanidi, Claudio Ferrari, Federico Nocentini, Stefano Berretti, Andrea Cavallaro, Birkan Tunç |
FG | 2 |
| 2025 | Mamba-ST: State Space Model for Efficient Style TransferabstractThe goal of style transfer is, given a content image and a style source, generating a new image preserving the content but with the artistic representation of the style source. Most of the state-of-the-art architectures use transformers or diffusion-based models to perform this task, despite the heavy computational burden that they require. In particular, transformers use self- and cross-attention layers which have large memory footprint, while diffusion models require high inference time. To overcome the above, this paper explores a novel design of Mamba, an emergent State-Space Model (SSM), called Mamba-ST, to perform style transfer. To do so, we adapt Mamba linear equation to simulate the behavior of cross-attention layers, which are able to combine two separate embeddings into a single output, but drastically reducing memory usage and time complexity. We modified the Mamba's inner equations so to accept inputs from, and combine, two separate data streams. To the best of our knowledge, this is the first attempt to adapt the equations of SSMs to a vision task like style transfer without requiring any other module like cross-attention or custom normalization layers. An extensive set of experiments demonstrates the superiority and efficiency of our method in performing style transfer compared to transformers and diffusion models. Results show improved quality in terms of both ArtFID and FID metrics. Code is available at https://github.com/FilippoBotti/MambaST. Filippo Botti, Alex Ergasti, Leonardo Rossi, Tomaso Fontanini, Claudio Ferrari, Massimo Bertozzi, Andrea Prati 0001 |
WACV | 5 |
| 2025 | EmoVOCA: Speech-Driven Emotional 3D Talking HeadsabstractA notable challenge in 3D talking head generation consists in blending speech-related motions with expression dynamics. This is primarily caused by the lack of comprehensive 3D datasets that combine diversity in spoken sentences with a variety of facial expressions. Some literature works attempted to overcome such lack of data by fitting parametric 3D models (3DMMs) to 2D videos, and using the reconstructed 3D faces as replacement. However, their underlying parametric space limits the precision required to accurately reproduce convincing lip motions and synching, which is crucial for the application at hand. In this work, we look at the problem from a different perspective, and developed a data-driven technique to combine inexpressive 3D talking heads with a set of 3D expressive sequences, which we used for creating a synthetic dataset, called EmoVOCA. We then designed and trained an emotional 3D talking head generator that accepts a 3D face, an audio file, an emotion label, and an intensity value as inputs, and learns to animate the audio-synchronized lip movements with expressive traits of the face. Comprehensive experiments, both quantitative and qualitative, using our data and generator evidence superior ability in synthesizing convincing animations, when compared with the best performing methods in the literature. Our code and pre-trained models are available at https://github.com/miccunifi/EmoVOCA. Federico Nocentini, Claudio Ferrari, Stefano Berretti |
WACV | 2 |
| 2025 | MARS: Paying More Attention to Visual Attributes for Text-Based Person SearchabstractText-Based Person Search (TBPS) is a problem that gained significant interest within the research community. The task is that of retrieving one or more images of a specific individual based on a textual description. The multi-modal nature of the task requires learning representations that bridge text and image data within a shared latent space. Existing TBPS systems face two major challenges. One is defined as inter-identity noise that is due to the inherent vagueness and imprecision of text descriptions, and it indicates how descriptions of visual attributes can be generally associated to different people; the other is the intra-identity variations, which are all those nuisances, e.g., pose, illumination, that can alter the visual appearance of the same textual attributes for a given subject. To address these issues, this article presents a novel TBPS architecture named Mae-Attribute-Relation-Sensitive (MARS), which enhances current state-of-the-art models by introducing two key components: a Visual Reconstruction Loss and an Attribute Loss. The former employs a Masked AutoEncoder trained to reconstruct randomly masked image patches with the aid of the textual description. In doing so the model is encouraged to learn more expressive representations and textual–visual relations in the latent space. The attribute loss, instead, balances the contribution of different types of attributes, defined as adjective–noun chunks of text. This loss ensures that every attribute is taken into consideration in the person retrieval process. Extensive experiments on three commonly used datasets, namely CUHK-PEDES, ICFG-PEDES, and RSTPReid, report performance improvements, with significant gains in the Mean Average Precision (mAP) metric w.r.t. the current state of the art. Code will be available at https://github.com/ErgastiAlex/MARS . Alex Ergasti, Tomaso Fontanini, Claudio Ferrari, Massimo Bertozzi, Andrea Prati 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | ScanTalk: 3D Talking Heads from Unregistered Scans
Federico Nocentini, Thomas Besnier, Claudio Ferrari, Sylvain Arguillère, Stefano Berretti, Mohamed Daoudi |
ECCV (29) | 3 |
| 2024 | IMEmo: An Interpersonal Relation Multi-Emotion DatasetabstractWhile engaged in a face-to-face conversation, being capable of understanding the attitude, emotion, and intention of another person allows one guiding his/her behavior establishing a comfortable communication either verbal and non-verbal (i.e., body and face language). This paper introduces the “The IMEmo Interpersonal Multi-Emotion video dataset”, a new in-the-wild dataset of face-to-face interaction, built from movies of romance and drama categories. We manually collected over 100 clips from different movies in different languages. The dataset consists of 79.3 minutes scenes, with a duration of each clip ranging between 0.20 and 2.13 minutes. Each clip contains two people communicating with each other both verbally and with expression and body pose and gestures. Currently, it includes age, gender, emotions, social relationships, actions and valence/arousal annotations for both individuals. Emotion recognition results using a baseline CNN approach are also reported to provide an estimation of the difficulty of the data also in comparison to existing benchmarks. Hajer Guerdelli, Claudio Ferrari, Stefano Berretti, Alberto Del Bimbo |
FG | 2 |
| 2024 | Generating Multiple 4D Expression Transitions by Learning Face Landmark TrajectoriesabstractIn this work, we address the problem of 4D facial expressions generation. This is usually addressed by animating a neutral 3D face to reach an expression peak, and then get back to the neutral state. In the real world though, people show more complex expressions, and switch from one expression to another. We thus propose a new model that generates transitions between different expressions, and synthesizes long and composed 4D expressions. This involves three sub-problems: (i) modeling the temporal dynamics of expressions, (ii) learning transitions between them, and (iii) deforming a generic mesh. We propose to encode the temporal evolution of expressions using the motion of a set of 3D landmarks, that we learn to generate by training a manifold-valued GAN (Motion3DGAN). To allow the generation of composed expressions, this model accepts two labels encoding the starting and the ending expressions. The final sequence of meshes is generated by a Sparse2Dense mesh Decoder (S2D-Dec) that maps the landmark displacements to a dense, per-vertex displacement of a known mesh topology. By explicitly working with motion trajectories, the model is totally independent from the identity. Extensive experiments on five public datasets show that our proposed approach brings significant improvements with respect to previous solutions, while retaining good generalization to unseen data. Naima Otberdout, Claudio Ferrari, Mohamed Daoudi, Stefano Berretti, Alberto Del Bimbo |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | The Florence 4D Facial Expression DatasetabstractHuman facial expressions change dynamically, so their recognition / analysis should be conducted by accounting for the temporal evolution of face deformations either in 2D or 3D. While abundant 2D video data do exist, this is not the case in 3D, where few 3D dynamic (4D) datasets were released for public use. The negative consequence of this scarcity of data is amplified by current deep learning based-methods for facial expression analysis that require large quantities of variegate samples to be effectively trained. With the aim of smoothing such limitations, in this paper we propose a large dataset, named Florence 4D, composed of dynamic sequences of 3D face models, where a combination of synthetic and real identities exhibit an unprecedented variety of 4D facial expressions, with variations that include the classical neutral-apex transition, but generalize to expression-to-expression. All these characteristics are not exposed by any of the existing 4D datasets and they cannot even be obtained by combining more than one dataset. We strongly believe that making such a data corpora publicly available to the community will allow designing and experimenting new applications that were not possible to investigate till now. To show at some extent the difficulty of our data in terms of different identities and varying expressions, we also report a baseline experimentation on the proposed dataset that can be used as baseline. Filippo Principi, Stefano Berretti, Claudio Ferrari, Naima Otberdout, Mohamed Daoudi, Alberto Del Bimbo |
FG | 3 |
| 2023 | Meta-evaluation for 3D Face Reconstruction Via Synthetic DataabstractThe standard benchmark metric for 3D face reconstruction is the geometric error between reconstructed meshes and the ground truth. Nearly all recent reconstruction methods are validated on real ground truth scans, in which case one needs to establish point correspondence prior to error computation, which is typically done with the Chamfer (i.e., nearest neighbor) criterion. However, a simple yet fundamental question have not been asked: Is the Chamfer error an appropriate and fair benchmark metric for 3D face reconstruction? More generally, how can we determine which error estimator is a better benchmark metric? We present a meta-evaluation framework that uses synthetic data to evaluate the quality of a geometric error estimator as a benchmark metric for face reconstruction. Further, we use this framework to experimentally compare four geometric error estimators. Results show that the standard approach not only severely underestimates the error, but also does so inconsistently across reconstruction methods, to the point of even altering the ranking of the compared methods. Moreover, although non-rigid ICP leads to a metric with smaller estimation bias, it could still not correctly rank all compared reconstruction methods, and is significantly more time consuming than Chamfer. In sum, we show several issues present in the current benchmarking and propose a procedure using synthetic data to address these issues. Evangelos Sariyanidi, Claudio Ferrari, Stefano Berretti, Robert T. Schultz, Birkan Tunç |
IJCB | 2 |
| 2023 | Learning graph-based features for relief patterns classification on mesh manifoldsabstractRelief patterns represent a surface characteristic that can be seen as the 3D counterpart of the texture concept in 2D images. Such characteristic is well distinct from the 3D object shape but represents a good information source to recognize the object itself. The majority of state-of-the-art techniques for 2D images rely on convolution-based filtering so, the idea of extending such techniques to the mesh manifold domain is quite intriguing as much as challenging. In this paper, we propose a novel approach based on Graph Neural Networks for 3D mesh relief pattern classification. To this end, we designed a bi-level architecture that learns on data structures computed thanks to a mesh resampling algorithm that allows us to represent local surface patches uniformly, while keeping a consistent points order. The local mesh structures are represented by SpiderPatches, that aim to capture local features of the 3D mesh surface, providing a fine-grained rich representation of the relief patterns; global structures are instead captured by MeshGraphs, whose nodes are SpiderPatches, representing the mesh at a macroscopic level. We tested our architecture using SpiderPatches and MeshGraphs on the original meshes of the SHREC’17 and SHREC’20 relief patterns track datasets, showing superior performance to that reported in the literature using a comparable experimental setting. Niccolò Guiducci, Claudio Tortorici, Claudio Ferrari, Stefano Berretti |
Comput. Graph. | 3 |
| 2023 | Interpersonal relation recognition: a surveyabstractAbstract People spend a considerable amount of their time in social activities, where person-to-person relations are of main relevance. Recently, there has been an increasing research interest in automatically analyzing interpersonal relations, for the social and behavioral implications, and the many practical applications it may have. However, to the best of our knowledge, there is not a systematic study providing a harmonized view of the literature in the field. On this ground, we summarize in our work interpersonal relation recognition datasets and methods aiming to help researchers to have a better understanding of the characteristics of the state-of-the-art. In the proposed study, we distinguish between methods that address objective relations that do not depend on behavior or emotional state, and methods that consider subjective ones that depend on emotions. It turns out quite evidently that aiming at the latter recognition task is more challenging, with the existing methods that provide convincing results only on limited and very specific cases. For both the broad categories, we discuss datasets and methods according to the different behavioural and psychological models used to annotate and classify the data. We conclude our review work, by providing a comprehensive discussion pointing out current limitations and future research perspectives. Hajer Guerdelli, Claudio Ferrari, Stefano Berretti |
Multim. Tools Appl. | 2 |
| 2023 | Correction to: Interpersonal relation recognition: a survey
Hajer Guerdelli, Claudio Ferrari, Stefano Berretti |
Multim. Tools Appl. | 2 |
| 2023 | The Florence multi-resolution 3D facial expression datasetabstractIn the literature, several 3D face datasets have been collected, aiming at advancing the field of 3D face analysis from different perspectives. Data collection generally follows specific research needs, and the existing 3D face datasets all have different characteristics that are tailored for investigating different tasks, encompassing face recognition, facial expressions and emotions analysis, 3D face reconstruction. However, the majority of these datasets are either collected with high-resolution scanners, or consumer level devices, such as the Kinect, the latter being motivated by the burdensome and costly process of collecting high-quality scans. Differently from 2D imagery, the difference in resolution in 3D data represents a non negligible problem that is under-investigated, and still prevents the successful development of methods that can work in real scenarios. In this paper, we propose a new 3D face dataset, named “Florence Multi-Resolution 3D Facial Expression” (Florence 3DMRE), which aims at bridging the gap between high- and low-resolution 3D face datasets. Its peculiarity consists in (1) including high-resolution (HR) models obtained with a HR scanner, and paired samples collected with a Kinect sensor, (2) LR and HR scans are synchronized and capture extreme and asymmetric facial deformations as used in facial rehabilitation exercises. In total, our dataset consists of 14 subjects, each performing 19 complex and asymmetric expressions. For each of them, we collected a high-resolution scan, and an RGB-D sequence. Finally, to highlight the value of the dataset and the challenges it introduces, we use the collected data to perform baseline experiments for cross-resolution 3D face recognition and reconstruction. The dataset is released for research purposes only, and complies to GDPR for data treatment. The dataset can be found at this link. Claudio Ferrari, Stefano Berretti, Pietro Pala, Alberto Del Bimbo |
Pattern Recognit. Lett. | 1 |
| 2023 | FrankenMask: Manipulating semantic masks with transformers for face parts editingabstractIn this paper, we propose FrankenMask, a novel framework that allows swapping and rearranging face parts in semantic masks for automatic editing of shape-related facial attributes. This is a novel yet challenging task as substituting face parts in a semantic mask requires to account for possible spatial misalignment and the adaptation of surrounding regions. We obtain such a feature by combining a Transformer encoder to learn the spatial relationships of facial parts, with an encoder–decoder architecture, which reconstructs a complete mask from the composition of local parts. Reconstruction and attribute classification results demonstrate the effective synthesis of facial images, while showing the generation of accurate and plausible facial attributes. Code is available at https://github.com/TFonta/FrankenMask_semantic. Tomaso Fontanini, Claudio Ferrari, Giuseppe Lisanti, Leonardo Galteri, Stefano Berretti, Massimo Bertozzi, Andrea Prati 0001 |
Pattern Recognit. Lett. | 2 |
| 2023 | (Compress and Restore)N: A Robust Defense Against Adversarial Attacks on Image ClassificationabstractModern image classification approaches often rely on deep neural networks, which have shown pronounced weakness to adversarial examples: images corrupted with specifically designed yet imperceptible noise that causes the network to misclassify. In this article, we propose a conceptually simple yet robust solution to tackle adversarial attacks on image classification. Our defense works by first applying a JPEG compression with a random quality factor; compression artifacts are subsequently removed by means of a generative model Artifact Restoration GAN. The process can be iterated ensuring the image is not degraded and hence the classification not compromised. We train different AR-GANs for different compression factors, so that we can change its parameters dynamically at each iteration depending on the current compression, making the gradient approximation difficult. We experiment with our defense against three white-box and two black-box attacks, with a particular focus on the state-of-the-art BPDA attack. Our method does not require any adversarial training, and is independent of both the classifier and the attack. Experiments demonstrate that dynamically changing the AR-GAN parameters is of fundamental importance to obtain significant robustness. Claudio Ferrari, Federico Becattini, Leonardo Galteri, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Learning Streamed Attention Network from Descriptor Images for Cross-Resolution 3D Face RecognitionabstractIn this article, we propose a hybrid framework for cross-resolution 3D face recognition which utilizes a Streamed Attention Network (SAN) that combines handcrafted features with Convolutional Neural Networks (CNNs). It consists of two main stages: first, we process the depth images to extract low-level surface descriptors and derive the corresponding Descriptor Images (DIs), represented as four-channel images. To build the DIs, we propose a variation of the 3D Local Binary Pattern (3DLBP) operator that encodes depth differences using a sigmoid function. Then, we design a CNN that learns from these DIs. The peculiarity of our solution consists in processing each channel of the input image separately, and fusing the contribution of each channel by means of both self- and cross-attention mechanisms. This strategy showed two main advantages over the direct application of Deep-CNN to depth images of the face; on the one hand, the DIs can reduce the diversity between high- and low-resolution data by encoding surface properties that are robust to resolution differences. On the other, it allows a better exploitation of the richer information provided by low-level features, resulting in improved recognition. We evaluated the proposed architecture in a challenging cross-dataset, cross-resolution scenario. To this aim, we first train the network on scanner-resolution 3D data. Next, we utilize the pre-trained network as feature extractor on low-resolution data, where the output of the last fully connected layer is used as face descriptor. Other than standard benchmarks, we also perform experiments on a newly collected dataset of paired high- and low-resolution 3D faces. We use the high-resolution data as gallery, while low-resolution faces are used as probe, allowing us to assess the real gap existing between these two types of data. Extensive experiments on low-resolution 3D face benchmarks show promising results with respect to state-of-the-art methods. João Baptista Cardia Neto, Claudio Ferrari, Aparecido Nilceu Marana, Stefano Berretti, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Sparse to Dense Dynamic 3D Facial Expression GenerationabstractIn this paper, we propose a solution to the task of generating dynamic 3D facial expressions from a neutral 3D face and an expression label. This involves solving two sub-problems: (i) modeling the temporal dynamics of expressions, and (ii) deforming the neutral mesh to obtain the expressive counterpart. We represent the temporal evolution of expressions using the motion of a sparse set of 3D landmarks that we learn to generate by training a manifold-valued GAN (Motion3DGAN). To better encode the expression-induced deformation and disentangle it from the identity information, the generated motion is represented as per-frame displacement from a neutral configuration. To generate the expressive meshes, we train a Sparse2Dense mesh Decoder (S2D-Dec) that maps the landmark displacements to a dense, per-vertex displacement. This allows us to learn how the motion of a sparse set of landmarks influences the deformation of the overall face surface, independently from the identity. Experimental results on the CoMA and D3DFACS datasets show that our solution brings significant improvements with respect to previous solutions in terms of both dynamic expression generation and mesh reconstruction, while retaining good generalization to unseen data. Code and models are available at https://github.com/CRISTAL-3DSAM/Sparse2Dense. Naima Otberdout, Claudio Ferrari, Mohamed Daoudi, Stefano Berretti, Alberto Del Bimbo |
CVPR | 2 |
| 2022 | Complete Verification via Multi-Neuron Relaxation Guided Branch-and-Bound
Claudio Ferrari, Mark Niklas Müller, Nikola Jovanovic 0001, Martin T. Vechev |
ICLR | 1 |
| 2022 | What makes you, you? Analyzing Recognition by Swapping Face PartsabstractDeep learning advanced face recognition to an unprecedented accuracy. However, understanding how local parts of the face affect the overall recognition performance is still mostly unclear. Among others, face swap has been experimented to this end, but just for the entire face. In this paper, we propose to swap facial parts as a way to disentangle the recognition relevance of different face parts, like eyes, nose and mouth. In our method, swapping parts from a source face to a target one is performed by fitting a 3D prior, which establishes dense pixels correspondence between parts, while also handling pose differences. Seamless cloning is then used to obtain smooth transitions between the mapped source regions and the shape and skin tone of the target face. We devised an experimental protocol that allowed us to draw some preliminary conclusions when the swapped images are classified by deep networks, indicating a prominence of the eyes and eyebrows region. Code available at https://github.com/clferrari/FacePartsSwap Claudio Ferrari, Matteo Serpentoni, Stefano Berretti, Alberto Del Bimbo |
ICPR | 1 |
| 2022 | Foreword to the Special Section on 3D Object Retrieval 2022 Symposium (3DOR2022)
Stefano Berretti, Theoharis Theoharis, Mohamed Daoudi, Claudio Ferrari, Remco C. Veltkamp |
Comput. Graph. | 4 |
| 2022 | A Sparse and Locally Coherent Morphable Face Model for Dense Semantic Correspondence Across Heterogeneous 3D FacesabstractThe 3D Morphable Model (3DMM) is a powerful statistical tool for representing 3D face shapes. To build a 3DMM, a training set of face scans in full point-to-point correspondence is required, and its modeling capabilities directly depend on the variability contained in the training data. Thus, to increase the descriptive power of the 3DMM, establishing a dense correspondence across heterogeneous scans with sufficient diversity in terms of identities, ethnicities, or expressions becomes essential. In this manuscript, we present a fully automatic approach that leverages a 3DMM to transfer its dense semantic annotation across raw 3D faces, establishing a dense correspondence between them. We propose a novel formulation to learn a set of sparse deformation components with local support on the face that, together with an original non-rigid deformation algorithm, allow the 3DMM to precisely fit unseen faces and transfer its semantic annotation. We extensively experimented our approach, showing it can effectively generalize to highly diverse samples and accurately establish a dense correspondence even in presence of complex facial expressions. The accuracy of the dense registration is demonstrated by building a heterogeneous, large-scale 3DMM from more than 9,000 fully registered scans obtained by joining three large datasets together. Claudio Ferrari, Stefano Berretti, Pietro Pala, Alberto Del Bimbo |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Measuring 3D face deformations from RGB images of expression rehabilitation exercisesabstractThe accurate (quantitative) analysis of face deformations in 3D is a problem of increasing interest for the many applications it may have. In particular, defining a 3D model of the face that can deform to a 2D target image, while capturing local and asymmetric deformations is still a challenge in the existing literature. Computing a measure of such local deformations may represent a relevant index for monitoring rehabilitation exercises that are used in Parkinson’s and Alzheimer’s disease or in recovering from a stroke. In this study, we present a complete framework that allows the construction of a 3D Morphable Shape Model (3DMM) of the face and its fitting to a target RGB image. The model has the specific characteristic of being based on localized components of deformation; the fitting transformation is performed from 3D to 2D and is guided by the correspondence between landmarks detected in the target image and landmarks manually annotated on the average 3DMM. The fitting has also the peculiarity of being performed in two steps, disentangling face deformations that are due to the identity of the target subject from those induced by facial actions. In the experimental validation of the method, we used the MICC-3D dataset that includes 11 subjects each acquired in one neutral pose plus 18 facial actions that deform the face in localized and asymmetric ways. For each acquisition, we fit the 3DMM to an RGB frame with an apex facial action and to the neutral frame, and computed the extent of the deformation. Results indicated that the proposed approach can accurately capture the face deformation even for localized and asymmetric ones. The proposed framework proved the idea of measuring the deformations of a reconstructed 3D face model to monitor the facial actions performed in response to a set of target ones. Interestingly, these results were obtained just using RGB targets without the need for 3D scans captured with costly devices. This opens the way to the use of the proposed tool for remote medical monitoring of rehabilitation. Claudio Ferrari, Stefano Berretti, Pietro Pala, Alberto Del Bimbo |
Virtual Real. Intell. Hardw. | 1 |
| 2021 | PLM-IPE: A Pixel-Landmark Mutual Enhanced Framework for Implicit Preference EstimationabstractIn this paper, we are interested in understanding how customers perceive fashion recommendations, in particular when observing a proposed combination of garments to compose an outfit. Automatically understanding how a suggested item is perceived, without any kind of active engagement, is in fact an essential block to achieve interactive applications. We propose a pixel-landmark mutual enhanced framework for implicit preference estimation, named PLM-IPE, which is capable of inferring the user’s implicit preferences exploiting visual cues, without any active or conscious engagement. PLM-IPE consists of three key modules: pixel-based estimator, landmark-based estimator and mutual learning based optimization. The former two modules work on capturing the implicit reaction of the user from the pixel level and landmark level, respectively. The last module serves to transfer knowledge between the two parallel estimators. Towards evaluation, we collected a real-world dataset, named SentiGarment, which contains 3,345 facial reaction videos paired with suggested outfits and human labeled reaction scores. Extensive experiments show the superiority of our model over state-of-the-art approaches. Federico Becattini, Xuemeng Song, Claudio Baecchi, Shi-Ting Fang, Claudio Ferrari, Liqiang Nie, Alberto Del Bimbo |
MMAsia | 5 |
| 2020 | Probability Guided MaxoutabstractIn this paper, we propose an original CNN training strategy that brings together ideas from both dropout-like regularization methods and solutions that learn discriminative features. We propose a dropping criterion that, differently from dropout and its variants, is deterministic rather than random. It grounds on the empirical evidence that feature descriptors with larger L2-norm and highly-active nodes are strongly correlated to confident class predictions. Thus, our criterion guides towards dropping a percentage of the most active nodes of the descriptors, proportionally to the estimated class probability. We simultaneously train a per-sample scaling factor to balance the expected output across training and inference. This further allows us to keep high the descriptor's L2-norm, which we show enforces confident predictions. The combination of these two strategies resulted in our “Probability Guided Maxout” solution that acts as a training regularizer. We prove the above behaviors by reporting extensive image classification results on the CIFAR10, CIFAR100, and Caltech256 datasets. Code is available at https://github.com/clferrari/probability-guided-maxout. Claudio Ferrari, Stefano Berretti, Alberto Del Bimbo |
ICPR | 1 |
| 2020 | Inner Eye Canthus Localization for Human Body Temperature ScreeningabstractIn this paper, we propose an automatic approach for localizing the inner eye canthus in thermal face images. We first coarsely detect 5 facial keypoints corresponding to the center of the eyes, the nosetip and the ears. Then we compute a sparse 2D-3D points correspondence using a 3D Morphable Face Model (3DMM). This correspondence is used to project the entire 3D face onto the image, and subsequently locate the inner eye canthus. Detecting this location allows to obtain the most precise body temperature measurement for a person using a thermal camera. We evaluated the approach on a thermal face dataset provided with manually annotated landmarks. However, such manual annotations are normally conceived to identify facial parts such as eyes, nose and mouth, and are not specifically tailored for localizing the eye canthus region. As additional contribution, we enrich the original dataset by using the annotated landmarks to deform and project the 3DMM onto the images. Then, by manually selecting a small region corresponding to the eye canthus, we enrich the dataset with additional annotations. By using the manual landmarks, we ensure the correctness of the 3DMM projection, which can be used as ground-truth for future evaluations. Moreover, we supply the dataset with the 3D head poses and per-point visibility masks for detecting self-occlusions. The data is publicly available at https://www.micc.unifi.it/resources/datasets/thermal-face/. Claudio Ferrari, Lorenzo Berlincioni, Marco Bertini 0001, Alberto Del Bimbo |
ICPR | 1 |
| 2019 | Discovering Identity Specific Activation Patterns in Deep Descriptors for Template Based Face RecognitionabstractThe majority of recent face recognition systems are based on Deep Convolutional Neural Networks (DCNNs). These networks are trained on massive amounts of face images so as to learn a compact representation (deep descriptor) aimed at capturing the identity information. Recognition is then performed by computing some similarity (or distance) measure between descriptors. However, in practice, descriptors encode also other intra-class variabilities such as pose and expressions. This well-known problem is usually addressed by designing specific loss-functions or metric learning modules such that the learned descriptors maximize the inter-class (identity) distances and minimize the intra-class differences in the feature space. We tackle this problem from a different perspective by observing that descriptors associated with images of the same subject, on average, share similar patterns in the highest activation units. We demonstrate this assumption by showing that improved accuracy can be obtained in a template-based recognition scenario by retaining the descriptor bins with the average highest activation, and dropping all the others to zero. These activation patterns are also employed to build identity-representative binary masks that are effectively used in place of the descriptors to match templates. We investigate this strategy by performing experiments on the IJB-A dataset, and show that it can significantly boost the recognition accuracy. Claudio Ferrari, Stefano Berretti, Alberto Del Bimbo |
FG | 1 |
| 2019 | Coarse to Fine 3D Face Reconstruction from Single ImageabstractIn this demo we propose a coarse to fine reconstruction pipeline, which takes a single RGB image as input and outputs a detailed 3D model of the face. The pipeline is composed by two main blocks, the coarse reconstruction block, which is based on a 3D Morphable Model, and the refinement block, which instead grounds on a Generative Adversarial Network (GAN). Leonardo Galteri, Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo |
FG | 2 |
| 2019 | 3D Face Reconstruction from RGB-D Data by Morphable Model to Point Cloud Dense Fittingabstract3D cameras for face capturing are quite common today thanks to their ease of use and affordable cost. The depth information they provide is mainly used to enhance face pose estimation and tracking, and face-background segmentation, while applications that require finer face details are usually not possible due to the low-resolution data acquired by such devices. In this paper, we propose a framework that allows us to derive high-quality 3D models of the face starting from corresponding low-resolution depth sequences acquired with a depth camera. To this end, we start by defining a solution that exploits temporal redundancy in a short-sequence of adjacent depth frames to remove most of the acquisition noise and produce an aggregated point cloud output with intermediate level details. Then, using a 3DMM specifically designed to support local and expression-related deformations of the face, we propose a two-steps 3DMM fitting solution: initially the model is deformed under the effect of landmarks correspondences; subsequently, it is iteratively refined using points closeness updating guided by a mean-square optimization. Preliminary results show that the proposed solution is able to derive 3D models of the face with high visual quality; quantitative results also evidence the superiority of our approach with respect to methods that use one step fitting based on landmarks. Claudio Ferrari, Stefano Berretti, Pietro Pala, Alberto Del Bimbo |
ICPRAM | 1 |
| 2019 | Deep 3D morphable model refinement via progressive growing of conditional Generative Adversarial Networks
Leonardo Galteri, Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo |
Comput. Vis. Image Underst. | 2 |
| 2018 | Extended YouTube Faces: a Dataset for Heterogeneous Open-Set Face IdentificationabstractIn this paper, we propose an extension of the famous YouTube Faces (YTF) dataset. In the YTF dataset, the goal was to state whether two videos contained the same subject or not (video-based face verification). We enrich YTF with still images and an identification protocol. In the classic face identification, given a probe image (or video), the correct identity has to be retrieved among the gallery ones; the main peculiarity of such protocol is that each probe identity has a correspondent in the gallery (closed-set). To resemble a realistic and practical scenario, we devised a protocol in which probe identities are not guaranteed to be in the gallery (open-set). Compared to a closed-set identification, the latter is definitely more challenging in as much as the system needs firstly to reject impostors (i.e., probe identities missing from the gallery), and subsequently, if the probe is accepted as genuine, retrieve the correct identity. In our case, the probe set is composed of full-length videos from the original dataset, while the gallery is composed of templates, i.e., sets of still images. To collect the images, an automatic application was developed. The main motivations behind this work can be found in both the lack of open-set identification protocols defined in the literature and the undeniable complexity of such. We also argued that extending an existing and widely used dataset could make its distribution easier and that data heterogeneity would make the problem even more challenging and realistic. We named the dataset Extended YTF (E-YTF). Finally, we report baseline recognition results using two well known DCNN architectures. Claudio Ferrari, Stefano Berretti, Alberto Del Bimbo |
ICPR | 1 |
| 2018 | Fairness Behind a Veil of Ignorance: A Welfare Analysis for Automated Decision MakingabstractWe draw attention to an important, yet largely overlooked aspect of evaluating fairness for automated decision making systems---namely risk and welfare considerations. Our proposed family of measures corresponds to the long-established formulations of cardinal social welfare in economics, and is justified by the Rawlsian conception of fairness behind a veil of ignorance. The convex formulation of our welfare-based measures of fairness allows us to integrate them as a constraint into any convex loss minimization pipeline. Our empirical analysis reveals interesting trade-offs between our proposal and (a) prediction accuracy, (b) group discrimination, and (c) Dwork et al's notion of individual fairness. Furthermore and perhaps most importantly, our work provides both heuristic justification and empirical evidence suggesting that a lower-bound on our measures often leads to bounded inequality in algorithmic outcomes; hence presenting the first computationally feasible mechanism for bounding individual-level inequality. Hoda Heidari, Claudio Ferrari, Krishna P. Gummadi, Andreas Krause 0001 |
NeurIPS | 2 |
| 2018 | Investigating Nuisances in DCNN-Based Face RecognitionabstractFace recognition "in the wild" has been revolutionized by the deployment of deep learning based approaches. In fact, it has been extensively demonstrated that Deep Convolutional Neural Networks (DCNNs) are powerful enough to overcome most of the limits that affected face recognition algorithms based on hand-crafted features. These include variations in illumination, pose, expression and occlusion, to mention some. The DCNNs discriminative power comes from the fact that low- and high-level representations are learned directly from the raw image data. As a consequence, we expect the performance of a DCNN to be influenced by the characteristics of the image/video data that are fed to the network, and their preprocessing. In this work, we present a thorough analysis of several aspects that impact on the use of DCNN for face recognition. The evaluation has been carried out from two main perspectives: the network architecture and the similarity measures used to compare deeply learned features; the data (source and quality) and their preprocessing (bounding box and alignment). Results obtained on the IJB-A, MegaFace, UMDFaces and YouTube Faces datasets indicate viable hints for designing, training and testing DCNNs. Taking into account the outcomes of the experimental evaluation, we show how competitive performance with respect to the state-of-the-art can be reached even with standard DCNN architectures and pipeline. Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo |
IEEE Trans. Image Process. | 1 |
| 2017 | A Dictionary Learning-Based 3D Morphable Shape ModelabstractFace analysis from 2D images and videos is a central task in many multimedia applications. Methods developed to this end perform either face recognition or facial expression recognition, and in both cases results are negatively influenced by variations in pose, illumination, and resolution of the face. Such variations have a lower impact on 3D face data, which has given the way to the idea of using a 3D morphable model as an intermediate tool to enhance face analysis on 2D data. In this paper, we propose a new approach for constructing a 3D morphable shape model (called DL-3DMM) and show our solution can reach the accuracy of deformation required in applications where fine details of the face are concerned. For constructing the model, we start from a set of 3D face scans with large variability in terms of ethnicity and expressions. Across these training scans, we compute a point-topoint dense alignment, which is accurate also in the presence of topological variations of the face. The DL-3DMM is constructed by learning a dictionary of basis components on the aligned scans. The model is then fitted to 2D target faces using an efficient regularized ridge-regression guided by 2D/3D facial landmark correspondences in order to generate pose-normalized face images. Comparison between the DL-3DMM and the standard PCA-based 3DMM demonstrates that in general a lower reconstruction error can be obtained with our solution. Application to action unit detection and emotion recognition from 2D images and videos shows competitive results with state of the art methods on two benchmark datasets. Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo |
IEEE Trans. Multim. | 1 |
| 2016 | Effective 3D based frontalization for unconstrained face recognitionabstractIn this paper, we propose a new and effective frontalization algorithm for frontal rendering of unconstrained face images, and experiment it for face recognition. Initially, a 3DMM is fit to the image, and an interpolating function maps each pixel inside the face region on the image to the 3D model's. Thus, we can render a frontal view without introducing artifacts in the final image thanks to the exact correspondence between each pixel and the 3D coordinate of the model. The 3D model is then back projected onto the frontalized image allowing us to localize image patches where to extract the feature descriptors, and thus enhancing the alignment between the same descriptor over different images. Our solution outperforms other frontalization techniques in terms of face verification. Results comparable to state-of-the-art on two challenging benchmark datasets are also reported, supporting our claim of effectiveness of the proposed face image representation. Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo |
ICPR | 1 |
| 2015 | Dictionary Learning Based 3D Morphable Model Construction for Face Recognition with Varying Expression and PoseabstractIn this paper, we propose a new approach for constructing a 3D morph able model (3DMM) and experiment its application to face recognition. Differently from existing solutions, the proposed 3DMM is constructed from a training set that includes a large spectrum of variability in terms of ethnicity and facial expressions. By exploiting annotated landmarks available in the training data, we are able of establishing dense correspondence across training scans also in the presence of strong facial expressions. The 3DMM is then constructed by learning a dictionary of basis components, instead of using the traditional approach based on PCA decomposition. Finally, we cast the proposed dictionary learning DL-3DMM to a rigid/non-rigid deformation framework, which includes pose estimation and regularized ridge-regression fitting to 2D images. Comparative results between the DL-3DMM and its PCA counterpart are reported, together with face recognition results for images with large pose and expression variations. Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo |
3DV | 1 |
| 2014 | Pose Independent Face Recognition by Localizing Local Binary Patterns via Deformation ComponentsabstractIn this paper we address the problem of pose independent face recognition with a gallery set containing one frontal face image per enrolled subject while the probe set is composed by just a face image undergoing pose variations. The approach uses a set of aligned 3D models to learn deformation components using a 3D Morph able Model (3DMM). This further allows fitting a 3DMM efficiently on an image using a Ridge regression solution, regularized on the face space estimated via PCA. Then the approach describes each profile face by computing Local Binary Pattern (LBP) histograms localized on each deformed vertex, projected on a rendered frontal view. In the experimental result we evaluate the proposed method on the CMU Multi-PIE to assess face recognition algorithm across pose. We show how our process leads to higher performance than regular baselines reporting high recognition rate considering a range of facial poses in the probe set, up to ±45°. Finally we remark that our approach can handle continuous pose variations and it is comparable with recent state-of-the-art approaches. Iacopo Masi, Claudio Ferrari, Alberto Del Bimbo, Gérard G. Medioni |
ICPR | 2 |