EDBT 2026 Demo / reviewers in the wild / expert
Stefano Berretti
dblp:41/561
· DBLP profile ↗
122ranked-venue papers
32as first author
47since 2021 · last 2026
0000-0003-1219-4386ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 77 · 23 first-author · 23 since 2021Artificial intelligence and machine learning · 56 · 7 first-author · 25 since 2021Computer networks · 8 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Security and privacy · 6 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Standardized Benchmark for Skeleton-Based Rehabilitation Assessment Using Deep LearningabstractAutomated assessment of human motion plays a vital role in rehabilitation, enabling objective evaluation of patient performance and progress. Unlike general human activity recognition, rehabilitation motion assessment focuses on analyzing the quality of movement within the same action class, requiring the detection of subtle deviations from ideal motion. Recent advances in deep learning and video-based skeleton extraction have opened new possibilities for accessible, scalable motion assessment using affordable devices such as smartphones or webcams. However, the field lacks standardized benchmarks, consistent evaluation protocols, and reproducible methodologies, limiting progress and comparability across studies. In this work, we address these gaps by (i) aggregating existing rehabilitation datasets into a unified archive called Rehab-Pile, (ii) proposing a general benchmarking framework for evaluating deep learning methods in this domain, and (iii) conducting extensive benchmarking of multiple architectures across classification and regression tasks. All datasets and implementations are released to the community to support transparency and reproducibility. This paper aims to establish a solid foundation for future research in automated rehabilitation assessment and foster the development of reliable, accessible, and personalized rehabilitation solutions. The datasets, source-code and results of this article are all publicly available. Ali Ismail-Fawaz, Maxime Devanne, Stefano Berretti, Jonathan Weber, Germain Forestier |
FG | 3 |
| 2026 | Polyglot: Multilingual Style Preserving Speech-Driven Facial Animation
Federico Nocentini, Kwanggyoon Seo, Qingju Liu, Claudio Ferrari, Stefano Berretti, David Ferman, Hyeongwoo Kim, Pablo Garrido 0001, Akin Caliskan |
FG | 5 |
| 2026 | SEDTalker: Emotion-Aware 3D Facial Animation Using Frame-Level Speech Emotion Diarization
Farzaneh Jafari, Stefano Berretti, Anup Basu |
ICPR (12) | 2 |
| 2026 | Patch-Based Reconstruction and Multimodal Residual Learning for Generalized Deepfake DetectionabstractDeepfake detectors often achieve near-perfect accuracy in-domain but fail under dataset or forgery shifts, especially when only a few manipulated samples are available for training. We propose a two-stage framework for generalized and data-efficient deepfake detection based on reconstruction residuals. In Stage I, we learn a real-face prior with PM-VAE, a masked patch reconstructor that augments a Masked Autoencoder with a lightweight variational bottleneck to regularize the patch latent space and reduce memorization. In Stage II, the generator is frozen and used to produce forensic evidence from partially observed inputs via block-wise masking on the patch grid, the resulting inpainted reconstructions yield residual cues that are stable and localized. We then train a multi-branch Transformer to fuse (i) RGB context, (ii) spatial residuals, and (iii) wavelet-domain residuals that explicitly capture high-frequency inconsistencies missed by spatial errors alone. Extensive experiments on FaceForensics++ dataset under cross-forgery protocols, on external benchmarks (Celeb-DF, DFD, DFDC) for cross-dataset evaluation, and on synthetic generation (StyleGAN family and Stable Diffusion) show improved robustness over reconstruction baselines, with consistent gains in low-data regimes down to a handful of fake frames per video. Niccolò Marini, Andrea Ciamarra, Roberto Caldelli, Stefano Berretti |
IH&MMSec | 4 |
| 2026 | No MoCap Needed: Post-Training Motion Diffusion Models with Reinforcement Learning using Only Textual PromptsabstractDiffusion models have recently advanced human motion generation, producing realistic and diverse animations from textual prompts. However, adapting these models to unseen actions or styles typically requires additional motion capture data and full retraining, which is costly and difficult to scale. We propose a post-training framework based on Reinforcement Learning that fine-tunes pretrained motion diffusion models using only textual prompts, without requiring any motion ground truth. Our approach employs a pretrained text–motion retrieval network as a reward signal and optimizes the diffusion policy with Denoising Diffusion Policy Optimization, effectively shifting the model’s generative distribution toward the target domain without relying on paired motion data. We evaluate our method on cross-dataset adaptation and leave-one-out motion experiments using the HumanML3D and KIT-ML datasets across both latent- and joint-space diffusion architectures. Results from quantitative metrics and user studies show that our approach consistently improves the quality and diversity of generated motions, while preserving performance on the original distribution. Our approach is a flexible, data-efficient, and privacy-preserving solution for motion adaptation. Code is available on GitHub at https://github.com/miccunifi/NoMoCapNeeded. Girolamo Macaluso, Lorenzo Mandelli, Mirko Bicchierai, Stefano Berretti, Andrew D. Bagdanov |
WACV | 4 |
| 2026 | A video dataset for Multi-Emotion and Interpersonal relation analysis
Hajer Guerdelli, Claudio Ferrari, Stefano Berretti, Walid Barhoumi, Alberto Del Bimbo |
Comput. Vis. Image Underst. | 3 |
| 2026 | Beyond Fixed Topologies: Unregistered Training and Comprehensive Evaluation Metrics for 3D Talking Heads
Federico Nocentini, Thomas Besnier, Claudio Ferrari, Sylvain Arguillère, Mohamed Daoudi, Stefano Berretti |
Int. J. Comput. Vis. | 6 |
| 2026 | DT-VNet: Deep Transformer-Based VNet Framework for 3D Prostate MRI SegmentationabstractMagnetic ResonanceImaging (MRI) is widely used in examining and diagnosing prostate diseases due to its high resolution. However, the diverse morphology of prostate tissue presents a significant challenge for precise gland segmentation. Convolutional Neural Networks have demonstrated effectiveness in segmenting prostate regions. Nevertheless, their limited capability in extracting global long-range semantic features often leads to unstable network segmentation performance. To address these challenges, we propose a Deep Transformer-based Vnet framework (DT-VNet), which consists of a symmetric encoder-decoder architecture that explores global contextual features and retains local feature information. To effectively learn global and local features, We propose the Deep Union Transformer (DU-Trans) as an encoding base module for capturing comprehensive information. Additionally, we introduce a Pool Fusion Attention (PFA) module for decoding, which emphasizes learning context dependencies and interaction relationships. PFA can also facilitate the fusion of deep and shallow features. To our knowledge, this is the first study about deep transformer-based Vnet framework for prostate segmentation. We validate and compare our method on several public datasets against current state-of-the-art methods. The results demonstrate the superior performance of our proposed method in segmenting 3D prostate MRI. Yunyao Cai, Hu Lu, Shengli Wu 0001, Stefano Berretti, Shaohua Wan 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2026 | Guest Editorial: Medical Information Security and Privacy Solution for Smart Healthcare Industriesabstract. Amit Kumar Singh 0001, Stefano Berretti |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | 3D Face Reconstruction Error Decomposed: A Modular Benchmark for Fair and Fast Method EvaluationabstractComputing the standard benchmark metric for 3D face reconstruction, namely geometric error, requires a number of steps, such as mesh cropping, rigid alignment, or point correspondence. Current benchmark tools are monolithic (they implement a specific combination of these steps), even though there is no consensus on the best way to measure error. We present a toolkit for a Modularized 3D Face reconstruction Benchmark (M3DFB), where the fundamental components of error computation are segregated and interchangeable, allowing one to quantify the effect of each. Furthermore, we propose a new component, namely correction, and present a computationally efficient approach that penalizes for mesh topology inconsistency. Using this toolkit, we test 16 error estimators with 10 reconstruction methods on two real and two synthetic datasets. Critically, the widely used ICP-based estimator provides the worst benchmarking performance, as it significantly alters the true ranking of the top- 5 reconstruction methods. Notably, the correlation of ICP with the true error can be as low as 0.41. Moreover, non-rigid alignment leads to significant improvement (correlation larger than 0.90), highlighting the importance of annotating 3D landmarks on datasets. Finally, the proposed correction scheme, together with non-rigid warping, leads to an accuracy on a par with the best non-rigid ICP-based estimators, but runs an order of magnitude faster. Our open-source codebase is designed for researchers to easily compare alternatives for each component, thus helping accelerating progress in benchmarking for 3D face reconstruction and, furthermore, supporting the improvement of learned reconstruction methods, which depend on accurate error estimation for effective training. Evangelos Sariyanidi, Claudio Ferrari, Federico Nocentini, Stefano Berretti, Andrea Cavallaro, Birkan Tunç |
FG | 4 |
| 2025 | Generation of Complex 3D Human Motion by Temporal and Spatial Composition of Diffusion ModelsabstractIn this paper, we address the challenge of generating re-alistic 3D human motions for action classes that were never seen during the training phase. Our approach involves de-composing complex actions into simpler movements, specifically those observed during training, by leveraging the knowledge of human motion contained in GPTs models. These simpler movements are then combined into a single, realistic animation using the properties of diffusion models. Our claim is that this decomposition and subsequent recombination of simple movements can synthesize an animation that accurately represents the complex input action. This method operates during the inference phase and can be in-tegrated with any pre-trained diffusion model, enabling the synthesis of motion classes not present in the training data. We evaluate our method by dividing two benchmark human motion datasets into basic and complex actions, and then compare its performance against the state-of-the-art. Our code and models are publicly available at our github page. Lorenzo Mandelli, Stefano Berretti |
WACV | 2 |
| 2025 | EmoVOCA: Speech-Driven Emotional 3D Talking HeadsabstractA notable challenge in 3D talking head generation consists in blending speech-related motions with expression dynamics. This is primarily caused by the lack of comprehensive 3D datasets that combine diversity in spoken sentences with a variety of facial expressions. Some literature works attempted to overcome such lack of data by fitting parametric 3D models (3DMMs) to 2D videos, and using the reconstructed 3D faces as replacement. However, their underlying parametric space limits the precision required to accurately reproduce convincing lip motions and synching, which is crucial for the application at hand. In this work, we look at the problem from a different perspective, and developed a data-driven technique to combine inexpressive 3D talking heads with a set of 3D expressive sequences, which we used for creating a synthetic dataset, called EmoVOCA. We then designed and trained an emotional 3D talking head generator that accepts a 3D face, an audio file, an emotion label, and an intensity value as inputs, and learns to animate the audio-synchronized lip movements with expressive traits of the face. Comprehensive experiments, both quantitative and qualitative, using our data and generator evidence superior ability in synthesizing convincing animations, when compared with the best performing methods in the literature. Our code and pre-trained models are available at https://github.com/miccunifi/EmoVOCA. Federico Nocentini, Claudio Ferrari, Stefano Berretti |
WACV | 3 |
| 2025 | Establishing a unified evaluation framework for human motion generation: A comparative analysis of metricsabstractThe development of generative artificial intelligence for human motion generation has expanded rapidly, necessitating a unified evaluation framework. This paper presents a detailed review of eight evaluation metrics for human motion generation, highlighting their unique features and shortcomings. We propose standardized practices through a unified evaluation setup to facilitate consistent model comparisons. Additionally, we introduce a novel metric that assesses diversity in temporal distortion by analyzing warping diversity, thereby enhancing the evaluation of temporal data. We also conduct experimental analyses of three generative models using two publicly available datasets, offering insights into the interpretation of each metric in specific case scenarios. Our goal is to offer a clear, user-friendly evaluation framework for newcomers, complemented by publicly accessible code: https://github.com/MSD-IRIMAS/Evaluating-HMG . Ali Ismail-Fawaz, Maxime Devanne, Stefano Berretti, Jonathan Weber, Germain Forestier |
Comput. Vis. Image Underst. | 3 |
| 2025 | A Deep-Learning-Based Traffic Classification Method for 5G Aerial Computing NetworksabstractWith the rapid progress made in aerial computing technology and the increased popularity of fifth-generation (5G) networks, uncrewed aerial vehicles (UAVs) have been playing a crucial role in real-time data collection, processing, and transmission. However, due to the diversity in traffic generated by UAVs in various mission scenarios, there is a significant challenge posed in traffic classification. Therefore, a novel traffic classification model is proposed in this article on the basis of the spatial attention-enhanced convolutional neural network (SAE-CNN). This model proves effective in improving classification accuracy and latency, particularly in the context of various 5G services, such as enhanced mobile broadband (eMBB), ultrareliable low-latency communication (URLLC), and Internet service. Also, a 5G heterogeneous network platform is built to collect UAV-related aerial computing data, with extensive experiments performed to verify the superior performance of the SAE-CNN model compared to other state-of-the-art methods. The experimental results demonstrate that the proposed approach enables effective traffic management and classification for the application of UAV in complex 5G environments. Chen Chen 0006, Ziye Liu, Yuejun Yu, Stefano Berretti, Lei Liu 0031, Qingqi Pei |
IEEE Internet Things J. | 6 |
| 2025 | Guest Editorial: Special Issue on Trends in Social Multimedia Computing: Models, Methodologies, and ApplicationsabstractAlong with the fast development of high-speed networks and advanced wearable and intelligent devices, a large number of multimedia contents are now widely used by social networking sites and content-sharing services as information carriers for various applications [1]. The integration of multimedia and social media, which we call social multimedia, supports new types of user interaction. Motivated by the tremendous growth of social media applications, social computing has emerged as a novel computing paradigm that involves studying and managing social behavior and organizational dynamics to produce intelligent applications [2]. However, the wide prevalence of social multimedia poses a significant challenge for social computing because many new issues involving social activity and interaction around multimedia must be addressed in a media-specific manner. Nevertheless, multimedia research still remains open, given the challenging nature of the research focus in this area. Social multimedia can help improve existing multimedia applications, so the term social multimedia computing to denote the more focused multidisciplinary research and application field between social sciences and multimedia technology. In computational social system, multimedia plays a vital role in comfortable communication and social interaction, and providing valued data for analysis. It enhances user experience, supports diverse content sharing, and drives the development of social multimedia computing. This special issue unites pioneering studies that confront these challenges and chart new directions in social multimedia computing research from quantitative and/or computational perspective. The contributions span innovative models, algorithmic approaches, and emerging applications. This issue focuses on the technical and practical challenges faced in analyzing social media, user-generated content, and multimedia data. Understanding the role of multimedia in social contexts, maintaining data reliability, privacy, and security, and developing artificial intelligence (AI)-based solutions are the key goals in this direction. Together, they advance both the theoretical foundations and practical implementations of social multimedia computing, fostering richer human–machine symbioses and promoting enhanced social well being in multimedia contexts. Amit Kumar Singh 0001, Jungong Han, Stefano Berretti |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | COCALITE: A Hybrid Model COmbining CAtch22 and LITE for Time Series ClassificationabstractTime series classification has achieved significant advancements through deep learning models; however, these models often suffer from high complexity and computational costs. To address these challenges while maintaining effectiveness, we introduce COCALITE, an innovative hybrid model that combines the efficient LITE model with an augmented version incorporating Catch22 features during training. COCALITE operates with only 4.7% of the parameters of the state-of-the-art Inception model, significantly reducing computational overhead. By integrating these complementary approaches, COCALITE leverages both effective feature engineering and deep learning techniques to enhance classification accuracy. Our extensive evaluation across 128 datasets from the UCR archive demonstrates that COCALITE achieves competitive performance, offering a compelling solution for resource-constrained environments. Oumaima Badi, Maxime Devanne, Ali Ismail-Fawaz, Javidan Abdullayev, Vincent Lemaire 0001, Stefano Berretti, Jonathan Weber, Germain Forestier |
IEEE Big Data | 6 |
| 2024 | ScanTalk: 3D Talking Heads from Unregistered Scans
Federico Nocentini, Thomas Besnier, Claudio Ferrari, Sylvain Arguillère, Stefano Berretti, Mohamed Daoudi |
ECCV (29) | 5 |
| 2024 | IMEmo: An Interpersonal Relation Multi-Emotion DatasetabstractWhile engaged in a face-to-face conversation, being capable of understanding the attitude, emotion, and intention of another person allows one guiding his/her behavior establishing a comfortable communication either verbal and non-verbal (i.e., body and face language). This paper introduces the “The IMEmo Interpersonal Multi-Emotion video dataset”, a new in-the-wild dataset of face-to-face interaction, built from movies of romance and drama categories. We manually collected over 100 clips from different movies in different languages. The dataset consists of 79.3 minutes scenes, with a duration of each clip ranging between 0.20 and 2.13 minutes. Each clip contains two people communicating with each other both verbally and with expression and body pose and gestures. Currently, it includes age, gender, emotions, social relationships, actions and valence/arousal annotations for both individuals. Emotion recognition results using a baseline CNN approach are also reported to provide an estimation of the difficulty of the data also in comparison to existing benchmarks. Hajer Guerdelli, Claudio Ferrari, Stefano Berretti, Alberto Del Bimbo |
FG | 3 |
| 2024 | Binary segmentation of relief patterns on point cloudsabstractAnalysis of 3D textures, also known as relief patterns is a challenging task that requires separating repetitive surface patterns from the underlying global geometry. Existing works classify entire surfaces based on one or a few patterns by extracting ad-hoc statistical properties. Unfortunately, these methods are not suitable for objects with multiple geometric textures and perform poorly on more complex shapes. In this paper, we propose a neural network for binary segmentation to infer per-point labels based on the presence of surface relief patterns. We evaluated the proposed architecture on a high resolution point cloud dataset, surpassing the state-of-the-art, while maintaining memory and computation efficiency. Gabriele Paolini, Claudio Tortorici, Stefano Berretti |
Comput. Graph. | 3 |
| 2024 | Generating Multiple 4D Expression Transitions by Learning Face Landmark TrajectoriesabstractIn this work, we address the problem of 4D facial expressions generation. This is usually addressed by animating a neutral 3D face to reach an expression peak, and then get back to the neutral state. In the real world though, people show more complex expressions, and switch from one expression to another. We thus propose a new model that generates transitions between different expressions, and synthesizes long and composed 4D expressions. This involves three sub-problems: (i) modeling the temporal dynamics of expressions, (ii) learning transitions between them, and (iii) deforming a generic mesh. We propose to encode the temporal evolution of expressions using the motion of a set of 3D landmarks, that we learn to generate by training a manifold-valued GAN (Motion3DGAN). To allow the generation of composed expressions, this model accepts two labels encoding the starting and the ending expressions. The final sequence of meshes is generated by a Sparse2Dense mesh Decoder (S2D-Dec) that maps the landmark displacements to a dense, per-vertex displacement of a known mesh topology. By explicitly working with motion trajectories, the model is totally independent from the identity. Extensive experiments on five public datasets show that our proposed approach brings significant improvements with respect to previous solutions, while retaining good generalization to unseen data. Naima Otberdout, Claudio Ferrari, Mohamed Daoudi, Stefano Berretti, Alberto Del Bimbo |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | A High Stability Clustering Scheme for the Internet of VehiclesabstractIn existing research on cluster head selection schemes in the Internet of Vehicles (IoV), designing a stable cluster structure poses a significant challenge. Choosing a centrally-located cluster head that can respond rapidly is crucial for meeting various requirements. To address the aforementioned challenges, this paper introduces a machine learning-based IoV cluster head selection scheme (HSCS). We introduce a new metric termed N-cycle Average Virtual Cluster Delay (XTn) for appropriate cluster head selection. To accommodate the high dynamism of vehicles, a machine learning model is integrated to predict cluster head selection metrics across different periods, and a set of cluster head selection guidelines is formulated. Experimental results demonstrate that our proposed HSCS ensures a relatively low average intra-cluster delay while maintaining a longer cluster head retention time, and it exhibits commendable robustness. Chen Chen 0006, Jiabao Si, Neeraj Kumar 0001, Stefano Berretti, Shaohua Wan 0001 |
IEEE Trans. Netw. Serv. Manag. | 6 |
| 2024 | MS-GDA: Improving Heterogeneous Recipe Representation via Multinomial Sampling Graph Data AugmentationabstractWe study the problem of classifying different cooking styles, based on the recipe. The difficulty is that the same food ingredients, seasoning, and the very similar instructions result in different flavors, with different cooking styles. Existing methods have limitations: they mainly focus on homogeneous data (e.g., instruction or image), ignoring heterogeneous data (e.g., flavor compound or ingredient), which certainly hurts the classification performance. This is because collecting enough available heterogeneous data of a recipe is a non-trivial task. In this paper, we present a new heterogeneous data augmentation method to improve classification performance. Specifically, we first construct a heterogeneous recipe graph network to represent heterogeneous data, which includes four main-stream types of heterogeneous data: ingredient, flavor compound, image, and instruction. Then, we draw a sequence of augmented graphs for Semi-Supervised learning through multinomial sampling. The probability distribution of sampling depends on the Cosine distance between the nodes of graph. In this way, we name our approach as Multinomial Sampling Graph Data Augmentation (MS-GDA). Extensive experiments demonstrate that MS-GDA significantly outperforms SOTA baselines on cuisine classification and region prediction with the recipe benchmark dataset. Code is available at https://github.com/LiangzheChen/MS-GDA . Liangzhe Chen, Wei Li 0121, Xiaohui Cui, Zhenyu Wang 0013, Stefano Berretti, Shaohua Wan 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | LITE: Light Inception with boosTing tEchniques for Time Series ClassificationabstractDeep learning models have been shown to be a powerful solution for Time Series Classification (TSC). State-of-the-art architectures, while conducting promising results on the UCR archive, present a high number of trainable parameters. This can lead to long training with a high CO2, Power consumption and possible increase in the number of FLoat-point Operation Per Second (FLOPS). In this paper, we present a new architecture for TSC, the Light Inception with boosTing tEchnique (LITE) with only 2.34% of the state-of-the-art model InceptionTime’s number of parameters, while preserving performance. This architecture, with only 9, 814 trainable parameters due to the usage of DepthWise Separable Convolutions (DWSC), is boosted by three techniques: multiplexing, custom filters, and dilated convolution. The LITE architecture, trained on the UCR, is 2.78 times faster than InceptionTime and consumes 2.79 times less CO2 and Power. Ali Ismail-Fawaz, Maxime Devanne, Stefano Berretti, Jonathan Weber, Germain Forestier |
DSAA | 3 |
| 2023 | The Florence 4D Facial Expression DatasetabstractHuman facial expressions change dynamically, so their recognition / analysis should be conducted by accounting for the temporal evolution of face deformations either in 2D or 3D. While abundant 2D video data do exist, this is not the case in 3D, where few 3D dynamic (4D) datasets were released for public use. The negative consequence of this scarcity of data is amplified by current deep learning based-methods for facial expression analysis that require large quantities of variegate samples to be effectively trained. With the aim of smoothing such limitations, in this paper we propose a large dataset, named Florence 4D, composed of dynamic sequences of 3D face models, where a combination of synthetic and real identities exhibit an unprecedented variety of 4D facial expressions, with variations that include the classical neutral-apex transition, but generalize to expression-to-expression. All these characteristics are not exposed by any of the existing 4D datasets and they cannot even be obtained by combining more than one dataset. We strongly believe that making such a data corpora publicly available to the community will allow designing and experimenting new applications that were not possible to investigate till now. To show at some extent the difficulty of our data in terms of different identities and varying expressions, we also report a baseline experimentation on the proposed dataset that can be used as baseline. Filippo Principi, Stefano Berretti, Claudio Ferrari, Naima Otberdout, Mohamed Daoudi, Alberto Del Bimbo |
FG | 2 |
| 2023 | Meta-evaluation for 3D Face Reconstruction Via Synthetic DataabstractThe standard benchmark metric for 3D face reconstruction is the geometric error between reconstructed meshes and the ground truth. Nearly all recent reconstruction methods are validated on real ground truth scans, in which case one needs to establish point correspondence prior to error computation, which is typically done with the Chamfer (i.e., nearest neighbor) criterion. However, a simple yet fundamental question have not been asked: Is the Chamfer error an appropriate and fair benchmark metric for 3D face reconstruction? More generally, how can we determine which error estimator is a better benchmark metric? We present a meta-evaluation framework that uses synthetic data to evaluate the quality of a geometric error estimator as a benchmark metric for face reconstruction. Further, we use this framework to experimentally compare four geometric error estimators. Results show that the standard approach not only severely underestimates the error, but also does so inconsistently across reconstruction methods, to the point of even altering the ranking of the compared methods. Moreover, although non-rigid ICP leads to a metric with smaller estimation bias, it could still not correctly rank all compared reconstruction methods, and is significantly more time consuming than Chamfer. In sum, we show several issues present in the current benchmarking and propose a procedure using synthetic data to address these issues. Evangelos Sariyanidi, Claudio Ferrari, Stefano Berretti, Robert T. Schultz, Birkan Tunç |
IJCB | 3 |
| 2023 | Learning graph-based features for relief patterns classification on mesh manifoldsabstractRelief patterns represent a surface characteristic that can be seen as the 3D counterpart of the texture concept in 2D images. Such characteristic is well distinct from the 3D object shape but represents a good information source to recognize the object itself. The majority of state-of-the-art techniques for 2D images rely on convolution-based filtering so, the idea of extending such techniques to the mesh manifold domain is quite intriguing as much as challenging. In this paper, we propose a novel approach based on Graph Neural Networks for 3D mesh relief pattern classification. To this end, we designed a bi-level architecture that learns on data structures computed thanks to a mesh resampling algorithm that allows us to represent local surface patches uniformly, while keeping a consistent points order. The local mesh structures are represented by SpiderPatches, that aim to capture local features of the 3D mesh surface, providing a fine-grained rich representation of the relief patterns; global structures are instead captured by MeshGraphs, whose nodes are SpiderPatches, representing the mesh at a macroscopic level. We tested our architecture using SpiderPatches and MeshGraphs on the original meshes of the SHREC’17 and SHREC’20 relief patterns track datasets, showing superior performance to that reported in the literature using a comparable experimental setting. Niccolò Guiducci, Claudio Tortorici, Claudio Ferrari, Stefano Berretti |
Comput. Graph. | 4 |
| 2023 | Interpersonal relation recognition: a surveyabstractAbstract People spend a considerable amount of their time in social activities, where person-to-person relations are of main relevance. Recently, there has been an increasing research interest in automatically analyzing interpersonal relations, for the social and behavioral implications, and the many practical applications it may have. However, to the best of our knowledge, there is not a systematic study providing a harmonized view of the literature in the field. On this ground, we summarize in our work interpersonal relation recognition datasets and methods aiming to help researchers to have a better understanding of the characteristics of the state-of-the-art. In the proposed study, we distinguish between methods that address objective relations that do not depend on behavior or emotional state, and methods that consider subjective ones that depend on emotions. It turns out quite evidently that aiming at the latter recognition task is more challenging, with the existing methods that provide convincing results only on limited and very specific cases. For both the broad categories, we discuss datasets and methods according to the different behavioural and psychological models used to annotate and classify the data. We conclude our review work, by providing a comprehensive discussion pointing out current limitations and future research perspectives. Hajer Guerdelli, Claudio Ferrari, Stefano Berretti |
Multim. Tools Appl. | 3 |
| 2023 | Correction to: Interpersonal relation recognition: a survey
Hajer Guerdelli, Claudio Ferrari, Stefano Berretti |
Multim. Tools Appl. | 3 |
| 2023 | The Florence multi-resolution 3D facial expression datasetabstractIn the literature, several 3D face datasets have been collected, aiming at advancing the field of 3D face analysis from different perspectives. Data collection generally follows specific research needs, and the existing 3D face datasets all have different characteristics that are tailored for investigating different tasks, encompassing face recognition, facial expressions and emotions analysis, 3D face reconstruction. However, the majority of these datasets are either collected with high-resolution scanners, or consumer level devices, such as the Kinect, the latter being motivated by the burdensome and costly process of collecting high-quality scans. Differently from 2D imagery, the difference in resolution in 3D data represents a non negligible problem that is under-investigated, and still prevents the successful development of methods that can work in real scenarios. In this paper, we propose a new 3D face dataset, named “Florence Multi-Resolution 3D Facial Expression” (Florence 3DMRE), which aims at bridging the gap between high- and low-resolution 3D face datasets. Its peculiarity consists in (1) including high-resolution (HR) models obtained with a HR scanner, and paired samples collected with a Kinect sensor, (2) LR and HR scans are synchronized and capture extreme and asymmetric facial deformations as used in facial rehabilitation exercises. In total, our dataset consists of 14 subjects, each performing 19 complex and asymmetric expressions. For each of them, we collected a high-resolution scan, and an RGB-D sequence. Finally, to highlight the value of the dataset and the challenges it introduces, we use the collected data to perform baseline experiments for cross-resolution 3D face recognition and reconstruction. The dataset is released for research purposes only, and complies to GDPR for data treatment. The dataset can be found at this link. Claudio Ferrari, Stefano Berretti, Pietro Pala, Alberto Del Bimbo |
Pattern Recognit. Lett. | 2 |
| 2023 | FrankenMask: Manipulating semantic masks with transformers for face parts editingabstractIn this paper, we propose FrankenMask, a novel framework that allows swapping and rearranging face parts in semantic masks for automatic editing of shape-related facial attributes. This is a novel yet challenging task as substituting face parts in a semantic mask requires to account for possible spatial misalignment and the adaptation of surrounding regions. We obtain such a feature by combining a Transformer encoder to learn the spatial relationships of facial parts, with an encoder–decoder architecture, which reconstructs a complete mask from the composition of local parts. Reconstruction and attribute classification results demonstrate the effective synthesis of facial images, while showing the generation of accurate and plausible facial attributes. Code is available at https://github.com/TFonta/FrankenMask_semantic. Tomaso Fontanini, Claudio Ferrari, Giuseppe Lisanti, Leonardo Galteri, Stefano Berretti, Massimo Bertozzi, Andrea Prati 0001 |
Pattern Recognit. Lett. | 5 |
| 2023 | Medical Image Encryption by Content-Aware DNA Computing for Secure HealthcareabstractThere exists a rising concern on security of healthcare data and service. Even small lost, stolen, displaced, hacked, or communicated in personal health data could bring huge damage to patients. Therefore, we propose a novel content-aware deoxyribonucleic acid (DNA) computing system to encrypt medical images, thus guaranteeing privacy and promoting secure healthcare environment. The proposed system consists of sender and receiver to perform tasks of encryption and decryption, respectively, where both contain the same structure design, but perform opposite operations. In either sender or receiver, we design a randomly DNA encoding and a content-aware permutation and diffusion module. Considering introducing random mechanism to increase difficulty of cracking, the former module builds a random encryption rule selector in DNA encoding process by randomly mapping quantity of medical image pixels to outputs. Meanwhile, the latter module constructs a permutation sequence, which not only encodes information of pixel values, but also involves redundant correlation between adjacent pixels located in a patch. Such design brings awareness property of medical image content to greatly increase complexity in cracking by embedding semantical information for encryption. We demonstrate that the proposed system successfully improve cybersecurity of medical images against various attacks in robustness and effectiveness when transmitting data in wireless broadcasting scenarios. Yirui Wu, Lilai Zhang, Stefano Berretti, Shaohua Wan 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Learning Streamed Attention Network from Descriptor Images for Cross-Resolution 3D Face RecognitionabstractIn this article, we propose a hybrid framework for cross-resolution 3D face recognition which utilizes a Streamed Attention Network (SAN) that combines handcrafted features with Convolutional Neural Networks (CNNs). It consists of two main stages: first, we process the depth images to extract low-level surface descriptors and derive the corresponding Descriptor Images (DIs), represented as four-channel images. To build the DIs, we propose a variation of the 3D Local Binary Pattern (3DLBP) operator that encodes depth differences using a sigmoid function. Then, we design a CNN that learns from these DIs. The peculiarity of our solution consists in processing each channel of the input image separately, and fusing the contribution of each channel by means of both self- and cross-attention mechanisms. This strategy showed two main advantages over the direct application of Deep-CNN to depth images of the face; on the one hand, the DIs can reduce the diversity between high- and low-resolution data by encoding surface properties that are robust to resolution differences. On the other, it allows a better exploitation of the richer information provided by low-level features, resulting in improved recognition. We evaluated the proposed architecture in a challenging cross-dataset, cross-resolution scenario. To this aim, we first train the network on scanner-resolution 3D data. Next, we utilize the pre-trained network as feature extractor on low-resolution data, where the output of the last fully connected layer is used as face descriptor. Other than standard benchmarks, we also perform experiments on a newly collected dataset of paired high- and low-resolution 3D faces. We use the high-resolution data as gallery, while low-resolution faces are used as probe, allowing us to assess the real gap existing between these two types of data. Extensive experiments on low-resolution 3D face benchmarks show promising results with respect to state-of-the-art methods. João Baptista Cardia Neto, Claudio Ferrari, Aparecido Nilceu Marana, Stefano Berretti, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Sparse to Dense Dynamic 3D Facial Expression GenerationabstractIn this paper, we propose a solution to the task of generating dynamic 3D facial expressions from a neutral 3D face and an expression label. This involves solving two sub-problems: (i) modeling the temporal dynamics of expressions, and (ii) deforming the neutral mesh to obtain the expressive counterpart. We represent the temporal evolution of expressions using the motion of a sparse set of 3D landmarks that we learn to generate by training a manifold-valued GAN (Motion3DGAN). To better encode the expression-induced deformation and disentangle it from the identity information, the generated motion is represented as per-frame displacement from a neutral configuration. To generate the expressive meshes, we train a Sparse2Dense mesh Decoder (S2D-Dec) that maps the landmark displacements to a dense, per-vertex displacement. This allows us to learn how the motion of a sparse set of landmarks influences the deformation of the overall face surface, independently from the identity. Experimental results on the CoMA and D3DFACS datasets show that our solution brings significant improvements with respect to previous solutions in terms of both dynamic expression generation and mesh reconstruction, while retaining good generalization to unseen data. Code and models are available at https://github.com/CRISTAL-3DSAM/Sparse2Dense. Naima Otberdout, Claudio Ferrari, Mohamed Daoudi, Stefano Berretti, Alberto Del Bimbo |
CVPR | 4 |
| 2022 | What makes you, you? Analyzing Recognition by Swapping Face PartsabstractDeep learning advanced face recognition to an unprecedented accuracy. However, understanding how local parts of the face affect the overall recognition performance is still mostly unclear. Among others, face swap has been experimented to this end, but just for the entire face. In this paper, we propose to swap facial parts as a way to disentangle the recognition relevance of different face parts, like eyes, nose and mouth. In our method, swapping parts from a source face to a target one is performed by fitting a 3D prior, which establishes dense pixels correspondence between parts, while also handling pose differences. Seamless cloning is then used to obtain smooth transitions between the mapped source regions and the shape and skin tone of the target face. We devised an experimental protocol that allowed us to draw some preliminary conclusions when the swapped images are classified by deep networks, indicating a prominence of the eyes and eyebrows region. Code available at https://github.com/clferrari/FacePartsSwap Claudio Ferrari, Matteo Serpentoni, Stefano Berretti, Alberto Del Bimbo |
ICPR | 3 |
| 2022 | SPEAKER VGG CCT: Cross-Corpus Speech Emotion Recognition with Speaker Embedding and Vision TransformersabstractIn recent years, Speech Emotion Recognition (SER) has been investigated mainly transforming the speech signal into spectrograms that are then classified using Convolutional Neural Networks pre-trained on generic images and fine tuned with spectrograms. In this paper, we start from the general idea above and develop a new learning solution for SER, which is based on Compact Convolutional Transformers (CCTs) combined with a speaker embedding. With CCTs, the learning power of Vision Transformers (ViT) is combined with a diminished need for large volume of data as made possible by the convolution. This is important in SER, where large corpora of data are usually not available. The speaker embedding allows the network to extract an identity representation of the speaker, which is then integrated by means of a self-attention mechanism with the features that the CCT extracts from the spectrogram. Overall, the solution is capable of operating in real-time showing promising results in a cross-corpus scenario, where training and test datasets are kept separate. Experiments have been performed on several benchmarks in a cross-corpus setting as rarely used in the literature, with results that are comparable or superior to those obtained with state-of-the-art network architectures. Our code is available at https://github.com/JabuMlDev/Speaker-VGG-CCT Alessandro Arezzo, Stefano Berretti |
MMAsia | 2 |
| 2022 | Foreword to the Special Section on 3D Object Retrieval 2022 Symposium (3DOR2022)
Stefano Berretti, Theoharis Theoharis, Mohamed Daoudi, Claudio Ferrari, Remco C. Veltkamp |
Comput. Graph. | 1 |
| 2022 | Deep learning for 3D visionabstractWith the rapid development of 3D imaging sensors, such as depth cameras and laser scanning systems, 3D data has become increasingly accessible. Meanwhile, the boost of various deep learning algorithms, such as convolutional neural networks and transformers, further increases the usability of 3D vision systems. Driven by these factors, 3D vision has become an emerging and core component for numerous applications, such as autonomous driving, augmented reality, virtual reality and robotics. Although remarkable progress has been achieved in this area during the last few years, there are still several challenges that need to be addressed, such as the noisy, sparse, and irregular nature of point clouds, the high cost to label 3D data and the necessity to integrate geometry-based and learning-based techniques. Besides, 3D data produced by different 3D imaging sensors (e.g. structured light, stereo, LiDAR and time-of-flight) can be highly different. It is, therefore, necessary to investigate general algorithms that can mitigate the domain gap between different types of 3D data. This special issue aims to collect and present the latest research development in learning-based 3D vision theories and their applications and to inspire future research in this area. In total, there are eight papers accepted for publication in this special issue through careful peer reviews and revisions. These accepted papers are broadly categorised into three topics, and the summary of each topic is given below. TOPIC A—OPTICAL FLOW AND DEPTH ESTIMATION Han et al., in their paper ‘DEMVSNet: Denoising and Depth Inference for Unstructured Multi-View Stereo on Noised Images’, proposed a DEMVSNet to simultaneously address the depth estimation and image denoising problems for unstructured multi-view stereo. The multi-scales feature maps for each image are wrapped to construct cost volumes containing both the depth and RGB information through differentiable homography and Gaussian probability mapping. The cost volume regularisation module is then adopted to predict the probability of depth and RGB. To avoid overfitting in multi-task learning, the gradient normalisation algorithm is utilised to dynamically fine-tune the weights between the depth prediction task and the denoising task. To evaluate the performance of proposed DEMVSNet, a noisy Technical University of Denmark dataset is generated by adding Gaussian-Poisson noise to each image, and the experimental results demonstrate the superiority of DEMVSNet on both the denoising and multi-view stereo reconstruction tasks. Lin et al., in their paper ‘EAGAN: Event-Based Attention Generative Adversarial Networks for Optical Flow and Depth Estimation’, proposed an event-based attention generative adversarial network named EAGAN to simultaneously deal with optical flow and depth estimation based on monocular event camera. The generator of EAGAN is similar to U-net except that a transformer structure is introduced between the encoder and decoder. The position-coding features learnt from the transformer is added to features learnt from the encoding layer, which helps to capture the correlation between sequence information. The discriminator of EAGAN is based on a fully convolutional network and aims to distinguish whether the depth image or the optical flow image is generated by the generator. Experimental results conducted on the multi-vehicle stereo event camera dataset demonstrate the effectiveness of EAGAN on both the depth and optical flow estimation tasks. TOPIC B—POSE ESTIMATION Gao et al., in their paper ‘Efficient 6D Object Pose Estimation based on Attentive Multi-Scale Contextual Information’, proposed an end-to-end 6D pose estimation network to utilise multi-scale contextual features learnt from two heterogeneous data. First, interesting objects are detected from an RGB-D image using an existing semantic segmentation method. Then, pixel-wise geometric and colour features are learnt from 3D point clouds and 2D images respectively. Next, three pixelwise feature attention mechanism modules are utilised to exploit the inter-channel relationship of multimodal features. Finally, multi-scale features are extracted at three different scales and 6D pose is estimated through a dense regression module. Experimental results conducted on the LineMOD and YCB-Video datasets demonstrate that the proposed method achieves state-of-the-art performances in terms of average point distance and average closest point distance. Liu et al., in their paper ‘Auto Calibration of Multi-Camera System for Human Pose Estimation’, proposed an iterative joint estimation of intrinsic and extrinsic parameters for a multi-camera system. Specifically, keypoints are detected with high confidence to estimate the essential matrix between two cameras, and the valid extrinsic parameters are estimated by assuming that the intrinsic parameters are known a priori. Then, the reconstructed 3D human body coordinates are projected into the pixel coordinate system, and the intrinsic parameters are estimated by minimising the projection errors. The experimental results show that the proposed method achieves better performance than commonly used calibration tools. TOPIC C—POINT CLOUD PROCESSING AND UNDERSTANDING Liu et al., in their paper ‘Point Cloud Completion by Dynamic Transformer with Adaptive Neighbourhood Feature Fusion’, utilised the adaptive neighbourhood feature extraction (ANE) module and genetic hierarchical point generation (GHG) module to accomplish the point cloud completion task. The ANE module selects k nearest points both in the spatial and feature spaces adaptively according to different target shapes. The GHG module generates finer point clouds hierarchically according to the local shape characteristics, and the shape information of current points is transferred to the next stage through a dynamic transformer structure. The experimental results conducted on the Point Completion Network and Completion3D datasets demonstrate the superiority of the proposed method. Wang et al., in their paper ‘PCCN-RE: Point Cloud Colourisation Network Based on Relevance Embedding’, proposed a highly authentic point cloud colorisation network based on conditional generative adversarial (cGANs) networks. The generator network predicts the colours from the coordinates of each point, while the discriminator utilises the coordinates and the generated colours to determine the reality of input colourised point clouds. Three key components are contained in the generator. Specifically, the relevance embedding structure captures the most related local information, the weighted pooling structure aggregates the local features based on the correlation values of the covariance matrix, and the enhanced spatial transform network keeps the point clouds invariant to the geometric transformations based on weighted pooling and maximal pooling. The experimental results show that the proposed method achieves the highest Peak Signal to Noise Ratio and Structural Similarity Index on the ShapeNetCore dataset. Fang et al., in their paper ‘Sparse Point-Voxel Aggregation Network for Efficient Point Cloud Semantic Segmentation’, proposed a sparse point-voxel aggregation network to overcome high computational costs in the point cloud semantic segmentation task. In the encoding layer, the local context features are learnt through a sparse convolutional network performed on the voxelised point cloud, and the individual point features are learnt through multi-layer perceptron (MLP)-based network performed on the original point cloud. In the decoding layer, these two kinds of features are aggregated at different encoding layers through simple MLP layers. The experimental results show that the proposed method achieves state-of-the-art performance on the SemanticKITTI and S3DIS datasets. Wang et al., in their paper ‘Scale Robust Point Matching-Net: End-to-End Scale Point Matching Using Lie Group’, proposed an end-to-end scale point cloud matching network named SRPM-Net based on Lie Group. The extracted pointwise features are composed of point absolute coordinates, relative coordinates and point pair features of neighbouring points, and the local context features are aggregated through an attentive pooling layer. The matching matrix is computed via the exponential map of Lie group, which represents the feature similarity of points in two point clouds. The final transformation estimation problem is transferred as estimating the coefficients of the Lie algebra optimisation problem and is optimised through an iterative linear optimisation approach. The experimental results show that SRPM-Net achieves the best performance on the ModelNet40 and Stanford 3D scanning datasets. SUMMARY/CONCLUSION The papers published in this Special Issue show that traditional topics, such as optical flow and depth estimation, pose estimation, and point cloud processing have developed very fast in recent years. In addition, many topics have emerged in deep learning-based 3D vision, such as multi-task joint learning and multimodality intelligence. Future research in this field is expected to boost the theoretical development and potential applications of 3D vision. Yulan Guo, Hanyun Wang, Ronald Clark, Stefano Berretti, Mohammed Bennamoun |
IET Comput. Vis. | 4 |
| 2022 | A Sparse and Locally Coherent Morphable Face Model for Dense Semantic Correspondence Across Heterogeneous 3D FacesabstractThe 3D Morphable Model (3DMM) is a powerful statistical tool for representing 3D face shapes. To build a 3DMM, a training set of face scans in full point-to-point correspondence is required, and its modeling capabilities directly depend on the variability contained in the training data. Thus, to increase the descriptive power of the 3DMM, establishing a dense correspondence across heterogeneous scans with sufficient diversity in terms of identities, ethnicities, or expressions becomes essential. In this manuscript, we present a fully automatic approach that leverages a 3DMM to transfer its dense semantic annotation across raw 3D faces, establishing a dense correspondence between them. We propose a novel formulation to learn a set of sparse deformation components with local support on the face that, together with an original non-rigid deformation algorithm, allow the 3DMM to precisely fit unseen faces and transfer its semantic annotation. We extensively experimented our approach, showing it can effectively generalize to highly diverse samples and accurately establish a dense correspondence even in presence of complex facial expressions. The accuracy of the dense registration is demonstrated by building a heterogeneous, large-scale 3DMM from more than 9,000 fully registered scans obtained by joining three large datasets together. Claudio Ferrari, Stefano Berretti, Pietro Pala, Alberto Del Bimbo |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Dynamic Facial Expression Generation on Hilbert Hypersphere With Conditional Wasserstein Generative Adversarial NetsabstractIn this work, we propose a novel approach for generating videos of the six basic facial expressions given a neutral face image. We propose to exploit the face geometry by modeling the facial landmarks motion as curves encoded as points on a hypersphere. By proposing a conditional version of manifold-valued Wasserstein generative adversarial network (GAN) for motion generation on the hypersphere, we learn the distribution of facial expression dynamics of different classes, from which we synthesize new facial expression motions. The resulting motions can be transformed to sequences of landmarks and then to images sequences by editing the texture information using another conditional Generative Adversarial Network. To the best of our knowledge, this is the first work that explores manifold-valued representations with GAN to address the problem of dynamic facial expression generation. We evaluate our proposed approach both quantitatively and qualitatively on two public datasets; Oulu-CASIA and MUG Facial Expression. Our experimental results demonstrate the effectiveness of our approach in generating realistic videos with continuous motion, realistic appearance and identity preservation. We also show the efficiency of our framework for dynamic facial expressions generation, dynamic facial expression transfer and data augmentation for training improved emotion recognition models. Naima Otberdout, Mohamed Daoudi, Anis Kacem 0001, Lahoucine Ballihi, Stefano Berretti |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Automatic Estimation of Self-Reported Pain by Trajectory Analysis in the Manifold of Fixed Rank Positive Semi-Definite MatricesabstractWe propose an automatic method to estimate self-reported pain based on facial landmarks extracted from videos. For each video sequence, we decompose the face into four different regions and the pain intensity is measured by modeling the dynamics of facial movement using the landmarks of these regions. A formulation based on Gram matrices is used for representing the trajectory of landmarks on the Riemannian manifold of symmetric positive semi-definite matrices of fixed rank. A curve fitting algorithm is used to smooth the trajectories and temporal alignment is performed to compute the similarity between the trajectories on the manifold. A Support Vector Regression classifier is then trained to encode extracted trajectories into pain intensity levels consistent with self-reported pain intensity measurement. Finally, a late fusion of the estimation for each region is performed to obtain the final predicted pain level. The proposed approach is evaluated on two publicly available datasets, the UNBCMcMaster Shoulder Pain Archive and the Biovid Heat Pain dataset. We compared our method to the state-of-the-art on both datasets using different testing protocols, showing the competitiveness of the proposed approach. Benjamin Szczapa, Mohamed Daoudi, Stefano Berretti, Pietro Pala, Alberto Del Bimbo, Zakia Hammal |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | A Psychologically Inspired Fuzzy Cognitive Deep Learning Framework to Predict Crowd BehaviorabstractIn an intelligent surveillance system, detecting and predicting diverse collective crowd behaviors has emerged as a challenging problem for efficient crowd management. In real-world scenarios, potential disasters and hazards can be averted by considering crowd psychology for predicting crowd behaviors. This article proposes an approach that exploits the psychological and cognitive aspects of human behavior in determining nine diverse crowd behaviors. The proposed approach is a combination of two cognitive deep learning frameworks and a psychological fuzzy computational model that utilizes OCC theory of emotions, OCEAN five-factor model of personality and visual attention for detecting crowd behaviors. Experiments are performed on different datasets and the results prove that our approach is successful in detecting and predicting crowd behavior in confronting situations and also outperforms the state-of-the-art methods. In particular, considering psychological aspects and cognition in determining crowd behavior is beneficial for rectifying the semantic ambiguity in identifying crowd behaviors. Elizabeth B. Varghese, Sabu M. Thampi, Stefano Berretti |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | Guest Editorial: Medical Data Security Solution for Healthcare IndustriesabstractSince smart healthcare systems are highly connected to advanced wearable devices, internet of things (IoT) and mobile internet, valuable patient information and other significant medical records are easily transmitted over the public network. The patient information and clinical records are also stored on the existing databases and local servers of hospitals and healthcare centres. These materials not only provide a reference for healthcare professionals to make correct decisions on the patients, but also provide a strong basis for other professionals to undertake effective treatment and develop plans for correct diagnosis. Furthermore, the databases may be used by various research communities for research, without any possibility of privacy violations. However, leaking of healthcare data is highly likely. Therefore, medical data security is becoming very important in smart healthcare. Amit Kumar Singh 0001, Huiyu Zhou 0001, Stefano Berretti |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | Guest Editorial Emerging IoT-Driven Smart Health: From Cloud to EdgeabstractThe papers in this special section focus on emerging Internet of Medical Things. Recent advances in advances in healthcare can be experienced with the development of smart sensorial things, Artificial Intelligence (AI), Machine Learning (ML), Deep Learning (DL), edge computing, Edge AI, 6G, cloud computing, and connected healthcare have attracted a great deal of attention and a wide range of views. However, the need to deliver real-time and accurate healthcare services to patients, while reducing costs is a challenging issue [1]. Especially, COVID-19 has recently demonstrated the importance of fast, comprehensive, and accurate intelligent healthcare involving different types of medical, physiological, and epidemiological investigation data to diagnose the virus. Smart health is a real-time, intelligent, ubiquitous healthcare service based on Internet of bioMedical Things (IoMT). With the rapid development of related technologies such as deep learning, edge computing and IoT, smart health is playing vital role in healthcare industry to increase the accuracy, reliability, and productivity of mobile sensory devices. Shaohua Wan 0001, Michele Nappi, Chen Chen 0001, Stefano Berretti |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Measuring 3D face deformations from RGB images of expression rehabilitation exercisesabstractThe accurate (quantitative) analysis of face deformations in 3D is a problem of increasing interest for the many applications it may have. In particular, defining a 3D model of the face that can deform to a 2D target image, while capturing local and asymmetric deformations is still a challenge in the existing literature. Computing a measure of such local deformations may represent a relevant index for monitoring rehabilitation exercises that are used in Parkinson’s and Alzheimer’s disease or in recovering from a stroke. In this study, we present a complete framework that allows the construction of a 3D Morphable Shape Model (3DMM) of the face and its fitting to a target RGB image. The model has the specific characteristic of being based on localized components of deformation; the fitting transformation is performed from 3D to 2D and is guided by the correspondence between landmarks detected in the target image and landmarks manually annotated on the average 3DMM. The fitting has also the peculiarity of being performed in two steps, disentangling face deformations that are due to the identity of the target subject from those induced by facial actions. In the experimental validation of the method, we used the MICC-3D dataset that includes 11 subjects each acquired in one neutral pose plus 18 facial actions that deform the face in localized and asymmetric ways. For each acquisition, we fit the 3DMM to an RGB frame with an apex facial action and to the neutral frame, and computed the extent of the deformation. Results indicated that the proposed approach can accurately capture the face deformation even for localized and asymmetric ones. The proposed framework proved the idea of measuring the deformations of a reconstructed 3D face model to monitor the facial actions performed in response to a set of target ones. Interestingly, these results were obtained just using RGB targets without the need for 3D scans captured with costly devices. This opens the way to the use of the proposed tool for remote medical monitoring of rehabilitation. Claudio Ferrari, Stefano Berretti, Pietro Pala, Alberto Del Bimbo |
Virtual Real. Intell. Hardw. | 2 |
| 2021 | Foreword to the Special Section on Smart Tool and Applications for Graphics (STAG 2020)
Ruggero Pintus, Silvia Biasotti, Stefano Berretti |
Comput. Graph. | 3 |
| 2021 | Representing and analyzing relief patterns using LBP variants on mesh manifold
Claudio Tortorici, Naoufel Werghi, Stefano Berretti |
Pattern Anal. Appl. | 3 |
| 2021 | Convolution operations for relief-pattern retrieval, segmentation and classification on mesh manifolds
Claudio Tortorici, Stefano Berretti, Ahmad Obeid 0001, Naoufel Werghi |
Pattern Recognit. Lett. | 2 |
| 2020 | Modelling the Statistics of Cyclic Activities by Trajectory Analysis on the Manifold of Positive-Semi-Definite MatricesabstractIn this paper, a model is presented to extract statistical summaries to characterize the repetition of a cyclic body action, for instance a gym exercise, for the purpose of checking the compliance of the observed action to a template one and highlighting the parts of the action that are not correctly executed (if any). The proposed system relies on a Riemannian metric to compute the distance between two poses in such a way that the geometry of the manifold where the pose descriptors lie is preserved; a model to detect the begin and end of each cycle; a model to temporally align the poses of different cycles so as to accurately estimate the cross-sectional mean and variance of poses across different cycles. The proposed model is demonstrated using gym videos taken from the Internet. Ettore Maria Celozzi, Luca Ciabini, Luca Cultrera, Pietro Pala, Stefano Berretti, Mohamed Daoudi, Alberto Del Bimbo |
FG | 5 |
| 2020 | Fused Geometry Augmented Images For Analyzing Textured MeshabstractIn this paper, we propose a multi-modal mesh surface representation by fusing texture and geometric data. Our approach defines an inverse mapping between different geometric descriptors computed on the mesh surface, and the corresponding 2D texture image of the mesh, allowing the construction of fused geometrically augmented images. This new fused modality enables us to learn feature representations from 3D data in a highly efficient manner by employing standard convolutional neural networks in a transfer-learning mode. In contrast to existing methods, the proposed approach is both computationally and memory efficient, preserves intrinsic geometric information and learns highly discriminative feature representations by effectively fusing shape and texture information at the data level. The efficacy is demonstrated on the task of facial expression classification, showing competitive performance with state-of-the-art methods. Bilal Taha, Munawar Hayat, Stefano Berretti, Naoufel Werghi |
ICIP | 3 |
| 2020 | CSIOR: An Algorithm For Ordered Triangular Mesh Regularizationabstract3D scanners generate irregularly distributed cloud of points in most of the cases. Dealing with such data, often in the form of triangular meshes, requires a pre-processing step to regularize the triangle facets shape and size. In this paper, we propose CSIOR, a novel mesh regularization technique which is capable of producing quasi-equilateral triangles, and distinguished by two novel features, namely, its intrinsic ordered aspect and its preservation of the geometric texture of the surface (relief patterns). We evidence the superiority of our technique over current methods through a series of experiments performed on a variety of geometric textured surfaces. Claudio Tortorici, Naoufel Werghi, Stefano Berretti |
ICIP | 3 |
| 2020 | Probability Guided MaxoutabstractIn this paper, we propose an original CNN training strategy that brings together ideas from both dropout-like regularization methods and solutions that learn discriminative features. We propose a dropping criterion that, differently from dropout and its variants, is deterministic rather than random. It grounds on the empirical evidence that feature descriptors with larger L2-norm and highly-active nodes are strongly correlated to confident class predictions. Thus, our criterion guides towards dropping a percentage of the most active nodes of the descriptors, proportionally to the estimated class probability. We simultaneously train a per-sample scaling factor to balance the expected output across training and inference. This further allows us to keep high the descriptor's L2-norm, which we show enforces confident predictions. The combination of these two strategies resulted in our “Probability Guided Maxout” solution that acts as a training regularizer. We prove the above behaviors by reporting extensive image classification results on the CIFAR10, CIFAR100, and Caltech256 datasets. Code is available at https://github.com/clferrari/probability-guided-maxout. Claudio Ferrari, Stefano Berretti, Alberto Del Bimbo |
ICPR | 2 |
| 2020 | Automatic Estimation of Self-Reported Pain by Interpretable Representations of Motion DynamicsabstractWe propose an automatic method for pain intensity measurement from video. For each video, pain intensity was measured using the dynamics of facial movement using 66 facial points. Gram matrices formulation was used for facial points trajectory representations on the Riemannian manifold of symmetric positive semi-definite matrices of fixed rank. Curve fitting and temporal alignment were then used to smooth the extracted trajectories. A Support Vector Regression model was then trained to encode the extracted trajectories into ten pain intensity levels consistent with the Visual Analogue Scale for pain intensity measurement. The proposed approach was evaluated using the UNBC McMaster Shoulder Pain Archive and was compared to the state-of-the-art on the same data. Using both 5-fold cross-validation and leave-one-subject-out cross-validation, our results are competitive with respect to state-of-the-art methods. Benjamin Szczapa, Mohamed Daoudi, Stefano Berretti, Pietro Pala, Alberto Del Bimbo, Zakia Hammal |
ICPR | 3 |
| 2020 | CSIOR: Circle-Surface Intersection Ordered Resampling
Claudio Tortorici, Mohamed Kamel Riahi, Stefano Berretti, Naoufel Werghi |
Comput. Aided Geom. Des. | 3 |
| 2020 | SHREC 2020: Retrieval of digital surfaces with similar geometric reliefs
Elia Moscoso Thompson, Silvia Biasotti, Andrea Giachetti 0001, Claudio Tortorici, Naoufel Werghi, Ahmad Obeid 0001, Stefano Berretti, Hoang-Phuc Nguyen-Dinh, Minh-Quan Le, Hai-Dang Nguyen, Minh-Triet Tran, Leonardo Gigli, Santiago Velasco-Forero, Beatriz Marcotegui, Ivan Sipiran, Benjamin Bustos, Ioannis Romanelis, Vlassis Fotis, Ramamoorthy Luxman |
Comput. Graph. | 7 |
| 2020 | A Novel Geometric Framework on Gram Matrix Trajectories for Human Behavior UnderstandingabstractIn this paper, we propose a novel space-time geometric representation of human landmark configurations and derive tools for comparison and classification. We model the temporal evolution of landmarks as parametrized trajectories on the Riemannian manifold of positive semidefinite matrices of fixed-rank. Our representation has the benefit to bring naturally a second desirable quantity when comparing shapes-the spatial covariance-in addition to the conventional affine-shape representation. We derived then geometric and computational tools for rate-invariant analysis and adaptive re-sampling of trajectories, grounding on the Riemannian geometry of the underlying manifold. Specifically, our approach involves three steps: (1) landmarks are first mapped into the Riemannian manifold of positive semidefinite matrices of fixed-rank to build time-parameterized trajectories; (2) a temporal warping is performed on the trajectories, providing a geometry-aware (dis-)similarity measure between them; (3) finally, a pairwise proximity function SVM is used to classify them, incorporating the (dis-)similarity measure into the kernel function. We show that such representation and metric achieve competitive results in applications as action recognition and emotion recognition from 3D skeletal data, and facial expression recognition from videos. Experiments have been conducted on several publicly available up-to-date benchmarks. Anis Kacem 0001, Mohamed Daoudi, Boulbaba Ben Amor, Stefano Berretti, Juan Carlos Álvarez Paiva |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Learned 3D Shape Representations Using Fused Geometrically Augmented Images: Application to Facial Expression and Action Unit DetectionabstractIn this paper, we propose an approach to learn generic multi-modal mesh surface representations using a novel scheme for fusing texture and geometric data. Our approach defines an inverse mapping between different geometric descriptors computed on the mesh surface or its down-sampled version, and the corresponding 2D texture image of the mesh, allowing the construction of fused geometrically augmented images (FGAI). This new fused modality enables us to learn feature representations from 3D data in a highly efficient manner by simply employing standard CNNs in a transfer-learning mode. The proposed approach is both computationally and memory efficient, preserves intrinsic geometric information and learns highly discriminative feature representations by effectively fusing shape and texture information at data level. The efficacy of our approach is demonstrated for the tasks of facial action unit detection and expression classification. The extensive experiments conducted on the Bosphorus and BU-4DFE datasets show that our method produces a significant boost in the performance when compared to state-of-the-art solutions. Bilal Taha, Munawar Hayat, Stefano Berretti, Dimitrios Hatzinakos, Naoufel Werghi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Automatic Analysis of Facial Expressions Based on Deep Covariance TrajectoriesabstractIn this article, we propose a new approach for facial expression recognition (FER) using deep covariance descriptors. The solution is based on the idea of encoding local and global deep convolutional neural network (DCNN) features extracted from still images, in compact local and global covariance descriptors. The space geometry of the covariance matrices is that of symmetric positive definite (SPD) matrices. By conducting the classification of static facial expressions using a support vector machine (SVM) with a valid Gaussian kernel on the SPD manifold, we show that deep covariance descriptors are more effective than the standard classification with fully connected layers and softmax. Besides, we propose a completely new and original solution to model the temporal dynamic of facial expressions as deep trajectories on the SPD manifold. As an extension of the classification pipeline of covariance descriptors, we apply SVM with valid positive definite kernels derived from global alignment for deep covariance trajectories classification. By performing extensive experiments on the Oulu-CASIA, CK+, static facial expression in the wild (SFEW), and acted facial expressions in the wild (AFEW) data sets, we show that both the proposed static and dynamic approaches achieve the state-of-the-art performance for FER outperforming many recent approaches. Naima Otberdout, Anis Kacem 0001, Mohamed Daoudi, Lahoucine Ballihi, Stefano Berretti |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2019 | Discovering Identity Specific Activation Patterns in Deep Descriptors for Template Based Face RecognitionabstractThe majority of recent face recognition systems are based on Deep Convolutional Neural Networks (DCNNs). These networks are trained on massive amounts of face images so as to learn a compact representation (deep descriptor) aimed at capturing the identity information. Recognition is then performed by computing some similarity (or distance) measure between descriptors. However, in practice, descriptors encode also other intra-class variabilities such as pose and expressions. This well-known problem is usually addressed by designing specific loss-functions or metric learning modules such that the learned descriptors maximize the inter-class (identity) distances and minimize the intra-class differences in the feature space. We tackle this problem from a different perspective by observing that descriptors associated with images of the same subject, on average, share similar patterns in the highest activation units. We demonstrate this assumption by showing that improved accuracy can be obtained in a template-based recognition scenario by retaining the descriptor bins with the average highest activation, and dropping all the others to zero. These activation patterns are also employed to build identity-representative binary masks that are effectively used in place of the descriptors to match templates. We investigate this strategy by performing experiments on the IJB-A dataset, and show that it can significantly boost the recognition accuracy. Claudio Ferrari, Stefano Berretti, Alberto Del Bimbo |
FG | 2 |
| 2019 | Coarse to Fine 3D Face Reconstruction from Single ImageabstractIn this demo we propose a coarse to fine reconstruction pipeline, which takes a single RGB image as input and outputs a detailed 3D model of the face. The pipeline is composed by two main blocks, the coarse reconstruction block, which is based on a 3D Morphable Model, and the refinement block, which instead grounds on a Generative Adversarial Network (GAN). Leonardo Galteri, Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo |
FG | 4 |
| 2019 | Extending LBP and Convolution-Like Operations on the MeshabstractExtending the concept of texture to the geometry of a mesh manifold surface is an emerging topic in image processing. This concept is different from gluing images to the surface, but rather indicates the presence of relief patterns that locally change the surface geometry, showing some regular and repetitive pattern. In this paper, we propose an efficient and effective framework to address this novel task, which encompasses the convolution operation and the casting of a variety of Local Binary Pattern on the mesh manifold. Results show that our technique outperforms the existing state-of-the-art methods in the challenging task of relief patterns classification. Claudio Tortorici, Naoufel Werghi, Stefano Berretti |
ICIP | 3 |
| 2019 | 3D Face Reconstruction from RGB-D Data by Morphable Model to Point Cloud Dense Fittingabstract3D cameras for face capturing are quite common today thanks to their ease of use and affordable cost. The depth information they provide is mainly used to enhance face pose estimation and tracking, and face-background segmentation, while applications that require finer face details are usually not possible due to the low-resolution data acquired by such devices. In this paper, we propose a framework that allows us to derive high-quality 3D models of the face starting from corresponding low-resolution depth sequences acquired with a depth camera. To this end, we start by defining a solution that exploits temporal redundancy in a short-sequence of adjacent depth frames to remove most of the acquisition noise and produce an aggregated point cloud output with intermediate level details. Then, using a 3DMM specifically designed to support local and expression-related deformations of the face, we propose a two-steps 3DMM fitting solution: initially the model is deformed under the effect of landmarks correspondences; subsequently, it is iteratively refined using points closeness updating guided by a mean-square optimization. Preliminary results show that the proposed solution is able to derive 3D models of the face with high visual quality; quantitative results also evidence the superiority of our approach with respect to methods that use one step fitting based on landmarks. Claudio Ferrari, Stefano Berretti, Pietro Pala, Alberto Del Bimbo |
ICPRAM | 2 |
| 2019 | Enhanced skeleton and face 3D data for person re-identification from depth cameras
Pietro Pala, Lorenzo Seidenari, Stefano Berretti, Alberto Del Bimbo |
Comput. Graph. | 3 |
| 2019 | Deep 3D morphable model refinement via progressive growing of conditional Generative Adversarial Networks
Leonardo Galteri, Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo |
Comput. Vis. Image Underst. | 4 |
| 2019 | Reconstructing 3D Face Models by Incremental Aggregation and Refinement of Depth FramesabstractFace recognition from two-dimensional (2D) still images and videos is quite successful even with “in the wild” conditions. Instead, less consolidated results are available for the cases in which face data come from non-conventional cameras, such as infrared or depth. In this article, we investigate this latter scenario assuming that a low-resolution depth camera is used to perform face recognition in an uncooperative context. To this end, we propose, first, to automatically select a set of frames from the depth sequence of the camera because they provide a good view of the face in terms of pose and distance. Then, we design a progressive refinement approach to reconstruct a higher-resolution model from the selected low-resolution frames. This process accounts for the anisotropic error of the existing points in the current 3D model and the points in a newly acquired frame so that the refinement step can progressively adjust the point positions in the model using a Kalman-like estimation. The quality of the reconstructed model is evaluated by considering the error between the reconstructed models and their corresponding high-resolution scans used as ground truth. In addition, we performed face recognition using the reconstructed models as probes against a gallery of reconstructed models and a gallery with high-resolution scans. The obtained results confirm the possibility to effectively use the reconstructed models for the face recognition task. Pietro Pala, Stefano Berretti |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2018 | Deep Covariance Descriptors for Facial Expression Recognition
Naima Otberdout, Anis Kacem 0001, Mohamed Daoudi, Lahoucine Ballihi, Stefano Berretti |
BMVC | 5 |
| 2018 | Extended YouTube Faces: a Dataset for Heterogeneous Open-Set Face IdentificationabstractIn this paper, we propose an extension of the famous YouTube Faces (YTF) dataset. In the YTF dataset, the goal was to state whether two videos contained the same subject or not (video-based face verification). We enrich YTF with still images and an identification protocol. In the classic face identification, given a probe image (or video), the correct identity has to be retrieved among the gallery ones; the main peculiarity of such protocol is that each probe identity has a correspondent in the gallery (closed-set). To resemble a realistic and practical scenario, we devised a protocol in which probe identities are not guaranteed to be in the gallery (open-set). Compared to a closed-set identification, the latter is definitely more challenging in as much as the system needs firstly to reject impostors (i.e., probe identities missing from the gallery), and subsequently, if the probe is accepted as genuine, retrieve the correct identity. In our case, the probe set is composed of full-length videos from the original dataset, while the gallery is composed of templates, i.e., sets of still images. To collect the images, an automatic application was developed. The main motivations behind this work can be found in both the lack of open-set identification protocols defined in the literature and the undeniable complexity of such. We also argued that extending an existing and widely used dataset could make its distribution easier and that data heterogeneity would make the problem even more challenging and realistic. We named the dataset Extended YTF (E-YTF). Finally, we report baseline recognition results using two well known DCNN architectures. Claudio Ferrari, Stefano Berretti, Alberto Del Bimbo |
ICPR | 2 |
| 2018 | Spontaneous Expression Detection from 3D Dynamic Sequences by Analyzing Trajectories on Grassmann ManifoldsabstractIn this paper, we propose a framework for online spontaneous emotion detection, such as happiness or physical pain, from depth videos. Our approach consists on mapping the video streams onto a Grassmann manifold (i.e., space of k-dimensional linear subspaces) to form time-parameterized trajectories. To this end, depth videos are decomposed into short-time subsequences, each approximated by a k-dimensional linear subspace, which is in turn a point on the Grassmann manifold. Then, the temporal evolution of subspaces gives rise to a precise mathematical representation of trajectories on the underlying manifold. In the final step, extracted spatio-temporal features based on computing the velocity vectors along the trajectories, termed Geometric Motion History (GMH), are fed to an early event detector based on Structured Output SVM, which enables online emotion detection from partially-observed data. Experimental results obtained on the publicly available Cam3D Kinect and BP4D-spontaneous databases validate the proposed solution. The first database has served to exemplify the proposed framework using depth sequences of the upper part of the body collected using depth-consumer cameras, while the second database allowed the application of the same framework to physical pain detection from high-resolution and long 3D-face sequences. Taleb Alashkar, Boulbaba Ben Amor, Mohamed Daoudi, Stefano Berretti |
IEEE Trans. Affect. Comput. | 4 |
| 2018 | Investigating Nuisances in DCNN-Based Face RecognitionabstractFace recognition "in the wild" has been revolutionized by the deployment of deep learning based approaches. In fact, it has been extensively demonstrated that Deep Convolutional Neural Networks (DCNNs) are powerful enough to overcome most of the limits that affected face recognition algorithms based on hand-crafted features. These include variations in illumination, pose, expression and occlusion, to mention some. The DCNNs discriminative power comes from the fact that low- and high-level representations are learned directly from the raw image data. As a consequence, we expect the performance of a DCNN to be influenced by the characteristics of the image/video data that are fed to the network, and their preprocessing. In this work, we present a thorough analysis of several aspects that impact on the use of DCNN for face recognition. The evaluation has been carried out from two main perspectives: the network architecture and the similarity measures used to compare deeply learned features; the data (source and quality) and their preprocessing (bounding box and alignment). Results obtained on the IJB-A, MegaFace, UMDFaces and YouTube Faces datasets indicate viable hints for designing, training and testing DCNNs. Taking into account the outcomes of the experimental evaluation, we show how competitive performance with respect to the state-of-the-art can be reached even with standard DCNN architectures and pipeline. Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo |
IEEE Trans. Image Process. | 3 |
| 2018 | Table of Contents: Online Supplement Volume 14, Number 1sabstracteditorial Free Access Share on Table of Contents: Online Supplement Volume 14, Number 1s Author: Stefano Berretti 0000-0003-1219-4386View Profile Authors Info & Claims ACM Transactions on Multimedia Computing, Communications, and ApplicationsVolume 14Issue 2May 2018 Article No.: 42pp 1–4https://doi.org/10.1145/3226039Published:22 May 2018Publication History 5citation206DownloadsMetricsTotal Citations5Total Downloads206Last 12 Months26Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Stefano Berretti |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | Introduction to the Special Issue on Representation, Analysis, and Recognition of 3D HumansabstractNo abstract available. Stefano Berretti, Mohamed Daoudi, Pavan Turaga, Anup Basu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | Representation, Analysis, and Recognition of 3D Humans: A SurveyabstractComputer Vision and Multimedia solutions are now offering an increasing number of applications ready for use by end users in everyday life. Many of these applications are centered for detection, representation, and analysis of face and body. Methods based on 2D images and videos are the most widespread, but there is a recent trend that successfully extends the study to 3D human data as acquired by a new generation of 3D acquisition devices. Based on these premises, in this survey, we provide an overview on the newly designed techniques that exploit 3D human data and also prospect the most promising current and future research directions. In particular, we first propose a taxonomy of the representation methods, distinguishing between spatial and temporal modeling of the data. Then, we focus on the analysis and recognition of 3D humans from 3D static and dynamic data, considering many applications for body and face. Stefano Berretti, Mohamed Daoudi, Pavan Turaga, Anup Basu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2017 | Motion segment decomposition of RGB-D sequences for human behavior understanding
Maxime Devanne, Stefano Berretti, Pietro Pala, Hazem Wannous, Mohamed Daoudi, Alberto Del Bimbo |
Pattern Recognit. | 2 |
| 2017 | A Dictionary Learning-Based 3D Morphable Shape ModelabstractFace analysis from 2D images and videos is a central task in many multimedia applications. Methods developed to this end perform either face recognition or facial expression recognition, and in both cases results are negatively influenced by variations in pose, illumination, and resolution of the face. Such variations have a lower impact on 3D face data, which has given the way to the idea of using a 3D morphable model as an intermediate tool to enhance face analysis on 2D data. In this paper, we propose a new approach for constructing a 3D morphable shape model (called DL-3DMM) and show our solution can reach the accuracy of deformation required in applications where fine details of the face are concerned. For constructing the model, we start from a set of 3D face scans with large variability in terms of ethnicity and expressions. Across these training scans, we compute a point-topoint dense alignment, which is accurate also in the presence of topological variations of the face. The DL-3DMM is constructed by learning a dictionary of basis components on the aligned scans. The model is then fitted to 2D target faces using an efficient regularized ridge-regression guided by 2D/3D facial landmark correspondences in order to generate pose-normalized face images. Comparison between the DL-3DMM and the standard PCA-based 3DMM demonstrates that in general a lower reconstruction error can be obtained with our solution. Application to action unit detection and emotion recognition from 2D images and videos shows competitive results with state of the art methods on two benchmark datasets. Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo |
IEEE Trans. Multim. | 3 |
| 2016 | Learning shape variations of motion trajectories for gait analysisabstractThe analysis of human gait is more and more investigated due to its large panel of potential applications in various domains, like rehabilitation, deficiency diagnosis, surveillance and movement optimization. In addition, the release of depth sensors offers new opportunities to achieve gait analysis in a non-intrusive context. In this paper, we propose a gait analysis method from depth sequences by analyzing separately each step so as to be robust to gait duration and incomplete cycles. We analyze the shape of the motion trajectory as signature of the gait and consider shape variations within a Riemannian manifold to learn step models. During classification, the derivation of each performed step is evaluated in an online manner to qualitatively analyze the gait. Experiments are carried out in the context of abnormal gait detection and person re-identification trough gait recognition. Results demonstrated the potential of the method in both scenarios. Maxime Devanne, Hazem Wannous, Mohamed Daoudi, Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
ICPR | 4 |
| 2016 | Effective 3D based frontalization for unconstrained face recognitionabstractIn this paper, we propose a new and effective frontalization algorithm for frontal rendering of unconstrained face images, and experiment it for face recognition. Initially, a 3DMM is fit to the image, and an interpolating function maps each pixel inside the face region on the image to the 3D model's. Thus, we can render a frontal view without introducing artifacts in the final image thanks to the exact correspondence between each pixel and the 3D coordinate of the model. The 3D model is then back projected onto the frontalized image allowing us to localize image patches where to extract the feature descriptors, and thus enhancing the alignment between the same descriptor over different images. Our solution outperforms other frontalization techniques in terms of face verification. Results comparable to state-of-the-art on two challenging benchmark datasets are also reported, supporting our claim of effectiveness of the proposed face image representation. Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo |
ICPR | 3 |
| 2016 | A Grassmann framework for 4D facial shape analysis
Taleb Alashkar, Boulbaba Ben Amor, Mohamed Daoudi, Stefano Berretti |
Pattern Recognit. | 4 |
| 2016 | Reconstructing High-Resolution Face Models From Kinect Depth SequencesabstractPerforming face recognition across 3D scans with different resolution is now attracting an increasing interest thanks to the introduction of a new generation of depth cameras, capable of acquiring color/depth images over time. In fact, these devices acquire and provide depth data with much lower resolution compared with the 3D high-resolution scanners typically used for face recognition applications. If data are acquired without user cooperation, the problem is even more challenging, and the gap of resolution between probe and gallery scans can yield to a severe loss in terms of recognition accuracy. Based on these premises, we propose a method to build a higher resolution 3D face model from 3D data acquired by a low-resolution scanner. This face model is built using data acquired when a person passes in front of the scanner, without assuming any particular cooperation. The 3D data are registered and filtered by combining a model of the expected distribution of the acquisition error with a variant of the lowess method to remove outliers and build the final face model. The proposed approach is evaluated in terms of accuracy of face reconstruction and face recognition. Enrico Bondi, Pietro Pala, Stefano Berretti, Alberto Del Bimbo |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2016 | Boosting 3D LBP-Based Face Recognition by Fusing Shape and Texture Descriptors on the MeshabstractIn this paper, we present a novel approach for fusing shape and texture local binary patterns (LBPs) on a mesh for 3D face recognition. Using a recently proposed framework, we compute LBP directly on the face mesh surface, then we construct a grid of the regions on the facial surface that can accommodate global and partial descriptions. Compared with its depth-image counterpart, our approach is distinguished by the following features: 1) inherits the intrinsic advantages of mesh surface (e.g., preservation of the full geometry); 2) does not require normalization; and 3) can accommodate partial matching. In addition, it allows early level fusion of texture and shape modalities. Through experiments conducted on the BU-3DFE and Bosphorus databases, we assess different variants of our approach with regard to facial expressions and missing data, also in comparison to the state-of-the-art solutions. Naoufel Werghi, Claudio Tortorici, Stefano Berretti, Alberto Del Bimbo |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2015 | Dictionary Learning Based 3D Morphable Model Construction for Face Recognition with Varying Expression and PoseabstractIn this paper, we propose a new approach for constructing a 3D morph able model (3DMM) and experiment its application to face recognition. Differently from existing solutions, the proposed 3DMM is constructed from a training set that includes a large spectrum of variability in terms of ethnicity and facial expressions. By exploiting annotated landmarks available in the training data, we are able of establishing dense correspondence across training scans also in the presence of strong facial expressions. The 3DMM is then constructed by learning a dictionary of basis components, instead of using the traditional approach based on PCA decomposition. Finally, we cast the proposed dictionary learning DL-3DMM to a rigid/non-rigid deformation framework, which includes pose estimation and regularized ridge-regression fitting to 2D images. Comparative results between the DL-3DMM and its PCA counterpart are reported, together with face recognition results for images with large pose and expression variations. Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo |
3DV | 3 |
| 2015 | Representing 3D texture on mesh manifolds for retrieval and recognition applicationsabstractIn this paper, we present and experiment a novel approach for representing texture of 3D mesh manifolds using local binary patterns (LBP). Using a recently proposed framework [37], we compute LBP directly on the mesh surface, either using geometric or photometric appearance. Compared to its depth-image counterpart, our approach is distinguished by the following features: a) inherits the intrinsic advantages of mesh surface (e.g., preservation of the full geometry); b) does not require normalization; c) can accommodate partial matching. In addition, it allows early-level fusion of the geometry and photometric texture modalities. Through experiments conducted on two application scenarios, namely, 3D texture retrieval and 3D face recognition, we assess the effectiveness of the proposed solution with respect to state of the art approaches. Naoufel Werghi, Claudio Tortorici, Stefano Berretti, Alberto Del Bimbo |
CVPR | 3 |
| 2015 | Boosting 3D LBP-based face recognition by fusing shape and texture descriptors on the meshabstractIn this paper, we present a novel approach for fusing shape and texture local binary patterns (LBP) for 3D face recognition. Using the framework proposed in [1], we compute LBP directly on the face mesh surface, then we construct a grid of the regions on the facial surface that can accommodate global and partial descriptions. Compared to its depth-image counterpart, our approach is distinguished by the following features: a) inherits the intrinsic advantages of mesh surface; b) does not require normalization; c) can accommodate partial matching. In addition, it allows early-level fusion of texture and shape modalities. Through experiments conducted on the BU-3DFE and Bosphorus databases, we assess different variants of our approach with regard to facial expressions and missing data. Claudio Tortorici, Naoufel Werghi, Stefano Berretti |
ICIP | 3 |
| 2015 | Local binary patterns on triangular meshes: Concept and applications
Naoufel Werghi, Claudio Tortorici, Stefano Berretti, Alberto Del Bimbo |
Comput. Vis. Image Underst. | 3 |
| 2015 | 3-D Human Action Recognition by Shape Analysis of Motion Trajectories on Riemannian ManifoldabstractRecognizing human actions in 3-D video sequences is an important open problem that is currently at the heart of many research domains including surveillance, natural interfaces and rehabilitation. However, the design and development of models for action recognition that are both accurate and efficient is a challenging task due to the variability of the human pose, clothing and appearance. In this paper, we propose a new framework to extract a compact representation of a human action captured through a depth sensor, and enable accurate action recognition. The proposed solution develops on fitting a human skeleton model to acquired data so as to represent the 3-D coordinates of the joints and their change over time as a trajectory in a suitable action space. Thanks to such a 3-D joint-based framework, the proposed solution is capable to capture both the shape and the dynamics of the human body, simultaneously. The action recognition problem is then formulated as the problem of computing the similarity between the shape of trajectories in a Riemannian manifold. Classification using k-nearest neighbors is finally performed on this manifold taking advantage of Riemannian geometry in the open curve shape space. Experiments are carried out on four representative benchmarks to demonstrate the potential of the proposed solution in terms of accuracy/latency for a low-latency action recognition. Comparative results with state-of-the-art methods are reported. Maxime Devanne, Hazem Wannous, Stefano Berretti, Pietro Pala, Mohamed Daoudi, Alberto Del Bimbo |
IEEE Trans. Cybern. | 3 |
| 2015 | The Mesh-LBP: A Framework for Extracting Local Binary Patterns From Discrete ManifoldsabstractIn this paper, we present a novel and original framework, which we dubbed mesh-local binary pattern (LBP), for computing local binary-like-patterns on a triangular-mesh manifold. This framework can be adapted to all the LBP variants employed in 2D image analysis. As such, it allows extending the related techniques to mesh surfaces. After describing the foundations, the construction and the main features of the mesh-LBP, we derive its possible variants and show how they can extend most of the 2D-LBP variants to the mesh manifold. In the experiments, we give evidence of the presence of the uniformity aspect in the mesh-LBP, similar to the one noticed in the 2D-LBP. We also report repeatability experiments that confirm, in particular, the rotation-invariance of mesh-LBP descriptors. Furthermore, we analyze the potential of mesh-LBP for the task of 3D texture classification of triangular-mesh surfaces collected from public data sets. Comparison with state-of-the-art surface descriptors, as well as with 2D-LBP counterparts applied on depth images, also evidences the effectiveness of the proposed framework. Finally, we illustrate the robustness of the mesh-LBP with respect to the class of mesh irregularity typical to 3D surface-digitizer scans. Naoufel Werghi, Stefano Berretti, Alberto Del Bimbo |
IEEE Trans. Image Process. | 2 |
| 2014 | 3D Face Recognition by Functional Data Analysis
Dania Porro-Muñoz, Francisco Silva-Mata, Anier Revilla-Eng, Isneri Talavera-Bustamante, Stefano Berretti |
CIARP | 5 |
| 2014 | Computing Local Binary Patterns on Discrete ManifoldsabstractIn this paper, we present a novel and original framework for computing Local Binary Pattern (LBP)-like patterns on a triangular mesh manifold. This framework, that we called mesh-LBP, can be adapted to all the LBP variants employed in 2D image analysis. As such, it allows extending the related techniques to mesh surfaces. First, we describe the foundations, the construction and the features of the mesh-LBP. In the experiments, we first show evidence of the presence of the "uniformity" aspect in the mesh-LBP patterns. Then, we report about the application of mesh-LBP to the problem of 3D texture-classification in comparison to standard 3D surface descriptors and show the mesh-LBP robustness to mesh irregularities. Naoufel Werghi, Stefano Berretti, Alberto Del Bimbo |
ICPR | 2 |
| 2014 | 4-D Facial Expression Recognition by Learning Geometric DeformationsabstractIn this paper, we present an automatic approach for facial expression recognition from 3-D video sequences. In the proposed solution, the 3-D faces are represented by collections of radial curves and a Riemannian shape analysis is applied to effectively quantify the deformations induced by the facial expressions in a given subsequence of 3-D frames. This is obtained from the dense scalar field, which denotes the shooting directions of the geodesic paths constructed between pairs of corresponding radial curves of two faces. As the resulting dense scalar fields show a high dimensionality, Linear Discriminant Analysis (LDA) transformation is applied to the dense feature space. Two methods are then used for classification: 1) 3-D motion extraction with temporal Hidden Markov model (HMM) and 2) mean deformation capturing with random forest. While a dynamic HMM on the features is trained in the first approach, the second one computes mean deformations under a window and applies multiclass random forest. Both of the proposed classification schemes on the scalar fields showed comparable results and outperformed earlier studies on facial expression recognition from 3-D video sequences. Boulbaba Ben Amor, Hassen Drira, Stefano Berretti, Mohamed Daoudi, Anuj Srivastava |
IEEE Trans. Cybern. | 3 |
| 2014 | Face Recognition by Super-Resolved 3D Models From Consumer Depth CamerasabstractFace recognition based on the analysis of 3D scans has been an active research subject over the last few years. However, the impact of the resolution of 3D scans on the recognition process has not been addressed explicitly, yet being an element of primal importance to enable the use of the new generation of consumer depth cameras for biometric purposes. In fact, these devices perform depth/color acquisition over time at standard frame-rate, but with a low resolution compared to the 3D scanners typically used for acquiring 3D faces in recognition applications. Motivated by these considerations, in this paper, we define a super-resolution approach for 3D faces by which a sequence of low-resolution 3D face scans is processed to extract a higher resolution 3D face model. The proposed solution relies on the scaled iterative closest point procedure to align the low-resolution scans with each other, and estimates the value of the high-resolution 3D model through a 2D box-spline functions approximation. To evaluate the approach, we built-and made it publicly available-the Florence Superface dataset that collects high-resolution and low-resolution data for about 50 different persons. Qualitative and quantitative results are reported to demonstrate the accuracy of the proposed solution, also in comparison with alternative techniques. Stefano Berretti, Pietro Pala, Alberto Del Bimbo |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2014 | Selecting stable keypoints and local descriptors for person identification using 3D face scans
Stefano Berretti, Naoufel Werghi, Alberto Del Bimbo, Pietro Pala |
Vis. Comput. | 1 |
| 2013 | Local descriptors matching for 3D face recognitionabstractAn original solution to 3D face recognition, which supports face matching also in the case of probes with varying expressions and missing parts is proposed in this work. Distinguishing traits of the face are captured by first extracting 3D keypoints of the face scan, then measuring how the face surface changes in the neighborhood of the keypoints using a local descriptor. To this end, an adaptation of the meshDOG detector to the case of 3D faces is proposed, together with a multi-ring geometric histogram descriptor. Face similarity is then evaluated by comparing local keypoint descriptors across inlier pairs of matching keypoints between probe and gallery scans. Experiments have been performed on the Bosphorus database, showing competitive results with respect to existing solutions for 3D face biometrics. Naoufel Werghi, Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
ICIP | 2 |
| 2013 | Matching 3D face scans using interest points and local histogram descriptorsabstractIn this work, we propose and experiment an original solution to 3D face recognition that supports face matching also in the case of probe scans with missing parts. In the proposed approach, distinguishing traits of the face are captured by first extracting 3D keypoints of the scan and then measuring how the face surface changes in the keypoints neighborhood using local shape descriptors. In particular: 3D keypoints detection relies on the adaptation to the case of 3D faces of the meshDOG algorithm that has been demonstrated to be effective for 3D keypoints extraction from generic objects; as 3D local descriptors we used the HOG descriptor and also proposed two alternative solutions that develop, respectively, on the histogram of orientations and the geometric histogram descriptors. Face similarity is evaluated by comparing local shape descriptors across inlier pairs of matching keypoints between probe and gallery scans. The face recognition accuracy of the approach has been first experimented on the difficult probes included in the new 2D/3D Florence face dataset that has been recently collected and released at the University of Firenze, and on the Binghamton University 3D facial expression dataset. Then, a comprehensive comparative evaluation has been performed on the Bosphorus, Gavab and UND/FRGC v2.0 databases, where competitive results with respect to existing solutions for 3D face biometrics have been obtained. Stefano Berretti, Naoufel Werghi, Alberto Del Bimbo, Pietro Pala |
Comput. Graph. | 1 |
| 2013 | Sparse Matching of Salient Facial Curves for Recognition of 3-D Faces With Missing PartsabstractIn this work, we propose and experiment a 3-D face recognition approach capable of performing accurate face matching also in the case where just parts of probe scans are available. This is obtained through an original face representation and matching solution that first extracts keypoints of the 3-D depth image of the face and then measures how the face depth changes along facial curves connecting pairs of keypoints. Face similarity is evaluated by sparse comparison of facial curves defined across inlier pairs of matching keypoints between probe and gallery scans. In doing so, a statistical model is also proposed to associate facial curves of the gallery scans with a saliency measure so that curves that model characterizing traits of some subjects are distinguished from curves that are frequently observed in the face of many different subjects. Following recent related work, the recognition accuracy of the approach is experimented using two datasets, both comprising scans with missing parts: the Face Recognition Grand Challenge v2.0 dataset combined with the University of Notre Dame probes; the Gavab dataset. Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2013 | Automatic facial expression recognition in real-time from dynamic sequences of 3D face scans
Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
Vis. Comput. | 1 |
| 2012 | 3D dynamic expression recognition based on a novel Deformation Vector Field and Random Forest
Hassen Drira, Boulbaba Ben Amor, Mohamed Daoudi, Anuj Srivastava, Stefano Berretti |
ICPR | 5 |
| 2012 | Distinguishing Facial Features for Ethnicity-Based 3D Face RecognitionabstractAmong different approaches for 3D face recognition, solutions based on local facial characteristics are very promising, mainly because they can manage facial expression variations by assigning different weights to different parts of the face. However, so far, a few works have investigated the individual relevance that local features play in 3D face recognition with very simple solutions applied in the practice. In this article, a local approach to 3D face recognition is combined with a feature selection model to study the relative relevance of different regions of the face for the purpose of discriminating between different subjects. The proposed solution is experimented using facial scans of the Face Recognition Grand Challenge dataset. Results of the experimentation are two-fold: they quantitatively demonstrate the assumption that different regions of the face have different relevance for face discrimination and also show that the relevance of facial regions changes for different ethnic groups. Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2011 | Shape analysis of local facial patches for 3D facial expression recognition
Ahmed Maalej, Boulbaba Ben Amor, Mohamed Daoudi, Anuj Srivastava, Stefano Berretti |
Pattern Recognit. | 5 |
| 2011 | 3D facial expression recognition using SIFT descriptors of automatically detected keypoints
Stefano Berretti, Boulbaba Ben Amor, Mohamed Daoudi, Alberto Del Bimbo |
Vis. Comput. | 1 |
| 2010 | A Set of Selected SIFT Features for 3D Facial Expression RecognitionabstractIn this paper, the problem of person-independent facial expression recognition is addressed on 3D shapes. To this end, an original approach is proposed that computes SIFT descriptors on a set of facial landmarks of depth images, and then selects the subset of most relevant features. Using SVM classification of the selected features, an average recognition rate of 77.5% on the BU-3DFE database has been obtained. Comparative evaluation on a common experimental setup, shows that our solution is able to obtain state of the art results. Stefano Berretti, Alberto Del Bimbo, Pietro Pala, Boulbaba Ben Amor, Mohamed Daoudi |
ICPR | 1 |
| 2010 | Local 3D Shape Analysis for Facial Expression RecognitionabstractWe investigate the problem of facial expression recognition using 3D face data. Our approach is based on local shape analysis of several relevant regions of a given face scan. These regions or patches from facial surfaces are extracted and represented by sets of closed curves. A Riemannian framework is used to derive the shape analysis of the extracted patches. The applied framework permits to calculate a similarity (or dissimilarity) distances between patches, and to compute the optimal deformation between them. Once calculated, these measures are employed as inputs to a commonly used classification techniques such as AdaBoost and Support Vector Machines (SVM). A quantitative evaluation of our novel approach is conducted on a subset of the publicly available BU-3DFE database. Ahmed Maalej, Boulbaba Ben Amor, Mohamed Daoudi, Anuj Srivastava, Stefano Berretti |
ICPR | 5 |
| 2010 | 3D Face Recognition Using Isogeodesic StripesabstractIn this paper, we present a novel approach to 3D face matching that shows high effectiveness in distinguishing facial differences between distinct individuals from differences induced by nonneutral expressions within the same individual. The approach takes into account geometrical information of the 3D face and encodes the relevant information into a compact representation in the form of a graph. Nodes of the graph represent equal width isogeodesic facial stripes. Arcs between pairs of nodes are labeled with descriptors, referred to as 3D Weighted Walkthroughs (3DWWs), that capture the mutual relative spatial displacement between all the pairs of points of the corresponding stripes. Face partitioning into isogeodesic stripes and 3DWWs together provide an approximate representation of local morphology of faces that exhibits smooth variations for changes induced by facial expressions. The graph-based representation permits very efficient matching for face recognition and is also suited to being employed for face identification in very large data sets with the support of appropriate index structures. The method obtained the best ranking at the SHREC 2008 contest for 3D face recognition. We present an extensive comparative evaluation of the performance with the FRGC v2.0 data set and the SHREC08 data set. Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2009 | 3D Mesh decomposition using Reeb graphs
Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
Image Vis. Comput. | 1 |
| 2008 | Face recognition by SVMS classification of 2D and 3D Radial GeodesicsabstractAn original approach to represent 2D and 3D faces using radial geodesic distances (RGDs) is proposed in this work. In 3D, the RGD of a generic point of the face surface is computed as the length of the geodesic connecting the point with a reference point along a radial direction. In 2D, the RGD of a pixel with respect to a reference pixel accounts for the difference of gray level intensities of the two pixels and the Euclidean distance between them. Support Vector Machines (SVMs) are used to perform face recognition using 2D- and 3D-RGDs. Due to the high dimensionality of face representations based on RGDs, embedding into lower-dimensional spaces is applied before SVMs classification. Experimental results are reported for 3D-3D and 2D-3D face recognition using the proposed approach. Stefano Berretti, Alberto Del Bimbo, Pietro Pala, Francisco Silva-Mata |
ICME | 1 |
| 2008 | SHREC'08 entry: 3D face recognition using integral shape informationabstractIn this work, we shortly describe an original 3D face recognition approach and its performance as resulted from the 3D Shape Retrieval Contest of 3D Face Scans organized by SHREC 2008 with the support of the Network of Excellence AIM@SHAPE. In particular, the evaluation shows that the proposed approach attains the highest performance on the SHREC 2008 data set. Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
Shape Modeling International | 1 |
| 2007 | Geodesic Distances for 3D-3D and 2D-3D Face RecognitionabstractIn this paper, we propose an original framework for representing 2D and 3D face information using geodesic distances. This aims to define a representation enabling the direct comparison between 2D face images of an individual against its 3D face model. This representation is extracted by measuring geodesic distances in 2D and 3D. In 3D, the geodesic distance between two points on a surface is computed as the length of the shortest path connecting the two points. In 2D, the geodesic distance between two pixels is computed based on the differences of gray level intensities along the segment connecting the two pixels. Experimental results are shown to demonstrate the viability of the proposed solution. Stefano Berretti, Alberto Del Bimbo, Pietro Pala, Francisco Silva-Mata |
ICME | 1 |
| 2006 | 3D Face Identification Based on Arrangement of Salient WrinklesabstractIn this paper, we propose an original framework for three dimensional face representation and matching for identification purposes. Basic traits of a face are encoded by extracting curves of salient ridges and ravines from the surface of a dense mesh. A compact graph representation is then extracted from these curves through an original modeling technique capable to quantitatively measure spatial relationships between curves in a three dimensional space. In this way, face recognition is obtained by matching 3D graph representations of faces. Experimental results on a 3D face database show that the proposed solution attains high recognition accuracy and is quite robust to facial expression and pose changes Gianni Antini, Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
ICME | 2 |
| 2005 | Retrieval of 3D objects using curvature correlogramsabstractAlong with images and videos, 3D models have raised a certain interest for a number of reasons, including advancements in 3D hardware and software technologies, their ever decreasing prices and increasing availability, affordable 3D authoring tools, and the establishment of open standards for 3D data interchange. The resulting proliferation of 3D models demands for tools supporting their effective and efficient management, including archival and retrieval. In order to support effective retrieval by content of 3D objects and enable retrieval by object parts, information about local object structure should be combined with spatial information on object surface. In this paper, as a solution to this requirement, we present a method relying on curvature correlograms to perform description and retrieval by content of 3D objects. Experimental results are presented both to show results of sample queries by content and to compare-in terms of precision/recall figures-the proposed solution to alternative techniques. Gianni Antini, Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
ICME | 2 |
| 2005 | 3D Mesh Partitioning for Retrieval by Parts ApplicationsabstractA solution for part segmentation of 3D objects is proposed in this paper. The approach is targeted to identify salient visual parts of a mesh by determining its main protrusions and discarding, at the same time, parts originated by un-relevant local properties. This is obtained by first breaking the 3D mesh into seed regions according to the sum of geodesic distances between vertices, then by using topological and curvature information to refine the number of regions and their boundaries. In so doing, effective segmentation is regarded as a prerequisite to enable retrieval of 3D objects based on similarity of parts. Experimental results show the applicability of the proposed solution to complex shapes and its effectiveness in the identification of object parts. Gianni Antini, Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
ICME | 2 |
| 2004 | Shape representation by spatial partitioning for content based retrieval applicationsabstractShape representations for classification or retrieval purposes, have been extensively investigated, but only a few methods have tried to represent shapes as extended entities without reducing them to their boundary profiles or to synthetic geometric descriptors. An original solution for shape representation is proposed; it relies on a modelling technique originally developed to express directional spatial relationships between extended spatial entities. The representation selects a discrete set of points with respect to which a relationship matrix is computed, accounting for the spatial distribution of shape pixels. This is accomplished at different levels of resolution by a tree based structural representation. Properties of the representation and a measure of shape similarity are discussed. The efficiency and effectiveness of the proposed solution have also been assessed in the context of content based retrieval applications, through an experimental evaluation using a shape collection. Stefano Berretti, Gianpaolo D'Amico, Alberto Del Bimbo |
ICME | 1 |
| 2004 | Merging Results for Distributed Content Based Image Retrieval
Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
Multim. Tools Appl. | 1 |
| 2003 | Merging results of distributed image librariesabstractExploitation of information repositories available on the Internet requires users to separately query each repository and manually gather retrieved results. Such a solution could be simplified by using a centralized server that acts as a gateway between the user and repositories: the centralized server forwards the user query to federated repositories and fuses retrieved documents for presentation to the user. To perform these tasks efficiently, the centralized server should perform two main functions: resource selection and data fusion. The former is required to forward the user query only to the repositories that are candidate to contain relevant documents. The latter is used to gather all retrieved documents and conveniently arrange them for presentation to the user. In the case of image repositories, data fusion is particularly challenging owing to the difficulty to normalize document scores returned by different repositories. In this paper a novel solution is presented for fusion of results returned by different image repositories. Experimental results are presented that show the potential of the proposed approach. Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
ICME | 1 |
| 2003 | MIND: resource selection and data fusion in multimedia distributed digital librariesabstractNo abstract available. Stefano Berretti, Jamie Callan, Henrik Nottelmann, Xiao Mang Shou, Shengli Wu 0001 |
SIGIR | 1 |
| 2003 | Weighted walkthroughs between extended entities for retrieval by spatial arrangementabstractIn the access to image databases, queries based on the appearing visual features of searched data reduce the gap between the user and the engineering representation. To support this access modality, image content can be modeled in terms of different types of features such as shape, texture, color, and spatial arrangement. An original framework is presented which supports quantitative nonsymbolic representation and comparison of the mutual positioning of extended nonrectangular spatial entities. Properties of the model are expounded to develop an efficient computation technique and to motivate and assess a metric of similarity for quantitative comparison of spatial relationships. Representation and comparison of binary relationships between entities is then embedded into a graph-theoretical framework supporting representation and comparison of the spatial arrangements of a picture. Two prototype applications are described. Stefano Berretti, Alberto Del Bimbo, Enrico Vicario |
IEEE Trans. Multim. | 1 |
| 2002 | Using indexing structures for resource descriptors extraction from distributed image repositoriesabstractContent based retrieval from distributed libraries raises new and challenging issues with respect to retrieval from a single repository. In particular, an effective management of distributed libraries develops upon three main processes: resource description (extraction of descriptors that qualify the content of a given archive), resource selection (given a user query, analyze resource descriptions and select the resources that contain relevant documents) and results merging (organize and present items returned by individual libraries). So far, these issues have been mainly addressed for text archives. We present a solution to resource descriptors extraction, developing on the use of techniques for multidimensional data indexing. In particular, we implement and compare the extraction of resource descriptors computed through two different indexing approaches; namely m-tree indexing and fuzzy clustering. Comparative results are presented for a test database of about 1000 images. Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
ICME (2) | 1 |
| 2002 | Spatial arrangement of color in retrieval by visual similarity
Stefano Berretti, Alberto Del Bimbo, Enrico Vicario |
Pattern Recognit. | 1 |
| 2001 | Description and retrieval of 3D cellular structuresabstractRecent advances in management of multimedia digital libraries enable effective retrieval of information in the form of audio, image and video. Many archives of 3D objects already exist and are expected to grow both in relevance and size. However, retrieval of information in the form of 3D objects has received limited attention. We address the problem of effective description and retrieval of 3D data representing intracellular structures. These structures are represented as image stacks, where an image stack is constituted by a set of 2D images representing sections of a cellular body at different heights. In the proposed approach 2D visual feature descriptors and hidden Markov models are combined to obtain a representation model which is able to distinguish such intracellular structures as Golgi, nucleus, endoplasmic reticulum and lysosomes. Preliminary results are presented to show the effectiveness of the proposed representation model. Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
ICIP (1) | 1 |
| 2001 | Content Based Retrieval Of 3d Cellular StructuresabstractRecent advances in management of multimedia digital libraries enable effective retrieval of information in the form of audio, image and video. However, retrieval of information in the form of 3D objects has received limited attention so far. Yet many archives of 3D objects already exist and are expected to grow both in relevance and size. In this paper, we address the problem of effective description and retrieval of 3D data representing intracellular structures. These structures are represented in the form of image stacks, being an image stack a set of 2D images representing planar sections of a cellular body at different heights. In the proposed method, 2D visual feature descriptors and Hidden Markov Models are combined to obtain a representation model which is able to distinguish such intracellular structures as Golgi, nucleus, endoplasmic reticulum and lysosomes. Preliminary results are presented to show the effectiveness of the proposed representation model. Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
ICME | 1 |
| 2001 | Modelling Spatial Relationships between Colour Clusters
Stefano Berretti, Alberto Del Bimbo, Enrico Vicario |
Pattern Anal. Appl. | 1 |
| 2001 | Efficient Matching and Indexing of Graph Models in Content-Based RetrievalabstractIn retrieval from image databases, evaluation of similarity, based both on the appearance of spatial entities and on their mutual relationships, depends on content representation based on attributed relational graphs. This kind of modeling entails complex matching and indexing, which presently prevents its usage within comprehensive applications. In this paper, we provide a graph-theoretical formulation for the problem of retrieval based on the joint similarity of individual entities and of their mutual relationships and we expound its implications on indexing and matching. In particular, we propose the usage of metric indexing to organize large archives of graph models, and we propose an original look-ahead method which represents an efficient solution for the (sub)graph error correcting isomorphism problem needed to compute object distances. Analytic comparison and experimental results show that the proposed lookahead improves the state-of-the-art in state-space search methods and that the combined use of the proposed matching and indexing scheme permits for the management of the complexity of a typical application of retrieval by spatial arrangement. Stefano Berretti, Alberto Del Bimbo, Enrico Vicario |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2000 | The Computational Aspect of Retrieval by Spatial ArrangementabstractImage retrieval by spatial arrangement underlies a matching problem for the interpretation of entities specified in the user query on the entities appearing in the image of the database, and for the joint comparison of their features and spatial relationships. In this paper, we provide a graph-theoretical formulation of the problem and discuss its implication on indexing and matching. We first identify an indexing scheme which may fit the characteristics of the problem, and then expound and evaluate an efficient graph matching technique which makes this indexing approach viable. Stefano Berretti, Alberto Del Bimbo, Enrico Vicario |
ICPR | 1 |
| 2000 | Retrieval by Shape Similarity with Perceptual Distance and Effective IndexingabstractAn important problem in accessing and retrieving visual information is to provide efficient similarity matching in large databases. Though much work is being done on the investigation of suitable perceptual models and the automatic extraction of features, little attention is given to the combination of useful representations and similarity models with efficient index structures. In this paper we propose retrieval by shape similarity using local descriptors and effective indexing. Shapes are partitioned into tokens in correspondence with their protrusions, and each token is modeled according to a set of perceptually salient attributes. Shape indexing is obtained by arranging shape tokens into a suitably modified M-tree index structure. Two distinct distance functions model respectively, token and shape perceptual similarity. Examples from a prototype system and computational experiences are reported for both retrieval accuracy and indexing efficiency. Shape retrieval has been tested under shape scaling, orientation changes, and partial shape occlusions. A comparative analysis of different indexing structures, for shape retrieval is presented. Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
IEEE Trans. Multim. | 1 |
| 1999 | Efficient shape retrieval by parts
Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
CAIP | 1 |
| 1997 | Sensations and Psychological Effects in Color Image DatabaseabstractThe development of a system supporting querying of image databases by color content tackles a major design choice about properties of colors which are referenced within user queries. On the one hand, low-level properties directly reflect numerical features and concepts tied to the machine representation of color information. On the other hand, high-level properties address concepts such as the perceptual quality of colors and the sensations that they convey. Color-induced sensations include warmth, accordance or contrast, harmony, excitement, depression, anguish etc. In particular, paintings are an example where the message is contained more in the high-level color qualities and spatial arrangements than in the physical properties of colors. Starting from this observation, Johannes Itten (1961) introduced a formalism to analyze the use of color in art and the effects that this induces on the user's psyche. In this paper, we present a system which translates the Itten theory into a formal language that allows us to express the semantics associated with the combination of chromatic properties of color images. Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
ICIP (1) | 1 |