Chiara Pero

dblp:257/2564 · DBLP profile ↗
← Back
27ranked-venue papers
0as first author
24since 2021 · last 2026
0000-0002-5517-2198ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 A modular augmented reality framework for real-time clinical data visualization and interaction
abstract
This paper presents a modular augmented reality (AR) framework designed to support healthcare professionals in the real-time visualization and interaction with clinical data. The system integrates biometric patient identification, large language models (LLMs) for multimodal clinical data structuring, and ontology-driven AR overlays for anatomy-aware spatial projection. Unlike conventional systems, the framework enables immersive, context-aware visualization that improves both the accessibility and interpretability of medical information. The architecture is fully modular and mobile-compatible, allowing independent refinement of its core components. Patient identification is performed through facial recognition, while clinical documents are processed by a vision-language pipeline that standardizes heterogeneous records into structured data. Body-tracking technology anchors these parameters to the corresponding anatomical regions, supporting intuitive and dynamic interaction during consultations. The framework has been validated through a diabetology case study and a usability assessment with five clinicians, achieving a System Usability Scale (SUS) score of 73.0, which indicates good usability. Experimental results confirm the accuracy of biometric identification (97.1%). The LLM-based pipeline achieved an exact match accuracy of 98.0% for diagnosis extraction and 86.0% for treatment extraction from unstructured clinical images, confirming its reliability in structuring heterogeneous medical content. The system is released as open source to encourage reproducibility and collaborative development. Overall, this work contributes a flexible, clinician-oriented AR platform that combines biometric recognition, multimodal data processing, and interactive visualization to advance next-generation digital healthcare applications. • Modular AR system for real-time clinical data visualization and interaction. • Clinical data mapped to anatomical regions via ontology-guided AR overlay. • Structured clinical data extraction via mobile-efficient multimodal LLMs.
Lucia Cascone, Lucia Cimmino, Michele Nappi, Chiara Pero
Comput. Vis. Image Underst.4
2026 TESA -Net: A Court-Aware Architecture for Flow-Free Basketball Action Recognition
abstract
ABSTRACT Human action recognition in sports videos is a challenging computer vision task due to fast motion, frequent occlusions and fine‐grained visual similarities among action classes. This work presents TESA‐Net (Temporal‐Efficient Spatial Attention Network), an efficient dual‐stream architecture for basketball action recognition that achieves state‐of‐the‐art performance while maintaining computational efficiency. Unlike existing methods that rely on expensive 3D convolutions or full spatio‐temporal attention mechanisms, TESA‐Net employs a pre‐trained 2D ResNet‐50 backbone with lightweight temporal aggregation. The key innovation is a novel Court Line Detection module that augments the appearance stream with edge‐based geometric features, enabling accurate discrimination between shot types that differ primarily in shooting distance. We evaluate TESA‐Net on two complementary benchmarks: Basketball‐51, which targets fine‐grained shot classification in professional broadcasts, and MultiSubjects, which addresses coarse‐grained action recognition in amateur gymnasium recordings. On Basketball‐51, TESA‐Net achieves a validation accuracy of 94.37%, surpassing the previous state‐of‐the‐art HAQT (92.76%) by 1.61 percentage points. On MultiSubjects, TESA‐Net reaches 96.12% accuracy, matching transformer‐based approaches while using significantly fewer parameters. Owing to its efficient design, TESA‐Net requires substantially less memory than competing 3D‐based methods, enabling practical deployment without high‐end hardware.
Andrea F. Abate, Michele Nappi, Chiara Pero, Gianluca Ronga
Expert Syst. J. Knowl. Eng.3
2026 A framework for bias-aware dataset evaluation in soft facial attribute recognition
abstract
Soft Facial Attribute Recognition (FAR) remains largely unexplored in terms of demographic fairness. To the best of our knowledge, this study presents one of the first comprehensive analyses of demographic bias in FAR, proposing a systematic framework to detect, quantify, and promote awareness of both representational and stereotypical biases, supporting their mitigation. Leveraging established taxonomies, we evaluate state-of-the-art datasets using a rigorous set of interpretable bias metrics to uncover hidden demographic imbalances. To support reliable fairness assessment, we first enrich the datasets with standardized demographic annotations using the FairFace model. We then address label inconsistencies through the integration of predictions from advanced Vision-Language Models (VLMs). Our analysis reveals substantial imbalances across gender, age, and racial categories-specifically White, Black, and Asian- affecting dataset composition. Furthermore, we show that conventional fairness metrics often yield divergent assessments, highlighting the importance of multi-metric evaluation. This study provides a replicable methodology and actionable insights to support bias-aware facial analysis.
Lucia Cascone, Michele Nappi, Chiara Pero, Xinggang Wang
Pattern Recognit.3
2025 Mel Spectrogram-Based CNN Framework for Explainable Audio Deepfake Detection
Muhammad Khurram Zahur Bajwa, Aniello Castiglione, Chiara Pero
AINA (8)3
2025 Advancements in basketball action recognition: Datasets, methods, explainability, and synthetic data applications
abstract
Basketball Action Recognition (BAR) has received increasing attention in the fields of computer vision and artificial intelligence, serving as a fundamental component in performance evaluation, automated game annotation, tactical analysis, and referee decision-making support. Despite notable advancements driven by deep learning approaches, BAR remains a challenging task due to the inherent complexity of basketball movements, frequent occlusions, and limited availability of standardized benchmark datasets. This survey provides a comprehensive and structured synthesis of current developments in BAR research, encompassing four principal dimensions: dataset curation, computational methodologies, synthetic data generation, and model explainability. A critical analysis of publicly available basketball-specific datasets is presented, delineating their modalities, annotation strategies, action taxonomies, and representational scope. Furthermore, the survey offers a structured classification of state-of-the-art action recognition methodologies, ranging from video-based and skeleton-based models to sensor-driven and multimodal fusion approaches, emphasizing architectural characteristics, evaluation protocols, and task-specific adaptations. The role of synthetic data is systematically examined as a means to address data scarcity, reduce annotation noise, and enhance model generalization through controlled variability and simulation-based augmentation. In parallel, the integration of explainable artificial intelligence (XAI) techniques is also analyzed, with a focus on post-hoc attribution methods, probabilistic reasoning models, and interpretable neural architectures, aimed at improving the transparency and accountability of decision-making processes. The survey identifies persisting research challenges, including dataset heterogeneity, limitations in cross-domain transferability, and the accuracy-interpretability trade-off in deep models. By delineating current limitations and prospective directions, this work provides a foundational reference to guide the development of robust, generalizable, and explainable BAR systems for deployment in real-world sports intelligence applications. • A systematic review of Basketball Action Recognition Datasets. • Integration of Explainable AI in action recognition models for basketball sport. • Applications and limitations of action recognition in basketball sport. • The role of synthetic data in addressing dataset limitations of state-of-the-art.
Marco Caruso, Lucia Cimmino, Fabio Narducci, Chiara Pero, Gianluca Ronga
Image Vis. Comput.4
2025 Integrating end-to-end multimodal deep learning and domain adaptation for robust facial expression recognition
Mahmoud Hassaballah, Chiara Pero, Ranjeet Kumar Rout, Saiyed Umer
Image Vis. Comput.2
2025 Novel vision transformer and data augmentation technique for efficient detection of monkeypox disease
Aisha Ahmed AlArfaj, Abeer Hakeem, Ebtisam Abdullah Alabdulqader, Chiara Pero, Shtwai Alsubai, Nisreen Innab, Imran Ashraf 0003
Multim. Tools Appl.5
2025 Integrating Post-Quantum Cryptography and Blockchain to Secure Low-Cost IoT Devices
abstract
In the contemporary era, the global proliferation of Internet of Things (IoT) devices exceeds 15 billion, serving functions from wearables to smart grid monitoring. These devices frequently manage sensitive data, underscoring the need for secure and reliable IoT networks leveraging blockchain technology. A key innovation of this study is an approach to mitigate vulnerabilities that quantum computing poses to blockchain-based IoT systems, which existing cryptographic methods cannot effectively address. Quantum computers could exploit these weaknesses to compromise key-pair generation and extract private keys from transaction signatures. To overcome this, the research introduces an optimized implementation of the post-quantum digital signature algorithm Dilithium-5, ensuring blockchain security and quantum readiness. These transaction signatures are designed for low-power, cost-effective microcontrollers, such as the ESP32, making the solution accessible for a wide range of IoT devices. In addition, the study includes a case study involving a post-quantum safe portable device for measuring blood oxygen levels and heart rate, illustrating the practical benefits and effectiveness of the proposed solution in enhancing IoT security against quantum threats. The results demonstrate that the proposed approach ensures quantum-resistant security while maintaining performance efficiency, making it suitable for real-world IoT applications.
Aniello Castiglione, Jacopo Gennaro Esposito, Vincenzo Loia, Michele Nappi, Chiara Pero, Matteo Polsinelli
IEEE Trans. Ind. Informatics5
2025 Enhancing trust of deep learning models with post-quantum digital signatures
abstract
Abstract High-performance computing (HPC) is crucial for artificial intelligence (AI) and deep learning (DL) but faces challenges related to scalability, data transfer costs, and security risks. Federated Learning (FL) enables collaborative model training without centralized data aggregation. However, FL introduces vulnerabilities, as exchanged models can be intercepted and manipulated, necessitating robust cryptographic protection. With the advent of quantum computing, traditional security mechanisms are at risk, requiring the adoption of Post-Quantum Cryptographic (PQC) algorithms. This study benchmarks three PQC digital signature algorithms: Falcon, SPHINCS+, and ML-DSA. Their execution time, memory usage, and computational efficiency are evaluated in a simulated FL setting. To extend the analysis, different cryptographic hash functions (SHA3-256, SHA3-512, and BLAKE3) are analyzed to assess hashing efficiency under varying computational loads. Both centralized and decentralized FL scenarios are simulated, incorporating PQC-based digital signatures at each phase of the communication pipeline to ensure model integrity and authenticity. The results provide insights into the trade-offs between security and computational overhead, guiding the selection of scalable cryptographic solutions for FL. Falcon and ML-DSA demonstrate minimal impact on computational performance, making them strong candidates for securing FL environments. Future research directions include the direct signing of DL models to enhance security and the integration of widely used FL libraries for more realistic evaluations. These advancements could improve the practical deployment of post-quantum security solutions in FL, ensuring resilience against emerging quantum threats.
Aniello Castiglione, Jacopo Gennaro Esposito, Vincenzo Loia, Michele Nappi, Chiara Pero, Matteo Polsinelli
J. Supercomput.5
2024 Acoustic features analysis for explainable machine learning-based audio spoofing detection
abstract
The rapid evolution of synthetic voice generation and audio manipulation technologies poses significant challenges, raising societal and security concerns due to the risks of impersonation and the proliferation of audio deepfakes. This study introduces a lightweight machine learning (ML)-based framework designed to effectively distinguish between genuine and spoofed audio recordings. Departing from conventional deep learning (DL) approaches, which mainly rely on image-based spectrogram features or learning-based audio features, the proposed method utilizes a diverse set of hand-crafted audio features – such as spectral, temporal, chroma, and frequency-domain features – to enhance the accuracy of deepfake audio content detection. Through extensive evaluation and experiments on three well-known datasets, ASVSpoof2019, FakeAVCelebV2, and an In-The-Wild database, the proposed solution demonstrates robust performance and a high degree of generalization compared to state-of-the-art methods. In particular, our method achieved 89% accuracy on ASVSpoof2019, 94.5% on FakeAVCelebV2, and 94.67% on the In-The-Wild database. Additionally, the experiments performed on explainability techniques clarify the decision-making processes within ML models, enhancing transparency and identifying crucial features essential for audio deepfake detection. • Enhanced spoof audio detection via multi-feature integration. • Employed a lightweight ML framework for real-time applications. • Adopted subject-independent protocols to mitigate biometric bias. • Utilized Explainable AI (XAI) for transparent decision-making.
Carmen Bisogni, Vincenzo Loia, Michele Nappi, Chiara Pero
Comput. Vis. Image Underst.4
2024 Walk as you feel: Privacy preserving emotion recognition from gait patterns
abstract
Emotion recognition from gait has gained significant interest due to its applicability in different fields such as healthcare, social cues, surveillance, and smart applications. Gait, as a biometric trait, offers unique advantages, allowing remote identification and robust recognition even in uncontrolled scenarios. Moreover, gait analysis can provide valuable insights into an individual’s emotional state. This work presents the “Walk-as-you-Feel” (WayF) framework, a novel approach for gait-based emotion recognition that does not rely on facial cues, ensuring user privacy. To address challenges with small and unbalanced datasets, a balancing procedure suitable for deep learning architecture is also developed. Adapted Inception-v3 and EfficientNet are employed for the feature extraction phase. Classification is performed using a Gated Recurrent Units network (GRUs) and Transformers-Encoder. Experimental results demonstrate the competitiveness of the proposed approach with respect to state-of-the-art works which also integrate facial cues. WayF reaches an average recognition rate of approximately 77% in its best configuration. Moreover, when excluding the neutral emotion, the proposed method achieves an outstanding overall accuracy of 83.3%.
Carmen Bisogni, Lucia Cimmino, Michele Nappi, Toni Pannese, Chiara Pero
Eng. Appl. Artif. Intell.5
2024 POSER: POsed vs Spontaneous Emotion Recognition using fractal encoding
abstract
Emotion recognition from facial expressions is a fundamental human ability that can be harnessed and transferred to machines. The ability to differentiate between spontaneous and posed emotions holds significant importance in various domains, including behavioral biometrics, forensics, and security. This paper introduces a novel method, called POsed vs Spontaneous Emotion Recognition (POSER), which leverages a modified version of the Partitioned Iterated Functions System (PIFS) to obtain a Fractal Encoding. This encoding is used for the first time as facial features to train a machine learning approach for the classification of emotions as either spontaneous or posed. Furthermore, by adapting the original architecture, we demonstrate the effectiveness of these features in distinguishing seven different emotions in controlled as well as wild environments, within a framework referred to as POSER-EMO. Experimental results are presented on the SPOS and DISFA + datasets for the first classification problem, where POSER outperforms the state of the art, and on the CK + and SFEW datasets for the second classification problem.
Carmen Bisogni, Lucia Cascone, Michele Nappi, Chiara Pero
Image Vis. Comput.4
2024 Head Pose Estimation Patterns as Deepfake Detectors
abstract
The capacity to create “fake” videos has recently raised concerns about the reliability of multimedia content. Identifying between true and false information is a critical step toward resolving this problem. On this issue, several algorithms utilizing deep learning and facial landmarks have yielded intriguing results. Facial landmarks are traits that are solely tied to the subject’s head posture. Based on this observation, we study how Head Pose Estimation (HPE) patterns may be utilized to detect deepfakes in this work. The HPE patterns studied are based on FSA-Net, SynergyNet, and WSM, which are among the most performant approaches on the state-of-the-art. Finally, using a machine learning technique based on K-Nearest Neighbor and Dynamic Time Warping, their temporal patterns are categorized as authentic or false. We also offer a set of experiments for examining the feasibility of using deep learning techniques on such patterns. The findings reveal that the ability to recognize a deepfake video utilizing an HPE pattern is dependent on the HPE methodology. On the contrary, performance is less dependent on the performance of the utilized HPE technique. Experiments are carried out on the FaceForensics++ dataset that presents both identity swap and expression swap examples. The findings show that FSA-Net is an effective feature extraction method for determining whether a pattern belongs to a deepfake or not. The approach is also robust in comparison to deepfake videos created using various methods or for different goals. In the mean the method obtain 86% of accuracy on the identity swap task and 86.5% of accuracy on the expression swap. These findings offer up various possibilities and future directions for solving the deepfake detection problem using specialized HPE approaches, which are also known to be fast and reliable.
Federico Becattini, Carmen Bisogni, Vincenzo Loia, Chiara Pero, Fei Hao 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2024 IoT-enabled Biometric Security: Enhancing Smart Car Safety with Depth-based Head Pose Estimation
abstract
Advanced Driver Assistance Systems (ADAS) are experiencing higher levels of automation, facilitated by the synergy among various sensors integrated within vehicles, thereby forming an Internet of Things (IoT) framework. Among these sensors, cameras have emerged as valuable tools for detecting driver fatigue and distraction. This study introduces HYDE-F, a Head Pose Estimation (HPE) system exclusively utilizing depth cameras. HYDE-F adeptly identifies critical driver head poses associated with risky conditions, thus enhancing the safety of IoT-enabled ADAS. The core of HYDE-F’s innovation lies in its dual-process approach: it employs a fractal encoding technique and keypoint intensity analysis in parallel. These two processes are then fused using an optimization algorithm, enabling HYDE-F to blend the strengths of both methods for enhanced accuracy. Evaluations conducted on a specialized driving dataset, Pandora, demonstrate HYDE-F’s competitive performance compared to existing methods, surpassing current techniques in terms of average Mean Absolute Error (MAE) by nearly 1 ∘ . Moreover, case studies highlight the successful integration of HYDE-F with vehicle sensors. Additionally, HYDE-F exhibits robust generalization capabilities, as evidenced by experiments conducted on standard laboratory-based HPE datasets, i.e., Biwi and ICT-3DHP databases, achieving an average MAE of 4.9 ∘ and 5 ∘ , respectively.
Carmen Bisogni, Lucia Cascone, Michele Nappi, Chiara Pero
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Visual and textual explainability for a biometric verification system based on piecewise facial attribute analysis
abstract
The decisions behind the mechanics of a biometric verification system based on Machine Learning (ML) are difficult to comprehend. Although there is now well-established research in various fields of application, such as health or justice, the use of ML-based methods is accompanied by a lack of confidence that results in their limited use. The explainability of a ML system and the comprehension of what lies behind its prediction is one of the numerous characteristics that define “trust” in these systems. Over the years, face-based biometric authentication has been the subject of extensive research in both academia and industry. However, existing biometric authentication systems still have problems regarding accuracy, robustness and, explainability. Still lacking in the literature is a comprehensive examination of the use of post-hoc explainability techniques for such systems. Cognitive neuroscience has always been interested in the method by which people perceive faces; local elements such as the nose, eyes, and mouth are critical to the perception and recognition of a face. In this work, starting from this assumption, we propose a framework of visual and textual explainability based on the parts of a face by analyzing them with respect to the facial attributes reported in the CelebA dataset. The primary objective is to be able to explain why two pictures of different subjects are distinct. This is done by sinthesizing pairs of images that illustrate how dissimilar the various parts of the face under investigation are and incisive and direct textual explanations of the distinguishing features are generated. A further study analyzes an interpretable mapping between the semantic space of the text and the space of the image.
Lucia Cascone, Chiara Pero, Hugo Proença 0001
Image Vis. Comput.2
2023 Face recognition system with hybrid template protection scheme for Cyber-Physical-Social Services
abstract
This paper presents a secure face recognition system with advanced template protection schemes for Cyber-Physical-Social Services (CPSS). The implementation of the proposed system consists of five components. The initial step performs image preprocessing, where it detects the facial region from the captured image using the Tree-Structured Part Model (TSPM). The second phase involves feature extraction, where it utilizes the Scale Invariant Feature Transform (SIFT) descriptor to extract features from small patches of the preprocessed images, forming a collection of feature descriptors. The collection of feature descriptors is then clustered using the K-means clustering algorithm, returning the centers of K-clusters that serve as the vocabulary of a dictionary. Finally, a histogram is generated using the vocabularies and frequencies, referred to as the “Bag of Visual Words (BoVW)”. Using this dictionary and a feature learning technique called Sparse Representation Coding (SRC), followed by Spatial Pyramid Mapping (SPM), the system generates feature vectors from training/testing image samples. In the third component, the modified FaceHashing technique is applied to the original feature vectors, generating cancelable feature vectors. The fourth component employs a Bio-Cryptographic technique to preserve the cancelable feature vectors in a database. Lastly, the fifth component utilizes a multi-class linear SVM classifier on the decrypted and query-cancellable feature vector to classify users. The system evaluates its performance using FERET and CASIA-FaceV5 benchmark databases, providing 100% identification accuracy for 200-dimensional cancelable feature vectors. The performance and security comparisons demonstrate the superiority of the proposed system over existing methods.
Alamgir Sardar, Saiyed Umer, Ranjeet Kumar Rout, Chiara Pero
Pattern Recognit. Lett.4
2023 Privacy Preserving Ear Recognition System Using Transfer Learning in Industry 4.0
abstract
This article presents an Industry 4.0 compliant ear biometric recognition technique using dense convolutional network (DenseNet), a well-known convolutional neural network model. Compared to other biometric traits, ear recognition has been a challenge due to the unavailability of a large number of images and, therefore, the improvements due to deep learning application are still unexplored. Additionally, ear biometrics has the natural advantage of privacy preservation through excellent feature encoding, which is not yet explored. In this article, the performance of DenseNet is initially tested on typically challenging benchmarks, such as street view house numbers, Canadian Institute for advanced research, and ImageNet, achieving state-of-the-art results and requiring minimal computation time and memory. All the experiments are performed on six popular ear databases namely mathematical analysis of images, annotated web ears (AWE), extended AWE (AWE-X), computer vision laboratory ear (CVLE), Indian Institute of Technology-Delhi, and West Pomeranian University of Technology, indicating that the proposed algorithm achieves a better performance over state-of-the-art. Due to less trainable parameters and fast processing, this Industry 4.0 compliant proposed recognition method can be widely used over Internet of Biometric Things, ensuring the privacy preservation.
Debbrota Paul Chowdhury, Sambit Bakshi, Chiara Pero, Gustavo Olague, Pankaj Kumar Sa
IEEE Trans. Ind. Informatics3
2022 Touch keystroke dynamics for demographic classification
Lucia Cascone, Michele Nappi, Fabio Narducci, Chiara Pero
Pattern Recognit. Lett.4
2021 Fostering secure cross-layer collaborative communications by means of covert channels in MEC environments
Aniello Castiglione, Michele Nappi, Fabio Narducci, Chiara Pero
Comput. Commun.4
2021 Partitioned iterated function systems by regression models for head pose estimation
abstract
Abstract Head pose estimation represents an important computer vision technique in different contexts where image acquisition cannot be controlled by an operator, making face recognition of unknown subjects more accurate and efficient. In this work, starting from partitioned iterated function systems to identify the pose, different regression models are adopted to predict the angular value errors (yaw, pitch and roll axes, respectively). This method combines the fractal image compression characteristics, such as self-similar structures in order to identify similar head rotation, with regression analysis prediction. The experimental evaluation is performed on widely used benchmark datasets, i.e., Biwi and AFLW2000, and the results are compared with many existing state-of-the-art methods, demonstrating the robustness of the proposed fusion approach and excellent performance.
Andrea F. Abate, Paola Barra, Chiara Pero, Maurizio Tucci
Mach. Vis. Appl.3
2021 Adversarial attacks through architectures and spectra in face recognition
Carmen Bisogni, Lucia Cascone, Jean-Luc Dugelay, Chiara Pero
Pattern Recognit. Lett.4
2021 User recognition based on periocular biometrics and touch dynamics
Andrea Casanova, Lucia Cascone, Aniello Castiglione, Weizhi Meng 0001, Chiara Pero
Pattern Recognit. Lett.5
2021 A method for user-customized compensation of metamorphopsia through video see-through enabled head mounted display
Lucia Cimmino, Chiara Pero, Stefano Ricciardi, Shaohua Wan 0001
Pattern Recognit. Lett.2
2021 FASHE: A FrActal Based Strategy for Head Pose Estimation
abstract
Head pose estimation (HPE) represents a topic central to many relevant research fields and characterized by a wide application range. In particular, HPE performed using a singular RGB frame is particular suitable to be applied at best-frame-selection problems. This explains a growing interest witnessed by a large number of contributions, most of which exploit deep learning architectures and require extensive training sessions to achieve accuracy and robustness in estimating head rotations on three axes. However, methods alternative to machine learning approaches could be capable of similar if not better performance. To this regard, we present FASHE, an approach based on partitioned iterated function systems (PIFS) to represent auto-similarities within face image through a contractive affine function transforming the domain blocks extracted only once by a single frontal reference image, in a good approximation of the range blocks which the target image has been partitioned into. Pose estimation is achieved by finding the closest match between fractal code of target image and a reference array by means of Hamming distance. The results of experiments conducted exceed the state of the art on both Biwi and Ponting'04 datasets as well as approaching those of the best performing methods on the challenging AFLW2000 database. In addition, the applications to GOTCHA Video Dataset demonstrate that FASHE successfully operates in-the-wild.
Carmen Bisogni, Michele Nappi, Chiara Pero, Stefano Ricciardi
IEEE Trans. Image Process.3
2020 AR Based User Adaptive Compensation of Metamorphopsia
abstract
The increasing diffusion of augmented reality applications, fostered by the commercial availability of see-through enabled head mounted displays, is opening new opportunities to exploit the potential of this technology to aid subjects with visual impairments in their everyday tasks. The capability of showing a corrected version of the visual field by means of real-time processing of a video stream representing the surrounding environment, is the foundation of the proposed method to compensate a serious visual deficit, known as metamorphopsia, resulting in a geometrical deformation of part of the subject's visus. To this regard, we describe an approach for interactive measurement of user's impaired visual biometrics and for real-time compensation or reduction of this deficiency. This goal is achieved by mapping the video streams acquired from the stereoscopic video see-through cameras each onto a 2D polygonal mesh and offsetting its vertices until the correct vision, for each eye, is restored.
Guido Bozzelli, Maurizio De Nino, Chiara Pero, Stefano Ricciardi
AVI3
2020 HP2IFS: Head Pose estimation exploiting Partitioned Iterated Function Systems
abstract
Estimating the actual head orientation from 2D images, with regard to its three degrees of freedom, is a well known problem that is highly significant for a large number of applications involving head pose knowledge. Consequently, this topic has been tackled by a plethora of methods and algorithms the most part of which exploits neural networks. Machine learning methods, indeed, achieve accurate head rotation values yet require an adequate training stage and, to that aim, a relevant number of positive and negative examples. In this paper we take a different approach to this topic by using fractal coding theory and particularly Partitioned Iterated Function Systems to extract the fractal code from the input head image and to compare this representation to the fractal code of a reference model through Hamming distance. According to experiments conducted on both the BIWI and the AFLW2000 databases, the proposed PIFS based head pose estimation method provides accurate yaw/pitch/roll angular values, with a performance approaching that of state of the art of machine-learning based algorithms and exceeding most of non-training based approaches.
Carmen Bisogni, Michele Nappi, Chiara Pero, Stefano Ricciardi
ICPR3
2020 Head pose estimation by regression algorithm
Andrea F. Abate, Paola Barra, Chiara Pero, Maurizio Tucci
Pattern Recognit. Lett.3