VLDB 2026 Research / reviewers in the wild / expert
Michele Nappi
dblp:35/2888
· DBLP profile ↗
146ranked-venue papers
13as first author
56since 2021 · last 2026
0000-0002-2517-2867ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 70 · 8 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 41 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 14 since 2021Human-computer interaction and ubiquitous computing · 14 · 2 first-author · 1 since 2021Computer networks · 8 · 7 since 2021Systems, architecture and hardware · 5 · 2 since 2021Security and privacy · 5 · 3 since 2021Software engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A modular augmented reality framework for real-time clinical data visualization and interactionabstractThis paper presents a modular augmented reality (AR) framework designed to support healthcare professionals in the real-time visualization and interaction with clinical data. The system integrates biometric patient identification, large language models (LLMs) for multimodal clinical data structuring, and ontology-driven AR overlays for anatomy-aware spatial projection. Unlike conventional systems, the framework enables immersive, context-aware visualization that improves both the accessibility and interpretability of medical information. The architecture is fully modular and mobile-compatible, allowing independent refinement of its core components. Patient identification is performed through facial recognition, while clinical documents are processed by a vision-language pipeline that standardizes heterogeneous records into structured data. Body-tracking technology anchors these parameters to the corresponding anatomical regions, supporting intuitive and dynamic interaction during consultations. The framework has been validated through a diabetology case study and a usability assessment with five clinicians, achieving a System Usability Scale (SUS) score of 73.0, which indicates good usability. Experimental results confirm the accuracy of biometric identification (97.1%). The LLM-based pipeline achieved an exact match accuracy of 98.0% for diagnosis extraction and 86.0% for treatment extraction from unstructured clinical images, confirming its reliability in structuring heterogeneous medical content. The system is released as open source to encourage reproducibility and collaborative development. Overall, this work contributes a flexible, clinician-oriented AR platform that combines biometric recognition, multimodal data processing, and interactive visualization to advance next-generation digital healthcare applications. • Modular AR system for real-time clinical data visualization and interaction. • Clinical data mapped to anatomical regions via ontology-guided AR overlay. • Structured clinical data extraction via mobile-efficient multimodal LLMs. Lucia Cascone, Lucia Cimmino, Michele Nappi, Chiara Pero |
Comput. Vis. Image Underst. | 3 |
| 2026 | TESA -Net: A Court-Aware Architecture for Flow-Free Basketball Action RecognitionabstractABSTRACT Human action recognition in sports videos is a challenging computer vision task due to fast motion, frequent occlusions and fine‐grained visual similarities among action classes. This work presents TESA‐Net (Temporal‐Efficient Spatial Attention Network), an efficient dual‐stream architecture for basketball action recognition that achieves state‐of‐the‐art performance while maintaining computational efficiency. Unlike existing methods that rely on expensive 3D convolutions or full spatio‐temporal attention mechanisms, TESA‐Net employs a pre‐trained 2D ResNet‐50 backbone with lightweight temporal aggregation. The key innovation is a novel Court Line Detection module that augments the appearance stream with edge‐based geometric features, enabling accurate discrimination between shot types that differ primarily in shooting distance. We evaluate TESA‐Net on two complementary benchmarks: Basketball‐51, which targets fine‐grained shot classification in professional broadcasts, and MultiSubjects, which addresses coarse‐grained action recognition in amateur gymnasium recordings. On Basketball‐51, TESA‐Net achieves a validation accuracy of 94.37%, surpassing the previous state‐of‐the‐art HAQT (92.76%) by 1.61 percentage points. On MultiSubjects, TESA‐Net reaches 96.12% accuracy, matching transformer‐based approaches while using significantly fewer parameters. Owing to its efficient design, TESA‐Net requires substantially less memory than competing 3D‐based methods, enabling practical deployment without high‐end hardware. Andrea F. Abate, Michele Nappi, Chiara Pero, Gianluca Ronga |
Expert Syst. J. Knowl. Eng. | 2 |
| 2026 | Image Inpainting in 30 Years: A SurveyabstractABSTRACT As a fundamental task in restoring continuous visual signals, image inpainting plays a critical role in autonomous driving perception, medical imaging, video editing and digital heritage preservation. Driven by deep learning and large‐scale generative models, the field has transitioned from low‐level texture synthesis to high‐level semantic generation, yielding major breakthroughs in structural fidelity and visual realism. Centring on the generative paradigm as the architectural trajectory, this survey systematically categorizes the 30‐year evolution of image inpainting into three distinct technological generations: traditional prior‐driven synthesis, deep learning data‐driven reconstruction and modern foundation model‐driven generation. Despite this progress, highly competitive methods still struggle with large‐scale missing regions, global consistency in complex scenes, fine‐grained micro‐details and alignment with human visual perception. To address these gaps, we critically evaluate the technical paradigms and main bottlenecks within each of these evolutionary stages. We categorize and compare mainstream breakthroughs across high‐resolution restoration, text‐guided synthesis and complex scene generation. Furthermore, we compile standard benchmarks, evaluation metrics and quantitative performance comparisons of representative algorithms. Finally, we dissect open challenges—focusing on cross‐scene generalization and evaluation metric alignment—and outline future trajectories, particularly the integration of inpainting with text‐guided foundation models, providing a definitive reference for future theoretical and engineering advancements. Hengxiang Zhao, Wenchao Zhang 0001, Yu Zheng 0021, Michele Nappi, Junxin Chen 0001 |
Expert Syst. J. Knowl. Eng. | 6 |
| 2026 | A framework for bias-aware dataset evaluation in soft facial attribute recognitionabstractSoft Facial Attribute Recognition (FAR) remains largely unexplored in terms of demographic fairness. To the best of our knowledge, this study presents one of the first comprehensive analyses of demographic bias in FAR, proposing a systematic framework to detect, quantify, and promote awareness of both representational and stereotypical biases, supporting their mitigation. Leveraging established taxonomies, we evaluate state-of-the-art datasets using a rigorous set of interpretable bias metrics to uncover hidden demographic imbalances. To support reliable fairness assessment, we first enrich the datasets with standardized demographic annotations using the FairFace model. We then address label inconsistencies through the integration of predictions from advanced Vision-Language Models (VLMs). Our analysis reveals substantial imbalances across gender, age, and racial categories-specifically White, Black, and Asian- affecting dataset composition. Furthermore, we show that conventional fairness metrics often yield divergent assessments, highlighting the importance of multi-metric evaluation. This study provides a replicable methodology and actionable insights to support bias-aware facial analysis. Lucia Cascone, Michele Nappi, Chiara Pero, Xinggang Wang |
Pattern Recognit. | 2 |
| 2026 | FedBayesMamba: Uncertainty-aware federated learning for multimodal and audio-visual sequential modeling with selective state space modelsabstractFederated learning has emerged as an effective paradigm for training machine learning models across distributed clients without sharing raw data. In many real-world applications, sequential data are inherently multimodal, involving heterogeneous streams such as audio, visual, and temporal signals. However, most existing federated approaches rely on deterministic neural networks, which often struggle to capture predictive uncertainty under heterogeneous data distributions, cross-modal inconsistencies, and dynamic client participation. In this paper, we propose FedBayesMamba , a Bayesian federated learning framework for multimodal sequential data modeling based on selective state space models. The proposed approach introduces Bayesian parameterization into the Mamba architecture to enable uncertainty-aware sequence modeling while preserving the computational efficiency of state space models. To effectively integrate uncertainty across distributed clients, we further develop a posterior aggregation strategy that combines client-level posterior distributions in a principled probabilistic manner. Extensive experiments on multiple benchmark datasets demonstrate that the proposed framework achieves competitive predictive performance and improved uncertainty estimation under Non-IID federated settings. The results also indicate that FedBayesMamba exhibits strong robustness and stability in challenging federated scenarios. These findings highlight the potential of combining Bayesian learning with state space models for multimodal temporal modeling, particularly in audio-visual perception and cross-modal sequence understanding tasks. Xianxun Zhu, Xiaosong E, Michele Nappi, Imad Rida, Hui Chen 0026 |
Pattern Recognit. | 3 |
| 2026 | Transferable online handwriting recognition via Siamese contrastive learning on inertial signals
Lucia Cascone, Michele Nappi, Giuseppe Placidi, Matteo Polsinelli |
Pattern Recognit. Lett. | 2 |
| 2026 | Explainable multimodal brain imaging through a multiple-branch neural networkabstract• The role of each imaging modality is evaluated in the identification and segmentation of the lesions. • Changes in the importance of imaging modalities are evaluated for different segmentation tasks. • The impact and challenges of using AI-generated data instead of real counterpart scans are evaluated to check how the network can explain potential differences. Brain studies require the use of several complementary imaging modalities. When some modality is unavailable, Artificial Intelligence (AI) has recently provided ways to estimate them. Radiologists modulate the use of the available modalities depending on the task they have to perform. We aim to trace artificially the radiological process through a multibranch neural network architecture, the StarNet. The goal is to explain how and where different imaging modalities, either really collected or artificially reconstructed, are used in different radiological tasks by reading inside the structure of the network. To do that, StarNet includes several satellite networks, one per source modality, connected at each layer by a central unit. This design enables us to assess the contribution of each imaging modality, identifying where the contribution occurs, and to quantify the variations if certain modalities are substituted with AI-generated counterparts. The ultimate goal is to enable data-related and task-related ablation studies through the complete explainability of StarNet, thus offering radiologists clear guidance on which imaging sequences contribute to the task, to what extent, and at which stages of the process. As an example, we applied the proposed architecture to the 2D slices extracted from 3D volumes acquired with multimodal magnetic resonance imaging (MRI), to assess: 1. The role of the used imaging modalities; 2. The change in role when the radiological task changes; 3. The effects of synthetic data on the process. The results are presented and discussed. Giuseppe Placidi, Alessia Cipriani, Michele Nappi, Matteo Polsinelli |
Pattern Recognit. Lett. | 3 |
| 2026 | OpenV2X: A Modular Cyber-Physical Framework for Vision-Driven Environmental Perception and Interactive Feedback in Software-Defined VehiclesabstractThe evolution of software-defined vehicles and intelligent transportation systems requires cyber-physical frameworks capable of integrating real-time perception, communication, and human interaction. In industrial automotive environments, enhancing situational awareness while maintaining system modularity, energy efficiency, and interpretability remains a critical challenge. This work presentsOpenV2X, a modular cyber-physical framework designed for real-time perception and human-in-the-loop feedback in software-defined vehicles. The primary goal is to provide a modular system for environment reconstruction, overcoming the limitations of closed industrial systems and research efforts focused on individual modules. OpenV2X combines onboard vision-based sensors with Message Queuing Telemetry Transport (MQTT)-enabled vehicle-to-everything (V2X) communication, extending environmental awareness beyond sensor line-of-sight and adhering to the principles of industrial informatics and resilient automation. Its architecture decouples object and lane detection modules, enabling dynamic model updates without system overhaul, and supports interactive feedback through a graphical interface, fostering end-user trust. Furthermore, the framework allows operation beyond simulated tests: the collected data can be sent to the original equipment manufacturer (OEM) to continuously improve software, achieving the main objective of the work. Experimental validation in real highway scenarios demonstrates 84.3% accuracy in orientation-aware object detection, 73% intersection-over-union for lane detection under adverse weather conditions, and submillisecond V2X latency with minimal energy overhead. Designed for embedded edge deployment, OpenV2X provides a reproducible open-source platform that integrates perception, communication, and user feedback in next-generation industrial vehicular systems. Aniello Castiglione, Lucia Cimmino, Michele Nappi, Luigi Emanuele Sica |
IEEE Trans. Ind. Informatics | 3 |
| 2026 | Plausible Deniable Medical Image Encryption by Large Language Models and Reversible Content-Aware StrategyabstractThere is a rising concern about healthcare system security, where data loss could bring lots of damages to patients and hospitals. As a promising encryption method for medical images, DNA encoding own characteristics of high speed, parallelism computation, minimal storage, and unbreakable cryptosystems. Inspired by the idea of involving Large Language Models(LLMs) to improve DNA encoding, we propose a medical image encryption method with LLM-enhanced DNA encoding, which consists of LLM enhancing module and content-aware permutation&diffusion module. Regarding medical images generally have plain backgrounds with low-entropy pixels, the first module compresses pixels into highly compact signals with features of probabilistic varying and plausibly deniability, serving as another LLM-based layer of defense against privacy breaches before DNA encoding. The second module not only adds permutation by randomly sampling from a redundant correlation between adjacent pixels to break the internal links between pixels but also performs a DNA-based diffusion process to greatly increase the complexity of cracking. Experiments on ChestXray-14, COVID-CT and fcon-1000 datasets show that the proposed method outperforms all comparative methods in sensitivity, correlation and entropy. Yirui Wu, Xinfu Liu 0001, Lucia Cascone, Michele Nappi, Shaohua Wan 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | A Context-Dependent CNN-Based Framework for Multiple Sclerosis Segmentation in MRIabstractDespite several automated strategies for identification/segmentation of Multiple Sclerosis (MS) lesions in Magnetic Resonance Imaging (MRI) being developed, they consistently fall short when compared to the performance of human experts. This emphasizes the unique skills and expertise of human professionals in dealing with the uncertainty resulting from the vagueness and variability of MS, the lack of specificity of MRI concerning MS, and the inherent instabilities of MRI. Physicians manage this uncertainty in part by relying on their radiological, clinical, and anatomical experience. We have developed an automated framework for identifying and segmenting MS lesions in MRI scans by introducing a novel approach to replicating human diagnosis, a significant advancement in the field. This framework has the potential to revolutionize the way MS lesions are identified and segmented, being based on three main concepts: (1) Modeling the uncertainty; (2) Use of separately trained Convolutional Neural Networks (CNNs) optimized for detecting lesions, also considering their context in the brain, and to ensure spatial continuity; (3) Implementing an ensemble classifier to combine information from these CNNs. The proposed framework has been trained, validated, and tested on a single MRI modality, the FLuid-Attenuated Inversion Recovery (FLAIR) of the MSSEG benchmark public data set containing annotated data from seven expert radiologists and one ground truth. The comparison with the ground truth and each of the seven human raters demonstrates that it operates similarly to human raters. At the same time, the proposed model demonstrates more stability, effectiveness and robustness to biases than any other state-of-the-art model though using just the FLAIR modality. Giuseppe Placidi, Luigi Cinque, Gian Luca Foresti, Francesca Galassi, Filippo Mignosi, Michele Nappi, Matteo Polsinelli |
Int. J. Neural Syst. | 6 |
| 2025 | Autonomous Driving: Integration of Segmentation and Depth Camera in a Curriculum Learning ApproachabstractAutonomous Driving (AD) entails vehicles that can perceive their surroundings and navigate without human intervention. This involves utilising a combination of sensors and algorithms to recognise obstacles, interpret traffic signals, and make driving decisions. While AD holds promise for transforming transportation by enhancing safety, reducing congestion, minimising pollution, and optimising efficiency, it poses technical challenges also. This work extends a novel approach to building an autonomous vehicle agent using Deep Reinforcement Learning (DRL) with Proximal Policy Optimisation (PPO) to navigate urban environments simulated by the CAR Learning to Act (CARLA) Simulator. The agent aims to maintain lane integrity and avoid collisions, even in adverse weather conditions. The proposed architecture integrates a 180-degree environmental view and various multimodal data inputs (RGB, segmentation, and depth camera inputs), extensively tested through experimentation. Notably, the integration of segmentation and depth data results in a 13% reduction in the collision rate, with the proposed agent achieving a total reward of 2510. This approach demonstrates significant progress over the previous framework, showcasing improved obstacle detection and collision avoidance accuracy. Moreover, these findings contribute to ongoing autonomous vehicle research, offering insights into effective strategies for developing robust and dependable driving agents capable of navigating urban environments and interacting with road infrastructure, contributing to advancements in Augmented Intelligence of Things (AIoT)-enabled autonomous driving. Silvio Barra, Lucia Cimmino, Vincenzo Loia, Michele Nappi, Matteo Polsinelli |
IEEE Internet Things J. | 4 |
| 2025 | Deep Customized Network Slicing and Efficient Routing for IoT Applications in B5G-Enabled Edge Computing NetworksabstractBeyond 5G-enabled edge computing networking (ECN) will further deploy computing and communication resources to the edge of the networks. Then, edge service demands for Internet of Things (IoT) applications are becoming more and more diverse, while the corresponding routing service capability is limited and not flexible enough to deal with the demands of ECN, which then leads to reducing the inherent routing capability of ECN. It becomes extremely difficult for ECN to support diversified demands and provide diverse IoT applications quickly and flexibly. In this article, we propose a novel and customized deep routing mechanism for IoT applications in ECN, in which the network slicing and deep learning methods are jointly applied and leveraged. First, we design a new ECN architecture that formulates four kinds of network slices to cope with various IoT scenarios, which are eMBB, uRLLC, mMTTC, and backup slices. Second, using these slices, we can customize the ECN environment flexibly, based on which we propose the corresponding routing method for the purpose of fast and efficient service delivery. In particular, the mapping between network slices and the infrastructure is established with the object of maximizing the resource utilization. Then, the routing is designed and customized by using the deep learning model. Lastly, the experimental results show that the deep customized mechanism designed in this article can reduce the average loss rate of the model, decrease the average delay, as well as improve the average resource utilization compared with the existing studies. Xingchi Chen, Bo Yi 0002, Qing Li 0006, Fa Zhu, Yingpu Nian, Achyut Shankar, Michele Nappi, Amr Tolba |
IEEE Internet Things J. | 7 |
| 2025 | Ricci curvature discretizations for head pose estimation from a single imageabstractHead pose estimation (HPE) is crucial in various real-world applications, like human–computer interaction and biometric framework enhancement. This research aims to leverage network curvature to predict head pose from a single image. In networks, certain groups of nodes fulfill significant functional roles. This study focuses on the interactions of facial landmarks, considered as vertices in a weighted graph. The experiments demonstrate that the underlying graph geometry and topology enable the detection of similarities among various head poses. Two independent notions of discrete Ricci curvature for graphs, namely Ollivier–Ricci and Forman–Ricci curvatures, are investigated. These two types of Ricci curvature, each reflecting distinct geometric properties of the network, serve as inputs to the regression model. The results from the BIWI, AFLW2000, and Pointing‘04 datasets reveal that the two discretizations of Ricci’s curvature are closely related and outperform state-of-the-art methods, including both landmark-based and image-only approaches. This demonstrates the effectiveness and promise of using network curvature for HPE in diverse applications. • The topology of the underlying graph can identify similarities across head poses. • Analyzes the performance differences between Ollivier and Forman Ricci curvatures. • Ollivier–Ricci-based method shows competitive results on three different datasets. • The Forman–Ricci-based method offers a computationally efficient alternative. Andrea F. Abate, Lucia Cascone, Michele Nappi |
Pattern Recognit. | 3 |
| 2025 | Cyberbullying Detection Using PCA Extracted GLOVE Features and RoBERTaNet Transformer Learning ModelabstractOnline platforms are nurturing social interactions, yet regrettably, they have also led to the proliferation of antisocial behaviors such as cyberbullying, trolling, and hate speech on a global scale. The identification of hate speech and aggression has become indispensable in the fight against cyberbullying and online harassment. Cyberbullying encompasses the use of aggressive and offensive language, including rude, insulting, hateful, and teasing comments, to inflict harm on individuals through social media platforms. Human moderation is both sluggish and costly, rendering it impractical in light of the exponential growth of data. Consequently, automated detection systems are imperative to effectively combat trolling. This study addresses the challenge of automatically discerning cyberbullying in tweets sourced from a publicly available cyberbullying dataset. The proposed methodology leverages the robustly optimized bidirectional encoder representations from transformers approach (RoBERTa), integrating principle component analysis (PCA) extracted global vectors for word representation (GLOVE) word embedding features. Furthermore, our proposed approach is benchmarked against state-of-the-art machine learning, deep learning, and transformer-based methods, utilizing the GLOVE word embedding technique. Statistical analyses reveal that our proposed model outperforms its counterparts, achieving a 0.98 accuracy and recall rate with 0.97 of precision and F1 score in detecting cyberbullying tweets. Results from$k$-fold cross validation further corroborate the superior performance of our proposed model. Muhammad Umer 0001, Ebtisam Abdullah Alabdulqader, Aisha Ahmed AlArfaj, Lucia Cascone, Michele Nappi |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2025 | Integrating Post-Quantum Cryptography and Blockchain to Secure Low-Cost IoT DevicesabstractIn the contemporary era, the global proliferation of Internet of Things (IoT) devices exceeds 15 billion, serving functions from wearables to smart grid monitoring. These devices frequently manage sensitive data, underscoring the need for secure and reliable IoT networks leveraging blockchain technology. A key innovation of this study is an approach to mitigate vulnerabilities that quantum computing poses to blockchain-based IoT systems, which existing cryptographic methods cannot effectively address. Quantum computers could exploit these weaknesses to compromise key-pair generation and extract private keys from transaction signatures. To overcome this, the research introduces an optimized implementation of the post-quantum digital signature algorithm Dilithium-5, ensuring blockchain security and quantum readiness. These transaction signatures are designed for low-power, cost-effective microcontrollers, such as the ESP32, making the solution accessible for a wide range of IoT devices. In addition, the study includes a case study involving a post-quantum safe portable device for measuring blood oxygen levels and heart rate, illustrating the practical benefits and effectiveness of the proposed solution in enhancing IoT security against quantum threats. The results demonstrate that the proposed approach ensures quantum-resistant security while maintaining performance efficiency, making it suitable for real-world IoT applications. Aniello Castiglione, Jacopo Gennaro Esposito, Vincenzo Loia, Michele Nappi, Chiara Pero, Matteo Polsinelli |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | HPA-UNet: A Hybrid Post-Processing Attention U-Net for Tongue SegmentationabstractTongue diagnosis is the kernel method of Traditional Chinese Medicine (TCM), and it has been proved that the condition of the tongue can serve as an indicator of a person's health status. To automatically recognize a person's latent diseases by computer vision technology, getting the tongue segmentation from a picture with high precision has significant importance. However, the precision of tongue segmentation images in most prior methods is not satisfactory, which will inevitably result in misjudging. In this paper, an effective method is proposed for highly precise tongue segmentation, which is combined with an improved U-shaped neural network and an edge refinement post-processing method. The contributions are three-fold. First, a carefully designed data augmentation strategy is imported to prevent the network from over-fitting. Second, an updated U-shaped neural network is designed to segment tongue images with high precision. Third, a post-processing method is imported to refine the edge of the tongue segmentation further. The proposed method achieves competitive performance in almost all experiments on two datasets. Furthermore, the proposed post-processing method can effectively improve all classic neural networks in tongue segmentation, which strongly proves the flexibility and generalization of the proposed method. Leiyue Yao, Yuchen Xu 0013, Jianying Xiong, Achyut Shankar, Mustufa Haider Abidi, Michele Nappi |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Enhancing trust of deep learning models with post-quantum digital signaturesabstractAbstract High-performance computing (HPC) is crucial for artificial intelligence (AI) and deep learning (DL) but faces challenges related to scalability, data transfer costs, and security risks. Federated Learning (FL) enables collaborative model training without centralized data aggregation. However, FL introduces vulnerabilities, as exchanged models can be intercepted and manipulated, necessitating robust cryptographic protection. With the advent of quantum computing, traditional security mechanisms are at risk, requiring the adoption of Post-Quantum Cryptographic (PQC) algorithms. This study benchmarks three PQC digital signature algorithms: Falcon, SPHINCS+, and ML-DSA. Their execution time, memory usage, and computational efficiency are evaluated in a simulated FL setting. To extend the analysis, different cryptographic hash functions (SHA3-256, SHA3-512, and BLAKE3) are analyzed to assess hashing efficiency under varying computational loads. Both centralized and decentralized FL scenarios are simulated, incorporating PQC-based digital signatures at each phase of the communication pipeline to ensure model integrity and authenticity. The results provide insights into the trade-offs between security and computational overhead, guiding the selection of scalable cryptographic solutions for FL. Falcon and ML-DSA demonstrate minimal impact on computational performance, making them strong candidates for securing FL environments. Future research directions include the direct signing of DL models to enhance security and the integration of widely used FL libraries for more realistic evaluations. These advancements could improve the practical deployment of post-quantum security solutions in FL, ensuring resilience against emerging quantum threats. Aniello Castiglione, Jacopo Gennaro Esposito, Vincenzo Loia, Michele Nappi, Chiara Pero, Matteo Polsinelli |
J. Supercomput. | 4 |
| 2024 | Modeling Conditional Relationships in the Management and Monitoring of Type 1 and Type 2 Diabetes through Bayesian NetworkabstractDiabetes mellitus is one of the most prevalent chronic diseases, affecting millions of people worldwide. Effective management of diabetes, particularly type 1 (T1DM) and type 2 diabetes (T2DM), requires a deep understanding of the complex interactions between clinical, behavioral, and socio-demographic factors. This study leverages Bayesian networks (BNs) to model these interactions, providing a transparent and interpretable visual framework that reveals how variables such as insulin use, BMI, and socio-economic status influence diabetes outcomes. However, constructing the graphical structure of a BN poses significant challenges due to the intricate and multifaceted relationships involved. To ensure a meaningful comparison between T1DM and T2DM, we utilized a cohort of subjects selected to be as demographically homogeneous as possible. This allowed us to reduce confounding effects and focus on the intrinsic differences between the two conditions. We compared network structures derived from three data samples (50%, 80%, and 100%) to explore how variable relationships evolve as the dataset size increases, ensuring that critical interactions were captured at different levels of data availability. The findings highlight key differences in the management of T1DM and T2DM, particularly with regard to behavioral and socio-economic factors. Lucia Cascone, Margherita Maria Napolitano, Michele Nappi, Severino Nappi, Genny Tortora |
BIBM | 3 |
| 2024 | Acoustic features analysis for explainable machine learning-based audio spoofing detectionabstractThe rapid evolution of synthetic voice generation and audio manipulation technologies poses significant challenges, raising societal and security concerns due to the risks of impersonation and the proliferation of audio deepfakes. This study introduces a lightweight machine learning (ML)-based framework designed to effectively distinguish between genuine and spoofed audio recordings. Departing from conventional deep learning (DL) approaches, which mainly rely on image-based spectrogram features or learning-based audio features, the proposed method utilizes a diverse set of hand-crafted audio features – such as spectral, temporal, chroma, and frequency-domain features – to enhance the accuracy of deepfake audio content detection. Through extensive evaluation and experiments on three well-known datasets, ASVSpoof2019, FakeAVCelebV2, and an In-The-Wild database, the proposed solution demonstrates robust performance and a high degree of generalization compared to state-of-the-art methods. In particular, our method achieved 89% accuracy on ASVSpoof2019, 94.5% on FakeAVCelebV2, and 94.67% on the In-The-Wild database. Additionally, the experiments performed on explainability techniques clarify the decision-making processes within ML models, enhancing transparency and identifying crucial features essential for audio deepfake detection. • Enhanced spoof audio detection via multi-feature integration. • Employed a lightweight ML framework for real-time applications. • Adopted subject-independent protocols to mitigate biometric bias. • Utilized Explainable AI (XAI) for transparent decision-making. Carmen Bisogni, Vincenzo Loia, Michele Nappi, Chiara Pero |
Comput. Vis. Image Underst. | 3 |
| 2024 | Walk as you feel: Privacy preserving emotion recognition from gait patternsabstractEmotion recognition from gait has gained significant interest due to its applicability in different fields such as healthcare, social cues, surveillance, and smart applications. Gait, as a biometric trait, offers unique advantages, allowing remote identification and robust recognition even in uncontrolled scenarios. Moreover, gait analysis can provide valuable insights into an individual’s emotional state. This work presents the “Walk-as-you-Feel” (WayF) framework, a novel approach for gait-based emotion recognition that does not rely on facial cues, ensuring user privacy. To address challenges with small and unbalanced datasets, a balancing procedure suitable for deep learning architecture is also developed. Adapted Inception-v3 and EfficientNet are employed for the feature extraction phase. Classification is performed using a Gated Recurrent Units network (GRUs) and Transformers-Encoder. Experimental results demonstrate the competitiveness of the proposed approach with respect to state-of-the-art works which also integrate facial cues. WayF reaches an average recognition rate of approximately 77% in its best configuration. Moreover, when excluding the neutral emotion, the proposed method achieves an outstanding overall accuracy of 83.3%. Carmen Bisogni, Lucia Cimmino, Michele Nappi, Toni Pannese, Chiara Pero |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Soft-orthogonal constrained dual-stream encoder with self-supervised clustering network for brain functional connectivity dataabstractIn many brain network studies, brain functional connectivity data is extracted from neuroimaging data and then used for disease prediction. For now, brain disease data not only has a small sample but also has the problem of high dimensional and nonlinear. Therefore, deep clustering on brain functional connectivity data is very challenging. To solve these problems, we propose a Soft-orthogonal Constrained Dual-stream Encoder with Self-supervised clustering network (SSCDE), which consists of a pretext task and downstream task, which can fully mine the effective information in brain disease data. In the pretext task, we use two brain disease data under the same category to do cross-domain learning to obtain effective information from the same dataset. In the downstream task, to reduce redundancy and avoid negative coding, we propose a soft-orthogonal constrained dual-stream encoder to encode features separately. At the same time, we use the pseudo labels given by the pretext task as prior information for self-supervised learning. We conduct validation on different brain disease recognition tasks, and the result have proved that the proposed framework has achieved good performance compared with the unsupervised clustering analysis algorithms. To our knowledge, this is the first cross-domain assisted recognition study on brain functional connectivity data. The code is available at https://github.com/hulu88/SSCDE . Hu Lu, Tingting Jin, Hui Wei 0001, Michele Nappi, Shaohua Wan 0001 |
Expert Syst. Appl. | 4 |
| 2024 | Label-aware Attention Network with Multi-scale Boosting for Medical Image Segmentation
Linbo Wang 0001, Peng Xu 0045, Xianfeng Cao, Michele Nappi, Shaohua Wan 0001 |
Expert Syst. Appl. | 4 |
| 2024 | Fuzzy-CNN: Improving personal human identification based on IRIS recognition using LBP features
Mashael Khayyat, Nuha Zamzami, Michele Nappi, Muhammad Umer 0001 |
J. Inf. Secur. Appl. | 4 |
| 2024 | POSER: POsed vs Spontaneous Emotion Recognition using fractal encodingabstractEmotion recognition from facial expressions is a fundamental human ability that can be harnessed and transferred to machines. The ability to differentiate between spontaneous and posed emotions holds significant importance in various domains, including behavioral biometrics, forensics, and security. This paper introduces a novel method, called POsed vs Spontaneous Emotion Recognition (POSER), which leverages a modified version of the Partitioned Iterated Functions System (PIFS) to obtain a Fractal Encoding. This encoding is used for the first time as facial features to train a machine learning approach for the classification of emotions as either spontaneous or posed. Furthermore, by adapting the original architecture, we demonstrate the effectiveness of these features in distinguishing seven different emotions in controlled as well as wild environments, within a framework referred to as POSER-EMO. Experimental results are presented on the SPOS and DISFA + datasets for the first classification problem, where POSER outperforms the state of the art, and on the CK + and SFEW datasets for the second classification problem. Carmen Bisogni, Lucia Cascone, Michele Nappi, Chiara Pero |
Image Vis. Comput. | 3 |
| 2024 | Gaze analysis: A survey on its applicationsabstractThe examination of ocular movements has a wide range of applications due to the current developments in sensors that are now able to collect this biometric. This type of investigation is known as “gaze analysis”. The gaze has successfully examined a subject's physical and mental status in the past. As a result, over the last few decades, a large and diverse amount of literature on this subject has been generated and presented. The aim of this study is to collect and debate current gaze analysis methods based on their application field. Due to the context-specific needs for performance and efficiency, the eye movements under research are frequently evaluated from completely distinct perspectives. As a result, a collection of data, methods, and discussions ranging from the medical community to virtual and augmented reality, as well as human computer interface and remote learning, has been produced. In addition to providing a peek of novel observation on the issue of gaze analysis, the gaps between and within areas are also discussed to provide points for researchers to pursue. Carmen Bisogni, Michele Nappi, Genny Tortora, Alberto Del Bimbo |
Image Vis. Comput. | 2 |
| 2024 | Student academic success prediction in multimedia-supported virtual learning system using ensemble learning approach
Oumaima Saidani, Muhammad Umer 0001, Amal Alshardan, Nazik Alturki, Michele Nappi, Imran Ashraf 0003 |
Multim. Tools Appl. | 5 |
| 2024 | CDT-CAD: Context-Aware Deformable Transformers for End-to-End Chest Abnormality Detection on X-Ray ImagesabstractDeep learning methods have achieved great success in medical image analysis domain. However, most of them suffer from slow convergency and high computing cost, which prevents their further widely usage in practical scenarios. Moreover, it has been proved that exploring and embedding context knowledge in deep network can significantly improve accuracy. To emphasize these tips, we present CDT-CAD, i.e., context-aware deformable transformers for end-to-end chest abnormality detection on X-Ray images. CDT-CAD firstly constructs an iterative context-aware feature extractor, which not only enlarges receptive fields to encode multi-scale context information via dilated context encoding blocks, but also captures unique and scalable feature variation patterns in wavelet frequency domain via frequency pooling blocks. Afterwards, a deformable transformer detector on the extracted context features is built to accurately classify disease categories and locate regions, where a small set of key points are sampled, thus leading the detector to focus on informative feature subspace and accelerate convergence speed. Through comparative experiments on Vinbig Chest and Chest Det 10 Datasets, CDT-CAD demonstrates its effectiveness in recognizing chest abnormities and outperforms 1.4% and 6.0% than the existing methods in$AP_{5}0$and$AR$on VinBig dateset, and 0.9% and 2.1% on Chest Det-10 dataset, respectively. Yirui Wu, Qiran Kong, Lilai Zhang, Aniello Castiglione, Michele Nappi, Shaohua Wan 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2024 | IoT-enabled Biometric Security: Enhancing Smart Car Safety with Depth-based Head Pose EstimationabstractAdvanced Driver Assistance Systems (ADAS) are experiencing higher levels of automation, facilitated by the synergy among various sensors integrated within vehicles, thereby forming an Internet of Things (IoT) framework. Among these sensors, cameras have emerged as valuable tools for detecting driver fatigue and distraction. This study introduces HYDE-F, a Head Pose Estimation (HPE) system exclusively utilizing depth cameras. HYDE-F adeptly identifies critical driver head poses associated with risky conditions, thus enhancing the safety of IoT-enabled ADAS. The core of HYDE-F’s innovation lies in its dual-process approach: it employs a fractal encoding technique and keypoint intensity analysis in parallel. These two processes are then fused using an optimization algorithm, enabling HYDE-F to blend the strengths of both methods for enhanced accuracy. Evaluations conducted on a specialized driving dataset, Pandora, demonstrate HYDE-F’s competitive performance compared to existing methods, surpassing current techniques in terms of average Mean Absolute Error (MAE) by nearly 1 ∘ . Moreover, case studies highlight the successful integration of HYDE-F with vehicle sensors. Additionally, HYDE-F exhibits robust generalization capabilities, as evidenced by experiments conducted on standard laboratory-based HPE datasets, i.e., Biwi and ICT-3DHP databases, achieving an average MAE of 4.9 ∘ and 5 ∘ , respectively. Carmen Bisogni, Lucia Cascone, Michele Nappi, Chiara Pero |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Introduction to Special Issue on "Recent Trends in Multimedia Forensics"abstractMultimedia forensics is a subject area which is the need of the hour in this modern era of media-manipulation and generation of fake images/videos assisted with artificial intelligence (AI) models. With the ubiquitous expansion of internet enabled devices, there is a humungous amount of data available to the perusal of forensic experts. This data comprises of audio, video, images, text or a mix of those. Hence multimedia forensics, which involves a set of scientific techniques to collect, scrutinize and analyze this digital content, becomes highly imperative. The increasing threat of compelling media manipulations through machine learning-based technologies is making the situation more alarming. The most common instances are generative adversarial networks (GANs) (to generate artificial yet realistic images/videos) and DeepFake algorithms (to swap faces and expressions in videos). Furthermore, the ease of getting these manipulations done has lowered the skill required from the attacker’s end, which has intensified the problem manifold. This special issue captures a few recent outstanding works beyond trivial research results in order to push the border of the state-of-the-art and record the developments on this subject of research. Ritesh Vyas, Michele Nappi, Alberto Del Bimbo, Sambit Bakshi |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Editorial to Special Issue on Multimedia Cognitive Computing for Intelligent Transportation SystemabstractNo abstract available. Shaohua Wan 0001, Yi Jin 0001, Guandong Xu, Michele Nappi |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Language Identification as Improvement for Lip-Based Biometric Visual SystemsabstractLanguage has always been one of humanity's defining characteristics. Visual Language Identification (VLI) is a relatively new field of research that is complex and largely understudied. In this paper, we present a preliminary study in which we use linguistic information as a soft biometric trait to enhance the performance of a visual (auditory-free) identification system based on lip movement. We report a significant improvement in the identification performance of the proposed visual system as a result of the integration of these data using a score-based fusion strategy. Methods of Deep and Machine Learning are considered and evaluated. To the experimentation purposes, the dataset called laBial Articulation for the proBlem of the spokEn Language rEcognition (BABELE), consisting of 8 different languages, has been created. It includes a collection of different features of which the spoken language represents the most relevant, while each sample is also manually labelled with gender and age of the subjects. Lucia Cascone, Michele Nappi, Fabio Narducci |
ICIP | 2 |
| 2023 | Face mask detection using deep convolutional neural network and multi-stage image processing
Muhammad Umer 0001, Saima Sadiq, Reemah M. Alhebshi, Shtwai Alsubai, Abdullah Al Hejaili, Alá Abdulmajid Eshmawi, Michele Nappi, Imran Ashraf 0003 |
Image Vis. Comput. | 7 |
| 2023 | Impact of convolutional neural network and FastText embedding on text classificationabstractAbstract Efficient word representation techniques (word embeddings) with modern machine learning models have shown reasonable improvement on automatic text classification tasks. However, the effectiveness of such techniques has not been evaluated yet in terms of insufficient word vector representation for training. Convolutional Neural Network has achieved significant results in pattern recognition, image analysis, and text classification. This study investigates the application of the CNN model on text classification problems by experimentation and analysis. We trained our classification model with a prominent word embedding generation model, Fast Text on publically available datasets, six benchmark datasets including Ag News, Amazon Full and Polarity, Yahoo Question Answer, Yelp Full, and Polarity. Furthermore, the proposed model has been tested on the Twitter US airlines non-benchmark dataset as well. The analysis indicates that using Fast Text as word embedding is a very promising approach. Muhammad Umer 0001, Zainab Imtiaz, Muhammad Ahmad 0002, Michele Nappi, Carlo Maria Medaglia, Gyu Sang Choi, Arif Mehmood |
Multim. Tools Appl. | 4 |
| 2023 | Segmentation of ultrasound image sequences by combing a novel deep siamese network with a deformable contour model
Bo Ni, Xiantao Cai, Michele Nappi, Shaohua Wan 0001 |
Neural Comput. Appl. | 4 |
| 2023 | Two Path Gland Segmentation Algorithm of Colon Pathological Image Based on Local Semantic GuidanceabstractColonic adenocarcinoma is a disease severely endangering human life caused by mucosal epidermal carcinogenesis. The segmentation of potentially cancerous glands is the key in the detection and diagnosis of colonic adenocarcinoma. The appearance of cancerous tissue is different in gland segmentation in colon pathological images, and it is impossible to accurately segment the changes of glands from benign to malignant using a single network. Given these issues, a two-path gland segmentation algorithm of colon pathological image based on local semantic guidance is proposed in this paper. The improved candidate region search algorithm is adopted to expand the original image data set and generate sub-datasets sensitive to specific features. Then, the semantic feature-guided model is employed to extract the local adenocarcinoma features and acts on the backbone network together with context feature extraction based on the attention mechanism. In this way, a larger receptive field and more local feature information are obtained, the learning ability of the network to the morphological features of glands is enhanced, and the performance of automatic gland segmentation is finally improved. The algorithm is verified on Warwick Qu-Dataset. Compared with the current popular segmentation algorithms, our algorithm has good performance in Dice coefficient, F1 score, and Hausdorff distance on different types of test sets. Songtao Ding, Hongyu Wang 0007, Hu Lu, Michele Nappi, Shaohua Wan 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Ollivier-Ricci Curvature for Head Pose Estimation from a Single ImageabstractHead pose estimation is not only a crucial challenge for many real-world applications, such as driver attention detection analysis, but it represents an interesting strategy to support biometric frameworks as well. This paper aims to estimate head pose from a single image by applying notions of network curvature. In the real world, many complex networks have groups of nodes that are well connected to each other with significant functional roles. Similarly, the interactions of facial landmarks can be represented as complex dynamic systems modeled by weighted graphs. The functionality of such a system is therefore intrinsically linked to the topology and geometry of the underlying graph. In this work, using the geometric notion of Ollivier-Ricci curvature (ORC) on weighted graphs as input to the XGBoost regression model, we show that the intrinsic geometric basis of ORC offers a natural approach to discovering underlying common structure within a pool of poses. Experiments on the BIWI, AFLW2000 and Pointing '04 datasets show that the ORC_XGB method performs well compared to state-of-the-art methods, both landmark-based and image-only. Andrea F. Abate, Lucia Cascone, Riccardo Distasi, Michele Nappi |
IJCB | 4 |
| 2022 | Guest Editorial Introduction to the Special Issue on "Biometrics Based Methods for Healthcare Applications"
Michele Nappi, Hugo Proença 0001, Sambit Bakshi, Vittorio Murino |
Comput. Vis. Image Underst. | 1 |
| 2022 | Head pose estimation: An extensive survey on recent techniques and applications
Andrea F. Abate, Carmen Bisogni, Aniello Castiglione, Michele Nappi |
Pattern Recognit. | 4 |
| 2022 | Touch keystroke dynamics for demographic classification
Lucia Cascone, Michele Nappi, Fabio Narducci, Chiara Pero |
Pattern Recognit. Lett. | 2 |
| 2022 | ETCNN: Extra Tree and Convolutional Neural Network-based Ensemble Model for COVID-19 Tweets Sentiment Classification
Muhammad Umer 0001, Saima Sadiq, Hanen Karamti, Alá Abdulmajid Eshmawi, Michele Nappi, Muhammad Usman Sana, Imran Ashraf 0003 |
Pattern Recognit. Lett. | 5 |
| 2022 | DTPAAL: Digital Twinning Pepper and Ambient Assisted LivingabstractPepper is a humanoid robot capable of expressing body language, perceiving, and interacting with its surrounding environment, thanks to a wide set of sensors and actuators and exposing capabilities and high-level interfaces for natural interaction with humans. In this article, we present the development of VPepper, the Pepper virtual replica, by describing experiences focused on the interaction of the digital twin with the replicas of the smart objects in a smart home. Pepper robot has been featured with arms and hands, but its motors and actuators cannot support intensive experimental sessions and training procedures to learn how safely touch objects. Here, digital twin metaphor plays a crucial role. By a virtual and reliable replica of the robot, machine learning procedures can be seamlessly moved to/from the digital twin with a significant speedup and preventing the physical robot from deterioration. As a practical application, the reported case study is inspired to ambient-assisted living in elderly assistance. The experience, as well as the entire design and development process, has revealed VPepper and the smart environment to offer interesting opportunities for the physical accuracy of the simulation and for the availability of machine learning instruments that may be converted and adopted for real settings. A final empirical evaluation, performed involving 25 volunteer caregivers, confirms the perceived value and the potential usefulness of the system. Lucia Cascone, Michele Nappi, Fabio Narducci, Ignazio Passero |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | Guest Editorial: Biometrics in Industry 4.0: Open Challenges and Future Perspectives
Zhiwei Gao 0001, Aniello Castiglione, Michele Nappi |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | Guest Editorial Emerging IoT-Driven Smart Health: From Cloud to EdgeabstractThe papers in this special section focus on emerging Internet of Medical Things. Recent advances in advances in healthcare can be experienced with the development of smart sensorial things, Artificial Intelligence (AI), Machine Learning (ML), Deep Learning (DL), edge computing, Edge AI, 6G, cloud computing, and connected healthcare have attracted a great deal of attention and a wide range of views. However, the need to deliver real-time and accurate healthcare services to patients, while reducing costs is a challenging issue [1]. Especially, COVID-19 has recently demonstrated the importance of fast, comprehensive, and accurate intelligent healthcare involving different types of medical, physiological, and epidemiological investigation data to diagnose the virus. Smart health is a real-time, intelligent, ubiquitous healthcare service based on Internet of bioMedical Things (IoMT). With the rapid development of related technologies such as deep learning, edge computing and IoT, smart health is playing vital role in healthcare industry to increase the accuracy, reliability, and productivity of mobile sensory devices. Shaohua Wan 0001, Michele Nappi, Chen Chen 0001, Stefano Berretti |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | An End-to-End Curriculum Learning Approach for Autonomous Driving ScenariosabstractIn this work, we combine Curriculum Learning with Deep Reinforcement Learning to learn without any prior domain knowledge, an end-to-end competitive driving policy for the CARLA autonomous driving simulator. To our knowledge, we are the first to provide consistent results of our driving policy on all towns available in CARLA. Our approach divides the reinforcement learning phase into multiple stages of increasing difficulty, such that our agent is guided towards learning an increasingly better driving policy. The agent architecture comprises various neural networks that complements the main convolutional backbone, represented by a ShuffleNet V2. Further contributions are given by (i) the proposal of a novel value decomposition scheme for learning the value function in a stable way and (ii) an ad-hoc function for normalizing the growth in size of the gradients. We show both quantitative and qualitative results of the learned driving policy. Luca Anzalone, Paola Barra, Silvio Barra, Aniello Castiglione, Michele Nappi |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | Reinforced Curriculum Learning For Autonomous Driving In CarlaabstractAutonomous Vehicles promise to transport people in a safer, accessible, and even efficient way. Nowadays, real-world autonomous vehicles are build by large teams from big companies with a tremendous amount of engineering effort. Deep Reinforcement Learning can be used instead, without domain experts, to learn end-to-end driving policies. Here, we combine Curriculum Learning with deep reinforcement learning, in order to learn without any prior domain knowledge, an end-to-end competitive driving policy for the CARLA autonomous driving simulator. To our knowledge, this is the first work which provides consistent results of our driving policy on all the town scenarios provided by CARLA. Moreover, we point out two important issues in reinforcement learning: the former is about learning the value function in a stable way, whereas the latter is related to normalizing the learned advantage function. A proposal of a solution to these problems is provided. Luca Anzalone, Silvio Barra, Michele Nappi |
ICIP | 3 |
| 2021 | Fostering secure cross-layer collaborative communications by means of covert channels in MEC environments
Aniello Castiglione, Michele Nappi, Fabio Narducci, Chiara Pero |
Comput. Commun. | 2 |
| 2021 | Discrepancy detection between actual user reviews and numeric ratings of Google App store using deep learning
Saima Sadiq, Muhammad Umer 0001, Saleem Ullah, Seyedali Mirjalili, Vaibhav Rupapara, Michele Nappi |
Expert Syst. Appl. | 6 |
| 2021 | Attention monitoring for synchronous distance learning
Andrea F. Abate, Lucia Cascone, Michele Nappi, Fabio Narducci, Ignazio Passero |
Future Gener. Comput. Syst. | 3 |
| 2021 | ECB2: A novel encryption scheme using face biometrics for signing blockchain transactions
Carmen Bisogni, Gerardo Iovane, Riccardo Emanuele Landi, Michele Nappi |
J. Inf. Secur. Appl. | 4 |
| 2021 | Futuristic person re-identification over internet of biometrics things (IoBT): Technical potential versus practical reality
Nayan Kumar Subhashis Behera, Tanmay Kumar Behera, Michele Nappi, Sambit Bakshi, Pankaj Kumar Sa |
Pattern Recognit. Lett. | 3 |
| 2021 | Introduction to the special issue on "Biometrics in Smart Cities: Techniques and Applications (BI_SCI)"
Michele Nappi, Silvio Barra, Aniello Castiglione, Fabio Narducci, Pandi Vijayakumar |
Pattern Recognit. Lett. | 1 |
| 2021 | Scientific papers citation analysis using textual features and SMOTE resampling techniquesabstractAscertaining the impact of research is significant for the research community and academia of all disciplines. The only prevalent measure associated with the quantification of research quality is the citation-count. Although a number of citations play a significant role in academic research, sometimes citations can be biased or made to discuss only the weaknesses and shortcomings of the research. By considering the sentiment of citations and recognizing patterns in text can aid in understanding the opinion of the peer research community and will also help in quantifying the quality of research articles. Efficient feature representation combined with machine learning classifiers has yielded significant improvement in text classification. However, the effectiveness of such combinations has not been analyzed for citation sentiment analysis. This study aims to investigate pattern recognition using machine learning models in combination with frequency-based and prediction-based feature representation techniques with and without using Synthetic Minority Oversampling Technique (SMOTE) on publicly available citation sentiment dataset. Sentiment of citation instances are classified into positive, negative or neutral. Results indicate that the Extra tree classifier in combination with Term Frequency-Inverse Document Frequency achieved 98.26% accuracy on the SMOTE-balanced dataset. Muhammad Umer 0001, Saima Sadiq, Malik Muhammad Saad Missen, Zahid Hameed, Zahid Aslam, Muhammad Abubakar Siddique, Michele Nappi |
Pattern Recognit. Lett. | 7 |
| 2021 | Trustworthy Method for Person Identification in IIoT Environments by Means of Facial DynamicsabstractIn industrial Internet of Things (IIoT) environments, dependability of a complex manufacturing process in which human operators play a key role can be improved by identity recognition/authentication of whoever is involved in various stages of a production process, according to where and when he/she is supposed to be. To this aim, we propose an approach that exploits the dynamic appearance and the time-dependent local features characterizing the face of an individual during speech utterance with regard to their spatial and temporal components. The proposed method models these dynamic facial patterns captured from edge Internet of Things devices by means of the Local Binary Pattern on Three Orthogonal Planes descriptor, which effectively extract both face's local features and movement at the fog level of the architecture. A deep feedforward network available in the cloud is trained and optimized to match the extracted features to a reference database. The achieved results highlight state-of-the-art performances of the proposed method with regard to robustness and trustworthiness of identification, especially for challenging IIoT scenarios. Aniello Castiglione, Michele Nappi, Stefano Ricciardi |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | COVID-19: Automatic Detection of the Novel Coronavirus Disease From CT Images Using an Optimized Convolutional Neural NetworkabstractIt is widely known that a quick disclosure of the COVID-19 can help to reduce its spread dramatically. Transcriptase polymerase chain reaction could be a more useful, rapid, and trustworthy technique for the evaluation and classification of the COVID-19 disease. Currently, a computerized method for classifying computed tomography (CT) images of chests can be crucial for speeding up the detection while the COVID-19 epidemic is rapidly spreading. In this article, the authors have proposed an optimized convolutional neural network model (ADECO-CNN) to divide infected and not infected patients. Furthermore, the ADECO-CNN approach is compared with pretrained convolutional neural network (CNN)-based VGG19, GoogleNet, and ResNet models. Extensive analysis proved that the ADECO-CNN-optimized CNN model can classify CT images with 99.99% accuracy, 99.96% sensitivity, 99.92% precision, and 99.97% specificity. Aniello Castiglione, Pandi Vijayakumar, Michele Nappi, Saima Sadiq, Muhammad Umer 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2021 | FASHE: A FrActal Based Strategy for Head Pose EstimationabstractHead pose estimation (HPE) represents a topic central to many relevant research fields and characterized by a wide application range. In particular, HPE performed using a singular RGB frame is particular suitable to be applied at best-frame-selection problems. This explains a growing interest witnessed by a large number of contributions, most of which exploit deep learning architectures and require extensive training sessions to achieve accuracy and robustness in estimating head rotations on three axes. However, methods alternative to machine learning approaches could be capable of similar if not better performance. To this regard, we present FASHE, an approach based on partitioned iterated function systems (PIFS) to represent auto-similarities within face image through a contractive affine function transforming the domain blocks extracted only once by a single frontal reference image, in a good approximation of the range blocks which the target image has been partitioned into. Pose estimation is achieved by finding the closest match between fractal code of target image and a reference array by means of Hamming distance. The results of experiments conducted exceed the state of the art on both Biwi and Ponting'04 datasets as well as approaching those of the best performing methods on the challenging AFLW2000 database. In addition, the applications to GOTCHA Video Dataset demonstrate that FASHE successfully operates in-the-wild. Carmen Bisogni, Michele Nappi, Chiara Pero, Stefano Ricciardi |
IEEE Trans. Image Process. | 2 |
| 2021 | Waiting for Tactile: Robotic and Virtual Experiences in the FogabstractSocial robots adopt an emotional touch to interact with users inducing and transmitting humanlike emotions. Natural interaction with humans needs to be in real time and well grounded on the full availability of information on the environment. These robots base their way of communicating on direct interaction (touch, listening, view), supported by a range of sensors on the surrounding environment that provide a radially central and partial knowledge on it. Over the past few years, social robots have been demonstrated to implement different features, going from biometric applications to the fusion of machine learning environmental information collected on the edge. This article aims at describing the experiences performed and still ongoing and characterizes a simulation environment developed for the social robot Pepper that aims to foresee the new scenarios and benefits that tactile connectivity will enable. Lucia Cascone, Aniello Castiglione, Michele Nappi, Fabio Narducci, Ignazio Passero |
ACM Trans. Internet Techn. | 3 |
| 2020 | DELEX: a DEep Learning Emotive eXperience: Investigating empathic HCIabstractRecent advances in Machine Learning have unveiled interesting possibilities for real-time investigating about user characteristics and expressions like, but not limited to, age, sex, body posture, emotions and moods. These new opportunities lay the foundations for new HCI tools for interactive applications that adopt user emotions as a communication channel. Andrea F. Abate, Aniello Castiglione, Michele Nappi, Ignazio Passero |
AVI | 3 |
| 2020 | HP2IFS: Head Pose estimation exploiting Partitioned Iterated Function SystemsabstractEstimating the actual head orientation from 2D images, with regard to its three degrees of freedom, is a well known problem that is highly significant for a large number of applications involving head pose knowledge. Consequently, this topic has been tackled by a plethora of methods and algorithms the most part of which exploits neural networks. Machine learning methods, indeed, achieve accurate head rotation values yet require an adequate training stage and, to that aim, a relevant number of positive and negative examples. In this paper we take a different approach to this topic by using fractal coding theory and particularly Partitioned Iterated Function Systems to extract the fractal code from the input head image and to compare this representation to the fractal code of a reference model through Hamming distance. According to experiments conducted on both the BIWI and the AFLW2000 databases, the proposed PIFS based head pose estimation method provides accurate yaw/pitch/roll angular values, with a performance approaching that of state of the art of machine-learning based algorithms and exceeding most of non-training based approaches. Carmen Bisogni, Michele Nappi, Chiara Pero, Stefano Ricciardi |
ICPR | 2 |
| 2020 | A Comparison of Neural Network Approaches for Melanoma ClassificationabstractMelanoma is the deadliest form of skin cancer and it is diagnosed mainly visually, starting from initial clinical screening and followed by dermoscopic analysis, biopsy and histopathological examination. A dermatologist's recognition of melanoma may be subject to errors and may take some time to diagnose it. In this regard, deep learning can be useful in the study and classification of skin cancer. In particular, by classifying images with Deep Neural Network methodologies, it is possible to obtain comparable or even superior results compared to those of dermatologists. In this paper, we propose a methodology for the classification of melanoma by adopting different deep learning techniques applied to a common dataset, composed of images from the ISIC dataset and consisting of different types of skin diseases, including melanoma on which we applied a specific pre-processing phase. In particular, a comparison of the results is performed in order to select the best effective neural network to be applied to the problem of recognition and classification of melanoma. Moreover, we also evaluate the impact of the preprocessing phase on the final classification. Different metrics such as accuracy, sensitivity, and specificity have been selected to assess the goodness of the adopted neural networks and compare them also with the manual classification of dermatologists. Maria Frasca, Michele Nappi, Michele Risi, Genny Tortora, Alessia Auriemma Citarella |
ICPR | 2 |
| 2020 | An attention recurrent model for human cooperation detection
David Freire-Obregón, Modesto Castrillón-Santana, Paola Barra, Carmen Bisogni, Michele Nappi |
Comput. Vis. Image Underst. | 5 |
| 2020 | Demographic classification through pupil analysis
Virginio Cantoni, Lucia Cascone, Michele Nappi, Marco Porta |
Image Vis. Comput. | 3 |
| 2020 | SAFFO: A SIFT based approach for digital anastylosis for fresco recOnstruction
Paola Barra, Silvio Barra, Michele Nappi, Fabio Narducci |
Pattern Recognit. Lett. | 3 |
| 2020 | Pupil size as a soft biometrics for age and gender classification
Lucia Cascone, Carlo Maria Medaglia, Michele Nappi, Fabio Narducci |
Pattern Recognit. Lett. | 3 |
| 2020 | Web-Shaped Model for Head Pose Estimation: An Approach for Best Exemplar SelectionabstractHead pose estimation is a sensitive topic in video surveillance/smart ambient scenarios since head rotations can hide/distort discriminative features of the face. Face recognition would often tackle the problem of video frames where subjects appear in poses making it quite impossible. In this respect, the selection of the frames with the best face orientation can allow triggering recognition only on these, therefore decreasing the possibility of errors. This paper proposes a novel approach to head pose estimation for smart cities and video surveillance scenarios, aiming at this goal. The method relies on a cascade of two models: the first one predicts the positions of 68 well-known face landmarks; the second one applies a web-shaped model over the detected landmarks, to associate each of them to a specific face sector. The method can work on detected faces at a reasonable distance and with a resolution that is supported by several present devices. Results of experiments executed over some classical pose estimation benchmarks, namely Point '04, Biwi, and AFLW datasets show good performance in terms of both pose estimation and computing time. Further results refer to noisy images that are typical of the addressed settings. Finally, examples demonstrate the selection of the best frames from videos captured in video surveillance conditions. Paola Barra, Silvio Barra, Carmen Bisogni, Maria De Marsico, Michele Nappi |
IEEE Trans. Image Process. | 5 |
| 2020 | An Encryption Approach Using Information Fusion Techniques Involving Prime Numbers and Face BiometricsabstractThe work shows a novel solution to create an access key which can be used within the transactions of electronic currencies, blockchain as well as in the field of computer security to guarantee a high level of secrecy, but also, with a high level of certainty, to provide a person identity through Information Fusion (IF) techniques and biometric data encryption. Specifically, two non-connected areas have been joined, Face Biometrics and Public-key Cryptography. This choice was taken in order to get through the limits these two approaches have found singularly and to give a suitable solution in the context of electronic and digital exchanges (electro-currencies, Internet of Things). An innovative and original algorithm has been developed, which can do fusion operations between Face Biometrics and numerical data, that is an algorithm of Hybrid Information Fusion, named FIF (Face Information Fusion). We decided to use a digital face as a biometric component, and the product of two prime numbers as a numerical component, that is the module in RSA algorithm. Gerardo Iovane, Carmen Bisogni, Luigi De Maio, Michele Nappi |
IEEE Trans. Sustain. Comput. | 4 |
| 2019 | Dynamic Facial Features for Inherently Safer Face RecognitionabstractAmong the many known type of intra-class variations, facial expressions are considered particularly challenging, as witnessed by the large number of methods that have been proposed to cope with them. The idea inspiring this work is that dynamic facial features (DFF) extracted from facial expressions while a sentence is pronounced, could possibly represent a salient and inherently safer biometric identifier, due to the greater difficulty in forging a time variable descriptor instead of a static one. We therefore investigated on how a set of geometrical features, defined as distances between landmarks located in the lower half of face, changes across time while a sentence is uttered to find the most effective yet compact representation. The features vectors built upon these time-series were used to train a deep feed-forward neural network on the OuluVS visual-speech database. Testing in identification modality resulted in 98.2% of average recognition accuracy, 0.64% of equal error rate and a remarkable robustness to how the sentence is pronounced. Davide Iengo, Michele Nappi, Stefano Ricciardi, Davide Vanore |
ICIP | 2 |
| 2019 | Biometric data on the edge for secure, smart and user tailored access to cloud services
Silvio Barra, Aniello Castiglione, Fabio Narducci, Maria De Marsico, Michele Nappi |
Future Gener. Comput. Syst. | 5 |
| 2019 | A hand-based biometric system in visible light for mobile environmentsabstractThe analysis of the shape and geometry of the human hand has long represented an attractive field of research to address the needs of digital image forensics. Over recent years, it has also turned out to be effective in biometrics, where several innovative research lines are pursued. Given the widespread diffusion of mobile and portable devices, the possibility of checking the owner identity and controlling the access to the device by the hand image looks particularly attractive and less intrusive than other biometric traits. This encourages new research to tacklethe present limitations. The proposed work implements the complete architecture of a mobile hand recognition system, which uses the camera of the mobile device for the acquisition of the hand in visible light spectrum. The segmentation of the hand starts from the detection of the convexities and concavities defined by the fingers, and allows extracting 57 different features from the hand shape. The main contributions of the paper develop along two directions. First, dimensionality reduction methods are investigated, in order to identify subsets of features including only the most discriminating and robust ones. Second, different matching strategies are compared. The proposed method is tested over a dataset of hands from 100 subjects. The best obtained Equal Error Rate is 0.52%. Results demonstrate that discarding features that are more prone to distortions allows lighter processing, but also produces better performance than using the full set of features. This confirms the feasibility of such an approach on mobile devices, and further suggests to adopt it even in more traditional settings. Silvio Barra, Maria De Marsico, Michele Nappi, Fabio Narducci, Daniel Riccio |
Inf. Sci. | 3 |
| 2019 | F-FID: fast fuzzy-based iris de-noising for mobile security applications
Silvio Barra, Carmen Bisogni, Michele Nappi, Stefano Ricciardi |
Multim. Tools Appl. | 3 |
| 2019 | Question action relevance and editing for visual question answering
Andeep S. Toor, Harry Wechsler, Michele Nappi |
Multim. Tools Appl. | 3 |
| 2019 | Biometric surveillance using visual question answering
Andeep S. Toor, Harry Wechsler, Michele Nappi |
Pattern Recognit. Lett. | 3 |
| 2019 | I-Am: Implicitly Authenticate Me - Person Authentication on Mobile Devices Through Ear Shape and Arm GestureabstractToday, identity verification is required in many common activities, and it is arguably true that most people would like to be authenticated in the easiest and most transparent way, without having to remember a personal identification number. To this regard, this paper presents a multibiometric system based on the observation that the instinctive gesture of responding to a phone call can be used to capture two different biometrics, namely ear and arm gesture, which are complementary due to their, respectively, physical and behavioral nature. We conducted a comprehensive set of experiments aimed at assessing the contribution of each of the two biometrics as well as the advantage in their fusion to the system's overall performance. Experiments also provide objective measurement of both saliency and correlation of data captured by each sensor involved (accelerometer, gyroscope, and camera) according to various features extraction, features matching, and data-fusion techniques. The reports provide evidences about the potential of the proposed system and method for user authentication “in-the-wild,” whilst its eventual usage for person identification is also investigated. All of the experiments have been carried out on a specifically built, publicly available ear-arm database, including multibiometric captures of more than 100 subjects performed during different sessions, that represents an additional contribution of this paper. Andrea F. Abate, Michele Nappi, Stefano Ricciardi |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2018 | What are you doing while answering your smartphone?abstractContext awareness is major component of Ambient Intelligence. In fact, Ambient Intelligent environments are designed to combine ubiquity, awareness, intelligence and natural interaction. Awareness is defined as the ability by the system to locate and recognize people and objects, and their intentions. Then, intelligence is the ability of the system to analyze the detected context and to adapt its behavior to people and situations, and to learn over time, in order to provide users with personalized services. These concepts date back to late'90, but nowadays the widespread and ubiquitous availability of mobile devices, equipped with several different sensors, allows to put them into practice in a number of unexpected ways. This work presents a preliminary investigation on the possibility to use some of the smartphone sensors, namely the accelerometer and the gyroscope, to identify the bodily context when the user lifts the device to answer a call. The arm gesture, i.e., the way it is performed, is classified into 4 different states: while standing, sitting, walking or running. This information can be used to trigger context-sensitive system actions. Andrea F. Abate, Michele Nappi, Silvio Barra, Maria De Marsico |
ICPR | 2 |
| 2018 | Fast QuadTree-Based Pose Estimation for Security Applications Using Face Biometrics
Paola Barra, Carmen Bisogni, Michele Nappi, Stefano Ricciardi |
NSS | 3 |
| 2018 | Visual Question Authentication Protocol (VQAP)
Andeep S. Toor, Harry Wechsler, Michele Nappi, Kim-Kwang Raymond Choo |
Comput. Secur. | 3 |
| 2018 | Context awareness in biometric systems and methods: State of the art and future scenarios
Michele Nappi, Stefano Ricciardi, Massimo Tistarelli |
Image Vis. Comput. | 1 |
| 2018 | Insights into the results of MICHE I - Mobile Iris CHallenge Evaluation
Maria De Marsico, Michele Nappi, Fabio Narducci, Hugo Proença 0001 |
Pattern Recognit. | 2 |
| 2018 | Introduction to the special issue on integrating biometrics and forensics
Michele Nappi, Nasir Memon, Daniel Riccio, Andreas Uhl |
Pattern Recognit. Lett. | 1 |
| 2017 | MOHAB: Mobile Hand-Based Biometric Recognition
Silvio Barra, Maria De Marsico, Michele Nappi, Fabio Narducci, Daniel Riccio |
GPC | 3 |
| 2017 | An Efficient Implementation of the Algorithm by Lukáš et al. on Hadoop
Giuseppe Cattaneo, Umberto Ferraro Petrillo, Michele Nappi, Fabio Narducci, Gianluca Roscigno |
GPC | 3 |
| 2017 | Secure User Authentication on Smartphones via Sensor and Face Recognition on Short Video Clips
Chiara Galdi, Michele Nappi, Jean-Luc Dugelay |
GPC | 2 |
| 2017 | MEG: Texture operators for multi-expert gender classification
Modesto Castrillón-Santana, Maria De Marsico, Michele Nappi, Daniel Riccio |
Comput. Vis. Image Underst. | 3 |
| 2017 | Special issue on Best of Biometrics 2015
Massimo Tistarelli, J. Ross Beveridge, Patrick J. Flynn, Michele Nappi |
Image Vis. Comput. | 4 |
| 2017 | Fusion of physiological measures for multimodal biometric systems
Silvio Barra, Andrea Casanova, Matteo Fraschini, Michele Nappi |
Multim. Tools Appl. | 4 |
| 2017 | Leveraging implicit demographic information for face recognition using a multi-expert system
Maria De Marsico, Michele Nappi, Daniel Riccio, Harry Wechsler |
Multim. Tools Appl. | 2 |
| 2017 | "Mobile Iris CHallenge Evaluation part II (MICHE II)"
Maria De Marsico, Michele Nappi, Hugo Proença 0001 |
Pattern Recognit. Lett. | 2 |
| 2017 | Results from MICHE II - Mobile Iris CHallenge Evaluation II
Maria De Marsico, Michele Nappi, Hugo Proença 0001 |
Pattern Recognit. Lett. | 2 |
| 2016 | Mobile Iris CHallenge Evaluation II: Results from the ICPR competitionabstractThe growing interest for mobile biometrics stems from the increasing need to secure personal data and services, which are often stored or accessed from there. Modern user mobile devices, with acquisition and computation resources to support related operations, are nowadays widely available. This makes this research topic very attracting and promising. Iris recognition plays a major role in this scenario. However, mobile biometrics still suffer from some hindering factors. The resolution of captured images and the computational power are not comparable to desktop systems yet. Furthermore, the acquisition setting is generally uncontrolled, with users who are not that expert to autonomously generate biometric samples of sufficient quality. Mobile Iris CHallenge Evaluation aims at providing a testbed to assess the progress of mobile iris recognition, and to evaluate the extent of its present limitations. This paper presents the results of the competition launched at the 2016 edition of the International Conference on Pattern Recognition (ICPR). Modesto Castrillón-Santana, Maria De Marsico, Michele Nappi, Fabio Narducci, Hugo Proença 0001 |
ICPR | 3 |
| 2016 | Smartphone enabled person authentication based on ear biometrics and arm gestureabstractSmartphones are arguably candidates to become the platform of choice for ubiquitous biometric-based identity verification, thanks to their embedded sensors, reasonably good computing power and widespread diffusion. While applications of the most established biometrics like face, fingerprint and even iris have already been proposed on mobile devices, other less exploited identifiers could also be worth investigating. To this regard, a novel multi-modal approach to person authentication based on ear biometrics and gesture analysis is proposed in this paper. The idea is to coupling the discriminant power of ear, captured during the act of responding to a phone call, with the user's arm dynamics affecting the smartphone motion pattern due to behavioral and anatomical characteristics involved in this gesture. According to experiments conducted on a specifically built multi-modal database comprising a hundred subjects, we confirm that the “responding gesture” has significant discriminating power and combined to ear features provides even greater robustness and accuracy in mobile authentication scenarios. Andrea F. Abate, Michele Nappi, Stefano Ricciardi |
SMC | 2 |
| 2016 | Multi indicator approach via mathematical inference for price dynamics in information fusion context
Gerardo Iovane, Antonino Amorosia, Marco Leone, Michele Nappi, Genny Tortora |
Inf. Sci. | 4 |
| 2016 | Deceiving faces: When plastic surgery challenges face recognition
Michele Nappi, Stefano Ricciardi, Massimo Tistarelli |
Image Vis. Comput. | 1 |
| 2016 | Guest Editorial: Augmented Reality Based Framework for Multimedia Training and Learning
Andrea F. Abate, Michele Nappi |
Multim. Tools Appl. | 2 |
| 2016 | WIRE: Watershed based iris recognition
Maria Frucci, Michele Nappi, Daniel Riccio, Gabriella Sanniti di Baja |
Pattern Recognit. | 2 |
| 2016 | Multimodal authentication on smartphones: Combining iris and sensor recognition for a double check of user identity
Chiara Galdi, Michele Nappi, Jean-Luc Dugelay |
Pattern Recognit. Lett. | 2 |
| 2016 | Eye movement analysis for human authentication: a critical survey
Chiara Galdi, Michele Nappi, Daniel Riccio, Harry Wechsler |
Pattern Recognit. Lett. | 2 |
| 2016 | Towards demographic categorization using gaze analysis
Chiara Galdi, Harry Wechsler, Virginio Cantoni, Marco Porta, Michele Nappi |
Pattern Recognit. Lett. | 5 |
| 2015 | Entropy-Based Automatic Segmentation and Extraction of Tumors from Brain MRI Images
Maria De Marsico, Michele Nappi, Daniel Riccio |
CAIP (2) | 2 |
| 2015 | Hybrid multi-sensor tracking system for field-deployable mixed reality environmentabstractMixed reality has been proposed for industrial applications such as AR-assisted manteinance and repair in which real equipment are augmented by visual aids. A possible complementary use of this technology for the same context can be achieved by performing virtual training in a “mixed reality environment” enabling to observe actual-size virtual replicas of real equipment within the physical space, providing an advantage in terms of perceptual impact and learning efficacy compared to traditional computer-based-training techniques. To this aim we present an effective hybrid tracking approach, exploiting both optical markers and inertial sensors to measure user's head position and rotation, providing both accuracy and robustness to fast head movements. A field-deployable tridimensionally-arranged marker-set, specifically designed for this kind of application, represents a further contribute of this paper, providing reliable close-to-medium distance optical tracking capabilities and increased robustness to challenging lighting conditions. First experiments confirm the potential of the proposed solution for this context and for many other applications as well. Andrea F. Abate, Virginio Cantoni, Michele Nappi, Fabio Narducci, Stefano Ricciardi |
INDIN | 3 |
| 2015 | GANT: Gaze analysis technique for human identification
Virginio Cantoni, Chiara Galdi, Michele Nappi, Marco Porta, Daniel Riccio |
Pattern Recognit. | 3 |
| 2015 | Robust face recognition after plastic surgery using region-based approaches
Maria De Marsico, Michele Nappi, Daniel Riccio, Harry Wechsler |
Pattern Recognit. | 2 |
| 2015 | Guest editorial introduction to the special executable issue on "Mobile Iris CHallenge Evaluation part I (MICHE I)"
Maria De Marsico, Michele Nappi, Hugo Proença 0001 |
Pattern Recognit. Lett. | 2 |
| 2015 | Mobile Iris Challenge Evaluation (MICHE)-I, biometric iris dataset and protocols
Maria De Marsico, Michele Nappi, Daniel Riccio, Harry Wechsler |
Pattern Recognit. Lett. | 2 |
| 2014 | IDEM: Iris DEtection on Mobile DevicesabstractIn this paper an iris detection scheme for noisy images acquired by means of mobile devices is presented. Iris segmentation is accomplished by exploiting the use of the watershed transform with the purpose of identifying the iris boundary as much precisely as possible. After a pre-processing step aimed at color/illumination correction, the watershed transform is computed and suitably binarized. Circle fitting is then accomplished to identify the limbus boundary by using curvature approximation and a cost function for circle scoring. The watershed transform is furthermore employed to distinguish, in the zone delimited by the best fitting circle, the regions actually belonging to the iris from those belonging to eyelids and sclera. Finally, pupil detection is accomplished by means of circle fitting and by using a voting function based on homogeneity and separability criteria. The suggested iris detection scheme has a positive impact on an the accuracy in computing the iris code, which has in turn a positive impact on the performance of iris recognition. Maria Frucci, Chiara Galdi, Michele Nappi, Daniel Riccio, Gabriella Sanniti di Baja |
ICPR | 3 |
| 2014 | Complex numbers as a Compact Way to Represent scores and their reliability in Recognition by Multi-Biometric FusionabstractMulti-biometric systems are a powerful solution to deal with limitations of single classifiers, therefore improving the final recognition accuracy. The sub-systems composing the final architecture often return supplementary indices of input quality and/or of response reliability, which further qualify each recognition score. These indices can enter different information fusion policies. First, they can be used as weights for the fusion of the corresponding scores, in such a way that less trustworthy responses have a lower influence. Alternatively, they can be used to drive the selection of a subset of systems actually enabled for each fusion operation. The present work discusses their appropriate combination with respective scores, to obtain single values which are easier to handle and compare. It is worth underlining the different nature of quality and reliability measures. The quality estimation of input samples requires a complex analysis of environmental conditions, including capture sensors, besides computations over acquired data. Reliability of a system estimates its ability to return a correct response. As an alternative to combination, some solutions rather estimate the joint distributions of conditional probabilities of the scores from the single subsystems. These solutions require training through a huge number of samples. Furthermore, they assume stable score distributions. Our unified representation of the recognition score and of the corresponding quality/reliability value into a single complex number provides simplification and speed up of fusion of multi-classifier results. It also allows to devise procedures to readily compare the performance of different modules in a multi-biometric system, given that there is no natural ordering of these pairs of values of different nature. Moreover our method achieves performance comparable to top performing schemes, yet does not require a prior estimation of (joint) score distributions. As a matter of fact, though representing an upper bound to the obtainable performance, Likelihood ratio has the limit to require an accurate estimation of score distributions, while our approach relies on the reliability of each single response. This feature is very interesting when the set of relevant subjects may present significant variations over time. Silvio Barra, Maria De Marsico, Michele Nappi, Daniel Riccio |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2014 | FIRME: Face and Iris Recognition for Mobile Engagement
Maria De Marsico, Chiara Galdi, Michele Nappi, Daniel Riccio |
Image Vis. Comput. | 3 |
| 2014 | ES-RU: an entropy based rule to select representative templates in face surveillance
Maria De Marsico, Michele Nappi, Daniel Riccio |
Multim. Tools Appl. | 2 |
| 2014 | Introduction to the Special Section on Biometric Systems and ApplicationsabstractNowadays, biometrics is an important technological area receiving continuously growing interest from academia, industry, government, and the general public, due to the criticality and the social impact of its applications. Biometric systems are in fact rapidly being adopted in a wide variety of applications such as security, ambient intelligence, electronic and physical access control, digital rights management, background checking and defense, medical diagnosis as well as for adaptive environments. Michele Nappi, Vincenzo Piuri, Tieniu Tan, David Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2013 | Fusion of Multi-biometric Recognition Results by Representing Score and Reliability as a Complex Number
Maria De Marsico, Michele Nappi, Daniel Riccio |
CIARP (2) | 2 |
| 2013 | Entropy based Biometric Template Clustering
Michele Nappi, Daniel Riccio, Maria De Marsico |
ICPRAM | 1 |
| 2013 | Robust Face Recognition for Uncontrolled Pose and Illumination ChangesabstractFace recognition has made significant advances in the last decade, but robust commercial applications are still lacking. Current authentication/identification applications are limited to controlled settings, e.g., limited pose and illumination changes, with the user usually aware of being screened and collaborating in the process. Among others, pose and illumination changes are limited. To address challenges from looser restrictions, this paper proposes a novel framework for real-world face recognition in uncontrolled settings named Face Analysis for Commercial Entities (FACE). Its robustness comes from normalization (“correction”) strategies to address pose and illumination variations. In addition, two separate image quality indices quantitatively assess pose and illumination changes for each biometric query, before submitting it to the classifier. Samples with poor quality are possibly discarded or undergo a manual classification or, when possible, trigger a new capture. After such filter, template similarity for matching purposes is measured using a localized version of the image correlation index. Finally, FACE adopts reliability indices, which estimate the “acceptability” of the final identification decision made by the classifier. Experimental results show that the accuracy of FACE (in terms of recognition rate) compares favorably, and in some cases by significant margins, against popular face recognition methods. In particular, FACE is compared against SVM, incremental SVM, principal component analysis, incremental LDA, ICA, and hierarchical multiscale local binary pattern. Testing exploits data from different data sets: CelebrityDB, Labeled Faces in the Wild, SCface, and FERET. The face images used present variations in pose, expression, illumination, image quality, and resolution. Our experiments show the benefits of using image quality and reliability indices to enhance overall accuracy, on one side, and to provide for individualized processing of biometric probes for better decision-making purposes, on the other side. Both kinds of indices, owing to the way they are defined, can be easily integrated within different frameworks and off-the-shelf biometric applications for the following: 1) data fusion; 2) online identity management; and 3) interoperability. The results obtained by FACE witness a significant increase in accuracy when compared with the results produced by the other algorithms considered. Maria De Marsico, Michele Nappi, Daniel Riccio, Harry Wechsler |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2012 | CABALA - Collaborative architectures based on biometric adaptable layers and activities
Maria De Marsico, Michele Nappi, Daniel Riccio |
Pattern Recognit. | 2 |
| 2012 | Novel pattern recognition-based methods for re-identification in biometric context
Mislav Grgic, Michele Nappi, Harry Wechsler |
Pattern Recognit. Lett. | 2 |
| 2012 | Noisy Iris Recognition Integrated Scheme
Maria De Marsico, Michele Nappi, Daniel Riccio |
Pattern Recognit. Lett. | 2 |
| 2012 | Robust re-identification using randomness and statistical learning: Quo vadis
Michele Nappi, Harry Wechsler |
Pattern Recognit. Lett. | 1 |
| 2011 | VIVIE: A video-surveillance indexer via identity extractionabstractVIVIE is a system for video sequence indexing. Video frames are annotated according to the identities of appearing subjects. Different interacting modules perform different processing steps, and each can be possibly substituted with a different one performing the same task using a different method. Classification and clustering are the most challenging activities. Differently from most existing similar systems, VIVIE accounts for the concomitant appearance of two identities in the same clip, and exploits such information for identity mapping. VIVIE was tested on 7 video clips and on a subset of the SCFace database to assess its performances. Andrea F. Abate, Maria De Marsico, Michele Nappi, Daniel Riccio |
ICME | 3 |
| 2011 | NABS: Novel Approaches for Biometric SystemsabstractResearch on biometrics has noticeably increased. However, no single bodily or behavioral feature is able to satisfy acceptability, speed, and reliability constraints of authentication in real applications. The present trend is therefore toward multimodal systems. In this paper, we deal with some core issues related to the design of these systems and propose a novel modular framework, namely, novel approaches for biometric systems (NABS) that we have implemented to address them. NABS proposal encompasses two possible architectures based on the comparative speeds of the involved biometries. It also provides a novel solution for the data normalization problem, with the new quasi-linear sigmoid (QLS) normalization function. This function can overcome a number of common limitations, according to the presented experimental comparisons. A further contribution is the system response reliability (SRR) index to measure response confidence. Its theoretical definition allows to take into account the gallery composition at hand in assigning a system reliability measure on a single-response basis. The unified experimental setting aims at evaluating such aspects both separately and together, using face, ear, and fingerprint as test biometries. The results provide a positive feedback for the overall theoretical framework developed herein. Since NABS is designed to allow both a flexible choice of the adopted architecture, and a variable compositions and/or substitution of its optional modules, i.e., QLS and SRR, it can support different operational settings. Maria De Marsico, Michele Nappi, Daniel Riccio, Genny Tortora |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2010 | Virtual-ICSI: a visual-haptic interface for virtual training in intra cytoplasmic sperm injectionabstractVirtual simulators have been used in the last twenty years for applications ranging from flight simulation to computer-based training, just to name a few. More recently a new level of simulation has been introduced thanks to haptic interfaces able to reproduce kinesthetic and/or tactile feedback, typically experimented during interaction with real-world objects. In this paper visual simulation and haptic interfaces are integrated in a novel training system for Intra Cytoplasmic Sperm Injection (ICSI), an in-vitro fertilization technique which is now a standard for the treatment of human infertility. We describe a virtual micromanipulation simulator made by two hand-based Cyberforce haptic devices (a synthetic replica of the actual manipulation gear) and a visual-haptic engine simulating the shape and the dynamic behavior of the main components in the artificial fertilization process: the human egg, the selected sperm and the micro needles required to inject the latter into the egg's cytoplasm. Our first tests, conducted so far, are encouraging. Andrea F. Abate, Michele Nappi, Stefano Ricciardi, Genny Tortora, Stefano Levialdi, Maria De Marsico |
AVI | 2 |
| 2010 | Face: face analysis for Commercial EntitiesabstractThough face recognition gained significant attention and credibility in the last decade, quite few commercial applications were able to benefit from this. In this paper we propose FACE (Face analysis for Commercial Entities), a robust framework to address face analysis, aiming at supporting the activities of various Commercial Entities. In particular, we present two case studies, with the related experimental results which sustain the presented approach. Maria De Marsico, Michele Nappi, Daniel Riccio |
ICIP | 2 |
| 2010 | IS_IS: Iris Segmentation for Identification SystemsabstractAdvances in processing procedures make the iris a realistic candidate to the role of biometry of the future. Precise detection and segmentation for such biometry are a crucial ongoing research area. We propose an iris segmentation technique and show that it is more reliable than existent ones. Maria De Marsico, Michele Nappi, Daniel Riccio |
ICPR | 2 |
| 2010 | FARO: FAce Recognition Against Occlusions and Expression VariationsabstractFace recognition is widely considered as one of the most promising biometric techniques, allowing high recognition rates without being too intrusive. Many approaches have been presented to solve this special pattern recognition problem, also addressing the challenging cases of face changes, mainly occurring in expression, illumination, or pose. On the other hand, less work can be found in literature that deals with partial occlusions (i.e., sunglasses and scarves). This paper presents face recognition against occlusions and expression variations (FARO) as a new method based on partitioned iterated function systems (PIFSs), which is quite robust with respect to expression changes and partial occlusions. In general, algorithms based on PIFSs compute a map of self-similarities inside the whole input image, searching for correspondences among small square regions. However, traditional algorithms of this kind suffer from local distortions such as occlusions. To overcome such limitation, information extracted by PIFS is made local by working independently on each face component (eyes, nose, and mouth). Distortions introduced by likely occlusions or expression changes are further reduced by means of anad hocdistance measure. In order to experimentally confirm the robustness of the proposed method to both lighting and expression variations, as well as to occlusions, FARO has been tested using AR-faces database, one of the main benchmarks for the scientific community in this context. A further validation of FARO performances is provided by the experimental results produced on Face Recognition Grand Challenge database. Maria De Marsico, Michele Nappi, Daniel Riccio |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2009 | Fine: Fractal indexing based on neighborhood estimationabstractThe structure of a fractal based image indexing system is described. FINE implies image space linearization, a custom clustering strategy, Peano-serialized spatial addressing, ad-hoc heuristics for improving search and specially defined distance functions. The resulting system is invariant, or at least robust, to a large class of typical variations that appear in natural images including rotations, scaling, and changes in color or illumination. The performance of FINE is illustrated, discussed and compared with other present alternatives using standard and custom-based image databases. Michele Nappi, Daniel Riccio, Maria De Marsico |
ICIP | 1 |
| 2009 | Advanced Maintenance Simulation by Means of Hand-Based Haptic Interfaces
Michele Nappi, Luca Paolino, Stefano Ricciardi, Monica Sebillo, Giuliana Vitiello |
INTERACT (2) | 1 |
| 2009 | A Self-tuning People Identification System from Split Face Components
Maria De Marsico, Michele Nappi, Daniel Riccio |
PSIVT | 2 |
| 2007 | Embedding Linear Transformations in Fractal Image Coding
Michele Nappi, Daniel Riccio |
ACIVS | 1 |
| 2007 | Fast 3D Face Alignment and Improved Recognition Through Pyramidal Normal map MetricabstractFace's tri-dimensional shape represents a highly discriminating yet challenging biometric identifier due to different issues, some of which related to capture, alignment and normalization. This paper presents an improved normal map based face recognition approach, which relies on a novel method to automatically align a captured 3D face mesh to a reference template, allowing a more precise face comparison. The alignment algorithm exploits pyramidal-normal-map metric, a coarse to finer measurement of angular distance between two surfaces computed through normal maps with progressively increasing resolution. After the registration has been performed, the normalized face can be rapidly compared to any other template in the gallery database for authentication or identification purposes using standard normal map metric. The alignment approach avoids the need for a rough or manual face pre-alignment and maximizes recognition precision, requiring a fraction of the time needed by the iterative closest point (ICP) method to operate. We show preliminary experimental results on a 3D dataset featuring 235 different subjects. Andrea F. Abate, Michele Nappi, Stefano Ricciardi, Gabriele Sabatino |
ICIP (1) | 2 |
| 2007 | Rbs: a Robust Bimodal System for Face RecognitionabstractDuring the last few years, many algorithms have been proposed in particular for face recognition using classical 2-D images. However, it is necessary to deal with occlusions when the subject is wearing sunglasses, scarves and such. In the same way, ear recognition is arising as a new promising biometric for people recognition, even if the related literature appears to be somewhat underdeveloped. In this paper, several hybrid face/ear recognition systems are investigated. The system is based on IFS (Iterated Function Systems) theory that are applied on both face and ear resulting in a bimodal architecture. One advantage is that the information used for the indexing and recognition task of face/ear can be made local, and this makes the method more robust to possible occlusions. The distribution of similarities in the input images is exploited as a signature for the identity of the subject. The amount of information provided by each component of the face and the ear image has been assessed, first independently and then jointly. At last, results underline that the system significantly outperforms the existing approaches in the state of the art. Andrea F. Abate, Michele Nappi, Daniel Riccio, Genny Tortora |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2007 | 2D and 3D face recognition: A survey
Andrea F. Abate, Michele Nappi, Daniel Riccio, Gabriele Sabatino |
Pattern Recognit. Lett. | 2 |
| 2006 | Multi-Modal Face Recognition by Means of Augmented Normal Map and PCAabstractFace represents a rich biometric identifier whose potential in term of discriminating power has not been fully exploited yet. This paper addresses face recognition through a multi-modal approach operating on face's 3D (geometry) and 2D (skin texture) features by means of two different metrics: augmented normal map and principal component analysis. Augmented normal map includes shape (surface normals represented as 24 bit colour pixels) and texture info (additional 8 bit for skin colour) into one 32 bit image. The proposed two-staged method firstly performs a fast one-to-many comparison of facial geometry exploiting normal map metric. Then, to further improve recognition precision and reliability, best rank faces are compared to probe by PCA resulting in a final score. Other advantages are robustness to facial expressions and the ability to selectively filter face's non-skin regions (beard, moustaches). We include preliminary experimental results on a dataset of 101 textured 3D faces. Andrea F. Abate, Michele Nappi, Stefano Ricciardi, Gabriele Sabatino |
ICIP | 2 |
| 2006 | A Semantic View for Flexible Communication Models between Humans, Sensors and ActuatorsabstractIn this work, we present last extensions done on H2ML, a new approach to model human interaction models inside intelligent environments at different abstraction levels. H2ML has been defined by considering the exigency to design Ambient Intelligence (Ami) applications, by assuring transparency, uniformity and abstractness in bridging multiple sensors properties to flexible and personalized actuators that interact with the human choosing appropriate media communication strategies. Within this aim, H2ML can be viewed as a methodological approach, based on markup languages and fuzzy theory, that provides semantic representation to human attitudes acquired by the intelligent environment. H2ML works in bi-directional way: if from an one hand it builds semantic models of human behavior, from the other hand enables to impact on the most suitable control strategy to perform at actuator side, conciliating the action with the most suitable media "facilitator". Giovanni Acampora, Vincenzo Loia, Michele Nappi, Stefano Ricciardi |
SMC | 3 |
| 2006 | Face authentication using speed fractal technique
Andrea F. Abate, Riccardo Distasi, Michele Nappi, Daniel Riccio |
Image Vis. Comput. | 3 |
| 2006 | A range/domain approximation error-based approach for fractal image compressionabstractFractals can be an effective approach for several applications other than image coding and transmission: database indexing, texture mapping, and even pattern recognition problems such as writer authentication. However, fractal-based algorithms are strongly asymmetric because, in spite of the linearity of the decoding phase, the coding process is much more time consuming. Many different solutions have been proposed for this problem, but there is not yet a standard for fractal coding. This paper proposes a method to reduce the complexity of the image coding phase by classifying the blocks according to an approximation error measure. It is formally shown that postponing range\slash domain comparisons with respect to a preset block, it is possible to reduce drastically the amount of operations needed to encode each range. The proposed method has been compared with three other fractal coding methods, showing under which circumstances it performs better in terms of both bit rate and/or computing time. Riccardo Distasi, Michele Nappi, Daniel Riccio |
IEEE Trans. Image Process. | 2 |
| 2005 | Fast 3D face recognition based on normal mapabstractThis paper presents a 3D face recognition method aimed to biometric applications. The proposed method compares any two faces represented as 3D polygonal surfaces through their corresponding normal map, a bidimensional array which stores local curvature (mesh normals) as the pixel's RGB components of a color image. The recognition approach, based on the computation of a difference map resulting from the comparison of normal maps, is simple yet fast and accurate. A weighting mask, automatically generated for each subject using a set of expression variations, improves the robustness to a broad range of facial expressions. First results show the effectiveness of the method on a database of 3D faces featuring different genders, ages and expressions. Andrea F. Abate, Michele Nappi, Stefano Ricciardi, Gabriele Sabatino |
ICIP (2) | 2 |
| 2005 | An IFS based approach for face recognitionabstractNowadays face recognition is gaining great attention from the researcher respect to other biometrics, because it represents a good compromise between reliability and people acceptance. While this growth largely is driven by growing application demands, such as identification for law enforcement and authentication, for banking and security system access. The recognition task is difficult because of image variation in terms of position, size, expression, and pose. In this paper a new IFS based recognition method is presented. It exploits the IFS theory, largely studied in still image compression and indexing, but not enough for the face recognition task. Andrea F. Abate, Michele Nappi, Daniel Riccio, Genny Tortora |
ICIP (2) | 2 |
| 2004 | ABI: analogy-based indexing for content image retrieval
Maurizio Cibelli, Michele Nappi, Maurizio Tucci |
Image Vis. Comput. | 2 |
| 2003 | FIRE: fractal indexing with robust extensions for image databasesabstractAs already documented in the literature, fractal image encoding is a family of techniques that achieves a good compromise between compression and perceived quality by exploiting the self-similarities present in an image. Furthermore, because of its compactness and stability, the fractal approach can be used to produce a unique signature, thus obtaining a practical image indexing system. Since fractal-based indexing systems are able to deal with the images in compressed form, they are suitable for use with large databases. We propose a system called FIRE, which is then proven to be invariant under three classes of pixel intensity transformations and under geometrical isometries such as rotations by multiples of /spl pi//2 and reflections. This property makes the system robust with respect to a large class of image transformations that can happen in practical applications: the images can be retrieved even in the presence of illumination and/or color alterations. Additionally, the experimental results show the effectiveness of FIRE in terms of both compression and retrieval accuracy. Riccardo Distasi, Michele Nappi, Maurizio Tucci |
IEEE Trans. Image Process. | 2 |
| 2002 | HEAT: Hierarchical Entropy Approach for Texture Indexing in Image DatabasesabstractThis paper illustrates a method, called HEAT, for image indexing based on texture information. The texture's partitioning element is first put into 1-D form and then its Hierarchical Entropy-based Representation is obtained. This representation is used to index the texture in the space of features. The same representation is well suited for contour data, and it has invariance and robustness properties that make it attractive for incorporation into larger systems. A comparison with another performing method is carried out, and the experiments show that the two techniques have slightly different strong points, suggesting different fields of application. In the experimental section, a case study involving over 2500 mammographies from different sources is presented and discussed. Riccardo Distasi, Michele Nappi, Sergio Vitulano |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2000 | Speed-up in fractal image coding: comparison of methodsabstractFractal image compression has received much attention from the research community because of some desirable properties like resolution independence, fast decoding, and very competitive rate-distortion curves. Despite the advances made, the long computing times in the encoding phase still remain the main drawback of this technique. So far, several methods have been proposed in order to speed-up fractal image coding. We address the problem of choosing the best speed-up techniques for fractal image coding, comparing some of the most effective classification and feature vector methods--namely Fisher, Hurtgen, and Saupe--and a new feature vector coding scheme based on the block's mass center. Furthermore, we introduce two new coding schemes combining Saupe with Fisher, and Saupe with mass center coding scheme. Experimental results demonstrate both the superiority of feature vector techniques on classification and the effectiveness of combining Saupe and the mass center coding scheme, an approach that exhibits the best time-distortion curves. Mario Polvere, Michele Nappi |
IEEE Trans. Image Process. | 2 |
| 1999 | IME: an image management environment with content-based access
Andrea F. Abate, Michele Nappi, Genny Tortora, Maurizio Tucci |
Image Vis. Comput. | 2 |
| 1999 | Linear prediction image coding using iterated function systems
Michele Nappi, Domenico Vitulano |
Image Vis. Comput. | 1 |
| 1998 | Edge Detection: Local and Global OperatorsabstractThe paper describes a technique called ISE for image segmentation using entropy. The relation between the entropy of an image domain and the entropy of its subdomains is explored as a uniformity predicate. Such entropy is obtained from the analysis of the image histogram associating a Gaussian distribution to the maximum frequency of gray levels. In order to implement the model, we have introduced a well-known technique of Problem Solving. In our model, the most important roles are played by the Evaluation Function (EF) and the Control Strategy. The EF is related to the ratio between the entropy of one region or zone of the picture and the entropy of the entire picture, while the Control Strategy determines the optimal path in the search tree (quadtree) so that the nodes in the optimal path have minimal entropy. The paper shows some comparisons between ISE and classical edge detection techniques. Sergio Vitulano, Michele Nappi, Cecilia Di Ruberto |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 1998 | FIRST: Fractal Indexing and Retrieval SysTem for Image Databases
Michele Nappi, Giuseppe Polese, Genny Tortora |
Image Vis. Comput. | 1 |
| 1997 | Different methods to segment biomedical images
Sergio Vitulano, Cecilia Di Ruberto, Michele Nappi |
Pattern Recognit. Lett. | 3 |
| 1997 | Color image coding combining linear prediction and iterated function systems
Laura Moltedo, Michele Nappi, Domenico Vitulano, Sergio Vitulano |
Signal Process. | 2 |
| 1997 | Image compression by B-tree triangular codingabstractThis paper describes an algorithm for still image compression called B-tree triangular coding (BTTC). The coding scheme is based on the recursive decomposition of the image domain into right-angled triangles arranged in a binary tree. The method is attractive because of its fast encoding, O(n log n), and decoding, /spl Theta/(n), where n is the number of pixels, and because it is easy to implement and to parallelize. Experimental studies indicate that BTTC produces images of satisfactory quality from a subjective and objective point of view, One advantage of BTTC over JPEG is its shorter execution time. Riccardo Distasi, Michele Nappi, Sergio Vitulano |
IEEE Trans. Commun. | 2 |
| 1996 | A B-tree based recursive technique for image codingabstractThis paper describes an algorithm for image compression called B-tree triangular coding (BTTC). An image is considered as a discrete 3D surface, which is then represented by a set of polyhedrons. The image domain is partitioned into right-angled triangles constituting the polyhedrons' bases. Each polyhedron is characterized by the vertices of its triangular upper face, near the approximated surface. During the approximation process the polyhedrons are organized into a B-tree that will be represented as a binary string; the B-tree's leaves are the polyhedrons needed to restore the image during decompression. The computing time is O(nlogn) for compression and /spl Theta/(n) for decompression, where n is the number of pixels. Especially in decompression, this is a very fast method if compared to the standard techniques (e.g. JPEG). Riccardo Distasi, Michele Nappi, Sergio Vitulano |
ICPR | 2 |
| 1996 | Edge detection using a new definition of entropyabstractThe paper describes a possible model of the human perceptive process. In this paper the relation between the entropy of an image domain and the entropy of its subdomains is explored as a uniformity predicate. Such entropy is obtained from the analysis of the image histogram associating a Gaussian distribution to the maximum frequency of grey levels. With the aim of implementing the model, we have introduced a well known technique of problem solving. The most important roles of our model are played by the evaluation function (EF) and the control strategy. So the EF is related to the ratio between the entropy of one region or zone of the picture and the entropy of the entire picture. The control strategy determines the optimal path in the quadtree so that the nodes of the optimal path have minimal entropy. The paper shows some comparisons between the method and classical edge detection techniques. Sergio Vitulano, Michele Nappi, Domenico Vitulano, C. Masuovito |
ICPR | 2 |