EDBT 2026 Demo / reviewers in the wild / expert
Carmen Bisogni
dblp:231/8694
· DBLP profile ↗
31ranked-venue papers
13as first author
25since 2021 · last 2026
0000-0003-1358-006XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 8 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 7 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explainable AI for multimodal stress detection: interpreting model decisions across physiological, video and audio modalities
Andrea F. Abate, Carmen Bisogni, Aniello Castiglione, Maddalena Migliaccio |
Multim. Tools Appl. | 2 |
| 2025 | OPD-Based Attribute-Oriented Concept Reduction for Cognitive DiagnosisabstractConcept reduction that preserves binary relations is an emerging reduction theory in the field of Formal Concept Analysis. Its core lies in reducing the number of concepts while ensuring that the original information is not lost, thereby significantly improving the efficiency of data processing. Based on Object Pictorial Diagram (OPD), this paper proposes a novel attribute-oriented concept reduction method that preserves complementary binary relations. First, this paper clarifies the definition of attribute-oriented concept reduction and presents a specific method for addressing it from the perspective of OPD. Against the backdrop of smart education's growing emphasis on data-driven decision-making, accurately diagnosing learners' knowledge states has become a core requirement for instructional reform and personalized tutoring. Practically, by integrating learners' response data to exercises, cognitive diagnosis is conducted by using the obtained attribute-oriented concept reduction results, enabling an in-depth analysis of learners' knowledge states and cognitive structures. Experimental results demonstrate that the proposed method exhibits high efficiency in both solving attribute-oriented concept reduction and performing cognitive diagnosis. The proposed method provides robust support for assessing learners' learning states and enhances the interpretability of various personalized learning applications. Fei Hao 0001, Qing Wan, Carmen Bisogni, Xu Zhang 0016, Lexi Xu |
HPCC | 4 |
| 2025 | Synthetic data sets for person Re-Identification: A critical analysis
Rita Delussu, Lorenzo Putzu, Fadi Boutros, Carmen Bisogni, Naser Damer, Giorgio Fumera |
Image Vis. Comput. | 4 |
| 2025 | Real time emotions recognition through facial expressions
Alisha Fida, Muhammad Umer 0001, Oumaima Saidani, Monia Hamdi, Khaled Alnowaiser, Carmen Bisogni, Andrea F. Abate, Imran Ashraf 0003 |
Multim. Tools Appl. | 6 |
| 2024 | Acoustic features analysis for explainable machine learning-based audio spoofing detectionabstractThe rapid evolution of synthetic voice generation and audio manipulation technologies poses significant challenges, raising societal and security concerns due to the risks of impersonation and the proliferation of audio deepfakes. This study introduces a lightweight machine learning (ML)-based framework designed to effectively distinguish between genuine and spoofed audio recordings. Departing from conventional deep learning (DL) approaches, which mainly rely on image-based spectrogram features or learning-based audio features, the proposed method utilizes a diverse set of hand-crafted audio features – such as spectral, temporal, chroma, and frequency-domain features – to enhance the accuracy of deepfake audio content detection. Through extensive evaluation and experiments on three well-known datasets, ASVSpoof2019, FakeAVCelebV2, and an In-The-Wild database, the proposed solution demonstrates robust performance and a high degree of generalization compared to state-of-the-art methods. In particular, our method achieved 89% accuracy on ASVSpoof2019, 94.5% on FakeAVCelebV2, and 94.67% on the In-The-Wild database. Additionally, the experiments performed on explainability techniques clarify the decision-making processes within ML models, enhancing transparency and identifying crucial features essential for audio deepfake detection. • Enhanced spoof audio detection via multi-feature integration. • Employed a lightweight ML framework for real-time applications. • Adopted subject-independent protocols to mitigate biometric bias. • Utilized Explainable AI (XAI) for transparent decision-making. Carmen Bisogni, Vincenzo Loia, Michele Nappi, Chiara Pero |
Comput. Vis. Image Underst. | 1 |
| 2024 | Multimodal Emotion Recognition via Convolutional Neural Networks: Comparison of different strategies on two multimodal datasetsabstractThe aim of this paper is to investigate emotion recognition using a multimodal approach that exploits convolutional neural networks (CNNs) with multiple input. Multimodal approaches allow different modalities to cooperate in order to achieve generally better performances because different features are extracted from different pieces of information. In this work, the facial frames, the optical flow computed from consecutive facial frames, and the Mel Spectrograms (from the word melody) are extracted from videos and combined together in different ways to understand which modality combination works better. Several experiments are run on the models by first considering one modality at a time so that good accuracy results are found on each modality. Afterward, the models are concatenated to create a final model that allows multiple inputs. For the experiments the datasets used are BAUM-1 ((Bahçeşehir University Multimodal Affective Database - 1) and RAVDESS (Ryerson Audio–Visual Database of Emotional Speech and Song), which both collect two distinguished sets of videos based on the different intensity of the expression, that is acted/strong or spontaneous/normal, providing the representations of the following emotional states that will be taken into consideration: angry, disgust, fearful, happy and sad. The performances of the proposed models are shown through accuracy results and some confusion matrices, demonstrating better accuracy than the compared proposals in the literature. The best accuracy achieved on BAUM-1 dataset is about 95%, while on RAVDESS it is about 95.5%. Umberto Bilotti, Carmen Bisogni, Maria De Marsico, S. Tramonte |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Walk as you feel: Privacy preserving emotion recognition from gait patternsabstractEmotion recognition from gait has gained significant interest due to its applicability in different fields such as healthcare, social cues, surveillance, and smart applications. Gait, as a biometric trait, offers unique advantages, allowing remote identification and robust recognition even in uncontrolled scenarios. Moreover, gait analysis can provide valuable insights into an individual’s emotional state. This work presents the “Walk-as-you-Feel” (WayF) framework, a novel approach for gait-based emotion recognition that does not rely on facial cues, ensuring user privacy. To address challenges with small and unbalanced datasets, a balancing procedure suitable for deep learning architecture is also developed. Adapted Inception-v3 and EfficientNet are employed for the feature extraction phase. Classification is performed using a Gated Recurrent Units network (GRUs) and Transformers-Encoder. Experimental results demonstrate the competitiveness of the proposed approach with respect to state-of-the-art works which also integrate facial cues. WayF reaches an average recognition rate of approximately 77% in its best configuration. Moreover, when excluding the neutral emotion, the proposed method achieves an outstanding overall accuracy of 83.3%. Carmen Bisogni, Lucia Cimmino, Michele Nappi, Toni Pannese, Chiara Pero |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | POSER: POsed vs Spontaneous Emotion Recognition using fractal encodingabstractEmotion recognition from facial expressions is a fundamental human ability that can be harnessed and transferred to machines. The ability to differentiate between spontaneous and posed emotions holds significant importance in various domains, including behavioral biometrics, forensics, and security. This paper introduces a novel method, called POsed vs Spontaneous Emotion Recognition (POSER), which leverages a modified version of the Partitioned Iterated Functions System (PIFS) to obtain a Fractal Encoding. This encoding is used for the first time as facial features to train a machine learning approach for the classification of emotions as either spontaneous or posed. Furthermore, by adapting the original architecture, we demonstrate the effectiveness of these features in distinguishing seven different emotions in controlled as well as wild environments, within a framework referred to as POSER-EMO. Experimental results are presented on the SPOS and DISFA + datasets for the first classification problem, where POSER outperforms the state of the art, and on the CK + and SFEW datasets for the second classification problem. Carmen Bisogni, Lucia Cascone, Michele Nappi, Chiara Pero |
Image Vis. Comput. | 1 |
| 2024 | Gaze analysis: A survey on its applicationsabstractThe examination of ocular movements has a wide range of applications due to the current developments in sensors that are now able to collect this biometric. This type of investigation is known as “gaze analysis”. The gaze has successfully examined a subject's physical and mental status in the past. As a result, over the last few decades, a large and diverse amount of literature on this subject has been generated and presented. The aim of this study is to collect and debate current gaze analysis methods based on their application field. Due to the context-specific needs for performance and efficiency, the eye movements under research are frequently evaluated from completely distinct perspectives. As a result, a collection of data, methods, and discussions ranging from the medical community to virtual and augmented reality, as well as human computer interface and remote learning, has been produced. In addition to providing a peek of novel observation on the issue of gaze analysis, the gaps between and within areas are also discussed to provide points for researchers to pursue. Carmen Bisogni, Michele Nappi, Genny Tortora, Alberto Del Bimbo |
Image Vis. Comput. | 1 |
| 2024 | Head Pose Estimation Patterns as Deepfake DetectorsabstractThe capacity to create “fake” videos has recently raised concerns about the reliability of multimedia content. Identifying between true and false information is a critical step toward resolving this problem. On this issue, several algorithms utilizing deep learning and facial landmarks have yielded intriguing results. Facial landmarks are traits that are solely tied to the subject’s head posture. Based on this observation, we study how Head Pose Estimation (HPE) patterns may be utilized to detect deepfakes in this work. The HPE patterns studied are based on FSA-Net, SynergyNet, and WSM, which are among the most performant approaches on the state-of-the-art. Finally, using a machine learning technique based on K-Nearest Neighbor and Dynamic Time Warping, their temporal patterns are categorized as authentic or false. We also offer a set of experiments for examining the feasibility of using deep learning techniques on such patterns. The findings reveal that the ability to recognize a deepfake video utilizing an HPE pattern is dependent on the HPE methodology. On the contrary, performance is less dependent on the performance of the utilized HPE technique. Experiments are carried out on the FaceForensics++ dataset that presents both identity swap and expression swap examples. The findings show that FSA-Net is an effective feature extraction method for determining whether a pattern belongs to a deepfake or not. The approach is also robust in comparison to deepfake videos created using various methods or for different goals. In the mean the method obtain 86% of accuracy on the identity swap task and 86.5% of accuracy on the expression swap. These findings offer up various possibilities and future directions for solving the deepfake detection problem using specialized HPE approaches, which are also known to be fast and reliable. Federico Becattini, Carmen Bisogni, Vincenzo Loia, Chiara Pero, Fei Hao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | IoT-enabled Biometric Security: Enhancing Smart Car Safety with Depth-based Head Pose EstimationabstractAdvanced Driver Assistance Systems (ADAS) are experiencing higher levels of automation, facilitated by the synergy among various sensors integrated within vehicles, thereby forming an Internet of Things (IoT) framework. Among these sensors, cameras have emerged as valuable tools for detecting driver fatigue and distraction. This study introduces HYDE-F, a Head Pose Estimation (HPE) system exclusively utilizing depth cameras. HYDE-F adeptly identifies critical driver head poses associated with risky conditions, thus enhancing the safety of IoT-enabled ADAS. The core of HYDE-F’s innovation lies in its dual-process approach: it employs a fractal encoding technique and keypoint intensity analysis in parallel. These two processes are then fused using an optimization algorithm, enabling HYDE-F to blend the strengths of both methods for enhanced accuracy. Evaluations conducted on a specialized driving dataset, Pandora, demonstrate HYDE-F’s competitive performance compared to existing methods, surpassing current techniques in terms of average Mean Absolute Error (MAE) by nearly 1 ∘ . Moreover, case studies highlight the successful integration of HYDE-F with vehicle sensors. Additionally, HYDE-F exhibits robust generalization capabilities, as evidenced by experiments conducted on standard laboratory-based HPE datasets, i.e., Biwi and ICT-3DHP databases, achieving an average MAE of 4.9 ∘ and 5 ∘ , respectively. Carmen Bisogni, Lucia Cascone, Michele Nappi, Chiara Pero |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Monocular Vision-aided Depth Measurement from RGB Images for Autonomous UAV NavigationabstractMonocular vision-based 3D scene understanding has been an integral part of many machine vision applications. Always, the objective is to measure the depth using a single RGB camera, which is at par with the depth cameras. In this regard, monocular vision-guided autonomous navigation of robots is rapidly gaining popularity among the research community. We propose an effective monocular vision-assisted method to measure the depth of an Unmanned Aerial Vehicle (UAV) from an impending frontal obstacle. This is followed by collision-free navigation in unknown GPS-denied environments. Our approach deals upon the fundamental principle of perspective vision that the size of an object relative to its field of view (FoV) increases as the center of projection moves closer towards the object. Our contribution involves modeling the depth followed by its realization through scale-invariant SURF features. Noisy depth measurements arising due to external wind, or the turbulence in the UAV, are rectified by employing a constant velocity-based Kalman filter model. Necessary control commands are then designed based on the rectified depth value to avoid the obstacle before collision. Rigorous experiments with SURF scale-invariant features reveal an overall accuracy of 88.6% with varying obstacles, in both indoor and outdoor environments. Ram Prasad Padhy, Pankaj Kumar Sa, Fabio Narducci, Carmen Bisogni, Sambit Bakshi |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Dual-LightGCN: Dual light graph convolutional network for discriminative recommendation
Wenqing Huang, Fei Hao 0001, Jiaxing Shang, Wangyang Yu 0001, Shengke Zeng, Carmen Bisogni, Vincenzo Loia |
Comput. Commun. | 6 |
| 2023 | Emotion recognition at a distance: The robustness of machine learning based on hand-crafted facial features vs deep learning modelsabstractEmotion estimation from face expression analysis is nowadays a widely-explored computer vision task. In turn, the classification of expressions relies on relevant facial features and their dynamics. Despite the promising accuracy results achieved in controlled and favorable conditions, the processing of faces acquired at a distance, entailing low-quality images, still suffers from a significant performance decrease. In particular, most approaches and related computational models become extremely unstable in the case of the very small amount of useful pixels that is typical in these conditions. Therefore, their behavior should be investigated more carefully. On the other hand, real-time emotion recognition at a distance may play a critical role in smart video surveillance, especially when controlling particular kinds of events, e.g., political meetings, to try to prevent adverse actions. This work compares facial expression recognition at a distance by: 1) a deep learning architecture based on state-of-the-art (SOTA) proposals, which exploits the whole images to autonomously learn the relevant embeddings; 2) a machine learning approach that relies on hand-crafted features, namely the facial landmarks preliminarily extracted using the popular Mediapipe framework. Instead of using either the complete sequence of frames or only the final still image of the expression, like current SOTA approaches, the two proposed methods are designed to use rich temporal information to identify three different stages of emotion. Expressions are time-split accordingly into four phases to better exploit their temporal-dependent dynamics. Experiments were conducted on the popular Extended Cohn-Kanade dataset (CK+). It was chosen for its wide use in related literature, and because it includes videos of facial expressions and not only still images. The results show that the approach relying on machine learning via hand-crafted features is more suitable for classifying the initial phases of the expression and does not decay in terms of accuracy when images are at a distance (only 0.08% of decay). On the contrary, deep learning not only has difficulties classifying the initial phases of the expressions but also suffers from relevant performance decay when considering images at a distance (52.68% accuracy decay). Carmen Bisogni, Lucia Cimmino, Maria De Marsico, Fei Hao 0001, Fabio Narducci |
Image Vis. Comput. | 1 |
| 2023 | Exploring invariance of concept stability for attribute reduction in three-way concept lattice
Fei Hao 0001, Carmen Bisogni, Vincenzo Loia, Zheng Pei 0001, Aziz Nasridinov |
Soft Comput. | 3 |
| 2022 | Imaging based cervical cancer diagnostics using small object detection - generative adversarial networks
R. Elakkiya, Kuppa Sai Sri Teja, L. Jegatha Deborah, Carmen Bisogni, Carlo Maria Medaglia |
Multim. Tools Appl. | 4 |
| 2022 | Head pose estimation: An extensive survey on recent techniques and applications
Andrea F. Abate, Carmen Bisogni, Aniello Castiglione, Michele Nappi |
Pattern Recognit. | 2 |
| 2022 | Impact of Deep Learning Approaches on Facial Expression Recognition in Healthcare IndustriesabstractA facial expression recognition system that can provide quick assistance to the healthcare system and exceptional services to the patients is proposed in this article. The implementation of this work is divided into three components. In the first component, landmark points on the facial region are detected; a fixed-sized rectangular box is obtained by normalizing the detected face region, and then, down sampled to its varying sizes producing multiresolution images. Different convolution neural network architectures are proposed in the second component for analyzing the textual information within the multiresolution facial images. To extract more discriminating features and enhance the proposed system’s performance, some amalgamation of transfer learning, progressive image resizing, data augmentation, and fine tuning of parameters are employed in the third component. For experimental purposes, three benchmark databases, static facial expressions in the wild, Cohn-Kanade, and Karolinska directed emotional faces, are employed with some existing methods concerning these databases. The comparison with these databases shows the superiority of the proposed system. Carmen Bisogni, Aniello Castiglione, Sanoar Hossain, Fabio Narducci, Saiyed Umer |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | Drowsiness Detection in the Era of Industry 4.0: Are We Ready?abstractInterconnectivity and smart automation of Internet of Things in recent times have led to the concept of Industry 4.0. Together with the improvement in productivity and new business models, employment conditions should take advantage of these new technologies. Safety in the workplace is one of the most sensitive topics on matters that needs targeted and accurate solutions. The safety can be guaranteed by investigating the attention states of the workers, and in particular, their drowsiness levels. Several technologies have faced this problem by using biometrics, but how many of them are applicable in a real-case-use scenario of Industry 4.0? This article aims to answer this question by discussing available data and methods that can be used in specific workplaces. We highlight their limitations and accuracy to sketch out the recent literature that may contribute to worker safety in Industry 4.0. Finally, we point out a gap that needs to be filled in order to implement these strategies on a large scale. Carmen Bisogni, Fei Hao 0001, Vincenzo Loia, Fabio Narducci |
IEEE Trans. Ind. Informatics | 1 |
| 2021 | ECB2: A novel encryption scheme using face biometrics for signing blockchain transactions
Carmen Bisogni, Gerardo Iovane, Riccardo Emanuele Landi, Michele Nappi |
J. Inf. Secur. Appl. | 1 |
| 2021 | Visual question answering: Which investigated applications?
Silvio Barra, Carmen Bisogni, Maria De Marsico, Stefano Ricciardi |
Pattern Recognit. Lett. | 2 |
| 2021 | Deep learning for emotion driven user experiences
Carmen Bisogni, Lucia Cascone, Aniello Castiglione, Ignazio Passero |
Pattern Recognit. Lett. | 1 |
| 2021 | Adversarial attacks through architectures and spectra in face recognition
Carmen Bisogni, Lucia Cascone, Jean-Luc Dugelay, Chiara Pero |
Pattern Recognit. Lett. | 1 |
| 2021 | Stability of three-way concepts and its application to natural language generation
Fei Hao 0001, Carmen Bisogni, Geyong Min, Vincenzo Loia, Carmen De Maio |
Pattern Recognit. Lett. | 3 |
| 2021 | FASHE: A FrActal Based Strategy for Head Pose EstimationabstractHead pose estimation (HPE) represents a topic central to many relevant research fields and characterized by a wide application range. In particular, HPE performed using a singular RGB frame is particular suitable to be applied at best-frame-selection problems. This explains a growing interest witnessed by a large number of contributions, most of which exploit deep learning architectures and require extensive training sessions to achieve accuracy and robustness in estimating head rotations on three axes. However, methods alternative to machine learning approaches could be capable of similar if not better performance. To this regard, we present FASHE, an approach based on partitioned iterated function systems (PIFS) to represent auto-similarities within face image through a contractive affine function transforming the domain blocks extracted only once by a single frontal reference image, in a good approximation of the range blocks which the target image has been partitioned into. Pose estimation is achieved by finding the closest match between fractal code of target image and a reference array by means of Hamming distance. The results of experiments conducted exceed the state of the art on both Biwi and Ponting'04 datasets as well as approaching those of the best performing methods on the challenging AFLW2000 database. In addition, the applications to GOTCHA Video Dataset demonstrate that FASHE successfully operates in-the-wild. Carmen Bisogni, Michele Nappi, Chiara Pero, Stefano Ricciardi |
IEEE Trans. Image Process. | 1 |
| 2020 | HP2IFS: Head Pose estimation exploiting Partitioned Iterated Function SystemsabstractEstimating the actual head orientation from 2D images, with regard to its three degrees of freedom, is a well known problem that is highly significant for a large number of applications involving head pose knowledge. Consequently, this topic has been tackled by a plethora of methods and algorithms the most part of which exploits neural networks. Machine learning methods, indeed, achieve accurate head rotation values yet require an adequate training stage and, to that aim, a relevant number of positive and negative examples. In this paper we take a different approach to this topic by using fractal coding theory and particularly Partitioned Iterated Function Systems to extract the fractal code from the input head image and to compare this representation to the fractal code of a reference model through Hamming distance. According to experiments conducted on both the BIWI and the AFLW2000 databases, the proposed PIFS based head pose estimation method provides accurate yaw/pitch/roll angular values, with a performance approaching that of state of the art of machine-learning based algorithms and exceeding most of non-training based approaches. Carmen Bisogni, Michele Nappi, Chiara Pero, Stefano Ricciardi |
ICPR | 1 |
| 2020 | An attention recurrent model for human cooperation detection
David Freire-Obregón, Modesto Castrillón-Santana, Paola Barra, Carmen Bisogni, Michele Nappi |
Comput. Vis. Image Underst. | 4 |
| 2020 | Web-Shaped Model for Head Pose Estimation: An Approach for Best Exemplar SelectionabstractHead pose estimation is a sensitive topic in video surveillance/smart ambient scenarios since head rotations can hide/distort discriminative features of the face. Face recognition would often tackle the problem of video frames where subjects appear in poses making it quite impossible. In this respect, the selection of the frames with the best face orientation can allow triggering recognition only on these, therefore decreasing the possibility of errors. This paper proposes a novel approach to head pose estimation for smart cities and video surveillance scenarios, aiming at this goal. The method relies on a cascade of two models: the first one predicts the positions of 68 well-known face landmarks; the second one applies a web-shaped model over the detected landmarks, to associate each of them to a specific face sector. The method can work on detected faces at a reasonable distance and with a resolution that is supported by several present devices. Results of experiments executed over some classical pose estimation benchmarks, namely Point '04, Biwi, and AFLW datasets show good performance in terms of both pose estimation and computing time. Further results refer to noisy images that are typical of the addressed settings. Finally, examples demonstrate the selection of the best frames from videos captured in video surveillance conditions. Paola Barra, Silvio Barra, Carmen Bisogni, Maria De Marsico, Michele Nappi |
IEEE Trans. Image Process. | 3 |
| 2020 | An Encryption Approach Using Information Fusion Techniques Involving Prime Numbers and Face BiometricsabstractThe work shows a novel solution to create an access key which can be used within the transactions of electronic currencies, blockchain as well as in the field of computer security to guarantee a high level of secrecy, but also, with a high level of certainty, to provide a person identity through Information Fusion (IF) techniques and biometric data encryption. Specifically, two non-connected areas have been joined, Face Biometrics and Public-key Cryptography. This choice was taken in order to get through the limits these two approaches have found singularly and to give a suitable solution in the context of electronic and digital exchanges (electro-currencies, Internet of Things). An innovative and original algorithm has been developed, which can do fusion operations between Face Biometrics and numerical data, that is an algorithm of Hybrid Information Fusion, named FIF (Face Information Fusion). We decided to use a digital face as a biometric component, and the product of two prime numbers as a numerical component, that is the module in RSA algorithm. Gerardo Iovane, Carmen Bisogni, Luigi De Maio, Michele Nappi |
IEEE Trans. Sustain. Comput. | 2 |
| 2019 | F-FID: fast fuzzy-based iris de-noising for mobile security applications
Silvio Barra, Carmen Bisogni, Michele Nappi, Stefano Ricciardi |
Multim. Tools Appl. | 2 |
| 2018 | Fast QuadTree-Based Pose Estimation for Security Applications Using Face Biometrics
Paola Barra, Carmen Bisogni, Michele Nappi, Stefano Ricciardi |
NSS | 2 |