Alessio Fagioli 0001

dblp:143/4807-1 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0002-8111-9120ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Thyroid Nodule Classification via Weak Self-Supervision and Transfer Learning
Alessio Fagioli 0001, Marco Cascio, Gian Luca Foresti, Luigi Cinque
ICPRAM1
2026 SAGE-networks: Shape-aware geometric embeddings for writer re-identification in historical manuscripts
Alessio Fagioli 0001, Nicola Follador, Marco Cascio, Emanuela Colombi, Gian Luca Foresti
Pattern Recognit.1
2024 Signal enhancement and efficient DTW-based comparison for wearable gait recognition
abstract
The popularity of biometrics-based user identification has significantly increased over the last few years. User identification based on the face, fingerprints, and iris, usually achieves very high accuracy only in controlled setups and can be vulnerable to presentation attacks, spoofing, and forgeries. To overcome these issues, this work proposes a novel strategy based on a relatively less explored biometric trait, i.e., gait, collected by a smartphone accelerometer, which can be more robust to the attacks mentioned above. According to the wearable sensor-based gait recognition state-of-the-art, two main classes of approaches exist: 1) those based on machine and deep learning; 2) those exploiting hand-crafted features. While the former approaches can reach a higher accuracy, they suffer from problems like, e.g., performing poorly outside the training data, i.e., lack of generalizability. This paper proposes an algorithm based on hand-crafted features for gait recognition that can outperform the existing machine and deep learning approaches. It leverages a modified Majority Voting scheme applied to Fast Window Dynamic Time Warping, a modified version of the Dynamic Time Warping (DTW) algorithm with relaxed constraints and majority voting, to recognize gait patterns. We tested our approach named MV-FWDTW on the ZJU-gaitacc, one of the most extensive datasets for the number of subjects, but especially for the number of walks per subject and walk lengths. Results set a new state-of-the-art gait recognition rate of 98.82% in a cross-session experimental setup. We also confirm the quality of the proposed method using a subset of the OU-ISIR dataset, another large state-of-the-art benchmark with more subjects but much shorter walk signals.
Danilo Avola, Luigi Cinque, Maria De Marsico, Alessio Fagioli 0001, Gian Luca Foresti, Maurizio Mancini, Alessio Mecca
Comput. Secur.4
2024 Spatio-Temporal Image-Based Encoded Atlases for EEG Emotion Recognition
abstract
Emotion recognition plays an essential role in human-human interaction since it is a key to understanding the emotional states and reactions of human beings when they are subject to events and engagements in everyday life. Moving towards human-computer interaction, the study of emotions becomes fundamental because it is at the basis of the design of advanced systems to support a broad spectrum of application areas, including forensic, rehabilitative, educational, and many others. An effective method for discriminating emotions is based on ElectroEncephaloGraphy (EEG) data analysis, which is used as input for classification systems. Collecting brain signals on several channels and for a wide range of emotions produces cumbersome datasets that are hard to manage, transmit, and use in varied applications. In this context, the paper introduces the Empátheia system, which explores a different EEG representation by encoding EEG signals into images prior to their classification. In particular, the proposed system extracts spatio-temporal image encodings, or atlases, from EEG data through the Processing and transfeR of Interaction States and Mappings through Image-based eNcoding (PRISMIN) framework, thus obtaining a compact representation of the input signals. The atlases are then classified through the Empátheia architecture, which comprises branches based on convolutional, recurrent, and transformer models designed and tuned to capture the spatial and temporal aspects of emotions. Extensive experiments were conducted on the Shanghai Jiao Tong University (SJTU) Emotion EEG Dataset (SEED) public dataset, where the proposed system significantly reduced its size while retaining high performance. The results obtained highlight the effectiveness of the proposed approach and suggest new avenues for data representation in emotion recognition from EEG signals.
Danilo Avola, Luigi Cinque, Angelo Di Mambro, Alessio Fagioli 0001, Marco Raoul Marini, Daniele Pannone, Bruno Fanini, Gian Luca Foresti
Int. J. Neural Syst.4
2022 Human Silhouette and Skeleton Video Synthesis Through Wi-Fi Signals
abstract
The increasing availability of wireless access points (APs) is leading toward human sensing applications based on Wi-Fi signals as support or alternative tools to the widespread visual sensors, where the signals enable to address well-known vision-related problems such as illumination changes or occlusions. Indeed, using image synthesis techniques to translate radio frequencies to the visible spectrum can become essential to obtain otherwise unavailable visual data. This domain-to-domain translation is feasible because both objects and people affect electromagnetic waves, causing radio and optical frequencies variations. In the literature, models capable of inferring radio-to-visual features mappings have gained momentum in the last few years since frequency changes can be observed in the radio domain through the channel state information (CSI) of Wi-Fi APs, enabling signal-based feature extraction, e.g. amplitude. On this account, this paper presents a novel two-branch generative neural network that effectively maps radio data into visual features, following a teacher-student design that exploits a cross-modality supervision strategy. The latter conditions signal-based features in the visual domain to completely replace visual data. Once trained, the proposed method synthesizes human silhouette and skeleton videos using exclusively Wi-Fi signals. The approach is evaluated on publicly available data, where it obtains remarkable results for both silhouette and skeleton videos generation, demonstrating the effectiveness of the proposed cross-modality supervision strategy.
Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti
Int. J. Neural Syst.4
2022 Affective Action and Interaction Recognition by Multi-View Representation Learning from Handcrafted Low-Level Skeleton Features
abstract
Human feelings expressed through verbal (e.g. voice) and non-verbal communication channels (e.g. face or body) can influence either human actions or interactions. In the literature, most of the attention was given to facial expressions for the analysis of emotions conveyed through non-verbal behaviors. Despite this, psychology highlights that the body is an important indicator of the human affective state in performing daily life activities. Therefore, this paper presents a novel method for affective action and interaction recognition from videos, exploiting multi-view representation learning and only full-body handcrafted characteristics selected following psychological and proxemic studies. Specifically, 2D skeletal data are extracted from RGB video sequences to derive diverse low-level skeleton features, i.e. multi-views, modeled through the bag-of-visual-words clustering approach generating a condition-related codebook. In this way, each affective action and interaction within a video can be represented as a frequency histogram of codewords. During the learning phase, for each affective class, training samples are used to compute its global histogram of codewords stored in a database and later used for the recognition task. In the recognition phase, the video frequency histogram representation is matched against the database of class histograms and classified as the closest affective class in terms of Euclidean distance. The effectiveness of the proposed system is evaluated on a specifically collected dataset containing 6 emotion for both actions and interactions, on which the proposed system obtains 93.64% and 90.83% accuracy, respectively. In addition, the devised strategy also achieves in line performances with other literature works based on deep learning when tested on a public collection containing 6 emotions plus a neutral state, demonstrating the effectiveness of the presented approach and confirming the findings in psychological and proxemic studies.
Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti
Int. J. Neural Syst.4
2022 SIRe-Networks: Convolutional neural networks architectural extension for information preservation via skip/residual connections and interlaced auto-encoders
Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti
Neural Networks3
2022 3D hand pose and shape estimation from RGB images for keypoint-based hand gesture recognition
Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti, Adriano Fragomeni, Daniele Pannone
Pattern Recognit.3
2022 Deep Temporal Analysis for Non-Acted Body Affect Recognition
abstract
In the field of body affect recognition, the majority of literature is based on experiments performed on datasets where trained actors simulate emotional reactions. These acted and unnatural expressions differ from the more challenging genuine emotions, thus leading to less valuable results. In this article, a solution for basic non-acted emotion recognition based on 3D skeleton and Deep Neural Networks (DNNs) is provided. The proposed work introduces three majors contributions. First, temporal local movements performed by subjects are examined frame-by-frame, unlike the current state-of-the-art in non-acted body affect recognition where only static or global body features are considered. Second, an original set of global and time-dependent features for body movement description is provided. Third, this is one of the first works to use deep learning methods in the current non-acted body affect recognition literature. Due to the novelty of the topic, only the UCLIC dataset is currently considered the benchmark for comparative tests. On the latter, the proposed method outperforms all the competitors.
Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti, Cristiano Massaroni
IEEE Trans. Affect. Comput.3
2022 Multimodal Feature Fusion and Knowledge-Driven Learning via Experts Consult for Thyroid Nodule Classification
abstract
Computer-aided diagnosis (CAD) is becoming a prominent approach to assist clinicians spanning across multiple fields. These automated systems take advantage of various computer vision (CV) procedures, as well as artificial intelligence (AI) techniques, to formulate a diagnosis of a given image, e.g., computed tomography and ultrasound. Advances in both areas (CV and AI) are enabling ever increasing performances of CAD systems, which can ultimately avoid performing invasive procedures such as fine-needle aspiration. In this study, a novel end-to-end knowledge-driven classification framework is presented. The system focuses on multimodal data generated by thyroid ultrasonography, and acts as a CAD system by providing a thyroid nodule classification into the benign and malignant categories. Specifically, the proposed system leverages cues provided by an ensemble of experts to guide the learning phase of a densely connected convolutional network (DenseNet). The ensemble is composed by various networks pretrained on ImageNet, including AlexNet, ResNet, VGG, and others. The previously computed multimodal feature parameters are used to create ultrasonography domain experts via transfer learning, decreasing, moreover, the number of samples required for training. To validate the proposed method, extensive experiments were performed, providing detailed performances for both the experts ensemble and the knowledge-driven DenseNet. As demonstrated by the results, the proposed system achieves relevant performances in terms of qualitative metrics for the thyroid nodule classification task, thus resulting in a great asset when formulating a diagnosis.
Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Sebastiano Filetti, Giorgio Grani, Emanuele Rodolà
IEEE Trans. Circuits Syst. Video Technol.3
2022 Person Re-Identification Through Wi-Fi Extracted Radio Biometric Signatures
abstract
Person re-identification (Re-ID) is a challenging task that tries to recognize a person across different cameras, and that can prove useful in video surveillance as well as in forensics and security applications. However, traditional Re-ID systems analyzing image or video sequences suffer from well-known issues such as illumination changes, occlusions, background clutter, and long-term re-identification. To simultaneously address all these difficult problems, we explore a Re-ID solution based on an alternative medium that is inherently not affected by them, i.e., the Wi-Fi technology. The latter, due to the widespread use of wireless communications, has grown rapidly and is already enabling the development of Wi-Fi sensing applications, such as human localization or counting. These sensing procedures generally exploit Wi-Fi signals variations that are a direct consequence, among other things, of human presence, and which can be observed through the channel state information (CSI) of Wi-Fi access points. Following this rationale, in this paper, for the first time in literature, we show how the pervasive Wi-Fi technology can also be directly exploited for person Re-ID. More accurately, Wi-Fi signals amplitude and phase are extracted from CSI measurements and analyzed through a two-branch deep neural network working in a siamese-like fashion. The designed pipeline can extract meaningful features from signals, i.e., radio biometric signatures, that ultimately allow the person Re-ID. The effectiveness of the proposed system is evaluated on a specifically collected dataset, where remarkable performances are obtained; suggesting that Wi-Fi signal variations differ between different people and can consequently be used for their re-identification.
Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Chiara Petrioli
IEEE Trans. Inf. Forensics Secur.4
2021 LieToMe: An Ensemble Approach for Deception Detection from Facial Cues
abstract
Deception detection is a relevant ability in high stakes situations such as police interrogatories or court trials, where the outcome is highly influenced by the interviewed person behavior. With the use of specific devices, e.g. polygraph or magnetic resonance, the subject is aware of being monitored and can change his behavior, thus compromising the interrogation result. For this reason, video analysis-based methods for automatic deception detection are receiving ever increasing interest. In this paper, a deception detection approach based on RGB videos, leveraging both facial features and stacked generalization ensemble, is proposed. First, a face, which is well-known to present several meaningful cues for deception detection, is identified, aligned, and masked to build video signatures. These signatures are constructed starting from five different descriptors, which allow the system to capture both static and dynamic facial characteristics. Then, video signatures are given as input to four base-level algorithms, which are subsequently fused applying the stacked generalization technique, resulting in a more robust meta-level classifier used to predict deception. By exploiting relevant cues via specific features, the proposed system achieves improved performances on a public dataset of famous court trials, with respect to other state-of-the-art methods based on facial features, highlighting the effectiveness of the proposed method.
Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti
Int. J. Neural Syst.4
2021 Automatic estimation of optimal UAV flight parameters for real-time wide areas monitoring
Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti, Daniele Pannone, Claudio Piciarelli
Multim. Tools Appl.3
2021 R-SigNet: Reduced space writer-independent feature learning for offline writer-dependent signature verification
Danilo Avola, Manoochehr Joodi Bigdello, Luigi Cinque, Alessio Fagioli 0001, Marco Raoul Marini
Pattern Recognit. Lett.4
2020 LieToMe: Preliminary study on hand gestures for deception detection via Fisher-LSTM
Danilo Avola, Luigi Cinque, Maria De Marsico, Alessio Fagioli 0001, Gian Luca Foresti
Pattern Recognit. Lett.4
2019 Master and Rookie Networks for Person Re-identification
Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti, Cristiano Massaroni
CAIP (2)4