Marco Cascio

dblp:247/5312 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0003-4370-8140ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
1 paper
Biometric security · 100%
Artificial intelligence
1 paper
Video understanding and tracking · 100%
Computer networks
1 paper
Wireless sensing and localization · 50% Physical-layer communications · 50%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Biometric security
person re-identification
0.612022
Person Re-Identification Through Wi-Fi Extracted Radio Biometric Signatures · IEEE Trans. Inf. Forensics Secur. 2022
Computer vision › Video understanding and tracking › action recognition
skeleton-based action recognition
0.412020
2-D Skeleton-Based Action Recognition via Two-Branch Stacked LSTM-RNNs · IEEE Trans. Multim. 2020
Physical-layer communications
channel state information
0.212022
Person Re-Identification Through Wi-Fi Extracted Radio Biometric Signatures · IEEE Trans. Inf. Forensics Secur. 2022
Wireless sensing and localization
wifi sensing
0.212022
Person Re-Identification Through Wi-Fi Extracted Radio Biometric Signatures · IEEE Trans. Inf. Forensics Secur. 2022
Computer vision › Video understanding and tracking
action recognition
0.112020
2-D Skeleton-Based Action Recognition via Two-Branch Stacked LSTM-RNNs · IEEE Trans. Multim. 2020

Methods — techniques the papers use, named apart from their topics

siamese neural network · 1.1deep learning · 1.1recurrent neural network · 0.4long short-term memory · 0.43d convolutional neural network · 0.4
YearPublicationVenuePosition
2026 Thyroid Nodule Classification via Weak Self-Supervision and Transfer Learning
Alessio Fagioli 0001, Marco Cascio, Gian Luca Foresti, Luigi Cinque
ICPRAM2
2026 SAGE-networks: Shape-aware geometric embeddings for writer re-identification in historical manuscripts
Alessio Fagioli 0001, Nicola Follador, Marco Cascio, Emanuela Colombi, Gian Luca Foresti
Pattern Recognit.3
2024 ReViT: Enhancing vision transformers feature diversity with attention residual connections
Anxhelo Diko, Danilo Avola, Marco Cascio, Luigi Cinque
Pattern Recognit.3
2022 Human Silhouette and Skeleton Video Synthesis Through Wi-Fi Signals
abstract
The increasing availability of wireless access points (APs) is leading toward human sensing applications based on Wi-Fi signals as support or alternative tools to the widespread visual sensors, where the signals enable to address well-known vision-related problems such as illumination changes or occlusions. Indeed, using image synthesis techniques to translate radio frequencies to the visible spectrum can become essential to obtain otherwise unavailable visual data. This domain-to-domain translation is feasible because both objects and people affect electromagnetic waves, causing radio and optical frequencies variations. In the literature, models capable of inferring radio-to-visual features mappings have gained momentum in the last few years since frequency changes can be observed in the radio domain through the channel state information (CSI) of Wi-Fi APs, enabling signal-based feature extraction, e.g. amplitude. On this account, this paper presents a novel two-branch generative neural network that effectively maps radio data into visual features, following a teacher-student design that exploits a cross-modality supervision strategy. The latter conditions signal-based features in the visual domain to completely replace visual data. Once trained, the proposed method synthesizes human silhouette and skeleton videos using exclusively Wi-Fi signals. The approach is evaluated on publicly available data, where it obtains remarkable results for both silhouette and skeleton videos generation, demonstrating the effectiveness of the proposed cross-modality supervision strategy.
Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti
Int. J. Neural Syst.2
2022 Affective Action and Interaction Recognition by Multi-View Representation Learning from Handcrafted Low-Level Skeleton Features
abstract
Human feelings expressed through verbal (e.g. voice) and non-verbal communication channels (e.g. face or body) can influence either human actions or interactions. In the literature, most of the attention was given to facial expressions for the analysis of emotions conveyed through non-verbal behaviors. Despite this, psychology highlights that the body is an important indicator of the human affective state in performing daily life activities. Therefore, this paper presents a novel method for affective action and interaction recognition from videos, exploiting multi-view representation learning and only full-body handcrafted characteristics selected following psychological and proxemic studies. Specifically, 2D skeletal data are extracted from RGB video sequences to derive diverse low-level skeleton features, i.e. multi-views, modeled through the bag-of-visual-words clustering approach generating a condition-related codebook. In this way, each affective action and interaction within a video can be represented as a frequency histogram of codewords. During the learning phase, for each affective class, training samples are used to compute its global histogram of codewords stored in a database and later used for the recognition task. In the recognition phase, the video frequency histogram representation is matched against the database of class histograms and classified as the closest affective class in terms of Euclidean distance. The effectiveness of the proposed system is evaluated on a specifically collected dataset containing 6 emotion for both actions and interactions, on which the proposed system obtains 93.64% and 90.83% accuracy, respectively. In addition, the devised strategy also achieves in line performances with other literature works based on deep learning when tested on a public collection containing 6 emotions plus a neutral state, demonstrating the effectiveness of the presented approach and confirming the findings in psychological and proxemic studies.
Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti
Int. J. Neural Syst.2
2022 Person Re-Identification Through Wi-Fi Extracted Radio Biometric Signatures
abstract
Person re-identification (Re-ID) is a challenging task that tries to recognize a person across different cameras, and that can prove useful in video surveillance as well as in forensics and security applications. However, traditional Re-ID systems analyzing image or video sequences suffer from well-known issues such as illumination changes, occlusions, background clutter, and long-term re-identification. To simultaneously address all these difficult problems, we explore a Re-ID solution based on an alternative medium that is inherently not affected by them, i.e., the Wi-Fi technology. The latter, due to the widespread use of wireless communications, has grown rapidly and is already enabling the development of Wi-Fi sensing applications, such as human localization or counting. These sensing procedures generally exploit Wi-Fi signals variations that are a direct consequence, among other things, of human presence, and which can be observed through the channel state information (CSI) of Wi-Fi access points. Following this rationale, in this paper, for the first time in literature, we show how the pervasive Wi-Fi technology can also be directly exploited for person Re-ID. More accurately, Wi-Fi signals amplitude and phase are extracted from CSI measurements and analyzed through a two-branch deep neural network working in a siamese-like fashion. The designed pipeline can extract meaningful features from signals, i.e., radio biometric signatures, that ultimately allow the person Re-ID. The effectiveness of the proposed system is evaluated on a specifically collected dataset, where remarkable performances are obtained; suggesting that Wi-Fi signal variations differ between different people and can consequently be used for their re-identification.
Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Chiara Petrioli
IEEE Trans. Inf. Forensics Secur.2
2021 LieToMe: An Ensemble Approach for Deception Detection from Facial Cues
abstract
Deception detection is a relevant ability in high stakes situations such as police interrogatories or court trials, where the outcome is highly influenced by the interviewed person behavior. With the use of specific devices, e.g. polygraph or magnetic resonance, the subject is aware of being monitored and can change his behavior, thus compromising the interrogation result. For this reason, video analysis-based methods for automatic deception detection are receiving ever increasing interest. In this paper, a deception detection approach based on RGB videos, leveraging both facial features and stacked generalization ensemble, is proposed. First, a face, which is well-known to present several meaningful cues for deception detection, is identified, aligned, and masked to build video signatures. These signatures are constructed starting from five different descriptors, which allow the system to capture both static and dynamic facial characteristics. Then, video signatures are given as input to four base-level algorithms, which are subsequently fused applying the stacked generalization technique, resulting in a more robust meta-level classifier used to predict deception. By exploiting relevant cues via specific features, the proposed system achieves improved performances on a public dataset of famous court trials, with respect to other state-of-the-art methods based on facial features, highlighting the effectiveness of the proposed method.
Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti
Int. J. Neural Syst.2
2020 2-D Skeleton-Based Action Recognition via Two-Branch Stacked LSTM-RNNs
abstract
Action recognition in video sequences is an interesting field for many computer vision applications, including behavior analysis, event recognition, and video surveillance. In this article, a method based on 2D skeleton and two-branch stacked Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM) cells is proposed. Unlike 3D skeletons, usually generated by RGB-D cameras, the 2D skeletons adopted in this article are reconstructed starting from RGB video streams, therefore allowing the use of the proposed approach in both indoor and outdoor environments. Moreover, any case of missing skeletal data is managed by exploiting 3D-Convolutional Neural Networks (3D-CNNs). Comparative experiments with several key works on KTH and Weizmann datasets show that the method described in this paper outperforms the current state-of-the-art. Additional experiments on UCF Sports and IXMAS datasets demonstrate the effectiveness of our method in the presence of noisy data and perspective changes, respectively. Further investigations on UCF Sports, HMDB51, UCF101, and Kinetics400 highlight how the combination between the proposed two-branch stacked LSTM and the 3D-CNN-based network can manage missing skeleton information, greatly improving the overall accuracy. Moreover, additional tests on KTH and UCF Sports datasets also show the robustness of our approach in the presence of partial body occlusions. Finally, comparisons on UT-Kinect and NTU-RGB+D datasets show that the accuracy of the proposed method is fully comparable to that of works based on 3D skeletons.
Danilo Avola, Marco Cascio, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni, Emanuele Rodolà
IEEE Trans. Multim.2
2019 Master and Rookie Networks for Person Re-identification
Danilo Avola, Marco Cascio, Luigi Cinque, Alessio Fagioli 0001, Gian Luca Foresti, Cristiano Massaroni
CAIP (2)2