Giuseppe Cascavilla

dblp:147/5165 · DBLP profile ↗
← Back
12ranked-venue papers in the field
5as first author
9since 2021 · last 2025
0000-0002-0724-3772ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 7 (2 first)Data Mining & Knowledge Discovery · 3 (2 first)Database Systems & Data Management · 1Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)
YearPublicationVenuePosition
2025 Mind the Shift: A Study on Transfer Learning and Domain Adaptation in Vehicular Intrusion Detection
abstract
This paper addresses the need for adaptable intrusion detection systems (IDS) in intra-vehicle networks. We evaluated two scalable IDS strategies-combined training and transfer learning-on novel traffic data to balance predictive accuracy and computational efficiency under distributional shifts. Previous IDS models for intra-vehicle attacks achieved high accuracy but relied heavily on the simple CarHacking dataset. To address this limitation, we integrate the newer CIC-IoV-2024 dataset, which reflects realistic vehicular traffic. In Strategy 1, we retrain models from scratch on a combined dataset. These models achieve strong classification across all classes, with accuracies ranging from 90.57% to 100% and F1 scores of 88.91% to 100%. However, training takes longer (61-184 minutes) and inference per packet is slower (30-80 ms). In Strategy 2, we apply transfer learning by fine-tuning pre-trained models while freezing earlier layers. This approach reduces training time (4-16 minutes) and improves latency (21.29-126.76 ms), but predictive performance declines. The models primarily distinguish between benign and malicious traffic, with F1 scores ranging from 82.99% to 89.32%, and exhibit high uncertainty in classifying diverse attack types. Our findings highlight a trade-off between predictive power and computational efficiency. These insights can guide the deployment of IDS frameworks in real-time vehicular environments.
Jakob Richard Proos, Giuseppe Cascavilla, Cristoffer Leite, Alfredo Cuzzocrea
IEEE Big Data2
2025 Uncovering the Darknet: A Hybrid Intelligence Framework for Drug Market Surveillance
Diogo Rio, Krishna Teja Atluri, Giuseppe Cascavilla, Jessica De Pascale
IEEE Big Data3
2024 Extremist discourse on alt-right sub-reddit: An Inferential Network Analysis Approach
abstract
The use of extreme language in social networks is a problem of increasing relevance since it might be a driver of radicalization. Reddit is one of the platforms that provide a place for extremist groups to express themselves. Groups that share (extreme) opinions often contain both passive members and activists. The goal of the activists is to shape the perceptions of the wider public and attract support for their ideology. The current study’s aims are twofold. First, it explores the engagement patterns of activists in the alt-right groups on a conservative subreddit. Second, this study explores the role played by anger and extremist wording in shaping the discussions that take place in the sub-reddit. The study employs inferential network analysis, namely conditional uniform graphs tests and exponential random graph model, in combination with sentiment analysis. Our findings suggest that a small group of users is engaged in the alt-right sub-reddit a lot more than the average user, defining a profile for activism. Moreover, we observe a relationship between engaging in alt-right discussions in the sub-reddit and the sentiment of the users deduced from their posts. These results are a first step toward finding strategies to prevent online radicalization.
Rodger Van Der Heijden, Ellen Mans, Tim Jongenelen, Claudia Zucca, Giuseppe Cascavilla
IEEE Big Data5
2024 Toxic language based echo chambers on the Incels.net community: A network analysis approach
abstract
This study examines interaction patterns on the web forum Incels.net through Social Network Analysis, focusing on the formation of echo chambers characterized by toxic language. We explore several hypotheses: (H2) users tend to engage in animated one-to-one interactions by quoting each other; (H4) frequent posters are less likely to be quoted by less active users, indicating a lack of clear leadership; and (H5) users sharing the same sentiment are more likely to participate in large discussions (echo chambers) but do not engage in one-to-one conversations. Our findings enhance the understanding of behaviors that may contribute to radicalization within fringe online communities and offer insights for future research into similar extreme hate forums.
Mathieu Janssen, Giuseppe Cascavilla, Claudia Zucca, Alfredo Cuzzocrea
IEEE Big Data2
2024 Copula-based Approaches for Anomaly Detection: a Case-study in Financial Forensics
abstract
Fraud detection is a critical challenge in various domains, necessitating accurate and reliable methods to distinguish between legitimate and fraudulent transactions. This work explores the application of copula-based models for anomaly detection in financial forensics. It focuses on their effectiveness in identifying fraudulent activities in a highly imbalanced dataset. Copula models are designed to capture the dependencies between continuous variables, providing a flexible framework for modeling the joint distribution of features. For instance, the variables in a dataset can follow different distributions and the copula is able to model how these variables jointly behave, particularly in extreme cases. In this work, we used a fraud dataset to calculate the copula-based probability of fraud and conditional Gaussian copula. Then, we derive the copula-based Generalized Linear Model (GLM) formula from the conditional copula which is essentially a GLM with probit link and transformed variables when the covariates are continuous. Finally, we compare the performance of these copula-based models with standard methods for predicting a binary variable like GLM with a probit link and logistic regression. Results indicate that copula-based probability and conditional copula formulas offer promising results, particularly in handling complex dependencies, but with a high computational time, while copula-based GLM, when combined with over-sampling, also outperforms traditional methods.
Vanessa Tenconi, Damian A. Tamburri, Giovanni Quattrocchi, Corrado Pellegrino, Giuseppe Cascavilla, Willem-Jan van den Heuvel
IEEE Big Data5
2024 Few Images, Many Insights: Illicit Content Detection Using a Limited Number of Images
abstract
The anonymity and untraceability benefits of the dark web increased its popularity exponentially. The cost of these technical benefits is that such anonymity has created a suitable womb for illicit activity. Hence—in collaboration with cybersecurity practitioners and law-enforcement agencies—the research community provided approaches for recognizing and classifying illicit activities. Most of these approaches exploit textual content from dark web markets, whereas few used images that originated from them. This article investigates alternative techniques for recognizing illegal activities from images. The significant contributions of our work are threefold: (a) We investigate label-agnostic learning techniques like One-Shot and Few-Shot learning that use Siamese Neural Networks. Our approach manages to handle small-scale datasets with promising accuracy. In particular, the Siamese Neural Network approach reaches 90.9% on 5-Shot experiments over a 10-class dataset. (b) This study’s satisfactory findings facilitate the creation of potent tools to assist authorities in identifying illicit content on the Web. Moreover, our proof-of-concept approach demonstrated the ability to recognize illegal images using a limited number of files, reducing the time constraint in collecting illegal images. (c) We provide a complete labeled dataset of 3,570 images from 55 different categories from dark web markets that can be used for future research activities.
Giuseppe Cascavilla, Gemma Catolino, Mauro Conti, Dimos Mellios, Damian A. Tamburri
ACM Trans. Intell. Syst. Technol.1
2023 BigData Fusion for Trajectory Prediction of Multi-Sensor Surveillance Information Systems
abstract
Video surveillance information systems assist forensics to examine and analyze the evidence from crime scenes to develop objective findings in the investigation of crime. Often, the existing surveillance information systems exploit an array of security cameras and IoT devices monitoring the same crime scene from different points of view while the crime unfolds over a range of time. However, none can automatically and selectively merge big data streams connected to such systems to provide a holistic, end-to-end safety picture.This work proposes a trajectory prediction architecture framework within a multi-sensor surveillance system. We developed a novel position measurement technique using monocular depth perception networks with multi-camera setup using triangulation. We tested and compared our technique with a single camera sensor in our first experiment and as the multi-camera setup determines the position of our target more accurately, we used our measurement function in our second experiment. In our second experiment, we employed the Unscented Kalman Filter (UKF) for predicting the trajectory of the target, and proved that UKF has good potential for being used in surveillance systems. Lastly, we designed a general architecture framework for big data analysis in multi-sensor surveillance systems consisting the four layers: the Sensor Layer, the Single Sensor Computation Layer, the Data Fusion and Interpretation Layer, and the Human Interaction Layer.
Giuseppe Cascavilla, Alfredo Cuzzocrea, Daniel De Pascale, Mandana Omidbakhsh, Damian A. Tamburri
IEEE Big Data1
2023 Real-world K-Anonymity applications: The KGen approach and its evaluation in fraudulent transactions
abstract
K-Anonymity is a property for the measurement, management, and governance of the data anonymization. Many implementations of k-anonymity have been described in state of the art, but most of them are not practically usable over a large number of attributes in a “Big” dataset, i.e., a dataset drawing from Big Data. To address this significant shortcoming, we introduce and evaluate KGen, an approach to K-anonymity featuring meta-heuristics, specifically, Genetic Algorithms to compute a permutation of the dataset which is both K-anonymized and still usable for further processing, e.g., for private-by-design analytics. KGen promotes such a meta-heuristic approach since it can solve the problem by finding a pseudo-optimal solution in a reasonable time over a considerable load of input. KGen allows the data manager to guarantee a high anonymity level while preserving the usability and preventing loss of information entropy over the data. Differently from other approaches that provide optimal global solutions compatible with smaller datasets, KGen works properly also over Big datasets while still providing a good-enough K-anonymized but still processable dataset. Evaluation results show how our approach can still work efficiently on a real world dataset, provided by Dutch Tax Authority, with 47 attributes (i.e., the columns of the dataset to be anonymized) and over 1.5K+ observations (i.e., the rows of that dataset), as well as on a dataset with 97 attributes and over 3942 observations.
Daniel De Pascale, Giuseppe Cascavilla, Damian A. Tamburri, Willem-Jan van den Heuvel
Inf. Syst.2
2022 Explaining IoT Attacks: An Effective and Efficient Semi-Supervised Learning Framework
abstract
Cyber-attacks targeting Internet-of-Things (IoT) devices are prevalent due to the limited security resources of the target devices and their often limited connectivity. Explaining such attacks is therefore greatly important to construct countermeasures. Current methods of automated IoT attack analysis require either large amounts of labelled data for classification, or use clustering methods which can be inaccurate. However, when a desired grouping of the data, as well as some prior knowledge about some observations in the data is available, approximate semi-supervised learning methods may be used to create accurate cluster arrangements. We therefore investigated the use of semi-supervised clustering approaches for creating accurate clusters of IoT attack sessions based on their goals and characteristic commonalities. We first manually created a ground-truth grouping of recent IoT attacks based on their goal. We differentiated the goal of each session according to the purpose of the used commands and the taken approach, resulting in a total of five classes. We then automatically constructed a feature set suitable for clustering similar IoT attack sessions using a method proposed in recent literature, and passed it to two different semi-supervised clustering algorithms using either labelled data (SeededKMeans) or pairwise constraints (PCKMeans) as prior knowledge. We found that both semi-supervised approaches were able to create accurate cluster arrangements using only small amounts of prior knowledge. Moreover, they outperformed an entirely unsupervised KMeans algorithm in terms of accuracy.
Giuseppe Cascavilla, Reinier Zwart, Damian A. Tamburri, Alfredo Cuzzocrea
IEEE Big Data1
2018 The insider on the outside: a novel system for the detection of information leakers in social networks
abstract
Confidential information is all too easily leaked by naive users posting comments. In this paper we introduce DUIL, a system for Detecting Unintentional Information Leakers. The value of DUIL is in its ability to detect those responsible for information leakage that occurs through comments posted on news articles in a public environment, when those articles have withheld material non-public information. DUIL is comprised of several artefacts, each designed to analyse a different aspect of this challenge: the information, the user(s) who posted the information, and the user(s) who may be involved in the dissemination of information. We present a design science analysis of DUIL as an information system artefact comprised of social, information, and technology artefacts. We demonstrate the performance of DUIL on real data crawled from several Facebook news pages spanning two years of news articles.
Giuseppe Cascavilla, Mauro Conti, David G. Schwartz, Inbal Yahav
Eur. J. Inf. Syst.1
2017 On the Influence of Emotional Valence Shifts on the Spread of Information in Social Networks
abstract
In this paper, we present a study on 4.4 million Twitter messages related to 24 systematically chosen real-world events. For each of the 4.4 million tweets, we first extracted sentiment scores based on the eight basic emotions according to Plutchik's wheel of emotions. Subsequently, we investigated the effects of shifts in the emotional valence on the spread of information. We found that in general OSN users tend to conform to the emotional valence of the respective real-world event. However, we also found empirical evidence that prospectively negative real-world events exhibit a significant amount of shifted emotions in the corresponding tweets (i.e. positive messages). To explain this finding, we use the theory of social connection and emotional contagion. To the best of our knowledge, this is the first study that provides empirical evidence for the undoing hypothesis in online social networks (OSNs). The undoing hypothesis postulates that positive emotions serve as an antidote during negative events.
Ema Kusen, Mark Strembeck, Giuseppe Cascavilla, Mauro Conti
ASONAM3
2015 Revealing Censored Information Through Comments and Commenters in Online Social Networks
abstract
In this work we study information leakage through discussions in online social networks. In particular, we focus on articles published by news pages, in which a person's name is censored, and we examine whether the person is identifiable (decensored) by analyzing comments and social network graphs of commenters. As a case study for our proposed methodology, in this paper we considered 48 articles (Israeli, military related) with censored content, followed by a threaded discussion. We qualitatively study the set of comments and identify comments (in this case referred as "leakers") and the commenter and the censored person. We denote these commenters as "leakers". We found that such comments are present for some 75% of the articles we considered. Finally, leveraging the social network graphs of the leakers, and specifically the overlap among the graphs of the leakers, we are able to identify the censored person. We show the viability of our methodology through some illustrative use cases.
Giuseppe Cascavilla, Mauro Conti, David G. Schwartz, Inbal Yahav
ASONAM1