Marco Mellia

dblp:18/6316 · DBLP profile ↗
← Back
21ranked-venue papers in the field
0as first author
9since 2021 · last 2025
0000-0003-1859-6693ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 8Big Data, Cloud & Distributed Data Systems · 8Data Mining & Knowledge Discovery · 4Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2025 Join the Chat: How Curiosity Sparks Participation in Telegram Groups
abstract
This study delves into the mechanisms that spark user curiosity driving active engagement within public Telegram groups. By analyzing approximately 6 million messages from 29,196 users across 409 groups, we identify and quantify the key factors that stimulate users to actively participate (i.e., send messages) in group discussions. These factors include social influence, novelty, complexity, uncertainty, and conflict, all measured through metrics derived from message sequences and user participation over time. After clustering the messages, we apply explainability techniques to assign meaningful labels to the clusters. This approach uncovers macro categories representing distinct curiosity stimulation profiles, each characterized by a unique combination of various stimuli. Social influence from peers and influencers drives engagement for some users, while for others, rare media types or a diverse range of senders and media sparks curiosity. Analyzing patterns, we found that user curiosity stimuli are mostly stable, but, as the time between the initial message increases, curiosity occasionally shifts. A graph-based analysis of influence networks reveals that users motivated by direct social influence tend to occupy more peripheral positions, while those who are not stimulated by any specific factors are often more central, potentially acting as initiators and conversation catalysts. These findings contribute to understanding information dissemination and spread processes on social media networks, potentially contributing to more effective communication strategies.
Giordano Paoletti, Jussara M. Almeida, Luca Vassio, Marcos André Gonçalves, Marco Mellia
ICWSM5
2025 MAD: Multicriteria Anomaly Detection of Suspicious Financial Accounts from Billions of Cash Transactions
abstract
This paper presents a real-world deployment case study on using unsupervised anomaly detection for Anti-Money Laundering (AML).Using more than 2 billion anonymized bank transactions that Intesa Sanpaolo, a primary Italian financial institution, registered over 8 months, we developed, tuned and deployed a machine learning pipeline in production.Experts from Intesa Sanpaolo validated the performance of our approach against the institution's traditional rule-based system and checked new real-world cases the system allowed them to identify.Besides increasing both precision and recall by a factor of 6 in the detection of high-risk cases, our pipeline raises 200+ additional alerts during the 8-month period, manually identified by branch managers, but missed by the rulebased system.More importantly, a manual inspection of 100 new unseen cases revealed 28 significant previously unreported cases.The pipeline, now fully deployed in Intesa Sanpaolo's Transaction Monitoring system, highlights the advantages of machine learning over traditional approaches typically adopted in this traditionally very conservative sector.
Giordano Paoletti, Flavio Giobergia, Danilo Giordano, Luca Cagliero, Silvia Ronchiadin, Dario Moncalvo, Marco Mellia, Elena Baralis
KDD (2)7
2025 Privacy Policies and Consent Management Platforms: Growth and Users' Interactions over Time
abstract
In response to growing concerns about user privacy, legislators have introduced new regulations and laws, such as the General Data Protection Regulation (GDPR) in the European Union and the California Consumer Privacy Act (CCPA), which force websites to obtain user consent before activating any personal data collection. The cornerstone of this consent-seeking process involves the use of Privacy Banners, the technical tools to collect users’ approval for data collection practices. Consent management platforms (CMPs) have emerged as practical solutions to simplify the configuration and management of such privacy banners for website administrators, allowing them to outsource the complexities of managing user consent and activating advertising features. This article presents a detailed and longitudinal analysis of the evolution of CMPs spanning 9 years. We take a twofold perspective: firstly, thanks to the HTTP Archive dataset, we provide insights into the growth, market share, and geographical spread of CMPs. Noteworthy observations include the substantial impact of the GDPR on the proliferation of CMPs in Europe, where more than 40% of websites currently adopt a CMP. Secondly, we analyse millions of user interactions with a medium-sized CMP present in thousands of websites worldwide. We observe how even small changes in the design of Privacy Banners have a critical impact on the user’s giving or denying one’s consent to data collection. For instance, over 60% of users do not consent when offered a simple “one-click reject-all” option. Conversely, when opting out requires more than one click, about 90% of users prefer to simply give their consent. This hints that their main objective is to eliminate the annoying privacy banner rather than make an informed decision. Curiously, we observe that iOS users exhibit a higher tendency to accept cookies compared with Android users, possibly indicating greater confidence in the privacy offered by Apple devices. We believe that the findings of this article contribute to a deeper understanding of the multifaceted interactions between privacy regulations, technological solutions and user choices in the evolving Web ecosystem. We also show that the availability of large open datasets, although not explicitly designed and collected for our goals, is fundamental to exploring different angles of the internet evolution over time. For this, we make the data and code used in this work available to the community. 1
Nikhil Jha, Martino Trevisan, Marco Mellia, Daniel Fernandez, Rodrigo Irarrazaval
ACM Trans. Web3
2025 CoDÆN: Benchmarks and Comparison of Evolutionary Community Detection Algorithms for Dynamic Networks
abstract
Web data are often modelled as complex networks in which entities interact and form communities. Nevertheless, web data evolves over time, and network communities change alongside it. This makes Community Detection (CD) in dynamic graphs a relevant problem, calling for evolutionary CD algorithms. The choice and evaluation of such algorithm performance is challenging because of the lack of a comprehensive set of benchmarks and specific metrics. To address these challenges, we propose CoDÆN—Community Detection Algorithms in Evolving Networks—a benchmarking framework for evolutionary CD algorithms in dynamic networks, that we offer as open source to the community. CoDÆN allows us to generate synthetic community-structured graphs with known ground truth and design evolving scenarios combining nine basic graph transformations that modify edges, nodes, and communities. We propose three complementary metrics (i.e., Correctness, Delay, and Stability) to compare evolutionary CD algorithms. Armed with CoDÆN, we consider three evolutionary modularity-based CD approaches, dissecting their performance to gauge the trade-off between the stability of the communities and their correctness. Next, we compare the algorithms in real Web-oriented datasets, confirming such a trade-off. Our findings reveal that algorithms that introduce memory in the graph maximise stability but add delay when abrupt changes occur. Conversely, algorithms that introduce memory by initialising the CD algorithms with the previous solution fail to identify the split and birth of new communities. These observations underscore the value of CoDÆN in facilitating the study and comparison of alternative evolutionary community detection algorithms.
Giordano Paoletti, Luca Gioacchini, Marco Mellia, Luca Vassio, Jussara M. Almeida
ACM Trans. Web3
2024 Re-Identification Attacks against the Topics API
abstract
Recently, Google proposed the Topics API framework as a privacy-friendly alternative for behavioural advertising as a possible solution to balance user’s privacy and advertisement effectiveness. Using the Topics API, the browser builds a user profile based on navigation history, which advertisers can access. The Topics API aim at becoming the new standard for behavioural advertising, thus it is necessary to fully understand its operation and find possible limitations. In this article, we evaluate the robustness of the Topics API to a re-identification attack. To build a user profile, we suppose an attacker accumulates over time the topics a user exposes to different websites. The attacker later re-identifies the same user matching the profiles of their audience. We leverage real traffic traces and realistic population models, and we present increasingly powerful attack threats. We find that the Topics API mitigates but cannot prevent re-identification from taking place, as there is a sizeable chance that a user’s profile remains unique within a website’s audience and the attacker successfully matches it with the profile of the same user on a second website. Depending on environmental factors, the probability of correct re-identification can reach 50%, considering a pool of 1,000 users. We offer the code and data we use in this work to stimulate further studies and the tuning of the Topic API parameters. 1
Nikhil Jha, Martino Trevisan, Emilio Leonardi, Marco Mellia
ACM Trans. Web4
2023 GLEm-Net: Unified Framework for Data Reduction with Categorical and Numerical Features
abstract
In the era of Big Data, effective data reduction through feature selection is of paramount importance for machine learning. This paper presents GLEm-Net (Grouped Lasso with Embeddings Network), a novel neural framework that seamlessly processes both categorical and numerical features to reduce the dimensionality of data while retaining as much information as possible. By integrating embedding layers, GLEm-Net effectively manages categorical features with high cardinality and compresses their information in a less dimensional space. By using a grouped Lasso penalty function in its architecture, GLEm-Net simultaneously processes categorical and numerical data, efficiently reducing high-dimensional data while preserving the essential information. We test GLEm-Net with a real-world application in an industrial environment where 6 million records exist and each is described by a mixture of 19 numerical and 7 categorical features with a strong class imbalance. A comparative analysis using state-of-the-art methods shows that despite the difficulty of building a high-performance model, GLEm-Net outperforms the other methods in both feature selection and classification, with a better balance in the selection of both numerical and categorical features.
Francesco De Santis, Danilo Giordano, Marco Mellia, Alessia Damilano
IEEE Big Data3
2023 Data driven scalability and profitability analysis in free floating electric car sharing systems
Alessandro Ciociola, Danilo Giordano, Luca Vassio, Marco Mellia
Inf. Sci.4
2022 Legal Entity Disambiguation for Financial Crime Detection
abstract
Transaction Monitoring is one of the main labor-intensive tasks of anti-financial crime and it requires to scrutinise billions of transactions per month against possible crimes. The first step in the process is the correct identification of the involved parties. This foundational step defines the focal entities on which transaction monitoring algorithms rely to spot suspicious events. Unfortunately, the loose syntax of protocols and the free text fields of inter-banking communications make party disambiguation particularly challenging. The first step of a fully automated data-driven strategy is thus the detection of the actual entity owning or using a given account.In this paper, we leverage data-driven techniques to identify and disambiguate the owners of accounts involved in cross-border international transactions when a Financial Institution only knows a minority fraction of such parties as its own customers. For this, we propose a data science pipeline relying on hierarchical clustering to capture similarities among names of parties involved in actual transactions. We test and tune the proposed approach using a large, real-world, multi-language, proprietary dataset of actual international transactions. Our highly parallel implementation completes the identification of parties that share an account and identifies all accounts owned by a party with f-score higher than 0.8.
Jacopo Fior, Thomas Favale, Luca Cagliero, Danilo Giordano, Marco Mellia, Elena Baralis, Silvia Ronchiadin, Paolo Baracco, Dario Moncalvo
IEEE Big Data5
2022 The Internet with Privacy Policies: Measuring The Web Upon Consent
abstract
To protect user privacy, legislators have regulated the use of tracking technologies, mandating the acquisition of users’ consent before collecting data. As a result, websites started showing more and more consent management modules–i.e., Consent Banners–the visitors have to interact with to access the website content. Since these banners change the content the browser loads, they challenge web measurement collection, primarily to monitor the extent of tracking technologies, but also to measure web performance. If not correctly handled, Consent Banners prevent crawlers from observing the actual content of the websites. In this paper, we present a comprehensive measurement campaign focusing on popular websites in Europe and the US, visiting both landing and internal pages from different countries around the world. We engineer Priv-Accept , a Web crawler able to accept the Consent Banners, as most users would do in practice. It lets us compare how webpages change before and after accepting such policies, if present. Our results show that all measurements performed ignoring the Consent Banners offer a biased and partial view of the Web. After accepting the privacy policies, web tracking is far more pervasive, and webpages are larger and slower to load.
Nikhil Jha, Martino Trevisan, Luca Vassio, Marco Mellia
ACM Trans. Web4
2020 z-anonymity: Zero-Delay Anonymization for Data Streams
abstract
With the advent of big data and the birth of the data markets that sell personal information, individuals' privacy is of utmost importance. The classical response is anonymization, i.e., sanitizing the information that can directly or indirectly allow users' re-identification. The most popular solution in the literature is the k-anonymity. However, it is hard to achieve k-anonymity on a continuous stream of data, as well as when the number of dimensions becomes high.In this paper, we propose a novel anonymization property called z-anonymity. Differently from k-anonymity, it can be achieved with zero-delay on data streams and it is well suited for high dimensional data. The idea at the base of z-anonymity is to release an attribute (an atomic information) about a user only if at least z - 1 other users have presented the same attribute in a past time window. z-anonymity is weaker than k-anonymity since it does not work on the combinations of attributes, but treats them individually. In this paper, we present a probabilistic framework to map the z-anonymity into the k-anonymity property. Our results show that a proper choice of the z-anonymity parameters allows the data curator to likely obtain a k-anonymized dataset, with a precisely measurable probability. We also evaluate a real use case, in which we consider the website visits of a population of users and show that z-anonymity can work in practice for obtaining the k-anonymity too.
Nikhil Jha, Thomas Favale, Luca Vassio, Martino Trevisan, Marco Mellia
IEEE BigData5
2019 The News We Like Are Not the News We Visit: News Categories Popularity in Usage Data
Zied Ben-Houidi, Giuseppe Scavo, Stefano Traverso, Renata Teixeira, Marco Mellia, Soumen Ganguly
ICWSM5
2019 Web Experience in Mobile Networks: Lessons from Two Million Page Visits
abstract
Measuring and characterizing web page performance is a challenging task. When it comes to the mobile world, the highly varying technology characteristics coupled with the opaque network configuration make it even more difficult. Aiming at reproducibility, we present a large scale empirical study of web page performance collected in eleven commercial mobile networks spanning four countries. By digging into measurement from nearly two million web browsing sessions, we shed light on the impact of different web protocols, browsers, and mobile technologies on the web performance. We find that the impact of mobile broadband access is sizeable. For example, the median page load time using mobile broadband increases by a third compared to wired access. Mobility clearly stresses the system, with handover causing the most evident performance penalties. Contrariwise, our measurements show that the adoption of HTTP/2 and QUIC has practically negligible impact. To understand the intertwining of all parameters, we adopt state-of-the-art statistical methods to identify the significance of different factors on the web performance. Our analysis confirms the importance of access technology and mobility context as well as webpage composition and browser. Our work highlights the importance of large-scale measurements. Even with our controlled setup, the complexity of the mobile web ecosystem is challenging to untangle. For this, we are releasing the dataset as open data for validation and further research.
Mohammad Rajiullah, Andra Lutu, Ali Safari Khatouni, Mah-Rukh Fida, Marco Mellia, Anna Brunström, Özgü Alay, Stefan Alfredsson, Vincenzo Mancuso
WWW5
2018 Achieving Horizontal Scalability in Density-based Clustering for URLs
abstract
Clustering has become an important means to analyze large datasets when labeled data is not available. The volume of data and its variety however challenge classical clustering algorithms, with density-based ones suffering from severe scalability issues.In this paper, we propose a way to perform density-based clustering efficiently by exploiting the horizontal scalability offered by big data solution such as Apache Spark. We are motivated by recent techniques for Internet monitoring that rely on clustering to group similar events and spot anomalies. We focus specifically on textual data, such as URLs or server logs. Computing the distance between points, here represented as strings, becomes a major issue. Indeed, when datasets become large, most of density-based clustering algorithms are bottlenecked by the computation of all the distances between any pairs of elements. To overcome this, we propose to decouple the distance computation, easily amenable to parallelization, from the algorithm execution. By using this approach, we can easily exploit the benefits of distributed platforms like Apache Spark or MapReduce. A faster execution of the algorithms is thus guaranteed, together with more flexibility in the choice of the clustering method.We make both the code and the dataset publicly available, to both guarantee the repeatability of the experiments, and possibly offering a new benchmark dataset.
Azadeh Faroughi, Reza Javidan, Marco Mellia, Andrea Morichetta 0002, Francesca Soro, Martino Trevisan
IEEE BigData3
2018 Mining Sensor Data for Predictive Maintenance in the Automotive Industry
abstract
Predictive maintenance is an ever-growing area of interest, spanning different fields and approaches. In the automotive industry faulty behaviors of the oxygen sensor are a key challenge to address. This paper presents OxyClog, a data-driven framework that, given a large number of time series collected from a vehicle's ECU (engine control unit), builds a model to predict if the oxygen sensor is currently unclogged, almost clogged (since the clogging of the sensor happens gradually), or clogged. OxyClog is characterized by a tailored preprocessing, which includes a custom and interpretable feature selection algorithm, along with a summarization strategy to transform a time-dependent problem into a time-independent one. Furthermore, a semi-supervised labeling methodology has been devised to use different data sources with different characteristics to define meaningful clogging labels. OxyClog integrates state-of-the-art classification algorithms - both interpretable and non-interpretable - to process real ECU data with good prediction performance.
Flavio Giobergia, Elena Baralis, Maria Camuglia, Tania Cerquitelli, Marco Mellia, Alessandra Neri, Davide Tricarico, Alessia Tuninetti
DSAA5
2018 You, the Web, and Your Device: Longitudinal Characterization of Browsing Habits
abstract
Understanding how people interact with the web is key for a variety of applications, e.g., from the design of effective web pages to the definition of successful online marketing campaigns. Browsing behavior has been traditionally represented and studied by means of clickstreams , i.e., graphs whose vertices are web pages, and edges are the paths followed by users. Obtaining large and representative data to extract clickstreams is, however, challenging. The evolution of the web questions whether browsing behavior is changing and, by consequence, whether properties of clickstreams are changing. This article presents a longitudinal study of clickstreams from 2013 to 2016. We evaluate an anonymized dataset of HTTP traces captured in a large ISP, where thousands of households are connected. We first propose a methodology to identify actual URLs requested by users from the massive set of requests automatically fired by browsers when rendering web pages. Then, we characterize web usage patterns and clickstreams, taking into account both the temporal evolution and the impact of the device used to explore the web. Our analyses precisely quantify various aspects of clickstreams and uncover interesting patterns, such as the typical short paths followed by people while navigating the web, the fast increasing trend in browsing from mobile devices, and the different roles of search engines and social networks in promoting content. Finally, we contribute a dataset of anonymized clickstreams to the community to foster new studies.1
Luca Vassio, Idilio Drago, Marco Mellia, Zied Ben-Houidi, Mohamed Lamine Lamali
ACM Trans. Web3
2017 Automatic detection of DNS manipulations
abstract
The DNS is a fundamental service that has been repeatedly attacked and abused. DNS manipulation is a prominent case: Recursive DNS resolvers are deployed to explicitly return manipulated answers to users' queries. While DNS manipulation is used for legitimate reasons too (e.g., parental control), rogue DNS resolvers support malicious activities, such as malware and viruses, exposing users to phishing and content injection. We introduce REMeDy, a system that assists operators to identify the use of rogue DNS resolvers in their networks. REMeDy is a completely automatic and parameter-free system that evaluates the consistency of responses across the resolvers active in the network. It operates by passively analyzing DNS traffic and, as such, requires no active probing of third-party servers. REMeDy is able to detect resolvers that manipulate answers, including resolvers that affect unpopular domains. We validate REMeDy using large-scale DNS traces collected in ISP networks where more than 100 resolvers are regularly used by customers. REMeDy automatically identifies regular resolvers, and pinpoint manipulated responses. Among those, we identify both legitimate services that offer additional protection to clients, and resolvers under the control of malwares that steer traffic with likely malicious goals.
Martino Trevisan, Idilio Drago, Marco Mellia, Maurizio M. Munafò
IEEE BigData3
2017 Mining and modeling web trajectories from passive traces
abstract
In modern web, users contact lots of services, identified by the domain name of the server. The temporal sequence and transitions of visited domains form a trajectory of the user on the web. In this work, we analyze 4 weeks of such trajectories, extracted from logs collected in our university network, and mine them via big data and machine learning methodologies to extract the interests of users. Our goal is to create a model of such trajectories and find similarities so to observe peculiarity of users' browsing. Thanks to the model, we propose a methodology to automatically group together the trajectories of single users and/or communities into highly descriptive environments which in turn allow the analyst to identify the topic of interest. We propose an automatic way to highlight differences in terms of popularity and content of environments. Lastly, we analyze the transition among environments, showing how people in smaller communities, e.g., in the same department, have a much more homogeneous behaviour than people at large, e.g., in the university.
Luca Vassio, Marco Mellia, Flavio Figueiredo, Ana Paula Couto da Silva, Jussara M. Almeida
IEEE BigData2
2016 WHAT: A big data approach for accounting of modern web services
abstract
HTTP(S) has become the main means to access the Internet. The web is a tangle, with (i) multiple services and applications co-located on the same infrastructure and (ii) several websites, services and applications embedding objects from CDN, ads and tracking platforms. Traditional solutions for traffic classification and metering fall short in providing visibility in users' activities. Service providers and corporate network administrators are left with huge amounts of measurements, which cannot immediately reveal the real impact of each web service on the network. Such visibility is key to dimension the network, charge users and policy traffic. This paper introduces the Web Helper Accounting Tool (WHAT), a system to uncover the overall traffic produced by specific web services. WHAT combines big data and machine learning approaches to process large volumes of network flow measurements and learn how to group traffic due to pre-defined services of interest. Our evaluation demonstrates WHAT effectiveness in enabling accurate accounting of the traffic associated to each service. WHAT illustrates the power of machine learning when applied to large datasets of network measurements, and allows network administrators to regain the lost visibility on network usage.
Martino Trevisan, Idilio Drago, Marco Mellia, Han Hee Song, Mario Baldi
IEEE BigData3
2014 Detecting malicious clients in ISP networks using HTTP connectivity graph and flow information
abstract
This paper considers an approach to identify previously undetected malicious clients in Internet Service Provider (ISP) networks by combining flow classification with a graph-based score propagation method. Our approach represents all HTTP communications between clients and servers as a weighted, near-bipartite graph, where the nodes correspond to the IP addresses of clients and servers while the links are their interconnections, weighted according to the output of a flow-based classifier. We employ a two-phase alternating score propagation algorithm on the graph to identify suspicious clients in a monitored network. Using a symmetrized weighted adjacency matrix as its input, we show that our score propagation algorithm is less vulnerable towards inflating the malicious scores of popular Web servers with high in-degrees compared to the normalization used in PageRank, a widely used graph-based method. Experimental results on a 4-hour network trace collected by a large Internet service provider showed that incorporating flow information into score propagation significantly improves the precision of the algorithm.
Sabyasachi Saha, Ruben Torres, Jianpeng Xu, Pang-Ning Tan, Antonio Nucci, Marco Mellia
ASONAM7
2014 Large-scale network traffic monitoring with DBStream, a system for rolling big data analysis
abstract
The complexity of the Internet has rapidly increased, making it more important and challenging to design scalable network monitoring tools. Network monitoring typically requires rolling data analysis, i.e., continuously and incrementally updating (rolling-over) various reports and statistics over highvolume data streams. In this paper, we describe DBStream, which is an SQL-based system that explicitly supports incremental queries for rolling data analysis. We also present a performance comparison of DBStream with a parallel data processing engine (Spark), showing that, in some scenarios, a single DBStream node can outperform a cluster of ten Spark nodes on rolling network monitoring workloads. Although our performance evaluation is based on network monitoring data, our results can be generalized to other Big Data problems with high volume and velocity.
Arian Bär, Alessandro Finamore, Pedro Casas, Lukasz Golab, Marco Mellia
IEEE BigData5
2013 TUCAN: Twitter user centric ANalyzer
abstract
Twitter has attracted millions of users that generate a humongous flow of information at constant pace. The research community has thus started proposing tools to extract meaningful information from tweets. In this paper, we take a different angle from the mainstream of previous works: we explicitly target the analysis of the timeline of tweets from "single users". We define a framework - named TUCAN - to compare information offered by the target users over time, and to pinpoint recurrent topics or topics of interest. First, tweets belonging to the same time window are aggregated into "bird songs". Several filtering procedures can be selected to remove stop-words and reduce noise. Then, each pair of bird songs is compared using a similarity score to automatically highlight the most common terms, thus highlighting recurrent or persistent topics. TUCAN can be naturally applied to compare bird song pairs generated from timelines of different users.
Luigi Grimaudo, Han Hee Song, Mario Baldi, Marco Mellia, Maurizio M. Munafò
ASONAM4