EDBT 2026 Demo / reviewers in the wild / expert
Martino Trevisan
dblp:186/8831
· DBLP profile ↗
44ranked-venue papers
12as first author
25since 2021 · last 2026
0000-0002-4258-4679ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 22 · 7 first-author · 14 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 2 since 2021Security and privacy · 5 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DRUID: Coordinating drone movements for compromised node identificationabstractIn recent years, Unmanned Aerial Vehicles (UAVs) (also called drones) networks have become increasingly popular in scenarios where rapid deployment, flexible mobility, and real-time data acquisition are crucial, such as disaster relief, environmental monitoring, military operations, and smart city infrastructure. However, due to their dynamic nature and dependence on wireless communication, they are intrinsically vulnerable to a variety of cyberattacks. In this work, we present DRUID , a decentralized scheme that silently identifies a compromised drone and selectively alters the messages it forwards. The scheme uses a combination of secret sharing and multipath routing to allow a pair of communicating drones, namely A and B , to detect the presence of a compromised drone along any route between them, thereby categorizing each route as either safe or compromised. The scheme operates iteratively and consists of three key modules: (i) an Information Retrieval Procedure that allows A to learn more about the topology, (ii) a binary search-like Identification Procedure, and (iii) if the previous module fails to identify the compromised drone, a Node Repositioning Procedure that relocates nodes closer to the compromised path. We validate DRUID on a large and diverse set of 178 731 graphs representing realistic UAV networks with different communication ranges. Comparing our scheme to previous work, experiments show that DRUID achieves a 97 % identification rate—up from the 54 % of the most recent alternative approach. We analyze the cost associated with the node repositioning procedure in terms of computation time and drone movement, and show that it generally takes a few seconds. Mauro Farina, Erica Salvato, Martino Trevisan, Alberto Bartoli |
Ad Hoc Networks | 3 |
| 2026 | Internet usage and performance in GEO satellite networks: A large-scale study across Europe and AfricaabstractSatellite Communication (SatCom) offers internet connectivity where traditional infrastructures are too expensive to deploy. When using satellites in a geostationary orbit, the distance from Earth forces a round-trip time of at least 550 ms. Coupled with the constrained capacity of the physical link, this challenges the traditional internet access quality we are used to. In this paper, we present a complete passive characterization of the traffic carried by an operational SatCom provider. With this unique vantage point, we observe the performance of the SatCom technology, as well as the usage habits of subscribers in different countries in Europe and Africa. We highlight the implications of such technology on Internet usage and functioning, and we pinpoint technical challenges due to the CDN and DNS resolution issues, while discussing possible optimizations that the ISP could implement to improve the service offered to SatCom subscribers. We complete the characterization of the adoption and performance of newer protocols with a focus on IPv6 and QUIC. Gabriele Merlach, Daniel Perdices, Gianluca Perna, Martino Trevisan, Danilo Giordano, Marco Mellia |
Comput. Networks | 4 |
| 2026 | DPMon: A differentially-private query engine for passive measurements
Martino Trevisan |
Comput. Networks | 1 |
| 2025 | Handling Large-Scale Network Flow Records: A Comparative Study on Lossy CompressionabstractFlow records, that summarize the characteristics of traffic flows, represent a practical and powerful way to monitor a network. While they already offer significant compression compared to full packet captures, their sheer volume remains daunting, especially for large Internet Service Providers (ISPs). In this paper, we investigate several lossy compression techniques to further reduce storage requirements while preserving the utility of flow records for key tasks, such as predicting the domain name of contacted servers. Our study evaluates scalar quantization, Principal Component Analysis (PCA), and vector quantization, applied to a real-world dataset from an operational campus network. Results reveal that scalar quantization provides the best tradeoff between compression and accuracy. PCA can preserve predictive accuracy but hampers subsequent entropic compression, and while vector quantization shows promise, it struggles with scalability due to the high-dimensional nature of the data. These findings result in practical strategies for optimizing flow record storage in large-scale monitoring scenarios. Gabriele Merlach, Martino Trevisan, Damiano Ravalico, Fabio Palmese, Giovanni Baccichet, Alessandro Redondi |
NOMS | 2 |
| 2025 | A Zero Trust Data-Driven Perspective on PKI Root StoresabstractSecurity and privacy on the Internet rely on the Public Key Infrastructure (PKI), which is based on unlimited trust in a set of predefined certification authorities included in the users’ root stores. However, the architecture of the PKI is no longer appropriate for the current threat landscape and security principles. Specifically, the implicit and permanent trust given to certification authorities collides with the rising zero trust approach, a cyber-security model that mandates that trust must never be granted implicitly or permanently to any entity. This work offers a zero trust perspective on the PKI and root store composition. Using navigation datasets collected from users’ browsers and passive monitors, we analyze their actual needs and identify the portion of root stores that are useful for their activity. We propose several zero trust policies to manage root stores that shrink the large perimeter of trust allowed by commercial root stores. Our experiments show that less than half of the root certificates included in the Mozilla root store are indeed used for navigation, while only 14 cover 99% of the traffic of our users. Moreover, implementing such policies requires little effort for a company, providing a practical way for managing root stores with up-to-date security principles. Mauro Farina, Damiano Ravalico, Martino Trevisan, Alberto Bartoli |
IEEE J. Sel. Areas Commun. | 3 |
| 2025 | Privacy Policies and Consent Management Platforms: Growth and Users' Interactions over TimeabstractIn response to growing concerns about user privacy, legislators have introduced new regulations and laws, such as the General Data Protection Regulation (GDPR) in the European Union and the California Consumer Privacy Act (CCPA), which force websites to obtain user consent before activating any personal data collection. The cornerstone of this consent-seeking process involves the use of Privacy Banners, the technical tools to collect users’ approval for data collection practices. Consent management platforms (CMPs) have emerged as practical solutions to simplify the configuration and management of such privacy banners for website administrators, allowing them to outsource the complexities of managing user consent and activating advertising features. This article presents a detailed and longitudinal analysis of the evolution of CMPs spanning 9 years. We take a twofold perspective: firstly, thanks to the HTTP Archive dataset, we provide insights into the growth, market share, and geographical spread of CMPs. Noteworthy observations include the substantial impact of the GDPR on the proliferation of CMPs in Europe, where more than 40% of websites currently adopt a CMP. Secondly, we analyse millions of user interactions with a medium-sized CMP present in thousands of websites worldwide. We observe how even small changes in the design of Privacy Banners have a critical impact on the user’s giving or denying one’s consent to data collection. For instance, over 60% of users do not consent when offered a simple “one-click reject-all” option. Conversely, when opting out requires more than one click, about 90% of users prefer to simply give their consent. This hints that their main objective is to eliminate the annoying privacy banner rather than make an informed decision. Curiously, we observe that iOS users exhibit a higher tendency to accept cookies compared with Android users, possibly indicating greater confidence in the privacy offered by Apple devices. We believe that the findings of this article contribute to a deeper understanding of the multifaceted interactions between privacy regulations, technological solutions and user choices in the evolving Web ecosystem. We also show that the availability of large open datasets, although not explicitly designed and collected for our goals, is fundamental to exploring different angles of the internet evolution over time. For this, we make the data and code used in this work available to the community. 1 Nikhil Jha, Martino Trevisan, Marco Mellia, Daniel Fernandez, Rodrigo Irarrazaval |
ACM Trans. Web | 2 |
| 2024 | Monitoring Web QoE in Satellite Networks from Passive MeasurementsabstractSatellite Communication (SatCom) is the only choice to access the Internet in remote regions and is characterized by extreme latency and constrained capacity. For SatCom operators, it is thus fundamental to monitor the Quality of Experience (QoE) of subscribers, to measure their satisfaction, spot anomalies and optimize the peculiar network setup. The Web has become the primary source of Internet content, and Web browsing is the main activity of internauts. This paper addresses the challenge of monitoring Web QoE in SatCom environments, proposing a tailored system that employs a supervised approach to predict Web QoE using passive measurements. The system collects training data through Test Agents that mimic real subscribers' traffic patterns and uses them to build Machine Learning (ML) models that predict performance metrics. The findings demonstrate the feasibility of monitoring Web QoE in SatCom environments, with limitations on website applicability and temporal stability. The need for periodic data generation and the development of a general machine learning model for unseen websites remain open challenges. This research contributes to enhancing web browsing experiences in SatCom and expanding understanding of Web QoE monitoring in diverse network settings. Gianluca Perna, Martino Trevisan, Danilo Giordano, Daniel Perdices, Marco Mellia |
CCNC | 2 |
| 2024 | TASP: Topic-based abstractive summarization of Facebook text posts
Irene Benedetto, Moreno La Quatra, Luca Cagliero, Luca Vassio, Martino Trevisan |
Expert Syst. Appl. | 5 |
| 2024 | Re-Identification Attacks against the Topics APIabstractRecently, Google proposed the Topics API framework as a privacy-friendly alternative for behavioural advertising as a possible solution to balance user’s privacy and advertisement effectiveness. Using the Topics API, the browser builds a user profile based on navigation history, which advertisers can access. The Topics API aim at becoming the new standard for behavioural advertising, thus it is necessary to fully understand its operation and find possible limitations. In this article, we evaluate the robustness of the Topics API to a re-identification attack. To build a user profile, we suppose an attacker accumulates over time the topics a user exposes to different websites. The attacker later re-identifies the same user matching the profiles of their audience. We leverage real traffic traces and realistic population models, and we present increasingly powerful attack threats. We find that the Topics API mitigates but cannot prevent re-identification from taking place, as there is a sizeable chance that a user’s profile remains unique within a website’s audience and the attacker successfully matches it with the profile of the same user on a second website. Depending on environmental factors, the probability of correct re-identification can reach 50%, considering a pool of 1,000 users. We offer the code and data we use in this work to stimulate further studies and the tuning of the Topic API parameters. 1 Nikhil Jha, Martino Trevisan, Emilio Leonardi, Marco Mellia |
ACM Trans. Web | 2 |
| 2023 | Measuring the Performance of iCloud Private Relay
Martino Trevisan, Idilio Drago, Paul Schmitt, Francesco Bronzino |
PAM | 1 |
| 2023 | Practical anonymization for data streams: z-anonymity and relation with k-anonymity
Nikhil Jha, Luca Vassio, Martino Trevisan, Emilio Leonardi, Marco Mellia |
Perform. Evaluation | 3 |
| 2023 | On the Robustness of Topics API to a Re-Identification AttackabstractWeb tracking through third-party cookies is considered a threat to users' privacy and is supposed to be abandoned in the near future. Recently, Google proposed the Topics API framework as a privacy-friendly alternative for behavioural advertising. Using this approach, the browser builds a user profile based on navigation history, which advertisers can access. The Topics API has the possibility of becoming the new standard for behavioural advertising, thus it is necessary to fully understand its operation and find possible limitations. This paper evaluates the robustness of the Topics API to a re-identification attack where an attacker reconstructs the user profile by accumulating user's exposed topics over time to later re-identify the same user on a different website. Using real traffic traces and realistic population models, we find that the Topics API mitigates but cannot prevent re-identification to take place, as there is a sizeable chance that a user's profile is unique within a website's audience. Consequently, the probability of correct re-identification can reach 15-17%, considering a pool of 1,000 users. We offer the code and data we use in this work to stimulate further studies and the tuning of the Topic API parameters. Nikhil Jha, Martino Trevisan, Emilio Leonardi, Marco Mellia |
Proc. Priv. Enhancing Technol. | 2 |
| 2023 | URLGEN - Toward Automatic URL Generation Using GANsabstractURLs play an essential role on the Internet, allowing access to Web resources. Automatically generating URLs is helpful in various tasks, such as application debugging, API testing, and blocklist creation for security applications. Current testing suites deeply embed experts’ domain knowledge to generate suitable URLs, resulting in an ad-hoc solution for each given application. These tools thus require heavy manual intervention, with the expensive coding of rules that are hard to maintain. We here introduce URLGEN, a system that uses Generative Adversarial Networks (GANs) to tackle the automatic URL generation problem. URLGEN is designed for web API testing and generates URL samples for an application without any system expertise, complementing the existing tools. It leverages Long Short-Term Memory (LSTM) and Convolutional Neural Network (CNN) architectures, augmented by an embedding layer that simplifies the URL learning and generation process. We show that URLGEN learns to generate new valid URLs from samples of real URLs without requiring any domain knowledge and following a purely data-driven approach. We compare the GAN architecture of URLGEN against other design options and show that the LSTM architecture can better capture the correlation among URL characters, outperforming previously proposed solutions. Finally, we show that the URLGEN approach can be extended to other scenarios, which we illustrate with two use cases, i.e., cybersquatting domain prediction and URL classification. Rodolfo V. Valentim, Idilio Drago, Martino Trevisan, Marco Mellia |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2023 | Attacking DoH and ECH: Does Server Name Encryption Protect Users' Privacy?abstractPrivacy on the Internet has become a priority, and several efforts have been devoted to limit the leakage of personal information. Domain names, both in the TLS Client Hello and DNS traffic, are among the last pieces of information still visible to an observer in the network. The Encrypted Client Hello extension for TLS, DNS over HTTPS or over QUIC protocols aim to further increase network confidentiality by encrypting the domain names of the visited servers. In this article, we check whether an attacker able to passively observe the traffic of users could still recover the domain name of websites they visit even if names are encrypted. By relying on large-scale network traces, we show that simplistic features and off-the-shelf machine learning models are sufficient to achieve surprisingly high precision and recall when recovering encrypted domain names. We consider three attack scenarios, i.e., recovering the per-flow name, rebuilding the set of visited websites by a user, and checking which users visit a given target website. We next evaluate the efficacy of padding-based mitigation, finding that all three attacks are still effective, despite resources wasted with padding. We conclude that current proposals for domain encryption may produce a false sense of privacy, and more robust techniques should be envisioned to offer protection to end users. Martino Trevisan, Francesca Soro, Marco Mellia, Idilio Drago, Ricardo Morla |
ACM Trans. Internet Techn. | 1 |
| 2022 | A first look at starlink performanceabstractWith new Low Earth Orbit satellite constellations such as Starlink, satellite-based Internet access is becoming an alternative to traditional fixed and wireless technologies with comparable throughputs and latencies. In this paper, we investigate the user-perceived performance of Starlink. Our measurements show that latency remains low and does not vary significantly under idle or lightly loaded links. Compared to another commercial Internet access using a geostationary satellite, Starlink achieves higher TCP throughput and provides faster web browsing. To avoid interference from performance enhancing proxies commonly used in satellite networks, we also use QUIC to assess performance under load and packet loss. Our results indicate that delay and packet loss increase slightly under load for both upload and download. François Michel, Martino Trevisan, Danilo Giordano, Olivier Bonaventure |
IMC | 2 |
| 2022 | When satellite is all you have: watching the internet from 550 msabstractSatellite Communication (SatCom) offers internet connectivity where traditional infrastructures are too expensive to deploy. When using satellites in a geostationary orbit, the distance from Earth forces a round trip time higher than 550 ms. Coupled with the limited and shared capacity of the physical link, this poses a challenge to the traditional internet access quality we are used to. Daniel Perdices, Gianluca Perna, Martino Trevisan, Danilo Giordano, Marco Mellia |
IMC | 3 |
| 2022 | Retina: An open-source tool for flexible analysis of RTC traffic
Gianluca Perna, Dena Markudova, Martino Trevisan, Paolo Garza, Michela Meo, Maurizio M. Munafò |
Comput. Networks | 3 |
| 2022 | A first look at HTTP/3 adoption and performance
Gianluca Perna, Martino Trevisan, Danilo Giordano, Idilio Drago |
Comput. Commun. | 2 |
| 2022 | Measuring Roaming in Europe: Infrastructure and Implications on Users' QoEabstract“Roam like Home” is the initiative of the European Commission (EC) to end the levy of extra charges when roaming within the European region. As a result, people can use data services more freely across Europe. However, the implications of roaming solutions on network performance have not been carefully examined yet. This paper provides an in-depth characterization of the implications of international data roaming within Europe. We build a unique roaming measurement platform using 16 different mobile networks deployed in six countries across Europe. Using this platform, we measure different aspects of international roaming in 4G networks in Europe, including mobile network configuration, performance characteristics, and quality of experience. We find that operators adopt a common approach to implement roaming called Home-routed roaming (HR). This results in additional latency penalties of 60 ms or more, depending on geographical distance. This leads to worse browsing performance, with an increase in the metrics related to Quality of Experience (QoE) of users (Page Load time and Speed Index) in the order of 15-20 percent. We further analyze in isolation the impact of latency on QoE metrics and find that the penalty imposed by HR leads to a degradation on QoE metrics up to 150 percent in case of intercontinental roaming. Anna Maria Mandalari, Andra Lutu, Ana Custura, Ali Safari Khatouni, Özgü Alay, Marcelo Bagnulo, Vaibhav Bajpai, Anna Brunström, Jörg Ott, Martino Trevisan, Marco Mellia, Gorry Fairhurst |
IEEE Trans. Mob. Comput. | 10 |
| 2022 | Real-Time Classification of Real-Time CommunicationsabstractReal-time communication (RTC) applications have become largely popular in the last decade with the spread of broadband and mobile Internet access. Nowadays, these platforms are a fundamental means for connecting people and supporting businesses that increasingly rely on forms of remote work. In this context, it is of paramount importance to operate at the network level to ensure adequate Quality of Experience (QoE) for users, and appropriate traffic management policies are essential to prioritize RTC traffic. This in turn requires the network to be able to identify RTC streams and the type of content they carry. In this paper, we propose a machine learning-based application to classify media streams generated by RTC applications encapsulated in Secure Real-Time Protocol (SRTP) flows in real-time. Using carefully tuned features extracted from packet characteristics, we train models to classify streams into a variety of classes, including media type (audio/video), video quality, and redundant streams. We validate our approach using traffic from over 62 hours of multi-party meetings conducted using two popular RTC applications, namely Cisco Webex Teams and Jitsi Meet. We achieve an overall accuracy of 96% for Webex and 95% for Jitsi, using a lightweight decision tree model that makes decisions based solely on 1 second of real-time traffic. Our results show that models trained for a particular meeting software have difficulty when used with another one, although domain adaptation techniques facilitate the transfer of pre-trained models. Gianluca Perna, Dena Markudova, Martino Trevisan, Paolo Garza, Michela Meo, Maurizio M. Munafò, Giovanna Carofiglio |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2022 | The Internet with Privacy Policies: Measuring The Web Upon ConsentabstractTo protect user privacy, legislators have regulated the use of tracking technologies, mandating the acquisition of users’ consent before collecting data. As a result, websites started showing more and more consent management modules–i.e., Consent Banners–the visitors have to interact with to access the website content. Since these banners change the content the browser loads, they challenge web measurement collection, primarily to monitor the extent of tracking technologies, but also to measure web performance. If not correctly handled, Consent Banners prevent crawlers from observing the actual content of the websites. In this paper, we present a comprehensive measurement campaign focusing on popular websites in Europe and the US, visiting both landing and internal pages from different countries around the world. We engineer Priv-Accept , a Web crawler able to accept the Consent Banners, as most users would do in practice. It lets us compare how webpages change before and after accepting such policies, if present. Our results show that all measurements performed ignoring the Consent Banners offer a biased and partial view of the Web. After accepting the privacy policies, web tracking is far more pervasive, and webpages are larger and slower to load. Nikhil Jha, Martino Trevisan, Luca Vassio, Marco Mellia |
ACM Trans. Web | 2 |
| 2021 | The stock exchange of influencers: a financial approach for studying fanbase variation trendsabstractIn many online social networks (OSNs), a limited portion of profiles emerges and reaches a large base of followers, i.e., the so-called social influencers. One of their main goals is to increase their fanbase to increase their visibility, engaging users through their content. In this work, we propose a novel parallel between the ecosystem of OSNs and the stock exchange market. Followers act as private investors, and they follow influencers, i.e., buy stocks, based on their individual preferences and on the information they gather through external sources. In this preliminary study, we show how the approaches proposed in the context of the stock exchange market can be successfully applied to social networks. Our case study focuses on 60 Italian Instagram influencers and shows how their followers short-term trends obtained through Bollinger bands become close to those found in external sources, Google Trends in our case, similarly to phenomena already observed in the financial market. Besides providing a strong correlation between these different trends, our results pose the basis for studying social networks with a new lens, linking them with a different domain. Fabio Bertone, Luca Vassio, Martino Trevisan |
ASONAM | 3 |
| 2021 | Online Classification of RTC TrafficabstractReal-time communication (RTC) platforms have become increasingly popular in the last decade, together with the spread of broadband Internet access. They are nowadays a fundamental means for connecting people and supporting the economy, which relies more and more on forms of remote working. In this context, it is particularly important to act at the network level to ensure adequate Quality of Experience (QoE) to users, where proper traffic management policies are essential to prioritize RTC traffic. This, in turn, requires in-network devices to identify RTC streams and the type of content they carry. In this paper, we propose a machine learning-based application to classify, in real-time, the media streams generated by RTC applications encapsulated in Secure Real Time Protocol (SRTP) flows. Using carefully tuned features extracted from packet characteristics, we train a model to classify streams into an ample set of classes, including media type (audio/video), video quality and redundant streams. To validate our approach, we use traffic from more than 88 hours of multi-party meeting calls made using the Cisco Webex Teams application. We reach an overall accuracy of 97% with a light-weight decision tree model, which makes decisions using only 1 second of traffic. Gianluca Perna, Dena Markudova, Martino Trevisan, Paolo Garza, Michela Meo, Maurizio M. Munafò, Giovanna Carofiglio |
CCNC | 3 |
| 2021 | Understanding web pornography usage from traffic analysis
Andrea Morichetta 0002, Martino Trevisan, Luca Vassio, Julia Krickl |
Comput. Networks | 2 |
| 2021 | α-MON: Traffic Anonymizer for Passive MonitoringabstractPacket measurements at scale are essential for several applications, such as cyber-security, accounting and troubleshooting. They, however, threaten users’ privacy by exposing sensitive information. Anonymization has been the answer to this challenge, i.e., replacing sensitive information with obfuscated copies. Anonymization of packet traces, however, comes with some challenges and drawbacks. First, it reduces the value of data. Second, it requires to consider diverse protocols because information may leak from many non-encrypted fields. Third, it must be performed at high speeds directly at the monitor, to prevent private data from leaking, calling for real-time solutions. We present$\alpha $-MON, a flexible tool for privacy-preserving packet monitoring. It replicates input packet streams to different consumers while anonymizing protocol fields according to flexible policies that cover all protocol layers. Beside classic anonymization mechanisms such as IP address obfuscation,$\alpha $-MON supports${z}$-anonymization, a novel solution to obfuscate rare values that can be uniquely traced back to limited sets of users. Differently from classic anonymization approaches,z-anonymityworks on a streaming fashion, with zero delay, operating at high-speed links on a packet-by-packet basis. We quantify the impact ofz-anonymityon traffic measurements, finding that it introduces minimal error when it comes to finding heavy-hitter services. We evaluate$\alpha $-MON performance using packet traces collected from an ISP network and show that it achieves a sustainable rate of 40 Gbit/s on a Commercial Off-the Shelf server.$\alpha $-MON is available to the community as an open-source project. Thomas Favale, Martino Trevisan, Idilio Drago, Marco Mellia |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2020 | z-anonymity: Zero-Delay Anonymization for Data StreamsabstractWith the advent of big data and the birth of the data markets that sell personal information, individuals' privacy is of utmost importance. The classical response is anonymization, i.e., sanitizing the information that can directly or indirectly allow users' re-identification. The most popular solution in the literature is the k-anonymity. However, it is hard to achieve k-anonymity on a continuous stream of data, as well as when the number of dimensions becomes high.In this paper, we propose a novel anonymization property called z-anonymity. Differently from k-anonymity, it can be achieved with zero-delay on data streams and it is well suited for high dimensional data. The idea at the base of z-anonymity is to release an attribute (an atomic information) about a user only if at least z - 1 other users have presented the same attribute in a past time window. z-anonymity is weaker than k-anonymity since it does not work on the combinations of attributes, but treats them individually. In this paper, we present a probabilistic framework to map the z-anonymity into the k-anonymity property. Our results show that a proper choice of the z-anonymity parameters allows the data curator to likely obtain a k-anonymized dataset, with a precisely measurable probability. We also evaluate a real use case, in which we consider the website visits of a population of users and show that z-anonymity can work in practice for obtaining the k-anonymity too. Nikhil Jha, Thomas Favale, Luca Vassio, Martino Trevisan, Marco Mellia |
IEEE BigData | 4 |
| 2020 | Realistic testing of RTC applications under mobile networksabstractThe increasing usage of Real-Time Communication (RTC) applications for leisure and remote working calls for realistic and reproducible techniques to test them. They are used under very different network conditions: from high-speed broadband networks, to noisy wireless links. As such, it is of paramount importance to assess the impact of the network on users' Quality of Experience (QoE), especially when it comes to the application's mechanisms such as video quality adjustment or transmission of redundant data. In this work, we pose the basis for a system in which a target RTC application is tested in an emulated mobile environment. To this end, we leverage ERRANT, a data-driven emulator which includes 32 distinct profiles modeling mobile network performance in different conditions. As a use case, we opt for Cisco Webex, a popular RTC application. We show how variable network conditions impact the packet loss, and, in turn, trigger video quality adjustments, impairing the users' QoE. Gianluca Perna, Martino Trevisan, Danilo Giordano |
CoNEXT | 2 |
| 2020 | A comparative study of RTC applicationsabstractReal-Time Communication (RTC) applications have become ubiquitous and are nowadays fundamental for people to communicate with friends and relatives, as well as for enterprises to allow remote working and save travel costs. Countless competing platforms differ in the ease of use, features they implement, supported user equipment and targeted audience (consumer of business). However, there is no standard protocol or interoperability mechanism. This picture complicates the traffic management, making it hard to isolate RTC traffic for prioritization or obstruction. Moreover, undocumented operation could result in the traffic being blocked at firewalls or middleboxes. In this paper, we analyze 13 popular RTC applications, from widespread consumer apps, like Skype and Whatsapp, to business platforms dedicated to enterprises - Microsoft Teams and Webex Teams. We collect packet traces under different conditions and illustrate similarities and differences in their use of the network. We find that most applications employ the well-known RTP protocol, but we observe a few cases of different (and even undocumented) approaches. The majority of applications allow peer-to-peer communication during calls with only two participants. Six of them send redundant data for Forward Error Correction or encode the user video at different bitrates. In addition, we notice that many of them are easy to identify by looking at the destination servers or the domain names resolved via DNS. The packet traces we collected, along with the metadata we extract, are made available to the community. Antonio Nisticò, Dena Markudova, Martino Trevisan, Michela Meo, Giovanna Carofiglio |
ISM | 3 |
| 2020 | Campus traffic and e-Learning during COVID-19 pandemic
Thomas Favale, Francesca Soro, Martino Trevisan, Idilio Drago, Marco Mellia |
Comput. Networks | 3 |
| 2020 | ERRANT: Realistic emulation of radio access networks
Martino Trevisan, Ali Safari Khatouni, Danilo Giordano |
Comput. Networks | 1 |
| 2020 | Five Years at the Edge: Watching Internet From the ISP NetworkabstractThe Internet and the way people use it are constantly changing. Knowing traffic is crucial for operating the network, understanding users' needs, and ultimately improving applications. Here, we provide an in-depth longitudinal view of Internet traffic during 5 years (from 2013 to 2017). We take the point of the view of a national-wide ISP and analyze rich flow-level measurements to pinpoint and quantify changes. We observe the traffic, both from a point of view of users and services. We show that an ordinary broadband subscriber downloaded in 2017 more than twice as much as they used to do 5 years before. Bandwidth hungry video services drove this change at the beginning, while recently social messaging applications contribute to increase of data consumption. We study how protocols and service infrastructures evolve over time, highlighting events that may challenge traffic management policies. In the rush to bring servers closer and closer to users, we witness the birth of the sub-millisecond Internet, with caches located directly at ISP edges. The picture we take shows a lively Internet that always evolves and suddenly changes. To support new analyses, we make anonymized data available at https://smartdata.polito.it/five-years-at-the-edge/. Martino Trevisan, Danilo Giordano, Idilio Drago, Maurizio M. Munafò, Marco Mellia |
IEEE/ACM Trans. Netw. | 1 |
| 2019 | Poster: On the Application of NLP to Discover Relationships between Malicious Network EntitiesabstractThe increase in network traffic volumes challenges the scalability of security analysis tools. In this paper, we present NetLearn, a solution to identify potentially malicious network entities from large amounts of network traffic data. NetLearn applies recently developed natural language processing algorithms to discover security-relevant relationships between the observed network entities, e.g., domain names and IP addresses, without requiring external sources of information for its analysis. Giuseppe Siracusano, Martino Trevisan, Roberto Gonzalez, Roberto Bifulco |
CCS | 2 |
| 2019 | Data-Driven Emulation of Mobile Access NetworksabstractNetwork monitoring is fundamental to understand network evolution and behavior. However, monitoring studies have the main limitation of running new experiments when the phenomenon under analysis is over e.g., congestion. To overcome this limitation, network emulation is of vital importance for network testing and research experiments either in wired and mobile networks. When it comes to mobile networks, the variety of technical characteristics, coupled with the opaque network configurations, make realistic network emulation a challenging task. In this paper, we address this issue leveraging a large scale dataset composed of 500M network latency measurements in Mobile BroadBand networks. By using this dataset, we create 51 different network latency profiles based on the Mobile BroadBand operator, the radio access technology and signal strength. These profiles are then processed to make them compatible with the tc-netem emulation tool. Finally, we show that, despite the limitation of current tc-netem emulation tool, Generative Adversarial Networks are a promising solution used to create realistic temporal emulation. We believe that this work could be the first step toward a comprehensive data-driven network emulation. For this, we make our profiles and codes available to foster further studies in these directions. Ali Safari Khatouni, Martino Trevisan, Danilo Giordano |
CNSM | 2 |
| 2019 | Are Darknets All The Same? On Darknet Visibility for Security MonitoringabstractDarknets are sets of IP addresses that are advertised but do not host any client or server. By passively recording the incoming packets, they assist network monitoring activities. Since packets they receive are unsolicited by definition, darknets help to spot misconfigurations as well as important security events, such as the appearance and spread of botnets, DDoS attacks using spoofed IP address, etc. A number of organizations worldwide deploys darknets, ranging from a few dozens of IP addresses to large /8 networks. We here investigate how similar is the visibility of different darknets. By relying on traffic from three darknets deployed in different contintents, we evaluate their exposure in terms of observed events given their allocated IP addresses. The latter is particularly relevant considering the shortage of IPv4 addresses on the Internet. Our results suggest that some well-known facts about darknet visibility seem invariant across deployments, such as the most commonly contacted ports. However, size and location matter. We find significant differences in the observed traffic from darknets deployed in different IP ranges as well as according to the size of the IP range allocated for the monitoring. Francesca Soro, Idilio Drago, Martino Trevisan, Marco Mellia, João M. Ceron, José Jair Santanna |
LANMAN | 3 |
| 2019 | Characterizing Web Pornography Consumption from Passive Measurements
Andrea Morichetta 0002, Martino Trevisan, Luca Vassio |
PAM | 2 |
| 2019 | PAIN: A Passive Web performance indicator for ISPs
Martino Trevisan, Idilio Drago, Marco Mellia |
Comput. Networks | 1 |
| 2019 | 4 Years of EU Cookie Law: Results and Lessons LearnedabstractAbstract Personalized advertisement has changed the web. It lets websites monetize the content they offer. The downside is the continuous collection of personal information with significant threats to personal privacy. In 2002, the European Union (EU) introduced a first set of regulations on the use of online tracking technologies. It aimed, among other things, to make online tracking mechanisms explicit to increase privacy awareness among users. Amended in 2009, the EU Directive mandates websites to ask for informed consent before using any kind of profiling technology, e.g., cookies. Since 2013, the ePrivacy Directive became mandatory, and each EU Member State transposed it in national legislation. Since then, most of European websites embed a “Cookie Bar”, the most visible effect of the regulation. In this paper, we run a large-scale measurement campaign to check the current implementation status of the EU cookie directive. For this, we use CookieCheck, a simple tool to automatically verify legislation violations. Results depict a shady picture: 49 % of websites do not respect the Directive and install profiling cookies before any user’s consent is given. Beside presenting a detailed picture, this paper casts lights on the difficulty of legislator attempts to regulate the troubled marriage between ad-supported web services and their users. In this picture, online privacy seems to be continuously at stake, and it is hard to reach transparency. Martino Trevisan, Stefano Traverso, Eleonora Bassi, Marco Mellia |
Proc. Priv. Enhancing Technol. | 1 |
| 2018 | Achieving Horizontal Scalability in Density-based Clustering for URLsabstractClustering has become an important means to analyze large datasets when labeled data is not available. The volume of data and its variety however challenge classical clustering algorithms, with density-based ones suffering from severe scalability issues.In this paper, we propose a way to perform density-based clustering efficiently by exploiting the horizontal scalability offered by big data solution such as Apache Spark. We are motivated by recent techniques for Internet monitoring that rely on clustering to group similar events and spot anomalies. We focus specifically on textual data, such as URLs or server logs. Computing the distance between points, here represented as strings, becomes a major issue. Indeed, when datasets become large, most of density-based clustering algorithms are bottlenecked by the computation of all the distances between any pairs of elements. To overcome this, we propose to decouple the distance computation, easily amenable to parallelization, from the algorithm execution. By using this approach, we can easily exploit the benefits of distributed platforms like Apache Spark or MapReduce. A faster execution of the algorithms is thus guaranteed, together with more flexibility in the choice of the clustering method.We make both the code and the dataset publicly available, to both guarantee the repeatability of the experiments, and possibly offering a new benchmark dataset. Azadeh Faroughi, Reza Javidan, Marco Mellia, Andrea Morichetta 0002, Francesca Soro, Martino Trevisan |
IEEE BigData | 6 |
| 2018 | Five years at the edge: watching internet from the ISP networkabstractThe Internet and the way people use it are constantly changing. Knowing traffic is crucial for operating the network, understanding users' need, and ultimately improving applications. Here, we provide an in-depth longitudinal view of Internet traffic in the last 5 years (from 2013 to 2017). We take the point of the view of a national-wide ISP and analyze flow-level rich measurements to pinpoint and quantify trends. We evaluate the providers' costs in terms of traffic consumption by users and services. We show that an ordinary broadband subscriber nowadays downloads more than twice as much as they used to do 5 years ago. Bandwidth hungry video services drive this change, while social messaging applications boom (and vanish) at incredible pace. We study how protocols and service infrastructures evolve over time, highlighting unpredictable events that may hamper traffic management policies. In the rush to bring servers closer and closer to users, we witness the birth of the sub-millisecond Internet, with caches located directly at ISP edges. The picture we take shows a lively Internet that always evolves and suddenly changes. Martino Trevisan, Danilo Giordano, Idilio Drago, Marco Mellia, Maurizio M. Munafò |
CoNEXT | 1 |
| 2018 | AWESoME: Big Data for Automatic Web Service Management in SDNabstractSoftware defined network (SDN) has enabled consistent and programmable management in computer networks. However, the explosion of cloud services and content delivery networks (CDNs)-coupled with the momentum of encryption-challenges the simple per-flow management and calls for a more comprehensive approach for managing Web traffic. We propose a new approach based on a “per service” management concept, which allows to identify and prioritize all traffic of important Web services, while segregating others, even if they are running on the same cloud platform, or served by the same CDN. We design and evaluate AWESoME, automatic Web service manager, a novel SDN application to address the above problem. On the one hand, it leverages big data algorithms to automatically build models describing the traffic of thousands of Web services. On the other hand, it uses the models to install rules in SDN switches to steer all flows related to the originating services. Using traffic traces from volunteers and operational networks, we provide extensive experimental results to show that AWESoME associates flows to the corresponding Web service in real-time and with high accuracy. AWESoME introduces a negligible load on the SDN controller and installs a limited number of rules on switches, hence scaling well in realistic deployments. Finally, for easy reproducibility, we release ground truth traces and scripts implementing AWESoME core components. Martino Trevisan, Idilio Drago, Marco Mellia, Han Hee Song, Mario Baldi |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2017 | Automatic detection of DNS manipulationsabstractThe DNS is a fundamental service that has been repeatedly attacked and abused. DNS manipulation is a prominent case: Recursive DNS resolvers are deployed to explicitly return manipulated answers to users' queries. While DNS manipulation is used for legitimate reasons too (e.g., parental control), rogue DNS resolvers support malicious activities, such as malware and viruses, exposing users to phishing and content injection. We introduce REMeDy, a system that assists operators to identify the use of rogue DNS resolvers in their networks. REMeDy is a completely automatic and parameter-free system that evaluates the consistency of responses across the resolvers active in the network. It operates by passively analyzing DNS traffic and, as such, requires no active probing of third-party servers. REMeDy is able to detect resolvers that manipulate answers, including resolvers that affect unpopular domains. We validate REMeDy using large-scale DNS traces collected in ISP networks where more than 100 resolvers are regularly used by customers. REMeDy automatically identifies regular resolvers, and pinpoint manipulated responses. Among those, we identify both legitimate services that offer additional protection to clients, and resolvers under the control of malwares that steer traffic with likely malicious goals. Martino Trevisan, Idilio Drago, Marco Mellia, Maurizio M. Munafò |
IEEE BigData | 1 |
| 2017 | Re-Designing Dynamic Content Delivery in the Light of a Virtualized InfrastructureabstractWe explore the opportunities and design options enabled by novel SDN and NFV technologies, by re-designing a dynamic content delivery network (CDN) service. Our system, named MOSTO, provides performance levels comparable to that of a regular CDN, but does not require the deployment of a large distributed infrastructure. In the process of designing the system, we identify relevant functions that could be integrated in the future Internet infrastructure. Such functions greatly simplify the design and effectiveness of services, such as MOSTO. We demonstrate our system using a mixture of simulation, emulation, testbed experiments, and by realizing a proof-of-concept deployment in a planet-wide commercial cloud system. Giuseppe Siracusano, Roberto Bifulco, Martino Trevisan, Tobias Jacobs, Simon Kuenzer, Stefano Salsano, Nicola Blefari-Melazzi, Felipe Huici |
IEEE J. Sel. Areas Commun. | 3 |
| 2016 | WHAT: A big data approach for accounting of modern web servicesabstractHTTP(S) has become the main means to access the Internet. The web is a tangle, with (i) multiple services and applications co-located on the same infrastructure and (ii) several websites, services and applications embedding objects from CDN, ads and tracking platforms. Traditional solutions for traffic classification and metering fall short in providing visibility in users' activities. Service providers and corporate network administrators are left with huge amounts of measurements, which cannot immediately reveal the real impact of each web service on the network. Such visibility is key to dimension the network, charge users and policy traffic. This paper introduces the Web Helper Accounting Tool (WHAT), a system to uncover the overall traffic produced by specific web services. WHAT combines big data and machine learning approaches to process large volumes of network flow measurements and learn how to group traffic due to pre-defined services of interest. Our evaluation demonstrates WHAT effectiveness in enabling accurate accounting of the traffic associated to each service. WHAT illustrates the power of machine learning when applied to large datasets of network measurements, and allows network administrators to regain the lost visibility on network usage. Martino Trevisan, Idilio Drago, Marco Mellia, Han Hee Song, Mario Baldi |
IEEE BigData | 1 |
| 2016 | Towards web service classification using addresses and DNSabstractThe identification of the services that generate traffic is crucial for ISPs and companies to plan and monitor the network. The widespread deployment of encryption and the convergence of the web services towards HTTP/HTTPS challenge traditional classification techniques. Algorithms to classify traffic are left with little information, such as server IP addresses, flow characteristics and queries performed at the DNS. Moreover, due to the usage of Content Delivery Networks and cloud infrastructure, it is unclear whether such coarse metadata is sufficient to differentiate the traffic. This paper studies to what extent basic information visible at flow-level measurements is useful for traffic classification on the web. By analyzing a large dataset of flow measurements, we quantify how often the same server IP address is used by different services, and how services use hostnames. Our results show that a very simple classifier that relies only on server IP addresses and on lists of hostnames can distinguish up to 55% of the traffic volume. Yet, collisions of names and addresses are common among popular services, calling for more ingenuity. This paper is a preliminary step in the evaluation of classification algorithms that are suitable for the modern Internet, where only minimal metadata collection will be possible in the network. Martino Trevisan, Idilio Drago, Marco Mellia, Maurizio M. Munafò |
IWCMC | 1 |