Idilio Drago

dblp:08/4676 · DBLP profile ↗
← Back
38ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0003-1932-1261ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 22 · 2 first-author · 10 since 2021Systems, architecture and hardware · 4 · 1 since 2021Security and privacy · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3Artificial intelligence and machine learning · 2Applied, interdisciplinary, general and emerging computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Building realistic cybersecurity datasets through a edge cloud testbed with cost-aware monitoring
abstract
This paper presents a distributed and autonomous experimental infrastructure designed to support the execution of cybersecurity experiments and the generation of flow and cloud monitoring datasets. The proposed system enables the execution of diverse and reproducible security experiments across geographically separated institutions with distinct physical and logical infrastructures. The infrastructure integrates real applications and networks to emulate both benign and malicious traffic, supporting the generation of flow-based and cloud-level datasets under varied monitoring configurations. Through collaborative deployment at two universities in Brazil, the proposed testbed shows its adaptability and scalability across multiple environments. The experimental results demonstrate that monitoring intervals ranging from 5 to 10 s achieve an effective balance between the detection performance of machine learning models for malicious activities in cloud services and the operational costs associated with network and cloud monitoring, maintaining high classification accuracy across diverse attack types. The generated datasets provide a consistent basis for evaluating monitoring strategies and developing data-driven detection models in cloud-native environments.
Willen Borges Coelho, Giovanni Comarela, Rodolfo V. Valentim, Idilio Drago, Rodolfo da Silva Villaça
Comput. Networks4
2026 FedScope - Federated Host Embeddings From Telescope Traffic: Design and Implementation
abstract
network telescope is a range of IP addresses that host no services. Millions of bots and scanners contact it to look for vulnerable systems, and the traffic it exposes is fundamental to understanding malicious activities. The visibility a telescope offers depends on its size and geolocation, and merging the information from multiple telescopes could help increase visibility and uncover more malicious activities. However, sharing raw telescope data is complicated, calling for solutions that allow one to directly share the knowledge rather than the data obtained from multiple deployments. In this paper, we explore the application of Federated Learning (FL) to create and share such global knowledge from the malicious activities seen in distributed telescopes. For that, we introduce FedScope, an FL-based solution for generatinghost embeddingsin a distributed way. We compare FedScope to local and distributed alternatives in downstream tasks, such as sender classification or coordinated activities detection. We show that FedScope (i) produces embeddings of equal or higher quality than those of a single telescope; (ii) increases coverage, allowing the global model to monitor more malicious actors; (iii) avoids the sharing of the raw data, limiting exchanged data.
Andrea Sordello, Rodolfo V. Valentim, Luca Vassio, Idilio Drago, Marco Mellia
IEEE Trans. Netw. Serv. Manag.5
2025 Energy-Efficient DNNs on FPGAs for Edge-Cloud Computer Vision
abstract
This research addresses the energy-efficient deployment of deep neural networks (DNNs) on FPGAs for edge-cloud computer vision applications. We evaluate and compare classic convolutional neural networks, Vision Transformers (ViTs), and alternative approaches that may be more suitable for FPGA deployment, such as differentiable boolean logic (DiffLogic) networks. Our methodology employs hardware-aware quantization techniques and specialized deployment frameworks to optimize performance on resource-constrained systems. The expected contributions include: i) energy characterization of DNN accelerators, ii) comparative analysis of FPGA toolchains, iii) optimization of modern networks with a focus on attention mechanisms, and iv) evaluation of DiffLogic networks as an alternative architecture. We demonstrate our approach through both toy cases (e.g., MNIST dataset) and a realistic Driver Distraction Detection (DDD) case study that requires real-time execution on edge devices. Preliminary results show that our optimized model achieves 66 % test accuracy on the DDD problem with very low power consumption on a PYNQ-Z2 board.
Qaisar Farooq, Idilio Drago
FPL2
2024 LogPrécis: Unleashing language models for automated malicious log analysis
abstract
Security logs are the key to understanding attacks and diagnosing vulnerabilities. Often coming in the form of text logs, their analysis remains a daunting challenge. Language Models (LMs) have demonstrated unmatched potential in understanding natural and programming languages. The question arises as to whether and how LMs could be also used to automatise the analysis of security logs. We here systematically study how to benefit from the state-of-the-art LM to support the analysis of text-like Unix shell attack logs automatically. For this, we thoroughly designed LogPrécis. LogPrécis receives as input malicious shell sessions. It then automatically identifies and assigns the attacker tactic to each portion of the session, i.e., unveiling the sequence of the attacker's goals. This creates a unique attack fingerprint. We demonstrate LogPrécis capability to support the analysis of two large datasets containing about 400,000 unique Unix shell attacks recorded in a 2-year-long honeypot deployment. LogPrécis reduces the analysis to about 3,000 unique fingerprints. Such abstraction lets us better understand attacks, extract attack prototypes, detect novelties, and track families and mutations. Overall, LogPrécis, released as open source, demonstrates the potential of adopting LMs for security analysis and paves the way for better and more responsive defence against cyberattacks.
Matteo Boffa, Idilio Drago, Marco Mellia, Luca Vassio, Danilo Giordano, Rodolfo V. Valentim, Zied Ben-Houidi
Comput. Secur.2
2024 X-squatter: AI Multilingual Generation of Cross-Language Sound-squatting
abstract
Sound-squatting is a squatting technique that exploits similarities in word pronunciation to trick users into accessing malicious resources. It is an understudied threat that has gained traction with the popularity of smart speakers and audio-only content, such as podcasts. The picture gets even more complex when multiple languages are involved. We here introduce X-squatter, a multi- and cross-language AI-based system that relies on a Transformer Neural Network for generating high-quality sound-squatting candidates. We illustrate the use of X-squatter by searching for domain name squatting abuse across hundreds of millions of issued TLS certificates, alongside other squatting types. Key findings unveil that approximately 15% of generated sound-squatting candidates have associated TLS certificates, well above the prevalence of other squatting types (7%). Furthermore, we employ X-squatter to assess the potential for abuse in PyPI packages, revealing the existence of hundreds of candidates within a 3-year package history. Notably, our results suggest that the current platform checks cannot handle sound-squatting attacks, calling for better countermeasures. We believe X-squatter uncovers the usage of multilingual sound-squatting phenomena on the Internet and it is a crucial asset for proactive protection against the threat.
Rodolfo V. Valentim, Idilio Drago, Marco Mellia, Federico Cerutti 0001
ACM Trans. Priv. Secur.2
2024 Cross-Network Embeddings Transfer for Traffic Analysis
abstract
Artificial Intelligence (AI) approaches have emerged as powerful tools to improve traffic analysis for network monitoring and management. However, the lack of large labeled datasets and the ever-changing networking scenarios make a fundamental difference compared to other domains where AI is thriving. We believe the ability to transfer the specific knowledge acquired in one network (or dataset) to a different network (or dataset) would be fundamental to speed up the adoption of AI-based solutions for traffic analysis and other networking applications (e.g., cybersecurity). We here propose and evaluate different options to transfer the knowledge built from a provider network, owning data and labels, to a customer network that desires to label its traffic but lacks labels. We formulate this problem as a domain adaptation problem that we solve with embedding alignment techniques and canonical transfer learning approaches. We present a thorough experimental analysis to assess the performance considering both supervised (e.g., classification) and unsupervised (e.g., novelty detection) downstream tasks related to darknet and honeypot traffic. Our experiments show the proper transfer techniques to use the models obtained from a network in a different network. We believe our contribution opens new opportunities and business models where network providers can successfully share their knowledge and AI models with customers.
Luca Gioacchini, Marco Mellia, Luca Vassio, Idilio Drago, Giulia Milan, Zied Ben-Houidi, Dario Rossi 0001
IEEE Trans. Netw. Serv. Manag.4
2023 Measuring the Performance of iCloud Private Relay
Martino Trevisan, Idilio Drago, Paul Schmitt, Francesco Bronzino
PAM2
2023 Enlightening the Darknets: Augmenting Darknet Visibility With Active Probes
abstract
Darknets collect unsolicited traffic reaching unused address spaces. They provide insights into malicious activities, such as the rise of botnets and DDoS attacks. However, darknets provide a shallow view, as traffic is never responded. Here we quantify how their visibility increases by responding to traffic with interactive responders with increasing levels of interaction. We consider four deployments: Darknets, simple, vertical bound to specific ports, and, a honeypot that responds to all protocols on any port. We contrast these alternatives by analyzing the traffic attracted by each deployment and characterizing how traffic changes throughout the responder lifecycle on the darknet. We show that the deployment of responders increases the value of darknet data by revealing patterns that would otherwise be unobservable. We measure Side-Scan phenomena where once a host starts responding, it attracts traffic to other ports and neighboring addresses. uncovers attacks that darknets and would not observe, e.g. large-scale activity on non-standard ports. And we observe how quickly senders can identify and attack new responders. The “enlightened” part of a darknet brings several benefits and offers opportunities to increase the visibility of sender patterns. This information gain is worth taking advantage of, and we, therefore, recommend that organizations consider this option.
Francesca Soro, Thomas Favale, Danilo Giordano, Idilio Drago, Tommaso Rescio, Marco Mellia, Zied Ben-Houidi, Dario Rossi 0001
IEEE Trans. Netw. Serv. Manag.4
2023 URLGEN - Toward Automatic URL Generation Using GANs
abstract
URLs play an essential role on the Internet, allowing access to Web resources. Automatically generating URLs is helpful in various tasks, such as application debugging, API testing, and blocklist creation for security applications. Current testing suites deeply embed experts’ domain knowledge to generate suitable URLs, resulting in an ad-hoc solution for each given application. These tools thus require heavy manual intervention, with the expensive coding of rules that are hard to maintain. We here introduce URLGEN, a system that uses Generative Adversarial Networks (GANs) to tackle the automatic URL generation problem. URLGEN is designed for web API testing and generates URL samples for an application without any system expertise, complementing the existing tools. It leverages Long Short-Term Memory (LSTM) and Convolutional Neural Network (CNN) architectures, augmented by an embedding layer that simplifies the URL learning and generation process. We show that URLGEN learns to generate new valid URLs from samples of real URLs without requiring any domain knowledge and following a purely data-driven approach. We compare the GAN architecture of URLGEN against other design options and show that the LSTM architecture can better capture the correlation among URL characters, outperforming previously proposed solutions. Finally, we show that the URLGEN approach can be extended to other scenarios, which we illustrate with two use cases, i.e., cybersquatting domain prediction and URL classification.
Rodolfo V. Valentim, Idilio Drago, Martino Trevisan, Marco Mellia
IEEE Trans. Netw. Serv. Manag.2
2023 i-DarkVec: Incremental Embeddings for Darknet Traffic Analysis
abstract
Darknets are probes listening to traffic reaching IP addresses that host no services. Traffic reaching a darknet results from the actions of internet scanners, botnets, and possibly misconfigured hosts. Such peculiar nature of the darknet traffic makes darknets a valuable instrument to discover malicious online activities, e.g., identifying coordinated actions performed by bots or scanners. However, the massive amount of packets and sources that darknets observe makes it hard to extract meaningful insights, calling for scalable tools to automatically identify and group sources that share similar behaviour. We here present i-DarkVec, a methodology to learn meaningful representations of Darknet traffic. i-DarkVec leverages Natural Language Processing techniques (e.g., Word2Vec) to capture the co-occurrence patterns that emerge when scanners or bots launch coordinated actions. As in NLP problems, the embeddings learned with i-DarkVec enable several new machine learning tasks on the darknet traffic, such as identifying clusters of senders engaged in similar activities. We extensively test i-DarkVec and explore its design space in a case study using real darknets. We show that with a proper definition of services , the learned embeddings can be used to (i) solve the classification problem to associate unknown sources’ IP addresses to the correct classes of coordinated actors and (ii) automatically identify clusters of previously unknown sources performing similar attacks and scans, easing the security analyst’s job. i-DarkVec leverages a novel incremental embedding learning approach that is scalable and robust to traffic changes, making it applicable to dynamic and large-scale scenarios.
Luca Gioacchini, Luca Vassio, Marco Mellia, Idilio Drago, Zied Ben-Houidi, Dario Rossi 0001
ACM Trans. Internet Techn.4
2023 Attacking DoH and ECH: Does Server Name Encryption Protect Users' Privacy?
abstract
Privacy on the Internet has become a priority, and several efforts have been devoted to limit the leakage of personal information. Domain names, both in the TLS Client Hello and DNS traffic, are among the last pieces of information still visible to an observer in the network. The Encrypted Client Hello extension for TLS, DNS over HTTPS or over QUIC protocols aim to further increase network confidentiality by encrypting the domain names of the visited servers. In this article, we check whether an attacker able to passively observe the traffic of users could still recover the domain name of websites they visit even if names are encrypted. By relying on large-scale network traces, we show that simplistic features and off-the-shelf machine learning models are sufficient to achieve surprisingly high precision and recall when recovering encrypted domain names. We consider three attack scenarios, i.e., recovering the per-flow name, rebuilding the set of visited websites by a user, and checking which users visit a given target website. We next evaluate the efficacy of padding-based mitigation, finding that all three attacks are still effective, despite resources wasted with padding. We conclude that current proposals for domain encryption may produce a false sense of privacy, and more robust techniques should be envisioned to offer protection to end users.
Martino Trevisan, Francesca Soro, Marco Mellia, Idilio Drago, Ricardo Morla
ACM Trans. Internet Techn.4
2022 A first look at HTTP/3 adoption and performance
Gianluca Perna, Martino Trevisan, Danilo Giordano, Idilio Drago
Comput. Commun.4
2021 DarkVec: automatic analysis of darknet traffic with word embeddings
abstract
Darknets are passive probes listening to traffic reaching IP addresses that host no services. Traffic reaching them is unsolicited by nature and often induced by scanners, malicious senders and misconfigured hosts. Its peculiar nature makes it a valuable source of information to learn about malicious activities. However, the massive amount of packets and sources that reach darknets makes it hard to extract meaningful insights. In particular, multiple senders contact the darknet while performing similar and coordinated tasks, which are often commanded by common controllers (botnets, crawlers, etc.). How to automatically identify and group those senders that share similar behaviors remains an open problem.
Luca Gioacchini, Luca Vassio, Marco Mellia, Idilio Drago, Zied Ben-Houidi, Dario Rossi 0001
CoNEXT4
2021 Modeling large-scale live video streaming client behavior
Thiago A. Guarnieri, Idilio Drago, Ítalo S. Cunha, Breno Almeida, Jussara M. Almeida, Alex Borges Vieira
Multim. Syst.2
2021 α-MON: Traffic Anonymizer for Passive Monitoring
abstract
Packet measurements at scale are essential for several applications, such as cyber-security, accounting and troubleshooting. They, however, threaten users’ privacy by exposing sensitive information. Anonymization has been the answer to this challenge, i.e., replacing sensitive information with obfuscated copies. Anonymization of packet traces, however, comes with some challenges and drawbacks. First, it reduces the value of data. Second, it requires to consider diverse protocols because information may leak from many non-encrypted fields. Third, it must be performed at high speeds directly at the monitor, to prevent private data from leaking, calling for real-time solutions. We present$\alpha $-MON, a flexible tool for privacy-preserving packet monitoring. It replicates input packet streams to different consumers while anonymizing protocol fields according to flexible policies that cover all protocol layers. Beside classic anonymization mechanisms such as IP address obfuscation,$\alpha $-MON supports${z}$-anonymization, a novel solution to obfuscate rare values that can be uniquely traced back to limited sets of users. Differently from classic anonymization approaches,z-anonymityworks on a streaming fashion, with zero delay, operating at high-speed links on a packet-by-packet basis. We quantify the impact ofz-anonymityon traffic measurements, finding that it introduces minimal error when it comes to finding heavy-hitter services. We evaluate$\alpha $-MON performance using packet traces collected from an ISP network and show that it achieves a sustainable rate of 40 Gbit/s on a Commercial Off-the Shelf server.$\alpha $-MON is available to the community as an open-source project.
Thomas Favale, Martino Trevisan, Idilio Drago, Marco Mellia
IEEE Trans. Netw. Serv. Manag.3
2020 Campus traffic and e-Learning during COVID-19 pandemic
Thomas Favale, Francesca Soro, Martino Trevisan, Idilio Drago, Marco Mellia
Comput. Networks4
2020 Five Years at the Edge: Watching Internet From the ISP Network
abstract
The Internet and the way people use it are constantly changing. Knowing traffic is crucial for operating the network, understanding users' needs, and ultimately improving applications. Here, we provide an in-depth longitudinal view of Internet traffic during 5 years (from 2013 to 2017). We take the point of the view of a national-wide ISP and analyze rich flow-level measurements to pinpoint and quantify changes. We observe the traffic, both from a point of view of users and services. We show that an ordinary broadband subscriber downloaded in 2017 more than twice as much as they used to do 5 years before. Bandwidth hungry video services drove this change at the beginning, while recently social messaging applications contribute to increase of data consumption. We study how protocols and service infrastructures evolve over time, highlighting events that may challenge traffic management policies. In the rush to bring servers closer and closer to users, we witness the birth of the sub-millisecond Internet, with caches located directly at ISP edges. The picture we take shows a lively Internet that always evolves and suddenly changes. To support new analyses, we make anonymized data available at https://smartdata.polito.it/five-years-at-the-edge/.
Martino Trevisan, Danilo Giordano, Idilio Drago, Maurizio M. Munafò, Marco Mellia
IEEE/ACM Trans. Netw.3
2019 Are Darknets All The Same? On Darknet Visibility for Security Monitoring
abstract
Darknets are sets of IP addresses that are advertised but do not host any client or server. By passively recording the incoming packets, they assist network monitoring activities. Since packets they receive are unsolicited by definition, darknets help to spot misconfigurations as well as important security events, such as the appearance and spread of botnets, DDoS attacks using spoofed IP address, etc. A number of organizations worldwide deploys darknets, ranging from a few dozens of IP addresses to large /8 networks. We here investigate how similar is the visibility of different darknets. By relying on traffic from three darknets deployed in different contintents, we evaluate their exposure in terms of observed events given their allocated IP addresses. The latter is particularly relevant considering the shortage of IPv4 addresses on the Internet. Our results suggest that some well-known facts about darknet visibility seem invariant across deployments, such as the most commonly contacted ports. However, size and location matter. We find significant differences in the observed traffic from darknets deployed in different IP ranges as well as according to the size of the IP range allocated for the monitoring.
Francesca Soro, Idilio Drago, Martino Trevisan, Marco Mellia, João M. Ceron, José Jair Santanna
LANMAN2
2019 Improving Performance of QUIC in WiFi
abstract
QUIC is a new transport protocol under standardization since 2016. Initially developed by Google as an experiment, the protocol is already deployed in large-scale, thanks to its support in Chromium and Google's servers. In this paper we experimentally analyze the performance of QUIC in WiFi networks. We perform experiments using both a controlled WiFi testbed and a production WiFi mesh network. In particular, we study how QUIC interplays with MAC layer features such as IEEE 802.11 frame aggregation. We show that the current implementation of QUIC in Chromium achieves sub-optimal throughput in wireless networks. Indeed, burstiness in modern WiFi standards may improve network performance, and we show that a Bursty QUIC (BQUIC), i.e., a customized version of QUIC that is targeted to increase its burstiness, can achieve better performance in WiFi. BQUIC outperforms the current version of QUIC in WiFi, with throughput gains ranging between 20% to 30%.
Jawad Manzoor, Llorenç Cerdà-Alabern, Ramin Sadre, Idilio Drago
WCNC4
2019 PAIN: A Passive Web performance indicator for ISPs
Martino Trevisan, Idilio Drago, Marco Mellia
Comput. Networks2
2019 A Survey on Big Data for Network Traffic Monitoring and Analysis
abstract
Network Traffic Monitoring and Analysis (NTMA) represents a key component for network management, especially to guarantee the correct operation of large-scale networks such as the Internet. As the complexity of Internet services and the volume of traffic continue to increase, it becomes difficult to design scalable NTMA applications. Applications such as traffic classification and policing require real-time and scalable approaches. Anomaly detection and security mechanisms require to quickly identify and react to unpredictable events while processing millions of heterogeneous events. At last, the system has to collect, store, and process massive sets of historical data for post-mortem analysis. Those are precisely the challenges faced by general big data approaches: Volume, Velocity, Variety, and Veracity. This survey brings together NTMA and big data. We catalog previous work on NTMA that adopt big data approaches to understand to what extent the potential of big data is being explored in NTMA. This survey mainly focuses on approaches and technologies to manage the big NTMA data, additionally briefly discussing big data analytics (e.g., machine learning) for the sake of NTMA. Finally, we provide guidelines for future work, discussing lessons learned, and research directions.
Alessandro D'Alconzo, Idilio Drago, Andrea Morichetta 0002, Marco Mellia, Pedro Casas
IEEE Trans. Netw. Serv. Manag.2
2018 HPC4AI: an AI-on-demand federated platform endeavour
abstract
In April 2018, under the auspices of the POR-FESR 2014-2020 program of Italian Piedmont Region, the Turin's Centre on High-Performance Computing for Artificial Intelligence (HPC4AI) was funded with a capital investment of 4.5M€ and it began its deployment. HPC4AI aims to facilitate scientific research and engineering in the areas of Artificial Intelligence and Big Data Analytics. HPC4AI will specifically focus on methods for the on-demand provisioning of AI and BDA Cloud services to the regional and national industrial community, which includes the large regional ecosystem of Small-Medium Enterprises (SMEs) active in many different sectors such as automotive, aerospace, mechatronics, manufacturing, health and agrifood.
Marco Aldinucci, Sergio Rabellino, Marco Pironti, Filippo Spiga, Paolo Viviani 0001, Maurizio Drocco, Marco Guerzoni, Guido Boella, Marco Mellia, Paolo Margara, Idilio Drago, Roberto Marturano, Guido Marchetto, Elio Piccolo, Stefano Bagnasco, Stefano Lusso, Sara Vallero, Giuseppe Attardi, Alex Barchiesi, Alberto Colla, Fulvio Galeazzi
CF11
2018 Five years at the edge: watching internet from the ISP network
abstract
The Internet and the way people use it are constantly changing. Knowing traffic is crucial for operating the network, understanding users' need, and ultimately improving applications. Here, we provide an in-depth longitudinal view of Internet traffic in the last 5 years (from 2013 to 2017). We take the point of the view of a national-wide ISP and analyze flow-level rich measurements to pinpoint and quantify trends. We evaluate the providers' costs in terms of traffic consumption by users and services. We show that an ordinary broadband subscriber nowadays downloads more than twice as much as they used to do 5 years ago. Bandwidth hungry video services drive this change, while social messaging applications boom (and vanish) at incredible pace. We study how protocols and service infrastructures evolve over time, highlighting unpredictable events that may hamper traffic management policies. In the rush to bring servers closer and closer to users, we witness the birth of the sub-millisecond Internet, with caches located directly at ISP edges. The picture we take shows a lively Internet that always evolves and suddenly changes.
Martino Trevisan, Danilo Giordano, Idilio Drago, Marco Mellia, Maurizio M. Munafò
CoNEXT3
2018 AWESoME: Big Data for Automatic Web Service Management in SDN
abstract
Software defined network (SDN) has enabled consistent and programmable management in computer networks. However, the explosion of cloud services and content delivery networks (CDNs)-coupled with the momentum of encryption-challenges the simple per-flow management and calls for a more comprehensive approach for managing Web traffic. We propose a new approach based on a “per service” management concept, which allows to identify and prioritize all traffic of important Web services, while segregating others, even if they are running on the same cloud platform, or served by the same CDN. We design and evaluate AWESoME, automatic Web service manager, a novel SDN application to address the above problem. On the one hand, it leverages big data algorithms to automatically build models describing the traffic of thousands of Web services. On the other hand, it uses the models to install rules in SDN switches to steer all flows related to the originating services. Using traffic traces from volunteers and operational networks, we provide extensive experimental results to show that AWESoME associates flows to the corresponding Web service in real-time and with high accuracy. AWESoME introduces a negligible load on the SDN controller and installs a limited number of rules on switches, hence scaling well in realistic deployments. Finally, for easy reproducibility, we release ground truth traces and scripts implementing AWESoME core components.
Martino Trevisan, Idilio Drago, Marco Mellia, Han Hee Song, Mario Baldi
IEEE Trans. Netw. Serv. Manag.2
2018 You, the Web, and Your Device: Longitudinal Characterization of Browsing Habits
abstract
Understanding how people interact with the web is key for a variety of applications, e.g., from the design of effective web pages to the definition of successful online marketing campaigns. Browsing behavior has been traditionally represented and studied by means of clickstreams , i.e., graphs whose vertices are web pages, and edges are the paths followed by users. Obtaining large and representative data to extract clickstreams is, however, challenging. The evolution of the web questions whether browsing behavior is changing and, by consequence, whether properties of clickstreams are changing. This article presents a longitudinal study of clickstreams from 2013 to 2016. We evaluate an anonymized dataset of HTTP traces captured in a large ISP, where thousands of households are connected. We first propose a methodology to identify actual URLs requested by users from the massive set of requests automatically fired by browsers when rendering web pages. Then, we characterize web usage patterns and clickstreams, taking into account both the temporal evolution and the impact of the device used to explore the web. Our analyses precisely quantify various aspects of clickstreams and uncover interesting patterns, such as the typical short paths followed by people while navigating the web, the fast increasing trend in browsing from mobile devices, and the different roles of search engines and social networks in promoting content. Finally, we contribute a dataset of anonymized clickstreams to the community to foster new studies.1
Luca Vassio, Idilio Drago, Marco Mellia, Zied Ben-Houidi, Mohamed Lamine Lamali
ACM Trans. Web2
2017 Automatic detection of DNS manipulations
abstract
The DNS is a fundamental service that has been repeatedly attacked and abused. DNS manipulation is a prominent case: Recursive DNS resolvers are deployed to explicitly return manipulated answers to users' queries. While DNS manipulation is used for legitimate reasons too (e.g., parental control), rogue DNS resolvers support malicious activities, such as malware and viruses, exposing users to phishing and content injection. We introduce REMeDy, a system that assists operators to identify the use of rogue DNS resolvers in their networks. REMeDy is a completely automatic and parameter-free system that evaluates the consistency of responses across the resolvers active in the network. It operates by passively analyzing DNS traffic and, as such, requires no active probing of third-party servers. REMeDy is able to detect resolvers that manipulate answers, including resolvers that affect unpopular domains. We validate REMeDy using large-scale DNS traces collected in ISP networks where more than 100 resolvers are regularly used by customers. REMeDy automatically identifies regular resolvers, and pinpoint manipulated responses. Among those, we identify both legitimate services that offer additional protection to clients, and resolvers under the control of malwares that steer traffic with likely malicious goals.
Martino Trevisan, Idilio Drago, Marco Mellia, Maurizio M. Munafò
IEEE BigData2
2017 Characterizing QoE in Large-Scale Live Streaming
abstract
Understanding the impact of performance degradation on users' QoE during live Internet streaming is key to maximize the audience and increase content providers' revenues. It is known that some problems have a strong correlation with low QoE--e.g., users experiencing video stalls tend to leave video sessions earlier. It is, however, mostly unknown whether such observations hold for live streaming of large-scale events (e.g., the FIFA World Cup). Such events are particular due to the widespread interest in the streamed content, reaching an impressively high audience worldwide. We study whether and to what extent performance degradation during live streaming of large-scale events affects users' QoE. We leverage a unique dataset collected from a major content provider in South America during the 2014 FIFA Soccer World Cup. We first extract performance metrics from the logs: stream bitrate, bitrate switches, playback stalls, and playback startup latency. We then correlate these performance metrics with session duration, which we use as a QoE indicator. We confirm the strong correlations between the metrics and QoE indicators; in particular, frequent stalls are often accompanied by higher probability of early session termination. Moreover, we quantify how such correlations vary according to broadcast matches and client terminals. Some of our findings challenge intuition--e.g., we find that PC users seem more tolerant to problems than users on mobile terminals. Our results provide better understanding of user QoE and are an important step towards user QoE models in large-scale events.
Thiago A. Guarnieri, Idilio Drago, Alex Borges Vieira, Ítalo S. Cunha, Jussara M. Almeida
GLOBECOM2
2017 Cost-Benefit Tradeoffs of Content Sharing in Personal Cloud Storage
abstract
Personal Cloud Storage (PCS) is a very popular Internet service. It allows users to backup data to the cloud as well as to perform collaborative work while sharing content. Notably, content sharing is a key feature for PCS users. It however comes with extra costs for service providers, as shared files must be synchronized to multiple user devices, generating more downloads from cloud servers. Despite the increasing interest in this type of service, a thorough investigation on the costs and benefits of PCS for service providers and end users has not been conducted yet. To that end, we propose a model to analyze cost-benefit tradeoffs for both parties. We develop utility functions that capture, in an abstract level, the satisfaction of the service provider and users in various scenarios. Then, we apply our model to evaluate alternative policies for content sharing in PCS. We consider two alternative policies for the current PCS sharing architecture, which count on user collaboration to reduce providers' costs. Our results show that such policies are advantageous for providers and users, leading to 39% utility improvements for both parties, while requiring low commitment of resources from participating users.
Glauber D. Gonçalves, Alex Borges Vieira, Idilio Drago, Ana Paula Couto da Silva, Jussara M. Almeida
MASCOTS3
2017 Personal Cloud Storage Benchmarks and Comparison
abstract
The large amount of space offered by personal cloud storage services (e.g., Dropbox and OneDrive), together with the possibility of synchronizing devices seamlessly, keep attracting customers to the cloud. Despite the high public interest, little information about system design and actual implications on performance is available when selecting a cloud storage service. Systematic benchmarks to assist in comparing services and understanding the effects of design choices are still lacking. This paper proposes a methodology to understand and benchmark personal cloud storage services. Our methodology unveils their architecture and capabilities. Moreover, by means of repeatable and customizable tests, it allows the measurement of performance metrics under different workloads. The effectiveness of the methodology is shown in a case study in which 11 services are compared under the same conditions. Our case study reveals interesting differences in design choices. Their implications are assessed in a series of benchmarks. Results show no clear winner, with all services having potential for improving performance. In some scenarios, the synchronization of the same files can take 20 times longer. In other cases, we observe a wastage of twice as much network capacity, questioning the design of some services. Our methodology and results are thus useful both as benchmarks and as guidelines for system design.
Enrico Bocchi, Idilio Drago, Marco Mellia
IEEE Trans. Cloud Comput.2
2016 WHAT: A big data approach for accounting of modern web services
abstract
HTTP(S) has become the main means to access the Internet. The web is a tangle, with (i) multiple services and applications co-located on the same infrastructure and (ii) several websites, services and applications embedding objects from CDN, ads and tracking platforms. Traditional solutions for traffic classification and metering fall short in providing visibility in users' activities. Service providers and corporate network administrators are left with huge amounts of measurements, which cannot immediately reveal the real impact of each web service on the network. Such visibility is key to dimension the network, charge users and policy traffic. This paper introduces the Web Helper Accounting Tool (WHAT), a system to uncover the overall traffic produced by specific web services. WHAT combines big data and machine learning approaches to process large volumes of network flow measurements and learn how to group traffic due to pre-defined services of interest. Our evaluation demonstrates WHAT effectiveness in enabling accurate accounting of the traffic associated to each service. WHAT illustrates the power of machine learning when applied to large datasets of network measurements, and allows network administrators to regain the lost visibility on network usage.
Martino Trevisan, Idilio Drago, Marco Mellia, Han Hee Song, Mario Baldi
IEEE BigData2
2016 The curious case of parallel connections in HTTP/2
abstract
Web pages and web-based services are becoming more and more complex. The average page size for the Alexa top 1000 websites in 2016 has reached 2.1 MB and fetching a page requires requests for 128 different objects. Although the bandwidth has been increasing exponentially in the last few years, the web experience is not improving at the same pace because of latency issues in HTTP/1. The HTTP/2 protocol aims to solve these issues by allowing clients and servers to multiplex HTTP requests and responses on a single TCP connection. If HTTP/2 is widely adopted, it can have enormous benefits not only for the user experience, but also for the servers and the network. Since clients do not have to open multiple parallel connections to avoid the problem of head-of-line blocking in HTTP/1.1, the number of concurrent TCP sessions can be significantly reduced. However, although multiplexing is one of the main features of HTTP/2, nothing actually prevents a client from opening multiple HTTP/2 connections to a server. In this paper we investigate the behavior of HTTP/2 traffic in the wild. We perform experiments to examine if web browsers use a single connection per domain over HTTP/2 in practice. Contrary to popular belief, our experiments on the traffic of a large university campus network and a residential network show that a significant number of HTTP/2 accesses are performed using parallel connections to a single domain on a server. We present two possible hypotheses for this behavior and discuss its implications for the future of the web.
Jawad Manzoor, Idilio Drago, Ramin Sadre
CNSM2
2016 Towards web service classification using addresses and DNS
abstract
The identification of the services that generate traffic is crucial for ISPs and companies to plan and monitor the network. The widespread deployment of encryption and the convergence of the web services towards HTTP/HTTPS challenge traditional classification techniques. Algorithms to classify traffic are left with little information, such as server IP addresses, flow characteristics and queries performed at the DNS. Moreover, due to the usage of Content Delivery Networks and cloud infrastructure, it is unclear whether such coarse metadata is sufficient to differentiate the traffic. This paper studies to what extent basic information visible at flow-level measurements is useful for traffic classification on the web. By analyzing a large dataset of flow measurements, we quantify how often the same server IP address is used by different services, and how services use hostnames. Our results show that a very simple classifier that relies only on server IP addresses and on lists of hostnames can distinguish up to 55% of the traffic volume. Yet, collisions of names and addresses are common among popular services, calling for more ingenuity. This paper is a preliminary step in the evaluation of classification algorithms that are suitable for the modern Internet, where only minimal metadata collection will be possible in the network.
Martino Trevisan, Idilio Drago, Marco Mellia, Maurizio M. Munafò
IWCMC2
2016 Detecting user actions from HTTP traces: Toward an automatic approach
abstract
Detecting explicit user actions, i.e., requests for web pages such as hyper-link clicks, from passive traces is fundamental for many applications, such as network forensics or content popularity estimation. Every URL explicitly visited by a user usually triggers further automatic URL requests to obtain all objects that compose the web page. HTTP traces provide a summary of all URLs requested by users, but no information that could be used to separate explicit from automatic requests. Previous works have targeted this problem and ad-hoc heuristics have been proposed. Validation has been typically done using synthetic traces. This paper investigates whether an approach based solely on machine learning can successfully detect user actions from HTTP traces. A machine learning approach would come with many advantages - e.g., it minimizes manual tuning of parameters and can easily adapt to page structure changes. We build both real and synthetic traces to assess the performance and gain insights on the features that bring most advantages in classification. Our results show that machine learning reaches similar or better performance as previous heuristics. Furthermore, we show that models built with machine learning algorithms are robust, presenting consistent performance in different scenarios.
Luca Vassio, Idilio Drago, Marco Mellia
IWCMC2
2016 Workload models and performance evaluation of cloud storage services
Glauber D. Gonçalves, Idilio Drago, Alex Borges Vieira, Ana Paula Couto da Silva, Jussara M. Almeida, Marco Mellia
Comput. Networks2
2014 Modeling the Dropbox client behavior
abstract
Cloud storage systems are currently very popular, generating a large amount of traffic. Indeed, many companies offer this kind of service, including worldwide providers such as Dropbox, Microsoft and Google. These companies, as well as new providers entering the market, could greatly benefit from knowing typical workload patterns that their services have to face in order to develop more cost-effective solutions. However, despite recent analyses of typical usage patterns and possible performance bottlenecks, no previous work investigated the underlying client processes that generate workload to the system. In this context, this paper proposes a hierarchical two-layer model for representing the Dropbox client behavior. We characterize the statistical parameters of the model using passive measurements gathered in 3 different network vantage points. Our contributions can be applied to support the design of realistic synthetic workloads, thus helping in the development and evaluation of new, well-performing personal cloud storage services.
Glauber D. Gonçalves, Idilio Drago, Ana Paula Couto da Silva, Alex Borges Vieira, Jussara M. Almeida
ICC2
2013 Benchmarking personal cloud storage
abstract
Personal cloud storage services are data-intensive applications already producing a significant share of Internet traffic. Several solutions offered by different companies attract more and more people. However, little is known about each service capabilities, architecture and -- most of all -- performance implications of design choices. This paper presents a methodology to study cloud storage services. We apply our methodology to compare 5 popular offers, revealing different system architectures and capabilities. The implications on performance of different designs are assessed executing a series of benchmarks. Our results show no clear winner, with all services suffering from some limitations or having potential for improvement. In some scenarios, the upload of the same file set can take seven times more, wasting twice as much capacity. Our methodology and results are useful thus as both benchmark and guideline for system design.
Idilio Drago, Enrico Bocchi, Marco Mellia, Herman Slatman, Aiko Pras
Internet Measurement Conference1
2013 Measurement Artifacts in NetFlow Data
Rick Hofstede, Idilio Drago, Anna Sperotto, Ramin Sadre, Aiko Pras
PAM2
2012 Inside dropbox: understanding personal cloud storage services
abstract
Personal cloud storage services are gaining popularity. With a rush of providers to enter the market and an increasing offer of cheap storage space, it is to be expected that cloud storage will soon generate a high amount of Internet traffic. Very little is known about the architecture and the performance of such systems, and the workload they have to face. This understanding is essential for designing efficient cloud storage systems and predicting their impact on the network.
Idilio Drago, Marco Mellia, Maurizio M. Munafò, Anna Sperotto, Ramin Sadre, Aiko Pras
Internet Measurement Conference1