Dario Rossi 0001

dblp:14/3340 · also Dario J. Rossi · DBLP profile ↗
← Back
112ranked-venue papers
16as first author
25since 2021 · last 2025
0000-0003-3936-8876ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 76 · 13 first-author · 16 since 2021Security and privacy · 11 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Databases, data management, data science and information retrieval · 6 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Systems, architecture and hardware · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Episodic Memories Generation and Evaluation Benchmark for Large Language Models
abstract
Episodic memory -- the ability to recall specific events grounded in time and space -- is a cornerstone of human cognition, enabling not only coherent storytelling, but also planning and decision-making. Despite their remarkable capabilities, Large Language Models (LLMs) lack a robust mechanism for episodic memory: we argue that integrating episodic memory capabilities into LLM is essential for advancing AI towards human-like cognition, increasing their potential to reason consistently and ground their output in real-world episodic events, hence avoiding confabulations. To address this challenge, we introduce a comprehensive framework to model and evaluate LLM episodic memory capabilities. Drawing inspiration from cognitive science, we develop a structured approach to represent episodic events, encapsulating temporal and spatial contexts, involved entities, and detailed descriptions. We synthesize a unique episodic memory benchmark, free from contamination, and release open source code and datasets to assess LLM performance across various recall and episodic reasoning tasks. Our evaluation of state-of-the-art models, including GPT-4 and Claude variants, Llama 3.1, and o1-mini, reveals that even the most advanced LLMs struggle with episodic memory tasks, particularly when dealing with multiple related events or complex spatio-temporal relationships -- even in contexts as short as 10k-100k tokens.
Alexis Huet, Zied Ben-Houidi, Dario Rossi 0001
ICLR3
2025 Localizing and Exploiting Concept Areas in LLMs for Downstream Classification Tasks
abstract
Localizing knowledge within Large Language Models (LLMs) is crucial for interpreting their mechanisms and outcomes. Whereas knowledge attribution has so far provided local sample-level explanations, in this work we argue that whenever LLMs are used for classification tasks, a class-level explanation is preferable. We therefore define broader concept areas, i.e., regions of the LLM comprising a small set of neurons that contains the most salient knowledge pertaining to each class and propose methods to identify such areas. We apply our methodology to BERT-based LLMs fine-tuned for downstream classification tasks such as sentiment analysis and attack classification: our results show that it is possible to (i) identify crucial sets of neurons that determine the behaviour of fine-tuned LLMs for explanation purposes, as well as (ii) exploit such concept areas to improve their classification outcomes-yielding up to 6% macro F1-Score improvement on sentiment analysis (public dataset) and 2% on attack classification (private dataset) without requiring further fine-tuning.
Alfredo Nascita, Jonatan Krolikowski, Valerio Persico, Antonio Pescapè, Dario Rossi 0001
IJCNN5
2025 Changepoint Detection via Subset Chains
Alexis Huet, José Manuel Navarro, Dario Rossi 0001
PAKDD (4)3
2024 Data Augmentation for Traffic Classification
Chao Wang 0103, Alessandro Finamore, Pietro Michiardi, Massimo Gallo, Dario Rossi 0001
PAM (1)5
2024 MEMENTO: A novel approach for class incremental learning of encrypted traffic
abstract
In the ever-changing digital environment, ensuring the ongoing effectiveness of traffic analysis and security measures is crucial. Therefore, Class Incremental Learning (CIL) in encrypted Traffic Classification (TC) is essential for adapting to evolving network behaviors and the rapid development of new applications. However, the application of CIL techniques in the TC domain is not straightforward, usually leading to unsatisfactory performance figures. Specifically, the improvement goal is to reduce forgetting on old apps and increase the capacity in learning new ones, in order to improve overall classification performance— reducing the drop from a model “trained-from-scratch”. The contribution of this work is the design of a novel fine-tuning approach called MEMENTO, which is obtained through the careful design of different building blocks: memory management, model training, and rectification strategies. In detail, we propose the application of traffic biflows augmentation strategies to better capitalize on old apps biflows, we introduce improvements in the distillation stage, and we design a general rectification strategy that includes several existing proposals. To assess our proposal, we leverage two publicly-available encrypted network traffic datasets, i.e., MIRAGE19 and CESNET-TLS22. As a result, on both datasets MEMENTO achieves a significant improvement in classifying new apps (w.r.t. the best-performing alternative, i.e., BiC) while maintaining stable performance on old ones. Equally important, MEMENTO achieves satisfactory overall TC performance, filling the gap toward a trained-from-scratch model and offering a considerable gain in terms of time (up to 10× speed-up) to obtain up-to-date and running classifiers. The experimental evaluation relies on a comprehensive performance evaluation workbench for CIL proposals, which is based on a wider set of metrics (as opposed to the existing literature in TC).
Francesco Cerasuolo, Alfredo Nascita, Giampaolo Bovenzi, Giuseppe Aceto, Domenico Ciuonzo, Antonio Pescapè, Dario Rossi 0001
Comput. Networks7
2024 Benchmarking Class Incremental Learning in Deep Learning Traffic Classification
abstract
Traffic Classification (TC) is experiencing a renewed interest, fostered by the growing popularity of Deep Learning (DL) approaches. In exchange for their proved effectiveness, DL models are characterized by a computationally-intensive training procedure that badly matches the fast-paced release of new (mobile) applications, resulting in significantly limited efficiency of model updates. To address this shortcoming, in this work we systematically explore Class Incremental Learning (CIL) techniques, aimed at adding new apps/services to pre-existing DL-based traffic classifiers without a full retraining, hence speeding up the model’s updates cycle. We investigate a large corpus of state-of-the-art CIL approaches for the DL-based TC task, and delve into their working principles to highlight relevant insight, aiming to understand if there is a case for CIL in TC. We evaluate and discuss their performance varying the number of incremental learning episodes, and the number of new apps added for each episode. Our evaluation is based on the publicly available$\mathtt {MIRAGE19}$dataset comprising traffic of 40 popular Android applications, fostering reproducibility. Despite our analysis reveals their infancy, CIL techniques are a promising research area on the roadmap towards automated DL-based traffic analysis systems.
Giampaolo Bovenzi, Alfredo Nascita, Lixuan Yang, Alessandro Finamore, Giuseppe Aceto, Domenico Ciuonzo, Antonio Pescapè, Dario Rossi 0001
IEEE Trans. Netw. Serv. Manag.8
2024 Cross-Network Embeddings Transfer for Traffic Analysis
abstract
Artificial Intelligence (AI) approaches have emerged as powerful tools to improve traffic analysis for network monitoring and management. However, the lack of large labeled datasets and the ever-changing networking scenarios make a fundamental difference compared to other domains where AI is thriving. We believe the ability to transfer the specific knowledge acquired in one network (or dataset) to a different network (or dataset) would be fundamental to speed up the adoption of AI-based solutions for traffic analysis and other networking applications (e.g., cybersecurity). We here propose and evaluate different options to transfer the knowledge built from a provider network, owning data and labels, to a customer network that desires to label its traffic but lacks labels. We formulate this problem as a domain adaptation problem that we solve with embedding alignment techniques and canonical transfer learning approaches. We present a thorough experimental analysis to assess the performance considering both supervised (e.g., classification) and unsupervised (e.g., novelty detection) downstream tasks related to darknet and honeypot traffic. Our experiments show the proper transfer techniques to use the models obtained from a network in a different network. We believe our contribution opens new opportunities and business models where network providers can successfully share their knowledge and AI models with customers.
Luca Gioacchini, Marco Mellia, Luca Vassio, Idilio Drago, Giulia Milan, Zied Ben-Houidi, Dario Rossi 0001
IEEE Trans. Netw. Serv. Manag.7
2023 Tree-based Kendall's τ Maximization for Explainable Unsupervised Anomaly Detection
abstract
We study the problem of building a regression tree with relatively small size, which maximizes the Kendall’s tau coefficient between the anomaly scores of a source anomaly detection algorithm and those predicted by our regression tree. We consider a labeling function which assigns to each leaf the inverse of its size, thereby providing satisfactory explanations when comparing examples with different anomaly scores. We show that our approach can be used as a post-hoc model, i.e. to provide global explanations for an existing anomaly detection algorithm. Moreover, it can be used as an in-model approach, i.e. the source anomaly detection algorithm can be replaced all together. This is made possible by leveraging the off-the-shelf transparency of tree-based approaches and from the fact that the explanations provided by our approach do not rely on the source anomaly detection algorithm. The main technical challenge to tackle is the efficient computation of the Kendall’s tau coefficients when determining the best split at each node of the regression tree. We show how such a coefficient can be computed incrementally, thereby making the running time of our algorithm almost linear (up to a logarithmic factor) in the size of the input. Our approach is completely unsupervised, which is appealing in the case when it is difficult to collect a large number of labeled examples. We complement our study with an extensive experimental evaluation against the state-of-the-art, showing the effectiveness of our approach.
Lanfang Kong, Alexis Huet, Dario Rossi 0001, Mauro Sozio
ICDM3
2023 Replication: Contrastive Learning and Data Augmentation in Traffic Classification Using a Flowpic Input Representation
abstract
Over the last years we witnessed a renewed interest toward Traffic Classification (TC) captivated by the rise of Deep Learning (DL). Yet, the vast majority of TC literature lacks code artifacts, performance assessments across datasets and reference comparisons against Machine Learning (ML) methods. Among those works, a recent study from IMC'22 [16] is worth of attention since it adopts recent DL methodologies (namely, few-shot learning, self-supervision via contrastive learning and data augmentation) appealing for networking as they enable to learn from a few samples and transfer across datasets. The main result of [16] on the UCDAVIS, ISCXVPN and ISCXTOR datasets is that, with such DL methodologies, 100 input samples are enough to achieve very high accuracy using an input representation called "flowpic'' (i.e., a per-flow 2d histograms of the packets size evolution over time).
Alessandro Finamore, Chao Wang 0103, Jonatan Krolikowski, José Manuel Navarro, Fuxing Chen, Dario Rossi 0001
IMC6
2023 A Lightweight, Efficient and Explainable-by-Design Convolutional Neural Network for Internet Traffic Classification
abstract
Traffic classification, i.e., the identification of the type of applications flowing in a network, is a strategic task for numerous activities (e.g., intrusion detection, routing). This task faces some critical challenges that current deep learning approaches do not address. The design of current approaches do not take into consideration the fact that networking hardware (e.g., routers) often runs with limited computational resources. Further, they do not meet the need for faithful explainability highlighted by regulatory bodies. Finally, these traffic classifiers are evaluated on small datasets which fail to reflect the diversity of applications in real-world settings.
Kevin Fauvel, Fuxing Chen, Dario Rossi 0001
KDD3
2023 Enlightening the Darknets: Augmenting Darknet Visibility With Active Probes
abstract
Darknets collect unsolicited traffic reaching unused address spaces. They provide insights into malicious activities, such as the rise of botnets and DDoS attacks. However, darknets provide a shallow view, as traffic is never responded. Here we quantify how their visibility increases by responding to traffic with interactive responders with increasing levels of interaction. We consider four deployments: Darknets, simple, vertical bound to specific ports, and, a honeypot that responds to all protocols on any port. We contrast these alternatives by analyzing the traffic attracted by each deployment and characterizing how traffic changes throughout the responder lifecycle on the darknet. We show that the deployment of responders increases the value of darknet data by revealing patterns that would otherwise be unobservable. We measure Side-Scan phenomena where once a host starts responding, it attracts traffic to other ports and neighboring addresses. uncovers attacks that darknets and would not observe, e.g. large-scale activity on non-standard ports. And we observe how quickly senders can identify and attack new responders. The “enlightened” part of a darknet brings several benefits and offers opportunities to increase the visibility of sender patterns. This information gain is worth taking advantage of, and we, therefore, recommend that organizations consider this option.
Francesca Soro, Thomas Favale, Danilo Giordano, Idilio Drago, Tommaso Rescio, Marco Mellia, Zied Ben-Houidi, Dario Rossi 0001
IEEE Trans. Netw. Serv. Manag.8
2023 i-DarkVec: Incremental Embeddings for Darknet Traffic Analysis
abstract
Darknets are probes listening to traffic reaching IP addresses that host no services. Traffic reaching a darknet results from the actions of internet scanners, botnets, and possibly misconfigured hosts. Such peculiar nature of the darknet traffic makes darknets a valuable instrument to discover malicious online activities, e.g., identifying coordinated actions performed by bots or scanners. However, the massive amount of packets and sources that darknets observe makes it hard to extract meaningful insights, calling for scalable tools to automatically identify and group sources that share similar behaviour. We here present i-DarkVec, a methodology to learn meaningful representations of Darknet traffic. i-DarkVec leverages Natural Language Processing techniques (e.g., Word2Vec) to capture the co-occurrence patterns that emerge when scanners or bots launch coordinated actions. As in NLP problems, the embeddings learned with i-DarkVec enable several new machine learning tasks on the darknet traffic, such as identifying clusters of senders engaged in similar activities. We extensively test i-DarkVec and explore its design space in a case study using real darknets. We show that with a proper definition of services , the learned embeddings can be used to (i) solve the classification problem to associate unknown sources’ IP addresses to the correct classes of coordinated actors and (ii) automatically identify clusters of previously unknown sources performing similar attacks and scans, easing the security analyst’s job. i-DarkVec leverages a novel incremental embedding learning approach that is scalable and robust to traffic changes, making it applicable to dynamic and large-scale scenarios.
Luca Gioacchini, Luca Vassio, Marco Mellia, Idilio Drago, Zied Ben-Houidi, Dario Rossi 0001
ACM Trans. Internet Techn.6
2022 STREamRHF: Tree-Based Unsupervised Anomaly Detection for Data Streams
abstract
We present STREAMRHF, an unsupervised anomaly detection algorithm for data streams. Our algorithm builds on some of the ideas of Random Histogram Forest (RHF) [1], a state-of-the-art algorithm for batch unsupervised anomaly detection. STREAMRHF constructs a forest of decision trees, where feature splits are determined according to the kurtosis score of every feature. It irrevocably assigns an anomaly score to data points, as soon as they arrive, by means of an incremental computation of its random trees and the kurtosis scores of the features. This allows efficient online scoring and concept drift detection altogether. Our approach is tree-based which boasts several appealing properties, such as explainability of the results [2]. We conduct an extensive experimental evaluation on multiple datasets from different real-world applications. Our evaluation shows that our streaming algorithm achieves comparable average precision to RHF while outperforming state-of-the-art streaming approaches for unsupervised anomaly detection with furthermore limited computational complexity.
Stefan Nesic, Andrian Putina, Maroua Bahri, Alexis Huet, José Manuel Navarro, Dario Rossi 0001, Mauro Sozio
AICCSA6
2022 Towards a systematic multi-modal representation learning for network data
abstract
Learning the right representations from complex input data is the key ability of successful machine learning (ML) models. The latter are often tailored to a specific data modality. For example, recurrent neural networks (RNNs) were designed having sequential data in mind, while convolutional neural networks (CNNs) were designed to exploit spatial correlation in images. Unlike computer vision (CV) and natural language processing (NLP), each of which targets a single well-defined modality, network ML problems often have a mixture of data modalities as input. Yet, instead of exploiting such abundance, practitioners tend to rely on sub-features thereof, reducing the problem to single modality for the sake of simplicity. In this paper, we advocate for exploiting all the modalities naturally present in network data. As a first step, we observe that network data systematically exhibits a mixture of quantities (e.g., measurements), and entities (e.g., IP addresses, names, etc.). Whereas the former are generally well exploited, the latter are often underused or poorly represented (e.g., with one-hot encoding). We propose to systematically leverage language models to learn entity representations, whenever significant sequences of such entities are historically observed. Through two diverse use-cases, we show that such entity encoding can benefit and naturally augment classic quantity-based features.
Zied Ben-Houidi, Raphaël Azorin, Massimo Gallo, Alessandro Finamore, Dario Rossi 0001
HotNets5
2022 Accelerating Deep Learning Classification with Error-controlled Approximate-key Caching
abstract
While Deep Learning (DL) technologies are a promising tool to solve networking problems that map to classification tasks, their computational complexity is still too high with respect to real-time traffic measurements requirements. To reduce the DL inference cost, we propose a novel caching paradigm, that we named approximate-key caching, which returns approximate results for lookups of selected input based on cached DL inference results. While approximate cache hits alleviate DL inference workload and increase the system throughput, they however introduce an approximation error. As such, we couple approximate-key caching with an error-correction principled algorithm, that we named auto-refresh. We analytically model our caching system performance for classic LRU and ideal caches, we perform a trace-driven evaluation of the expected performance, and we compare the benefits of our proposed approach with the state-of-the-art similarity caching – this testifies the practical interest of our proposal.
Alessandro Finamore, James Roberts, Massimo Gallo, Dario Rossi 0001
INFOCOM4
2022 Local Evaluation of Time Series Anomaly Detection Algorithms
abstract
In recent years, specific evaluation metrics for time series anomaly detection algorithms have been developed to handle the limitations of the classical precision and recall. However, such metrics are heuristically built as an aggregate of multiple desirable aspects, introduce parameters and wipe out the interpretability of the output. In this article, we first highlight the limitations of the classical precision/recall, as well as the main issues of the recent event-based metrics -- for instance, we show that an adversary algorithm can reach high precision and recall on almost any dataset under weak assumption. To cope with the above problems, we propose a theoretically grounded, robust, parameter-free and interpretable extension to precision/recall metrics, based on the concept of "affiliation'' between the ground truth and the prediction sets. Our metrics leverage measures of duration between ground truth and predictions, and have thus an intuitive interpretation. By further comparison against random sampling, we obtain a normalized precision/recall, quantifying how much a given set of results is better than a random baseline prediction. By construction, our approach keeps the evaluation local regarding ground truth events, enabling fine-grained visualization and interpretation of algorithmic results. We compare our proposal against various public time series anomaly detection datasets, algorithms and metrics. We further derive theoretical properties of the affiliation metrics that give explicit expectations about their behavior and ensure robustness against adversary strategies.
Alexis Huet, José Manuel Navarro, Dario Rossi 0001
KDD3
2022 Human readable network troubleshooting based on anomaly detection and feature scoring
José Manuel Navarro, Alexis Huet, Dario Rossi 0001
Comput. Networks3
2022 Neural language models for network configuration: Opportunities and reality check
Zied Ben-Houidi, Dario Rossi 0001
Comput. Commun.2
2022 Landing AI on Networks: An Equipment Vendor Viewpoint on Autonomous Driving Networks
abstract
The tremendous achievements of Artificial Intelligence (AI) in computer vision, natural language processing, games and robotics, has extended the reach of the AI hype to other fields: in telecommunication networks, the long term vision is to let AI fully manage, and autonomously drive, all aspects of network operation. In this industry vision paper, we discuss challenges and opportunities of Autonomous Driving Network (ADN) driven by AI technologies. To understand how AI can be successfully landed in current and future networks, we start by outlining challenges that are specific to the networking domain, putting them in perspective with advances that AI has achieved in other fields. We then present a system view, clarifying how AI can be fitted in the network architecture. We finally discuss current achievements as well as future promises of AI in networks, mentioning a roadmap to avoid bumps in the road that leads to true large-scale deployment of AI technologies in networks.
Dario Rossi 0001
IEEE Trans. Netw. Serv. Manag.1
2021 DarkVec: automatic analysis of darknet traffic with word embeddings
abstract
Darknets are passive probes listening to traffic reaching IP addresses that host no services. Traffic reaching them is unsolicited by nature and often induced by scanners, malicious senders and misconfigured hosts. Its peculiar nature makes it a valuable source of information to learn about malicious activities. However, the massive amount of packets and sources that reach darknets makes it hard to extract meaningful insights. In particular, multiple senders contact the darknet while performing similar and coordinated tasks, which are often commanded by common controllers (botnets, crawlers, etc.). How to automatically identify and group those senders that share similar behaviors remains an open problem.
Luca Gioacchini, Luca Vassio, Marco Mellia, Idilio Drago, Zied Ben-Houidi, Dario Rossi 0001
CoNEXT6
2021 FENXI: Deep-learning Traffic Analytics at the edge
Massimo Gallo, Alessandro Finamore, Gwendal Simon, Dario Rossi 0001
SEC4
2021 Real-Time Channel Management in WLANs: Deep Reinforcement Learning versus Heuristics
abstract
Today's WLANs rely on a centralized Access Controller (AC) entity for managing distributed wireless Access Points (APs) to which user devices connect. The availability of real-time analytics at the AC opens the possibility to automate the allocation of scarce radio resources, continuously adapting to changes in traffic demands. Often, the allocation problem is formulated in terms of weighted graph coloring, which is NP-hard, and custom heuristics are used to find satisfactory solutions. In this paper, we contrast solutions that are based on (and even improve) state of the art heuristics to a data-driven solution that leverages Deep Reinforcement Learning (DRL). Based on both simulation results as well as experiments in a real deployment, we show that our DRL-based scheme not only learns to solve the complex combinatorial problem in bounded time, outperforming heuristics, but it also exhibits appealing generalization properties, e.g. to different network sizes and densities.
Ovidiu Iacoboaiea, Jonatan Krolikowski, Zied Ben-Houidi, Dario Rossi 0001
Networking4
2021 Deployable Models for Approximating Web QoE Metrics From Encrypted Traffic
abstract
Being on endpoints, Content Providers can easily evaluate end users' Web browsing quality of experience (Web QoE) by accessing in-browser computed application-level metrics. Because of end-to-end traffic encryption, it is becoming considerably harder for Internet Service Providers (ISPs) to evaluate the Web QoE of their customers, which is important for management purposes. In this paper, we propose data-driven machine learning techniques and exact flow-level algorithmic methods to infer well-known application-level Web performance metrics (such as SpeedIndex and Page Load Time) from raw encrypted streams of network traffic. We prove the efficiency of our approach taking as input a unique dataset of more than 200,000 experiments, targeting a large set of popular pages (Alexa top-500), from probes from several ISPs networks, with different browsers (Chrome, Firefox) and viewport combinations. Results show that our data-driven models are not only accurate for several Web performance metrics, but also feature the ability to generalize to previously unseen conditions. Furthermore, we discuss how our extremely lightweight flow-level method has a provable accuracy on a specific metric, and is thus of particular appeal from a deployment viewpoint.
Alexis Huet, Antoine Saverimoutou, Zied Ben-Houidi, Hao Shi 0002, Shengming Cai, Jinchun Xu, Bertrand Mathieu, Dario Rossi 0001
IEEE Trans. Netw. Serv. Manag.8
2021 Online Anomaly Detection Leveraging Stream-Based Clustering and Real-Time Telemetry
abstract
Recent technology evolution allows network equipment to continuously stream a wealth of “telemetry” information, which pertains to multiple protocols and layers of the stack, at a very fine spatial-grain and high-frequency. This deluge of telemetry data clearly offers new opportunities for network control and troubleshooting, but also poses a serious challenge for what concerns its real-time processing. We tackle this challenge by applying streaming machine-learning techniques to the continuous flow of control and data-plane telemetry data, with the purpose of real-time detection of anomalies. In particular, we implement an anomaly detection engine that leverages DenStream, an unsupervised clustering technique, and apply it to features collected from a large-scale testbed comprising tens of routers traversed up to 3Terabit/s worth of real application traffic. We contrast DenStream with offline algorithms such as DBScan and Local Outlier Factor (LOF), as well as online algorithms such as the windowed version of DBScan, ExactSTORM, Continuous Outlier Detection (COD) and Robust Random Cut Forest (RRCF). Our experimental campaign compares these seven algorithms under both accuracy and computational complexity viewpoints: results testify that DenStream (i) achieves detection results on par with RRCF, the best performing algorithm and (ii) is significantly faster than other approaches, notably over two orders of magnitude faster than RRCF. In spirit with the recent trend toward reproducibility of results, we make our code available as open source to the scientific community.
Andrian Putina, Dario Rossi 0001
IEEE Trans. Netw. Serv. Manag.2
2021 Deep Learning and Zero-Day Traffic Classification: Lessons Learned From a Commercial-Grade Dataset
abstract
The increasing success of Machine Learning (ML) and Deep Learning (DL) has recently re-sparked interest towards traffic classification. While supervised techniques provide satisfactory performance when classifyingknowntraffic, the detection ofzero-day(i.e., unknown) traffic is a more challenging task. At the same time, zero-day detection, generally tackled with unsupervised techniques such as clustering, received less coverage by the traffic classification literature which focuses more on deriving DL models via supervised techniques. Moreover, the combination of supervised and unsupervised techniques poses challenges not fully covered by the traffic classification literature. In this paper, we share our experience on a commercial-grade DL traffic classification engine that combines supervised and unsupervised techniques to identify known and zero-day traffic. In particular, we rely on a dataset withhundredsof very fine grained application labels, and perform a thorough assessment of two state of the art traffic classifiers in commercial-grade settings. This pushes the boundaries of traffic classifiers evaluation beyond thefew tensof classes typically used in the literature. Our main contribution is the design and evaluation of GradBP, a novel technique for zero-day applications detection. Based on gradient backpropagation and tailored for DL models, GradBP yields superior performance with respect to state of the art alternatives, in both accuracy and computational cost. Overall, while ML and DL models are both equally able to provide excellent performance for the classification of known traffic, the non-linear feature extraction process of DL models backbone provides sizable advantages for the detection of unknown classes over classical ML models.
Lixuan Yang, Alessandro Finamore, Jun Feng 0009, Dario Rossi 0001
IEEE Trans. Netw. Serv. Manag.4
2020 Detecting Degradation of Web Browsing Quality of Experience
abstract
Quality of Experience (QoE) inference, and particularly the detection of its degradation is an important management tool for ISPs. Yet, this task is made difficult due to widespread use of encryption on the data-plane on the one hand so that measuring QoE is hard, and to the ephemeral properties of the web content on the other hand so that changes in QoE indicators may be rooted in changes in properties of the content itself, more than being caused by network-related events. In this paper, we phrase the QoE degradation detection issue as a change point detection problem, that we tackle by leveraging a unique dataset consisting on several hundreds thousands browsing sessions spanning multiple months. Our results, beyond showing feasibility, warn about the exclusive use of QoE indicators that are very close to content, as changes in the content space can lead to false alarms that are not tied to network-related problems.
Alexis Huet, Zied Ben-Houidi, Bertrand Mathieu, Dario Rossi 0001
CNSM4
2020 Are you on Mobile or Desktop? On the Impact of End-User Device on Web QoE Inference from Encrypted Traffic
abstract
Web browsing is one of the key applications of the Internet, if not the most important one. We address the problem of Web Quality-of-Experience (QoE) monitoring from the ISP perspective, relying on in-network, passive measurements. As a proxy to Web QoE, we focus on the analysis of the well-known SpeedIndex (SI) metric. Given the lack of application-level-data visibility introduced by the wide adoption of end-to-end encryption, we resort to machine-learning models to infer the SI and the QoE level of individual web-page loading sessions, using as input only packet- and flow-level data. In this paper, we study the impact of different end-user device types (e.g., smartphone, desktop, tablet) on the performance of such models. Empirical evaluations on a large, multi-device, heterogeneous corpus of Web-QoE measurements for the most popular websites demonstrate that the proposed solution can infer the SI as well as estimate QoE ranges with high accuracy, using either packet-level or flow-level measurements. In addition, we show that the device type adds a strong bias to the feasibility of these Web-QoE models, putting into question the applicability of previously conceived approaches on single-device measurements. To improve the state of the art, we conceive cross-device generalizable models operating at both packet and flow levels, offering a feasible solution for Web-QoE monitoring in operational, multi-device networks. To the best of our knowledge, this is the first study tackling the analysis of Web QoE from encrypted network traffic in multi-device scenarios.
Sarah Wassermann, Pedro Casas, Zied Ben-Houidi, Alexis Huet, Michael Seufert, Nikolas Wehner, Joshua Schüler, Shengming Cai, Hao Shi 0002, Jinchun Xu, Tobias Hoßfeld, Dario Rossi 0001
CNSM12
2020 Random Histogram Forest for Unsupervised Anomaly Detection
abstract
Roughly speaking, anomaly detection consists of identifying instances whose features significantly deviate from the rest of input data. It is one of the most widely studied problems in unsupervised machine learning, boasting applications in network intrusion detection, healthcare and many others. Several methods have been developed in recent years, however, a satisfactory solution is still missing to the best of our knowledge. We present Random Histogram Forest an effective approach for unsupervised anomaly detection. Our approach is probabilistic, which has been proved to be effective in identifying anomalies. Moreover, it employs the fourth central moment (aka kurtosis), so as to identify potential anomalous instances. We conduct an extensive experimental evaluation on 38 datasets including all benchmarks for anomaly detection, as well as the most successful algorithms for unsupervised anomaly detection, to the best of our knowledge. We evaluate all the approaches in terms of the average precision of the area under the precision-recall curve (AP). Our evaluation shows that our approach significantly outperforms all other approaches in terms of AP while boasting linear running time.
Andrian Putina, Mauro Sozio, Dario Rossi 0001, José Manuel Navarro
ICDM3
2020 Revealing QoE of Web Users from Encrypted Network Traffic
Alexis Huet, Antoine Saverimoutou, Zied Ben-Houidi, Hao Shi 0002, Shengming Cai, Jinchun Xu, Bertrand Mathieu, Dario Rossi 0001
Networking8
2020 Analyzing Wikipedia Users' Perceived Quality of Experience: A Large-Scale Study
abstract
The Web is one of the most successful Internet applications. Yet, the quality of Web users' experience is still largely impenetrable. Whereas Web performance is typically studied with controlled experiments, in this work we perform a large-scale study of a real site, Wikipedia, explicitly asking (a small fraction of its) users for feedback on the browsing experience. The analysis of the collected feedback reveals that 85% of users are satisfied, along with both expected (e.g., the impact of browser and network connectivity) and surprising findings (e.g., absence of day/night, weekday/weekend seasonality) that we detail in this paper. Also, we leverage user responses to build supervised data-driven models to predict user satisfaction which, despite including state-of-the art quality of experience metrics, are still far from achieving accurate results (0.62 recall of negative answers). Finally, we make our dataset publicly available, hopefully contributing in enriching and refining the scientific community knowledge on Web users' QoE.
Flavia Salutari, Diego N. da Hora, Gilles Dubuc, Dario Rossi 0001
IEEE Trans. Netw. Serv. Manag.4
2019 Discrete-Time Modeling of NFV Accelerators that Exploit Batched Processing
abstract
Network Functions Virtualization (NFV) is among the latest network revolutions, bringing flexibility and avoiding network ossification. At the same time, all-software NFV implementations on commodity hardware raise performance issues with respect to ASIC solutions. To address these issues, numerous software acceleration frameworks for packet processing have appeared in the last few years. Common among these frameworks is the use of batching techniques. In this context, packets are processed in groups as opposed to individually, which is required at high-speed to minimize the framework overhead, reduce interrupt pressure, and leverage instruction-level cache hits. Whereas several system implementations have been proposed and experimentally benchmarked, the scientific community has so far only to a limited extent attempted to model the system dynamics of modern NFV routers exploiting batching acceleration. In this paper, we fill this gap by proposing a simple generic model for such batching-based mechanisms, which allows a very detailed prediction of highly relevant performance indicators. These include the distribution of the processed batch size as well as queue size, which can be used to identify loss-less operational regimes or quantify the packet loss probability in high-load scenarios. We contrast the model prediction with experimental results gathered in a high-speed testbed including an NFV router, showing that the model not only correctly captures system performance under simple conditions, but also in more realistic scenarios in which traffic is processed by a mixture of functions.
Stanislav Lange, Leonardo Linguaglossa, Stefan Geißler, Dario Rossi 0001, Thomas Zinner
INFOCOM4
2019 A Large-scale Study of Wikipedia Users' Quality of Experience
abstract
The Web is one of the most successful Internet application. Yet, the quality of Web users' experience is still largely impenetrable. Whereas Web performances are typically gathered with controlled experiments, in this work we perform a large-scale study of one of the most popular websites,namely Wikipedia, explicitly asking (a small fraction of its) users for feedback on the browsing experience. We leverage user survey responses to build a data-driven model of user satisfaction which, despite including state-of-the art quality of experience metrics, is still far from achieving accurate results, and discuss directions to move forward. Finally, we aim at making our dataset publicly available, which hopefully contributes in enriching and refining the scientific community knowledge on Web users' quality of experience (QoE).
Flavia Salutari, Diego N. da Hora, Gilles Dubuc, Dario Rossi 0001
WWW4
2019 High-speed data plane and network functions virtualization by vectorizing packet processing
Leonardo Linguaglossa, Dario Rossi 0001, Salvatore Pontarelli, David Barach, Damjan Marjon, Pierre Pfister
Comput. Networks2
2019 Survey of Performance Acceleration Techniques for Network Function Virtualization
abstract
The ongoing network softwarization trend holds the promise to revolutionize network infrastructures by making them more flexible, reconfigurable, portable, and more adaptive than ever. Still, the migration from hard-coded/hard-wired network functions toward their software-programmable counterparts comes along with the need for tailored optimizations and acceleration techniques so as to avoid or at least mitigate the throughput/latency performance degradation with respect to fixed function network elements. The contribution of this paper is twofold. First, we provide a comprehensive overview of the host-based network function virtualization (NFV) ecosystem, covering a broad range of techniques, from low-level hardware acceleration and bump-in-the-wire offloading approaches to high-level software acceleration solutions, including the virtualization technique itself. Second, we derive guidelines regarding the design, development, and operation of NFV-based deployments that meet the flexibility and scalability requirements of modern communication networks.
Leonardo Linguaglossa, Stanislav Lange, Salvatore Pontarelli, Gábor Rétvári, Dario Rossi 0001, Thomas Zinner, Roberto Bifulco, Michael Jarschel, Giuseppe Bianchi 0001
Proc. IEEE5
2019 FloWatcher-DPDK: Lightweight Line-Rate Flow-Level Monitoring in Software
abstract
In the last few years, several software-based solutions have been proved to be very efficient for high-speed packet processing, traffic generation, and monitoring, and can be considered valid alternatives to expensive and non-flexible hardware-based solutions. In this paper, we first benchmark heterogeneous design choices for software-based packet monitoring systems in terms of achievable performance and required resources (i.e., the number of CPU cores). Building on this extensive analysis we design FloWatcher-DPDK, a DPDK-based high-speed software traffic monitor we provide to the community as an open source project. In a nutshell, FloWatcher-DPDK provides tunable fine-grained statistics at packet and flow levels. Experimental results demonstrate that FloWatcher-DPDK sustains per-flow statistics with 5-nines precision at high-speed (e.g., 14.88 Mpps) using a limited amount of resources. Finally, we showcase the usage of FloWatcher-DPDK by configuring it to analyze the performance of two open source prototypes for stateful flow-level end-host and in-network packet processing.
Tianzhu Zhang 0002, Leonardo Linguaglossa, Massimo Gallo, Paolo Giaccone, Dario Rossi 0001
IEEE Trans. Netw. Serv. Manag.5
2019 TupleMerge: Fast Software Packet Processing for Online Packet Classification
abstract
Packet classification is an important part of many networking devices, such as routers and firewalls. Software-defined networking (SDN) heavily relies on online packet classification which must efficiently process two different streams: incoming packets to classify and rules to update. This rules out many offline packet classification algorithms that do not support fast updates. We propose a novel online classification algorithm, TupleMerge (TM), derived from tuple space search (TSS), the packet classifier used by Open vSwitch (OVS). TM improves upon TSS by combining hash tables which contain rules with similar characteristics. This greatly reduces classification time preserving similar performance in updates. We validate the effectiveness of TM using both simulation and deployment in a full-fledged software router, specifically within the vector packet processor (VPP). In our simulation results, which focus solely on the efficiency of the classification algorithm, we demonstrate that TM outperforms all other state of the art methods, including TSS, PartitionSort (PS), and SAX-PAC. For example, TM is 34% faster at classifying packets and 30% faster at updating rules than PS. We then experimentally evaluate TM deployed within the VPP framework comparing TM against linear search and TSS, and also against TSS within the OVS framework. This validation of deployed implementations is important as SDN frameworks have several optimizations such as caches that may minimize the influence of a classification algorithm. Our experimental results clearly validate the effectiveness of TM. VPP TM classifies packets nearly two orders of magnitude faster than VPP TSS and at least one order of magnitude faster than OVS TSS.
James Daly, Valerio Bruschi, Leonardo Linguaglossa, Salvatore Pontarelli, Dario Rossi 0001, Jerome Tollet, Eric Torng, Andrew Yourtchenko
IEEE/ACM Trans. Netw.5
2018 Per-Flow Fairness in the Datacenter Network
abstract
Datacenter network (DCN) design has been actively researched for over a decade. Solutions proposed range from end-to-end transport protocol redesign to more intricate, monolithic and cross-layer architectures. Despite this intense activity, to date we remark the absence of DCN proposals based on simple fair scheduling strategies. In this paper, we evaluate the effectiveness of FQ-CoDel in the DCN environment. Our results show, (i) that average throughput is greater than that attained with DCN tailored protocols like DCTCP, and (ii) the completion time of short flows is close to that of state-of-art DCN proposals like pFabric. Good enough performance and striking simplicity make FQ-CoDel a serious contender in the DCN arena.
YiXi Gong, James Roberts, Dario Rossi 0001
HPSR3
2018 A simple yet effective network-assisted signal for enhanced DASH quality of experience
abstract
We propose and evaluate simple signals coming from in-network telemetry that are effective to enhance the quality of DASH streaming. Specifically, in-network caching is known to positively affect DASH streaming quality but at the same time negatively affect the controller stability, increasing the quality switch ratio. Our contributions are to first (i) consider the broad spectrum of interaction between the network and the application, and then (ii) to devise how to effectively exploit in a DASH controller a very simple signal (i.e., per-quality hit ratio) that can be exported by framework such as Server and Network Assisted DASH (SAND) at fairly low rate (i.e., a timescale of 10s of seconds). Our thorough experimental campaign confirms the soundness of the approach (that significantly ameliorate performance with respect to network-blind DASH), as well as its robustness (i.e., tuning is not critical) and practical appeal (i.e., due to its simplicity and compatibility with SAND).
Jacques Samain, Giovanna Carofiglio, Michele Tortelli, Dario Rossi 0001
NOSSDAV4
2018 Leveraging Inter-domain Stability for BGP Dynamics Analysis
Thomas Green, Anthony Lambert, Cristel Pelsser, Dario Rossi 0001
PAM4
2018 Narrowing the Gap Between QoS Metrics and Web QoE Using Above-the-fold Metrics
Diego N. da Hora, Alemnew Sheferaw Asrese, Vassilis Christophides, Renata Teixeira, Dario Rossi 0001
PAM5
2018 A Closer Look at IP-ID Behavior in the Wild
Flavia Salutari, Danilo Cicalese, Dario Rossi 0001
PAM3
2018 Speed Index: Relating the Industrial Standard for User Perceived Web Performance to web QoE
abstract
In 2012, Google introduced the Speed Index (SI) metric to quantify the speed of the Web page visual completeness for the actually displayed above-the-fold (ATF) portion of a Web page. In Web browsing a page might appear to the user to be already fully rendered, even though further content may still be retrieved, resulting in the Page Load Time (PLT). This happens due to the browser progressively rendering all objects, part of which can also be located below the browser window's current viewport. The SI metric (and variants) thereof have since established themselves as a de facto standard in Web page and browser testing. While SI is a step in the direction of including the user experience into Web metrics, the actual meaning of the metric and especially its relationship between Speed Index and Web QoE is however far from being clear. The contributions of this paper are thus to first develop an understanding of the SI based on a theoretical analysis and second, to analyze the interdependency between SI and MOS values from an existing public dataset. Specifically, our analysis is based on two well established models that map the user waiting time to a user ACR-rating of the QoE. The analysis show that ATF-based metrics are more appropriate than pure PLT as input to Web QoE models.
Tobias Hoßfeld, Florian Metzger, Dario Rossi 0001
QoMEX3
2018 Parallel Simulation of Very Large-Scale General Cache Networks
abstract
In this paper, we propose a methodology for the study of general cache networks, which is intrinsically scalable and amenable to parallel execution. We contrast two techniques: one that slices the network and another that slices the content catalog. In the former, each core simulates requests for the whole catalog on a subgraph of the original topology, whereas in the latter each core simulates requests for a portion of the original catalog on a replica of the whole network. Interestingly, we find out that when the number of cores increases (and so the split ratio of the network topology), the overhead of message passing required to keeping consistency among nodes actually offsets any benefit from the parallelization: this is strictly due to the correlation among neighboring caches, meaning that requests arriving at one cache allocated on one core may depend on the status of one or more caches allocated on different cores. Even more interestingly, we find out that the newly proposed catalog slicing, on the contrary, achieves an ideal speedup in the number of cores. Overall, our system, which we make available as open source software, enables performance assessment of large-scale general cache networks, i.e., comprising hundreds of nodes, trillions contents, and complex routing and caching algorithms, in minutes of CPU time and with exiguous amounts of memory.
Michele Tortelli, Dario Rossi 0001, Emilio Leonardi
IEEE J. Sel. Areas Commun.2
2018 Caching Encrypted Content Via Stochastic Cache Partitioning
abstract
In-network caching is an appealing solution to cope with the increasing bandwidth demand of video, audio, and data transfer over the Internet. Nonetheless, in order to protect consumer privacy and their own business, content providers (CPs) increasingly deliver encrypted content, thereby preventing Internet service providers (ISPs) from employing traditional caching strategies, which require the knowledge of the objects being transmitted. To overcome this emerging tussle between security and efficiency, in this paper we propose an architecture in which the ISP partitions the cache space into slices, assigns each slice to a different CP, and lets the CPs remotely manage their slices. This architecture enables transparent caching of encrypted content and can be deployed in the very edge of the ISP's network (i.e., base stations and femtocells), while allowing CPs to maintain exclusive control over their content. We propose an algorithm, called SDCP, for partitioning the cache storage into slices so as to maximize the bandwidth savings provided by the cache. A distinctive feature of our algorithm is that ISPs only need to measure the aggregated miss rates of each CP, but they need not know the individual objects that are requested. We prove that the SDCP algorithm converges to a partitioning that is close to the optimal, and we bound its optimality gap. We use simulations to evaluate SDCP's convergence rate under stationary and nonstationary content popularity. Finally, we show that SDCP significantly outperforms traditional reactive caching techniques, considering both CPs with perfect and with imperfect knowledge of their content popularity.
Andrea Araldo, György Dán, Dario Rossi 0001
IEEE/ACM Trans. Netw.3
2017 The Web, the Users, and the MOS: Influence of HTTP/2 on User Experience
Enrico Bocchi, Luca De Cicco, Marco Mellia, Dario Rossi 0001
PAM4
2017 Exploiting parallelism in hierarchical content stores for high-speed ICN routers
Rodrigo B. Mansilha, Marinho P. Barcellos, Emilio Leonardi, Dario Rossi 0001
Comput. Networks4
2017 A hybrid methodology for the performance evaluation of Internet-scale cache networks
Michele Tortelli, Dario Rossi 0001, Emilio Leonardi
Comput. Networks2
2017 Dynamic Adaptive Video Streaming: Towards a Systematic Comparison of ICN and TCP/IP
abstract
Streaming of video content over the Internet is experiencing an unprecedented growth. While video permeates every application, it also puts tremendous pressure in the network-to support users having heterogeneous accesses and expecting a high quality of experience, in a furthermore cost-effective manner. In this context, future internet paradigms, such as information centric networking (ICN), are particularly well suited to not only enhance video delivery at the client (as in the dynamic adaptive streaming over HTTP (DASH) approach), but to also naturally and seamlessly extend video support deeper in the network functions. In this paper, we contrast ICN and transmission control protocol/internet protocol (TCP/IP) with an experimental approach, where we employ several state-of-the-art DASH controllers (PANDA, AdapTech, and BOLA) on an ICN versus TCP/IP network stack. Our campaign, based on tools that we developed and made available as open-source software, includes multiple clients (homogeneous vesrus heterogeneous mixture and synchronous vesrus asynchronous arrivals), videos (up to 4k resolution), channels (e.g., DASH profiles, emulated WiFi and LTE, and real 3G/4G traces), and levels of integration with an ICN network (i.e., vanilla named data networking (NDN), wireless loss detection and recovery at the access point, and load balancing). Our results clearly illustrate, as well as quantitatively assess, the benefits of ICN-based streaming, warning about potential pitfalls that are however easy to avoid.
Jacques Samain, Giovanna Carofiglio, Luca Muscariello, Michele Papalini, Mauro Sardara, Michele Tortelli, Dario Rossi 0001
IEEE Trans. Multim.7
2016 Statistical network monitoring: Methodology and application to carrier-grade NAT
Enrico Bocchi, Ali Safari Khatouni, Stefano Traverso, Alessandro Finamore, Maurizio M. Munafò, Marco Mellia, Dario Rossi 0001
Comput. Networks7
2016 Framework, models and controlled experiments for network troubleshooting
François Espinet, Diana Joumblatt, Dario Rossi 0001
Comput. Networks3
2016 Latency-Based Anycast Geolocation: Algorithms, Software, and Data Sets
abstract
Use of IP-layer anycast has increased in the last few years beyond the DNS realm. Existing measurement techniques to identify and enumerate anycast replicas exploit specifics of the DNS protocol, which limits their applicability to this particular service. In this paper, we propose and thoroughly validate a protocol-agnostic technique for anycast replicas discovery and geolocation. Furthermore, we also provide the community with open-source software and data sets allowing others to replicate our experimental results, potentially facilitating the development of new techniques such as ours. In particular, our proposed method achieves thorough enumeration and city-level geolocalization of anycast instances from a set of known vantage points. The algorithm features an iterative workflow, pipelining enumeration (an optimization problem using latency as an input), and geolocalization (a classification problem using side channel information, such as city population) of anycast replicas. Results of a thorough validation campaign show our algorithm to be robust to measurement noise, and very lightweight as it requires only a handful of latency measurements.
Danilo Cicalese, Diana Joumblatt, Dario Rossi 0001, Marc-Olivier Buob, Jordan Augé, Timur Friedman
IEEE J. Sel. Areas Commun.3
2016 Cost-Aware Caching: Caching More (Costly Items) for Less (ISPs Operational Expenditures)
abstract
Albeit an important goal of caching is traffic reduction, a perhaps even more important aspect follows from the above achievement: the reduction of internet service provider (ISP) operational costs that comes as a consequence of the reduced load on transit and provider links. Surprisingly, to date this crucial aspect has not been properly taken into account in cache design. In this paper, we show that the classic caching efficiency indicator, i.e., the hit ratio, conflicts with cost. We therefore propose a mechanism whose goal is the reduction of cost and, in particular, we design a cost-aware (CoA) cache decision policy that, leveraging price heterogeneity among external links, tends to store with more probability the objects that the ISP has to retrieve through the most expensive links. We provide a model of our mechanism, based on Che's approximation, and, by means of a thorough simulation campaign, we contrast it with traditional cost-blind schemes, showing that CoA yields a significant cost saving, that is furthermore consistent over a wide range of scenarios. We show that CoA is easy to implement and robust, making the proposal of practical relevance.
Andrea Araldo, Dario Rossi 0001, Fabio Martignon
IEEE Trans. Parallel Distributed Syst.2
2015 Characterizing IPv4 anycast adoption and deployment
abstract
This paper provides a comprehensive picture of IP-layer anycast adoption in the current Internet. We carry on multiple IPv4 anycast censuses, relying on latency measurement from PlanetLab. Next, we leverage our novel technique for anycast detection, enumeration, and geolocation [17] to quantify anycast adoption in the Internet. Our technique is scalable and, unlike previous efforts that are bound to exploiting DNS, is protocol-agnostic. Our results show that major Internet companies (including tier-1 ISPs, over-the-top operators, Cloud providers and equipment vendors) use anycast: we find that a broad range of TCP services are offered over anycast, the most popular of which include HTTP and HTTPS by anycast CDNs that serve websites from the top-100k Alexa list. Additionally, we complement our characterization of IPv4 anycast with a description of the challenges we faced to collect and analyze large-scale delay measurements, and the lessons learned.
Danilo Cicalese, Jordan Augé, Diana Joumblatt, Timur Friedman, Dario Rossi 0001
CoNEXT5
2015 Snooping Wikipedia vandals with MapReduce
abstract
In this paper, we present and validate an algorithm able to accurately identify anomalous behaviors on online and collaborative social networks, based on their interaction with other fellows. We focus on Wikipedia, where accurate ground truth for the classification of vandals can be reliably gathered by manual inspection of the page edit history. We develop a distributed crawler and classifier tasks, both implemented in MapReduce, with whom we are able to explore a very large dataset, consisting of over 5 millions articles collaboratively edited by 14 millions authors, resulting in over 8 billion pairwise interactions. We represent Wikipedia as a signed network, where positive arcs imply constructive interaction between editors. We then isolate a set of high reputation editors (i.e., nodes having many positive incoming links) and classify the remaining ones based on their interactions with high reputation editors. We demonstrate our approach not only to be practically relevant (due to the size of our dataset), but also feasible (as it requires few MapReduce iteration) and accurate (over 95% true positive rate). At the same time, we are able to classify only about half of the dataset editors (recall of 50%) for which we outline some solution under study.
Michele Spina, Dario Rossi 0001, Mauro Sozio, Silviu Maniu, Bogdan Cautis
ICC2
2015 A fistful of pings: Accurate and lightweight anycast enumeration and geolocation
abstract
Use of IP-layer anycast has increased in the last few years: once relegated to DNS root and top-level domain servers, anycast is now commonly used to assist distribution of general purpose content by CDN providers. Yet, the measurement techniques for discovering anycast replicas have been designed around DNS, limiting their usefulness to this particular service. This raises the need for protocol agnostic methodologies, that should additionally be as lightweight as possible in order to scale up anycast service discovery. This is precisely the aim of this paper, which proposes a new method for exhaustive and accurate enumeration and city-level geolocation of anycast instances, requiring only a handful of latency measurements from a set of known vantage points. Our method exploits an iterative workflow that enumerates (an optimization problem) and geolocates (a classification problem) anycast replicas. We thoroughly validate our methodology on available ground truth (several DNS root servers), using multiple measurement infrastructures (PlanetLab, RIPE), obtaining extremely accurate results (even with simple algorithms, that we compare with the global optimum), that we make available to the scientific community. Compared to the state of the art work that appeared in INFOCOM 2013 and IMC 2013, our technique (i) is not bound to a specific protocol, (ii) requires 1000 times fewer vantage points, not only (iii) achieves over 50% recall but also (iv) accurately identifies the city-level geolocation for over 78% of the enumerated servers, with (v) a mean geolocation error of 361 km for all enumerated servers.
Danilo Cicalese, Diana Joumblatt, Dario Rossi 0001, Marc-Olivier Buob, Jordan Augé, Timur Friedman
INFOCOM3
2015 Impact of Carrier-Grade NAT on web browsing
abstract
Public IPv4 addresses are a scarce resource. While IPv6 adoption is lagging, Network Address Translation (NAT) technologies have been deployed over the last years to alleviate IPv4 exiguity and their high rental cost. In particular, Carrier-Grade NAT (CGN) is a well known solution to mask a whole ISP network behind a limited amount of public IP addresses, significantly reducing expenses.
Enrico Bocchi, Ali Safari Khatouni, Stefano Traverso, Alessandro Finamore, Valeria Di Gennaro, Marco Mellia, Maurizio M. Munafò, Dario Rossi 0001
IWCMC8
2015 Design and analysis of an improved bitmessage anti-spam mechanism
abstract
The BitMessage protocol offers privacy to its anonymous users. It is a completely decentralized messaging system, enabling users to exchange messages preventing accidental eavesdropping - a nice features in the Post-Snowden Internet Era. Not only messages are sent to every node on the network (making it impossible to understand the intended recipient), but their content is encrypted with the intended recipient public key (so that s/he only can decipher it). As these two properties combined might facilitate spamming, a proof-of-work (PoW) mechanism has been designed to mitigate this threat: only messages exhibiting properties of the PoW are forwarded on the network: since PoW is based on computationally heavy cryptographic functions, this slows down the rate at which spammers can introduce unsolicited messages in the network on the one hand, but also makes it harder to send legitimate messages for regular users on the other hand. In this paper, we (i) carry on an analysis of the current PoW mechanism, (ii) propose a very simple, yet very effective, generalization of the formula that decouples spammers vs legitimate users penalty showing that (iii) at the optimum, our proposal halves the harm spammers can do, avoiding by definition any impact for legitimate users.
Alexander Schaub 0001, Dario Rossi 0001
P2P2
2014 Cost-aware caching: Optimizing cache provisioning and object placement in ICN
abstract
Caching is frequently used by Internet Service Providers as a viable technique to reduce the latency perceived by end users, while jointly offloading network traffic. While the cache hit-ratio is generally considered in the literature as the dominant performance metric for such type of systems, in this paper we argue that a critical missing piece has so far been neglected. Adopting a radically different perspective, in this paper we explicitly account for the cost of content retrieval, i.e. the cost associated to the external bandwidth needed by an ISP to retrieve the contents requested by its customers. Interestingly, we discover that classical cache provisioning techniques that maximize cache efficiency (i.e., the hit-ratio), lead to suboptimal solutions with higher overall cost. To show this mismatch, we propose two optimization models that either minimize the overall costs or maximize the hit-ratio, jointly providing cache sizing, object placement and path selection. We formulate a polynomial-time greedy algorithm to solve the two problems and analytically prove its optimality. We provide numerical results and show that significant cost savings are attainable via a cost-aware design.
Andrea Araldo, Michele Mangili, Fabio Martignon, Dario Rossi 0001
GLOBECOM4
2014 Multi-Terabyte and multi-Gbps information centric routers
abstract
One of the main research directions along which the future Internet is evolving can be identified in the paradigmatic shift from a network of hosts toward a network of caches. Yet, several questions remain concerning the scalability of individual algorithms (e.g., name based lookup and routing) and components (e.g., caches) of these novel Information Centric Networking (ICN) architectures. Exploiting a peculiar characteristics of ICN (i.e., the fact that contents are split in chunks), and the nature of video streaming (which dominates Internet traffic), this paper proposes a novel two-layers caching scheme that allows multi-Terabyte caches to sustain content streaming at multi-Gbps speed. We model the system as an extension, to the case of chunked contents, of the well known Che approximation, that has the advantage of being very simple and accurate at the same time. Simulations under synthetic and realistic trace-driven traffic confirm the accuracy of the analysis and the feasibility of the proposed architecture.
Giuseppe Rossini, Dario Rossi 0001, Michele Garetto, Emilio Leonardi
INFOCOM2
2014 A per-application account of bufferbloat: Causes and impact on users
abstract
We propose a methodology to gauge the extent of queueing delay (aka bufferbloat) in the Internet, based on purely passive measurement of TCP traffic. We implement our methodology in Tstat and make it available as open source software. We leverage Deep Packet Inspection (DPI) and behavioral classification of Tstat to breakdown the queueing delay across different applications, in order to evaluate the impact of bufferbloat on user experience. We show that there is no correlation between the ISP traffic load and the queueing delay, thus confirming that bufferbloat is related only to the traffic of each single user (or household). Finally, we use frequent itemset mining techniques to associate the amount of queueing delay seen by each host with the set of its active applications, with the goal of investigating the root cause of bufferbloat.
Andrea Araldo, Dario Rossi 0001
IWCMC2
2014 Distributed Active Measurement of Internet Queuing Delays
Pellegrino Casoria, Dario Rossi 0001, Jordan Augé, Marc-Olivier Buob, Timur Friedman, Antonio Pescapè
PAM2
2014 Scalable Accurate Consolidation of Passively Measured Statistical Data
Silvia Colabrese, Dario Rossi 0001, Marco Mellia
PAM2
2014 Violation of Interdomain Routing Assumptions
Riad Mazloum, Marc-Olivier Buob, Jordan Augé, Bruno Baynat, Dario Rossi 0001, Timur Friedman
PAM5
2014 Pedestrian crossing: The long and winding road toward fair cross-comparison of ICN quality
abstract
While numerous Information Centric Networking (ICN) architectures have been proposed over the last years, the community has so far only timidly attempted at a quantitative assessment of the relative quality of service level that users are expected to enjoy in each of them. This paper starts a journey toward the cross comparison of ICN alternatives, making several contributions along this road. Specifically, a census of 20 ICN software tools reveals that about 10 are dedicated to a specific architecture, about half of which are simulators. Second, we survey ICN research papers using simulation to gather information concerning the used simulator, finding that a large fraction either uses custom proprietary and unavailable software, or even plainly fails to mention any information on this regard, which is deceiving. Third, we cross-compare some of the available simulators, finding that they achieve consistent results, which is instead encouraging. Fourth, we propose a methodology to increase and promote cross-comparison, which is within reach but requires community-wide agreement, promotion and enforcement.
Michele Tortelli, Dario Rossi 0001, Gennaro Boggia, Luigi Alfredo Grieco
QSHINE2
2014 Fighting the bufferbloat: On the coexistence of AQM and low priority congestion control
abstract
Nowadays, due to excessive queuing, delays on the Internet can grow longer than the round trip time between the Moon and the Earth – for which the “bufferbloat” term was recently coined. Some point to active queue management (AQM) as the solution. Others propose end-to-end low-priority congestion control techniques (LPCC). Under both approaches, promising advances have been made in recent times: notable examples are CoDel for AQM, and LEDBAT for LPCC. In this paper, we warn of a potentially fateful interaction when AQM and LPCC techniques are combined: namely, AQM resets the relative level of priority between best-effort and low-priority congestion control protocols. We validate the generality of our findings by an extended set of experiments with packet-level ns2 simulation, considering 5 AQM techniques and 3 LPCC protocols, and carry on a thorough sensitivity analysis varying several parameters of the networking scenario. We complete the simulation via an experimental campaign conducted on both controlled testbeds and on the Internet, confirming the reprioritization issue to hold in the real world at least under all combination of AQM policies and LPCC protocols available in the Linux kernel. To promote cross-comparison, we make our scripts and dataset available to the research community.
YiXi Gong, Dario Rossi 0001, Claudio Testa, Silvio Valenti, M. Dave Taht
Comput. Networks2
2014 Delay-based congestion control: Flow vs. BitTorrent swarm perspectives
Claudio Testa, Dario Rossi 0001
Comput. Networks2
2013 Passive bufferbloat measurement exploiting transport layer information
abstract
“Bufferbloat” is the growth in buffer size that has led Internet delays to occasionally exceed the light propagation delay from the Earth to the Moon. Manufacturers have built in large buffers to prevent losses on Wi-Fi, cable and ADSL links. But the combination of some links' limited bandwidth with TCP's tendency to saturate that bandwidth results in excessive queuing delays. In response, new congestion control protocols such as BitTorrent's uTP/LEDBAT aim at explicitly limiting the delay that they add at the bottleneck link. This work proposes a methodology to monitor the upstream queuing delay experienced by remote hosts, both those using LEDBAT, through LEDBAT's native one-way delay measurements, and those using TCP, through the Timestamp Option. We report preliminary findings on bufferbloat-related queuing delays on an Internet measurement campaign involving a few thousand hosts.
Chiara Chirichella, Dario Rossi 0001, Claudio Testa, Timur Friedman, Antonio Pescapè
GLOBECOM2
2013 ccnSim: An highly scalable CCN simulator
abstract
Research interest about Information Centric Networking (ICN) has grown at a very fast pace over the last few years, especially after the 2009 seminal paper of Van Jacobson et al. describing a Content Centric Network (CCN) architecture. While significant research effort has been produced in terms of architectures, algorithms, and models, the scientific community currently lacks common tools and scenarios to allow a fair cross-comparison among the different proposals. The situation is particularly complex as the commonly used general-purpose simulators cannot cope with the expected system scale: thus, many proposals are currently evaluated over small and unrealistic scale, especially in terms of dominant factors like catalog and cache sizes. As such, there is need of a scalable tool under which different algorithms can be tested and compared. Over the last years, we have developed and optimized ccnSim, an highly scalable chunk-level simulator especially suitable for the analysis of caching performance of CCN network. In this paper, we briefly describe the tool, and present an extensive benchmark of its performance. To give an idea of ccnSim scalability, a common off-the-shelf PC equipped with 8GB of RAM memory is able to simulate 2-hours of a 50-nodes CCN network, where each nodes is equipped with 10GB caches, serving a 1PB catalog in about 20 min CPU time.
Raffaele Chiocchetti, Dario Rossi 0001, Giuseppe Rossini
ICC2
2013 To the Moon and back: Are Internet bufferbloat delays really that large?
abstract
Recently, the “bufferbloat” term has been coined to describe very large queuing delays (up to several seconds) experienced by Internet users. This problem has pushed protocol designer to deploy alternative (delay-based) models to the standard (lossbased) TCP best effort congestion control. In this work, we exploit timestamp information carried in the LEDBAT header, a protocol proposed by BitTorrent as replacement for TCP data transfer, to infer the queuing delay suffered by remote hosts. We conduct a thorough measurement campaign, that let us conclude that (i) LEDBAT delay-based congestion control is effective in keeping the queuing delay low for the bulk of the peers, (ii) yet about 1% of peers often experience queuing delay in excess of 1s, and (iii) not only the network access type, but also the BitTorrent client and the operating system concurr in determining the bufferbloat magnitude.
Chiara Chirichella, Dario Rossi 0001
INFOCOM2
2013 Fighting the bufferbloat: On the coexistence of AQM and low priority congestion control
abstract
Nowadays, due to excessive queuing, delays on the Internet can grow longer than several round trips between the Moon and the Earth - for which the “bufferbloat” term was recently coined. Some point to active queue management (AQM) as the solution. Others propose end-to-end low-priority congestion control techniques (LPCC). Under both approaches, promising advances have been made in recent times: notable examples are CoDel for AQM, and LEDBAT for LPCC. In this paper, we warn of a potentially fateful interaction when AQM and LPCC techniques are combined: namely (i) AQM resets the relative level of priority between best effort and low-priority congestion control protocols; (ii) while reprioritization generally equalizes the priority of LPCC and TCP, we also find that some AQM settings may actually lead best effort TCP to starvation. By an extended set of experiments conducted on both controlled testbeds and on the Internet, we show the problem to hold in the real world for all tested combination of AQM policies and LPCC protocols. To further validate the generality of our findings, we complement our experiments with packet-level simulation, to cover cases of other popular AQM and LPCC that are not available in the Linux kernel. To promote cross-comparison, we make our scripts and dataset available to the research community.
YiXi Gong, Dario Rossi 0001, Claudio Testa, Silvio Valenti, M. Dave Taht
INFOCOM2
2013 I Tube, YouTube, P2PTube: Assessing ISP benefits of peer-assisted caching of YouTube content
abstract
The last few years have seen an explosion of video on demand traffic carried over the Internet infrastructure. While P2P applications have been proposed to carry VoD and TV content, they have so far encountered limited adoption except in Asian countries. Part of why this happens is explained with the fact that (i) the current asymmetric network infrastructure does not offer enough system capacity needed to let a fully P2P-VoD/TV to be self-sustainable, (ii) that the actual capacity at nominal peers is often smaller than the available one due to inefficiency in NAT punching[1], and (iii) the very same nonelastic nature of the service, that makes the system inherently less robust w.r.t elastic file-sharing to dynamic changes in the istantaneously available bandwidth. The other part of the story can be summarized with the success of CDN-managed services such as Netflix, Hulu and especially YouTube - according to [2], about 3 billion YouTube videos are viewed and 100's of thousand videos are uploaded every day, with independent research confirming YouTube to represent 20-30% of ISPs incoming traffic[3].
Yann Nicolas, Daniel Wolff, Dario Rossi 0001, Alessandro Finamore
P2P3
2013 Data plane throughput vs control plane delay: Experimental study of BitTorrent performance
abstract
In this paper, we address the trade-off between the data plane efficiency and the control plane timeliness for the BitTorrent performance. We argue that loss-based congestion control protocols can fill large buffers, leading to a higher end-to-end delay, unlike low-priority or delay-based congestion control protocols. We perform experiments for both the uTorrent and mainline BitTorrent clients, and we study the impact of uTP (a novel transport protocol proposed by BitTorrent) and several TCP congestion control algorithms (Cubic, New Reno, LP, Vegas and Nice) on the download completion time. Briefly, in case peers in the swarm all use the same congestion control algorithm, we observe that the specific algorithm has only a limited impact on the swarm performance. Conversely, when a mix of TCP congestion control algorithms coexists, peers employing a delay-based low-priority algorithm exhibit shorter completion time.
Claudio Testa, Dario Rossi 0001, Ashwin Rao, Arnaud Legout
P2P2
2013 Remotely Gauging Upstream Bufferbloat Delays
Chiara Chirichella, Dario Rossi 0001, Claudio Testa, Timur Friedman, Antonio Pescapè
PAM2
2013 Rethinking the Low Extra Delay Background Transport (LEDBAT) Protocol
Giovanna Carofiglio, Luca Muscariello, Dario Rossi 0001, Claudio Testa, Silvio Valenti
Comput. Networks3
2013 FIB Aplasia through probabilistic routing and autoforwarding
Giuseppe Rossini, Dario Rossi 0001, Christophe Betoule, Remi Clavier, Gilles Thouénon
Comput. Networks2
2013 Evaluating CCN multi-path interest forwarding strategies
Giuseppe Rossini, Dario Rossi 0001
Comput. Commun.2
2013 ModelNet-TE: An emulation tool for the study of P2P and traffic engineering interaction dynamics
Dario Rossi 0001, Paolo Veglia, Matteo Sammarco, Federico Larroca
Peer-to-Peer Netw. Appl.1
2013 Performance evaluation of P2P-TV diffusion algorithms under realistic settings
Paolo Veglia, Dario Rossi 0001
Peer-to-Peer Netw. Appl.2
2012 Wire-speed statistical classification of network traffic on commodity hardware
abstract
In this paper we present a software-based traffic classification engine running on commodity multi-core hardware, able to process in real-time aggregates of up to 14.2 Mpps over a single 10 Gbps interface -- i.e., the maximum possible packet rate over a 10 Gbps Ethernet links given the minimum frame size of 64 Bytes.
Pedro M. Santiago del Río, Dario Rossi 0001, Francesco Gringoli, Lorenzo Nava, Luca Salgarelli, Javier Aracil 0001
Internet Measurement Conference2
2012 Inferring the buffering delay of remote BitTorrent peers under LEDBAT vs TCP
abstract
Nowadays, due to excessive queuing, Internet delays grow sometimes as large as the propagation delay from moon to earth - for which the bufferbloat term was recently coined. Some points to active queue management (AQM) as its solution, others propose end-to-end congestion control techniques - like BitTorrent that recently replaced TCP with the LEDBAT transport protocol. In this demo, we implement a methodology to monitor the upstream queuing delay experienced by remote hosts, both those using LEDBAT, through LEDBAT's native one-way delay measurements, and those using TCP, through the timestamp option. By actively taking part into torrent downloads as leechers, our software is able to infer (and visualize) the amount of access delay suffered by the remote peers.
Chiara Chirichella, Dario Rossi 0001, Claudio Testa, Timur Friedman, Antonio Pescapè
P2P2
2011 Identifying Key Features for P2P Traffic Classification
abstract
Many researchers have recently dealt with P2P traffic classification, mainly because P2P applications are continuously growing in number as well as in traffic volume. Additionally, in response to the shift of the operational community from packet-level to flow-level monitoring, witnessed by the widespread use of NetFlow, a number of behavioral classifiers have been proposed. These techniques, usually having P2P applications as their main target, base the classification on the analysis of the pattern of traffic generated by a host and proved accurate even when using only flow-level data. Yet, all these approaches are very specific and the community lacks a broader view of the actual amount of information of behavioral features derived by flow-level data. The preliminary results presented in this paper try to fill this gap. First of all we define a comprehensive framework by means of which we systematically explore the space of behavioral properties and build a large set of potentially expressive features. Thanks to our general approach, most features already used by existing classifiers fall into this set. Then, by employing tools from information theory and data from packet-level traces captured on real networks, we evaluate the amount of information conveyed by each feature, ranking them according to their usefulness for application identification. Finally we show the classification performance of these set of features, using a supervised machine learning algorithm.
Silvio Valenti, Dario Rossi 0001
ICC2
2011 Fine-grained behavioral classification in the core: the issue of flow sampling
abstract
This work studies the impact of flow sampling on the accuracy of behavioral traffic classification. More precisely, we consider the case where the traffic classification engine is located at different vantage points of the network. Usually, behavioral classification is performed close to the user access network - where all the traffic exchanged by an endpoint can be observed. In this work instead we take into account the case of a classifier placed deeper in the aggregation network - where, due to load balancing or routing issues, only part of the traffic is generally observed. We use the Abacus behavioral classification engine as our case study, as it has been shown to provide accurate classification of P2P applications, by relying only on the count of packets and bytes peers exchange during fixed-length time-windows. We further consider multiple policies of flow sampling, that either reflect real router forwarding table, or allow to assess parameter impact under controlled settings. An accurate measurement campaign shows that, provided that the signature definition does not rely on absolute counters of the traffic volume (e.g., which can be achieved by means of normalization), even when signatures are computed over a minority of the traffic (e.g., about 10%) the classification is still reliable (e.g, accuracy exceeds 75%).
Silvio Valenti, Dario Rossi 0001
IWCMC2
2011 On the impact of uTP on BitTorrent completion time
abstract
BitTorrent, one of the most widespread file sharing P2P applications, recently introduced uTP, an application-level congestion control protocol which aims to efficiently use the available link capacity, while avoiding to interfere with the rest of user traffic (e.g., Web, VoIP and gaming) sharing the same access bottleneck. Research on uTP has so far focused on the investigation of the congestion control behavior on rather simple settings (i.e., single bottleneck, few backlogged flows, etc.), that are fairly far from the P2P settings in which the protocol is deployed. Moreover, prior work typically addressed questions, such as fairness and efficiency, that are natural from a congestion control context perspective, but are not directly related with the performance of the overall P2P system. In this work, we refine the understanding of uTP, by gauging its impact on the primary BitTorrent user-centric metric, namely the torrent download time, by means of packet level simulation. Results of our initial investigations show that: (i) in case uTP clients fully substitute TCP clients, no performance difference arise; (ii) in case of heterogeneous swarms, comprising peers using uTP and TCP congestion control, completion time of uTP peers can possibly benefit of lower uplink queuing delays, as signaling traffic (e.g., chunk requests) are not slowed down by long waits in the ADSL buffers.
Claudio Testa, Dario Rossi 0001
Peer-to-Peer Computing2
2011 Abacus: Accurate behavioral classification of P2P-TV traffic
Paola Bermolen, Marco Mellia, Michela Meo, Dario Rossi 0001, Silvio Valenti
Comput. Networks4
2011 Black-box analysis of Internet P2P applications
Dario Rossi 0001, Elisa Sottile, Paolo Veglia
Peer-to-Peer Netw. Appl.1
2011 On the quality of broadcast services in vehicular ad hoc networks
abstract
Abstract We investigate the broadcast problem in suburban and highway inter‐vehicular networks, aiming at providing a definitive comparison of two antipodean broadcast algorithm classes: the first one makes use of someinstantaneous informationlocally available at the vehicles (such as vehicle position and speed), while the second one relies onlong‐term knowledgegained through a beaconing procedure. Using a realistic microscopic model to represent the vehicular traffic flow, we investigate the performance of the above broadcast algorithm classes by simulation, considering different classes of network services (e.g., Critical, Normal, and Low‐priority). In order to explore a very large algorithmic design space, we devise a convex hull framework that allows us to effectively compare and compactly present the boundaries of the solution space for each algorithm class. By the use of such framework, we show that the beaconless performance encompasses a wider spectrum with respect to the beaconed one, with lower complexity and overhead. Copyright © 2010 John Wiley & Sons, Ltd.
Dario Rossi 0001, Roberta Fracchia, Michela Meo
Secur. Commun. Networks1
2010 The Quest for LEDBAT Fairness
abstract
BitTorrent developers have recently introduced a new application layer congestion control algorithm based on UDP framing at transport layer and currently under definition at the IETF ledbat Working Group. Ledbat is a delay-based protocol which aims at offering a ``lower than Best Effort'''' data transfer service, with a lower priority with respect to elastic TCP and interactive traffic (e.g., VoIP, game). However, in its current specification, ledbat is affected by a late-comer advantage: indeed the last flow arriving at the bottleneck is more aggressive due to a wrong estimation of the base delay and finally takes over all resources. In this work, we study several solutions to the late-comer problem by means of packet level simulations and simple analysis: in the investigation process, we individuate the root cause for LEDBAT unfairness and propose effective countermeasures.
Giovanna Carofiglio, Luca Muscariello, Dario Rossi 0001, Silvio Valenti
GLOBECOM3
2010 Stochastic Packet Inspection for TCP Traffic
abstract
In this paper, we extend the concept of Stochastic Packet Inspection (SPI) to support TCP traffic classification. SPI is a method based on the statistical fingerprint of the application-layer headers: by characterizing the frequencies of observed symbols, SPI can identify application protocol formats by automatically recognizing group of bits that take e.g., constant values, or random values, or are part of a counter. To correctly characterize symbol frequencies, SPI needs volumes of traffic to obtain statistically significant signatures. Earlier proposed for UDP traffic, SPI has to be modified to cope with the connection oriented service offered by TCP, in which application-layer headers are only found at the beginning of a TCP connection. In this paper, we extend SPI to support TCP traffic, and analyze its performance on real network data. The key idea is to move the classification target from single flows to endpoints, which aggregates all traffic sent/received by the same IP address and TCP port pair. The first few packets of flows sent from (or destined to) the same endpoint are then aggregated to yield a single SPI signature. Results show that SPI is able to achieve remarkably good results, with an average true positive rate of about 98%.
Gianluca La Mantia, Dario Rossi 0001, Alessandro Finamore, Marco Mellia, Michela Meo
ICC2
2010 LEDBAT: The New BitTorrent Congestion Control Protocol
abstract
A few months ago, BitTorrent developers announced that the transfer of torrent data in the official client was about to switch to a new application-layer congestion-control protocol using UDP at the transport-layer. This announcement immediately raised an unmotivated buzz about a new, imminent congestion collapse of the whole Internet. As the new congestion control aims at offering a lower than best effort transport service, this reaction was not built on solid technical foundation. Nevertheless, a legitimate question remains: whether this new protocol is a necessary building block for future Internet applications, or whether it may result in an umpteenth addition to the already well populated world of Internet congestion control algorithms. To tackle this issue, we implement the novel congestion control algorithm and investigate its performance by means of packet-level simulations. Considering a simple bottleneck scenario, where the new BitTorrent competes against either TCP or other BitTorrent flows, we evaluate the fairness of resource share as well as the protocol efficiency. Our results show that the new protocol successfully meets some of its design goals, as for instance the efficiency one. At the same time, we also identify some potential fairness issues, that need to be dealt with. Finally, we point out that end-users will be the final judges of the new protocol: therefore, further research should evaluate the effects of its adoption on the performance of the applications ultimately relying on it.
Dario Rossi 0001, Claudio Testa, Silvio Valenti, Luca Muscariello
ICCCN1
2010 Fine-grained traffic classification with netflow data
abstract
Nowadays Cisco Netflow is the de facto standard tool used by network operators and administrators for monitoring large edge and core networks. Implemented by all major vendors and recently a IETF standard, Netflow reports aggregated information about traffic traversing the routers in the form of flow-records. While this kind of data is already effectively used for accounting, monitoring and anomaly detection, the limited amount of information it conveys has until now hindered its employment for traffic classification purposes. In this paper, we present a behavioral algorithm which successfully exploits Netflow records for traffic classification. Since our classifier identifies an application by means of the simple counts of received packets and bytes, Netflow records contain all information required. We test our classification engine, based on a machine learning algorithm, over an extended set of traces containing a heterogeneous mix of applications ranging from P2P file-sharing and P2P live-streaming to traditional client-server services. Results show that our methodology correctly identifies the byte-wise traffic volume with an accuracy of 90% in the worst case, thus representing a first step towards the use of Netflow data for fine-grained classification of network traffic.
Dario Rossi 0001, Silvio Valenti
IWCMC1
2010 A hands-on assessment of transport protocols with lower than best effort priority
abstract
Last year, the official BitTorrent client switched to LEDBAT, a new congestion control algorithm targeting a lower-than Best Effort transport service. In this paper, we study this new protocol through packet-level simulations, with a special focus on a performance comparison with other lower-than Best Effort protocols such as TCP-LP and TCP-NICE: our aim is indeed to quantify and relatively weight the level of Low-priority provided by such protocols. Our results show that LEDBAT transport generally achieves the lowest possible level of priority, with the default configurations of TCP-NICE and TCP-LP representing increasing levels of aggressiveness. In addition, we perform a careful sensitivity analysis of LEDBAT performance, by tuning its main parameters in both an inter-protocol (against TCP) and intra-protocol (against LEDBAT itself) scenarios. In the inter-protocol case, even in case of misconfiguration LEDBAT competes as aggressively as TCP, but we show that it is not possible to achieve an arbitrary level of low-priority by merely tuning its parameters. In the intra-protocol case, we show that coexistence of legacy flows with slightly dissimilar settings, or experiencing different network conditions, can result in significant unfairness.
Giovanna Carofiglio, Luca Muscariello, Dario Rossi 0001, Claudio Testa
LCN3
2010 Yes, We LEDBAT: Playing with the New BitTorrent Congestion Control Algorithm
Dario Rossi 0001, Claudio Testa, Silvio Valenti
PAM1
2010 Network Awareness of P2P Live Streaming Applications: A Measurement Study
abstract
Early P2P-TV systems have already attracted millions of users, and many new commercial solutions are entering this market. Little information is however available about how these systems work, due to their closed and proprietary design. In this paper, we present large scale experiments to compare three of the most successful P2P-TV systems, namely PPLive, SopCast and TVAnts. Our goal is to assess what level of "network awareness" has been embedded in the applications. We first define a general framework to quantify which network layer parameters leverage application choices, i.e., what parameters mainly drive the peer selection and data exchange. We then apply the methodology to a large dataset, collected during a number of experiments where we deployed about 40 peers in several European countries. From analysis of the dataset, we observe that TVAnts and PPLive exhibit a mild preference to exchange data among peers in the same autonomous system the peer belongs to, while this clustering effect is less intense in SopCast. However, no preference versus country, subnet or hop count is shown. Therefore, we believe that next-generation P2P live streaming applications definitively need to improve the level of network-awareness, so to better localize the traffic in the network and thus increase their network-friendliness as well.
Delia Ciullo, M.-A. Garcia da Rocha Neta, Ákos Horváth 0004, Emilio Leonardi, Marco Mellia, Dario Rossi 0001, Miklós Telek, Paolo Veglia
IEEE Trans. Multim.6
2010 KISS: Stochastic Packet Inspection Classifier for UDP Traffic
abstract
This paper proposes KISS, a novel Internet classification engine. Motivated by the expected raise of UDP traffic, which stems from the momentum of Peer-to-Peer (P2P) streaming applications, we propose a novel classification framework that leverages on statistical characterization of payload. Statistical signatures are derived by the means of a Chi-Square (χ2)-like test, which extracts the protocol “format,” but ignores the protocol “semantic” and “synchronization” rules. The signatures feed a decision process based either on the geometric distance among samples, or on Support Vector Machines. KISS is very accurate, and its signatures are intrinsically robust to packet sampling, reordering, and flow asymmetry, so that it can be used on almost any network. KISS is tested in different scenarios, considering traditional client-server protocols, VoIP, and both traditional and new P2P Internet applications. Results are astonishing. The average True Positive percentage is 99.6%, with the worst case equal to 98.1,% while results are almost perfect when dealing with new P2P streaming applications.
Alessandro Finamore, Marco Mellia, Michela Meo, Dario Rossi 0001
IEEE/ACM Trans. Netw.4
2009 Do Next Generation Networks Need Path Diversity?
abstract
We have currently reached a phase where big shifts in the network traffic might impose to rethink the design of current architectures, and where new technologies, being pushed into market, will act as enabler of such changes. Taking into account the current scenario and its likely evolution as well, in this paper we examine the case for multi-path routing within the metropolitan access network. Through an optimization framework, we undertake the analysis of several interesting aspects of the problem, such as (i) the user access technology, (ii) the topology of the access network and (iii) the traffic locality ratio within the access. By numerical solution of the problem we quantify the potential gain given by path-diversity: our results confirm the appeal of multi-path routing strategies both from the user and the network perspectives.
Luca Muscariello, Diego Perino, Dario Rossi 0001
ICC3
2009 Evidences Behind Skype Outage
abstract
Skype is one of the most successful VoIP application in the current Internet spectrum. One of the most peculiar characteristics of Skype is that it relies on a P2P infrastructure for the exchange of signaling information amongst active peers. During August 2007, an unexpected outage hit the Skype overlay, yielding to a service blackout that lasted for more than two days: this paper aims at throwing light to this event. Leveraging on the use of an accurate Skype classification engine, we carry on an experimental study of Skype signaling during the outage. In particular, we focus on the signaling traffic before, during and after the outage, in the attempt to quantify interesting properties of the event. While it is very difficult to gather clear insights concerning the root causes of the breakdown itself, the collected measurement allow nevertheless to quantify several interesting aspects of the outage: for instance, measurements show that the outage caused, on average, a 3-fold increase of signaling traffic and a 10-fold increase of number of contacted peers, topping to more than 11 million connections for the most active node in our network - which immediately gives the feeling of the extent of the phenomenon.
Dario Rossi 0001, Marco Mellia, Michela Meo
ICC1
2009 Network awareness of P2P live streaming applications
abstract
Early P2P-TV systems have already attracted millions of users, and many new commercial solutions are entering this market. Little information is however available about how these systems work. In this paper we present large scale sets of experiments to compare three of the most successful P2P-TV systems, namely PPLive, SopCast and TVAnts. Our goal is to assess what level of "network awareness" has been embedded in the applications, i.e., what parameters mainly drive the peer selection and data exchange. By using a general framework that can be extended to other systems and metrics, we show that all applications largely base their choices on the peer bandwidth, i.e., they prefer high-bandwidth users, which is rather intuitive. Moreover, TVAnts and PPLive exhibits also a preference to exchange data among peers in the same autonomous system the peer belongs to. However, no evidence about preference versus peers in the same subnet or that are closer to the considered peer emerges. We believe that next-generation P2P live streaming applications definitively need to improve the level of network-awareness, so to better localize the traffic in the network and thus increase their network-friendliness as well.
Delia Ciullo, M.-A. Garcia da Rocha Neta, Ákos Horváth 0004, Emilio Leonardi, Marco Mellia, Dario Rossi 0001, Miklós Telek, Paolo Veglia
IPDPS6
2009 Sherlock: A Framework for P2P Traffic Analysis
abstract
After P2P file-sharing and VoIP telephony applications, VoD and live-streaming P2P applications have finally gained a large Internet audience as well. In this work, we define a framework for the comparison of these applications, based on the measurement and analysis of the traffic they generate. In order for the framework to be descriptive for all P2P applications, we first define the observable of interest: such metrics either pertain to different layers of the protocol stack (from network up to the application), or convey cross-layer information (such as the degree of awareness, at overlay layer, of properties characterizing the underlying physical network). The framework is compact (as it allows to represent all the above information at once), general (as is can be extended to consider metrics different from the one reported in this work), and flexible in both space and time (as it allows different levels of spatial aggregation, and also to represent the temporal evolution of the quantities of interest). Based on this framework, we analyze some of the most popular P2P application nowadays, highlighting their main similarities and differences.
Dario Rossi 0001, Elisa Sottile
Peer-to-Peer Computing1
2009 Support vector regression for link load prediction
Paola Bermolen, Dario Rossi 0001
Comput. Networks2
2009 Understanding Skype signaling
Dario Rossi 0001, Marco Mellia, Michela Meo
Comput. Networks1
2009 Detailed Analysis of Skype Traffic
abstract
Skype is beyond any doubt the VoIP application in the current Internet application spectrum. Its amazing success has drawn the attention of telecom operators and the research community, both interested in knowing its internal mechanisms, characterizing its traffic, understanding its users' behavior. In this paper, we investigate the characteristics of traffic streams generated by voice and video communications, and the signaling traffic generated by Skype. Our approach is twofold, as we make use of both active and passive measurement techniques to gather a deep understanding on the traffic Skype generates. From extensive testbed experiments, we devise a source model which takes into account: i) the service type, i.e., SkypeOut calls or calls between two Skype clients, ii) the selected source Codec, iii) the adopted transport layer protocol, and iv) network conditions. Leveraging on the use of an accurate Skype classification engine that we recently proposed, we study and characterize Skype traffic based on extensive passive measurements collected from our campus LAN.
Dario Bonfiglio, Marco Mellia, Michela Meo, Dario Rossi 0001
IEEE Trans. Multim.4
2008 VANETs: Why Use Beaconing at All?
abstract
We investigate the broadcast problem in intervehicular networks, aiming at assessing a definitive comparison of two antipodean algorithm classes: the first one makes use of instantaneous information, while the second one relies on longer- term knowledge gained through a beaconing procedure. Using a realistic microscopic model to represent the vehicular traffic flow, we investigate the performance of the above broadcast algorithm classes by simulation. In order to explore a very large algorithmic design space, we devise a Convex Hull framework that allows us to effectively compare and compactly present the boundaries of the solution space for each algorithm class. By the use of such framework we show that the beaconing approach is not justified for broadcast in suburban and highway VANETs, as there is no performance gain that justifies the complexity entailed by the beaconing procedure.
Dario Rossi 0001, Roberta Fracchia, Michela Meo
ICC1
2008 Tracking Down Skype Traffic
abstract
Skype is beyond any doubt the most popular VoIP application in the current Internet application spectrum. Its amazing success drawn the attention of telecom operators and the research community, both interested in knowing Skype's internal mechanisms, characterizing traffic and understanding users' behavior. We dissect the following fundamental components: data traffic generated by voice and video communication, and signaling traffic generated by Skype. We use both active and passive measurement techniques to gather a deep understanding on the traffic Skype generates. From extensive testbed experiments, we devise a source model which takes into account: (i) the service type, i.e., voice or video calls (ii) the selected source Codec, (iii) the adopted transport-layer protocol, and (iv) network conditions. Furthermore, leveraging on the use of an accurate Skype classification engine that we recently proposed, we study and characterize Skype traffic based on extensive passive measurements collected from our campus LAN.
Dario Bonfiglio, Marco Mellia, Michela Meo, Nicolo Ritacca, Dario Rossi 0001
INFOCOM5
2008 Passive analysis of TCP anomalies
Marco Mellia, Michela Meo, Luca Muscariello, Dario Rossi 0001
Comput. Networks4
2007 Understanding VoIP from Backbone Measurements
abstract
VoIP has widely been addressed as the technology that will change the Telecommunication model opening the path for convergence. Still today this revolution is far from being complete, since the majority of telephone calls are originated by circuit-oriented networks. In this paper for the first time to the best of our knowledge, we present a large dataset of measurements collected from the FastWeb backbone, which is one of the first worldwide Telecom operator to offer VoIP and high-speed data access to the end-user. Traffic characterization will focus on several layers, focusing on both end-user and ISP perspective. In particular, we highlight that, among loss, delay and jitter, only the first index may affect the VoIP call quality. Results show that the technology is mature to make the final step, allowing the integration of data and real-time services over the Internet.
Robert Birke, Marco Mellia, Michele Petracca, Dario Rossi 0001
INFOCOM4
2007 Revealing skype traffic: when randomness plays with you
abstract
Skype is a very popular VoIP software which has recently attracted the attention of the research community and network operators. Following a closed source and proprietary design, Skype protocols and algorithms are unknown. Moreover, strong encryption mechanisms are adopted by Skype, making it very difficult to even glimpse its presence from a traffic aggregate. In this paper, we propose a framework based on two complementary techniques to reveal Skypetraffic in real time. The first approach, based on Pearson'sChi-Square test and agnostic to VoIP-related trafficcharacteristics, is used to detect Skype's fingerprint from the packet framing structure, exploiting the randomness introduced at the bit level by the encryption process. Conversely, the second approach is based on a stochastic characterization of Skype traffic in terms of packet arrival rate and packet length, which are used as features of a decision process based on Naive Bayesian Classifiers.In order to assess the effectiveness of the above techniques, we develop an off-line cross-checking heuristic based on deep-packet inspection and flow correlation, which is interesting per se. This heuristic allows us to quantify the amount of false negatives and false positives gathered by means of the two proposed approaches: results obtained from measurements in different networks show that the technique is very effective in identifying Skype traffic. While both Bayesian classifier and packet inspection techniques are commonly used, the idea of leveraging on randomness to reveal traffic is novel. We adopt this to identify Skype traffic, but the same methodology can be applied to other classification problems as well.
Dario Bonfiglio, Marco Mellia, Michela Meo, Dario Rossi 0001, Paolo Tofanelli
SIGCOMM4
2006 Passive Identification and Analysis of TCP Anomalies
abstract
In this paper we focus on passive measurements of TCP traffic, main component of nowadays traffic. We propose a heuristic technique for the classification of the anomalies that may occur during the lifetime of a TCP flow, such as out-of-sequence and duplicate segments. Since TCP is a closed-loop protocol that infers network conditions by means of losses and reacts accordingly, the possibility of carefully distinguishing the causes of anomalies in TCP traffic is very appealing, since it may be instrumental to the deep understanding of TCP behavior in real environments and to protocol engineering as well. We apply the proposed heuristic to traffic traces collected at both networks edges and backbone links. By studying the statistical properties of TCP anomalies, we find that their aggregate exhibits Long Range Dependence phenomena, but that anomalies suffered by individual long-lived flows are on the contrary uncorrelated. Interestingly, no dependence to the actual link load is observed.
Marco Mellia, Michela Meo, Luca Muscariello, Dario Rossi 0001
ICC4
2006 Real-Time TCP/IP Analysis with Common Hardware
abstract
Traffic measurement represents an indispensable and valuable tool for the analysis of nowadays telecommunication networks. Moreover, it is desirable for traffic measurement and analysis to be both continuous and persistent, since only these joint requirements allow to track important changes on the traffic pattern. On the other hand, transmission links bandwidth keep improving, at a seemingly inexorable rate: therefore, the analysis of the traffic is becoming more complex than ever. This paper focuses on the description and the benchmarking of a network traffic analyzer, called Tstat, able to process real-time traffic further providing i) several advanced measurement indexes of transport layer protocols and ii) ever-lasting monitoring capabilities. Particularly, our aim is to assess what kind of links, and under which load, can be continuously and persistently monitored without compromising the complexity of the traffic analysis that has to be performed.
Dario Rossi 0001, Marco Mellia
ICC1
2005 Gambling heuristic on a chord ring
abstract
Chord routing is greedy and non-symmetric, and is based on a skiplist-like data structure, whose entries are known as fingers. This work explores the benefits arising from a modified greedy lookup strategy that, without introducing any additional communication overhead, simply exploits the implicit symmetry knowledge intrinsic to the highly structured Chord ring. Through extensive simulation on a dynamic peer environment, we show a practical and feasible solution that actually boosts DHT lookup performance under a wide range of scenarios.
Dario Rossi 0001, Ion Stoica
GLOBECOM1
2004 On the properties of TCP flow arrival process
abstract
We study the TCP flow arrival process, starting from the aggregated measurement at the TCP flow level taken from our campus network. In particular, we analyze the statistical properties of the TCP flow arrival process. We define the different traffic aggregates by splitting the original trace, such that i) each of them is constituted by all the TCP flows belonging to the same traffic relation, i.e., with the same source/destination IP addresses and ii) each traffic aggregate has, bytewise, the same amount of traffic. To induce a divisions of TCP-elephants and TCP-mice into different traffic aggregates, the used algorithm packs the largest traffic relations in the first traffic aggregates, so that subsequently generated aggregates are constituted by an increasing number of smaller traffic relations. The long range dependency (LRD) characteristics are presented, showing the possible causes of the LRD of TCP flow arrival process in i) the heavy tailed distribution of the number of flows in a traffic aggregate, and ii) the presence of TCP-elephants within them.
Dario Rossi 0001, Luca Muscariello, Marco Mellia
ICC1
2003 User patience and the Web: a hands-on investigation
abstract
We present a study of Web user behavior when network performance decreases causing an increase of page transfer times. Real traffic measurements are analyzed to infer whether worsening network conditions translate into greater impatience by the user, which translates in early interruption of TCP connections. Several parameters are studied in order to gather their impact on the interruption probability on Web transfers: time of day, file size, throughput and time elapsed since the beginning of the download. From the results presented, we try to paint a picture of the complex interactions between user perception of the Web and network-level events.
Dario Rossi 0001, Marco Mellia, Claudio Casetti
GLOBECOM1
2002 A simulation study of Web traffic over DiffServ networks
abstract
We present a simulation study of HTTP traffic crossing a DiffServ domain. We consider both the cases where the reserved bandwidth is not exceeded by the offered traffic (overprovisioning) and where the assured traffic competes with the classic best effort class (underprovisioning). The reported simulation shows that the DiffServ approach is able to protect the assured flows in the first case, while the performance benefits are tighter in the second case, in which fairness issues arise between long and short-lived flows.
Dario Rossi 0001, Claudio Casetti, Marco Mellia
GLOBECOM1