VLDB 2026 Research / reviewers in the wild / expert
Damiano Carra
dblp:72/4280
· DBLP profile ↗
49ranked-venue papers
23as first author
12since 2021 · last 2026
0000-0002-3467-1166ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 26 · 16 first-author · 6 since 2021Systems, architecture and hardware · 13 · 6 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A multimodal and perturbation-aware learning approach for robust traffic classificationabstractTraffic Classification (TC) is pivotal for network management, cybersecurity, and Quality of Experience (QoE) monitoring. However, while Deep Learning (DL) has significantly advanced TC, most existing works assume static, idealized conditions, overlooking key challenges of real-world deployments—such as traffic variability, routing asymmetries, out-of-order packet arrivals, and partial visibility at the Vantage Points (VPs). This motivates the need for robustness evaluations under such scenarios. In this work, we investigate the robustness of state-of-the-art (SOTA) TC models under realistic, yet controlled, perturbation scenarios. Specifically, we introduce novel, model-agnostic traffic perturbations—simulating time jitter, retransmissions, and partial visibility—to reflect conditions commonly encountered in live network traffic. We evaluate our approach on three public datasets—i.e., VPN-16 , MIRAGE-19 , and MIRAGE-24 —and show how Mimetic-Enhanced , a multimodal model, tends to outperform two representative single-modal counterparts both in terms of TC effectiveness on clean traffic and robustness under perturbations. Nonetheless, our analysis also reveals that multimodal models remain vulnerable under specific perturbation settings. To address this limitation, we propose a model-agnostic perturbation-aware training framework based on Supervised Data Augmentation ( Aug ) and Contrastive Learning ( CL )—considering both self-supervised and supervised variants. Unlike architecture-specific solutions, our approach operates at the learning strategy level , allowing it to be seamlessly applied to diverse classifiers without requiring structural modifications. Adopting Mimetic-Enhanced as a primary multimodal case study, we integrate the proposed strategies into its two-stage training pipeline. Experimental results demonstrate that perturbation-aware training not only improves TC effectiveness on clean (i.e., unperturbed) traffic—particularly when applied across both training stages—but also significantly strengthens the model’s robustness under diverse and realistic perturbation scenarios. Furthermore, we investigate Out-of-Distribution (OOD) detection, model calibration, and TC effectiveness in low-data regimes. Finally, we explicitly demonstrate the framework’s generalizability by validating it on other SOTA architectures, spanning both single- and multi-modal approaches. Idio Guarino, Giampaolo Bovenzi, Alfredo Nascita, Domenico Ciuonzo, Damiano Carra, Antonio Pescapè |
Comput. Networks | 5 |
| 2026 | A survey on CSI-based Wi-Fi sensing datasets and models with a focus on reproducibilityabstractWi-Fi sensing based on Channel State Information (CSI) has witnessed considerable research activity in recent years. However, a critical literature analysis reveals that only a limited amount of proposals are potentially reproducible, with many works lacking essential experimental details, publicly available datasets, or accessible analysis code. This may impede the research progress and the subsequent transition of promising findings into practical applications. The objective of this work is to identify CSI-based sensing proposals that are potentially reproducible based on the published information. Our goal is to provide a focused review of resources that can serve as a concrete starting point for researchers and practitioners seeking to experiment with and advance the field of Wi-Fi sensing. We perform a comprehensive analysis of publicly available datasets (encompassing both the collection methodologies and the environmental characteristics) and existing sensing models, accompanied by their code, pre-processing steps, and evaluation procedures. Finally, we discuss what are the minimum requirements for truly verifiable contributions in this field, and outline the best practices for creating and sharing reproducible CSI-based sensing datasets and models. Idio Guarino, Damiano Carra, Marco Cominelli, Francesco Gringoli, Renato Lo Cigno |
Comput. Commun. | 2 |
| 2025 | Probabilistic Resource Sharing in Cloud Caches with Heterogeneous Object SizesabstractIn-memory key-value stores are critical caching infrastructure for numerous cloud services. Unlike traditional CPU caches that often assume uniform item sizes, cloud caches frequently handle heterogeneous-sized objects, introducing significant challenges in cache management, particularly in shared multi-tenant environments. Existing cache sharing solutions designed for uniform-sized objects are often not optimized for these scenarios.This paper presents a probabilistic cache sharing scheme that dynamically adapts eviction probabilities across different traffic classes based on their performance. Our approach redistributes cache space at eviction events, reclaiming space probabilistically from one class to serve the needs of another. Our scheme operates independently of the underlying per-class eviction policies and adapts to time-varying traffic patterns. We evaluate our approach using real-world traces with heterogeneous object sizes, demonstrating its effectiveness in dynamically allocating cache space and improving overall cache performance in shared environments. Lorenzo Marini, Damiano Carra |
IC2E | 2 |
| 2025 | Low-Complexity online learning for caching
Damiano Carra, Giovanni Neglia |
Comput. Networks | 1 |
| 2024 | Can Miss Ratio Curves Predict Cache Performance?abstractMiss Ratio Curves (MRCs) serve as a widely recognized tool for cache profiling, offering a comprehensive visualization of the relationship between cache size and miss ratio. This enables users to assess the cost of storage and estimate the impact of cache misses effectively. However, MRCs are constructed based on past requests, raising questions about their suitability for what-if analysis in predicting future cache occupancy. In this work, we aim to evaluate the effectiveness of MRCs in predicting cache performance. To achieve this, we explore the influence of the interval during which requests are collected and establish metrics for quantifying the error between predictions and actual outcomes. Our findings highlight that predictive whatif analysis remains an open problem that necessitates thorough investigation and exploration. Lorenzo Marini, Damiano Carra |
IC2E | 2 |
| 2023 | DNN Split Computing: Quantization and Run-Length Coding are EnoughabstractSplit computing, a recently developed paradigm, capitalizes on the computational resources of end devices to enhance the inference efficiency in machine learning (ML) applications. This approach involves the end device processing input data and transmitting intermediate results to a cloud server, which then completes the inference computation. While the main goals of split computing are to reduce latency, minimize energy consumption, and decrease data transfer overhead, minimizing data transmission time remains a challenge. Many existing strategies involve modifying the ML model architecture which ultimately requires resource-intensive retraining. In our work, we explore lossless and lossy techniques to encode intermediate results without modifying the ML model. Concentrating on image classification and object detection-two prevalent ML applications-we assess the advantages and limitations of each technique. Our findings indicate that simple tools, such as linear quantization and run-length encoding, already accomplish considerable information reduction, which is on par with more complex state-of-the-art techniques that necessitate model retraining. These tools are computationally efficient and do not burden the end device. Damiano Carra, Giovanni Neglia |
GLOBECOM | 1 |
| 2023 | A Flexible Heuristic to Schedule Distributed Analytic Applications in Compute ClustersabstractThis work addresses the problem of scheduling user-defined analytic applications, which we define as high-level compositions of frameworks, their components, and the logic necessary to carry out work. The key idea in our application definition, is to distinguish classes of components, including core and elastic types: the first being required for an application to make progress, the latter contributing to reduced execution times. We show that the problem of scheduling such applications poses new challenges, which existing approaches address inefficiently. Francesco Pace, Daniele Venzano, Damiano Carra, Pietro Michiardi |
IEEE Trans. Cloud Comput. | 3 |
| 2023 | Ascent Similarity Caching With Approximate IndexesabstractSimilarity search is a key operation in multimedia retrieval systems and recommender systems, and it will play an important role also for future machine learning and augmented reality applications. When these systems need to serve large objects with tight delay constraints, edge servers close to the end-user can operate as similarity caches to speed up the retrieval. In this paper we present AÇAI, a new similarity caching policy which improves on the state of the art by using (i) an (approximate) index for the whole catalog to decide which objects to serve locally and which to retrieve from the remote server, and (ii) a mirror ascent algorithm to update the set of local objects with strong guarantees even when the request process does not exhibit any statistical regularity. Tareq Si Salem, Giovanni Neglia, Damiano Carra |
IEEE/ACM Trans. Netw. | 3 |
| 2022 | On the Impact of Transport Times in Flexible Job Shop Scheduling ProblemsabstractManufacturing systems require a careful scheduling of the resource usage to maximize the production efficiency. In a completely automated environment, the transport system should be orchestrated to work smoothly with the other resources. While the impact of job characteristics, such as fixed or variable processing times of the tasks composing the jobs, or task dependencies, has been extensively studied, the role of the transport system has received less attention.In this paper we consider a conveyor belt as a mean of transportation among a set of production machines. In this scenario, there is no input or output buffer at the machines, and the transport times depend on the availability of the machines. We propose a heuristic based on randomization, called SCHED-T, which is able to find a near optimal joint schedule for job processing and transfer in few seconds. We test our solution on known benchmarks, along with real-world instances, showing that our scheduler is able to predict accurately the overall processing time of a production line. Sebastiano Gaiardelli, Damiano Carra, Stefano Spellini, Franco Fummi |
ETFA | 2 |
| 2022 | I-SPLIT: Deep Network Interpretability for Split ComputingabstractThis work makes a substantial step in the field of split computing, i.e., how to split a deep neural network to host its early part on an embedded device and the rest on a server. So far, potential split locations have been identified exploiting uniquely architectural aspects, i.e., based on the layer sizes. Under this paradigm, the efficacy of the split in terms of accuracy can be evaluated only after having performed the split and retrained the entire pipeline, making an exhaustive evaluation of all the plausible splitting points prohibitive in terms of time. Here we show that not only the architecture of the layers does matter, but the importance of the neurons contained therein too. A neuron is important if its gradient with respect to the correct class decision is high. It follows that a split should be applied right after a layer with a high density of important neurons, in order to preserve the information flowing until then. Upon this idea, we propose Interpretable Split (I-SPLIT): a procedure that identifies the most suitable splitting points by providing a reliable prediction on how well this split will perform in terms of classification accuracy, beforehand of its effective implementation. As a further major contribution of I-SPLIT, we show that the best choice for the splitting point on a multiclass categorization problem depends also on which specific classes the network has to deal with. Exhaustive experiments have been carried out on two networks, VGG16 and ResNet-50, and three datasets, Tiny-Imagenet-200, notMNIST, and Chest X-Ray Pneumonia. The source code is available at https://github.com/vips4/I-Split. Federico Cunico, Luigi Capogrosso, Francesco Setti, Damiano Carra, Franco Fummi, Marco Cristani |
ICPR | 4 |
| 2022 | Sequence recommendations for groups: A dynamic approach to balance preferences
Sara Migliorini 0001, Elisa Quintarelli, Mauro Gambini, Alberto Belussi, Damiano Carra |
Inf. Syst. | 5 |
| 2021 | Taking two Birds with one k-NN Cacheabstractk-Nearest Neighbors aims at efficiently finding items close to a query in a large collection of objects, and it is used in different applications, from image retrieval to recommendation. These applications achieve high throughput combining two different elements: 1) approximate nearest neighbours searches that reduce the complexity at the cost of providing inexact answers and 2) caches that store the most popular items. In this paper we propose to combine the approximate index for the whole catalog with a more precise index for the items stored in the cache. Our experiments on realistic traces show that this approach is doubly advantageous as it 1) improves the quality of the final answer provided to a query, 2) additionally reduces the service latency. Damiano Carra, Giovanni Neglia |
GLOBECOM | 1 |
| 2020 | A Context-based Approach for Partitioning Big Data
Sara Migliorini 0001, Alberto Belussi, Elisa Quintarelli, Damiano Carra |
EDBT | 4 |
| 2020 | Efficient Miss Ratio Curve Computation for Heterogeneous Content Popularity
Damiano Carra, Giovanni Neglia |
USENIX ATC | 1 |
| 2020 | Elastic Provisioning of Cloud Caches: A Cost-Aware TTL ApproachabstractWe consider elastic resource provisioning in the cloud, focusing on in-memory key-value stores used as caches. Our goal is to dynamically scale resources to the traffic pattern minimizing the overall cost, which includes not only the storage cost, but also the cost due to misses. In fact, a small variation of the cache miss ratio may have a significant impact on user perceived performance in modern web services, which in turn has an impact on the overall revenues for the content provider using such services. We propose and study a dynamic algorithm for TTL caches, which is able to obtain close-to-minimal costs. Since high-throughput caches require low complexity operations, we discuss a practical implementation of such a scheme requiring constant overhead per request independently from the cache size. We evaluate our solution with real-world traces collected from Akamai, and show that the TTL approach is able to track the optimal cache configuration and achieve significant cost savings specially in highly dynamic settings that are likely to require elastic cloud services. Damiano Carra, Giovanni Neglia, Pietro Michiardi |
IEEE/ACM Trans. Netw. | 1 |
| 2019 | TTL-based Cloud CachesabstractWe consider in-memory key-value stores used as caches, and their elastic provisioning in the cloud. The cost associated to such caches not only includes the storage cost, but also the cost due to misses: in fact, the cache miss ratio has a direct impact on the performance perceived by end users, and this directly affects the overall revenues for content providers. Our aim is to adapt dynamically the number of caches based on the traffic pattern, to minimize the overall costs.We present a dynamic algorithm for TTL caches whose goal is to obtain close-to-minimal costs. We then propose a practical implementation with limited computational complexity: our scheme requires constant overhead per request independently from the cache size. Using real-world traces collected from the Akamai content delivery network, we show that our solution achieves significant cost savings specially in highly dynamic settings that are likely to require elastic cloud services. Damiano Carra, Giovanni Neglia, Pietro Michiardi |
INFOCOM | 1 |
| 2019 | Memory Partitioning and Management in MemcachedabstractMemcached is a popular component of modern Web architectures, which allows fast response times-a fundamental performance index for measuring the Quality of Experience of end-users-for serving popular objects. In this work, we study how memory partitioning in Memcached works and how it affects system performance in terms of hit ratio. We first present a cost-based memory partitioning and management mechanism for Memcached that is able to dynamically adapt to user requests and manage the memory according to both object sizes and costs. We present a comparative analysis of the vanilla memory management scheme of Memcached and our approach, using real traces from a major content delivery network operator. We show that our proposed memory management scheme achieves near-optimal performance, striking a good balance between the performance perceived by end-users and the pressure imposed on back-end servers. We then consider the problem known as “calcification”: Memcached divides the memory into different classes proportionally to the percentage of requests for objects of different sizes. Once all the available memory has been allocated, reallocation is not possible or limited. Using synthetic traces, we show the negative impact of calcification on the hit ratio with Memcached, while our scheme, thanks to its adaptivity, is able to solve the calcification problem, achieving near-optimal performance. Damiano Carra, Pietro Michiardi |
IEEE Trans. Serv. Comput. | 1 |
| 2018 | Elastic Provisioning of Cloud Caches: a Cost-aware TTL ApproachabstractNo abstract available. Damiano Carra, Giovanni Neglia, Pietro Michiardi |
SoCC | 1 |
| 2018 | Data-Driven Resource Shaping for Compute ClustersabstractNo abstract available. Francesco Pace, Dimitrios Milios, Damiano Carra, Daniele Venzano, Pietro Michiardi |
SoCC | 3 |
| 2018 | Cache Policies for Linear Utility MaximizationabstractCache policies to minimize the content retrieval cost have been studied through competitive analysis when the miss costs are additive and the sequence of content requests is arbitrary. More recently, a cache utility maximization problem has been introduced, where contents have stationary popularities and utilities are strictly concave in the hit rates. This paper bridges the two formulations, considering linear costs and content popularities. We show that minimizing the retrieval cost corresponds to solving an online knapsack problem, and we propose new dynamic policies inspired by simulated annealing, including DynqLRU, a variant of qLRU. We prove that DynqLRU asymptotically asymptotic converges to the optimum under the characteristic time approximation. In a real scenario, popularities vary over time and their estimation is very difficult. DynqLRU does not require popularity estimation, and our realistic, trace-driven evaluation shows that it significantly outperforms state-of-the-art policies, with up to 45% cost reduction. Giovanni Neglia, Damiano Carra, Pietro Michiardi |
IEEE/ACM Trans. Netw. | 2 |
| 2017 | Flexible Scheduling of Distributed Analytic ApplicationsabstractThis work addresses the problem of scheduling user-defined analytic applications, which we define as high-level compositions of frameworks, their components, and the logic necessary to carry out work. The key idea in our application definition, is to distinguish classes of components, including rigid and elastic types: the first being required for an application to make progress, the latter contributing to reduced execution times. We show that the problem of scheduling such applications poses new challenges, which existing approaches address inefficiently. Thus, we present the design and evaluation of a novel, flexible heuristic to schedule analytic applications, that aims at high system responsiveness, by allocating resources efficiently. Our algorithm is evaluated using trace-driven simulations, with large-scale real system traces: our flexible scheduler outperforms a baseline approach across a variety of metrics, including application turnaround times, and resource allocation efficiency. We also present the design and evaluation of a full-fledged system, which we have called Zoe, that incorporates the ideas presented in this paper, and report concrete improvements in terms of efficiency and performance, with respect to prior generations of our system. Francesco Pace, Daniele Venzano, Damiano Carra, Pietro Michiardi |
CCGrid | 3 |
| 2017 | Cache policies for linear utility maximizationabstractCache policies to minimize the content retrieval cost have been studied through competitive analysis when the miss costs are additive and the sequence of content requests is arbitrary. More recently, a cache utility maximization problem has been introduced, where contents have stationary popularities and utilities are strictly concave in the hit rates. This paper bridges the two formulations, considering linear costs and content popularities. We show that minimizing the retrieval cost corresponds to solving an online knapsack problem, and we propose new dynamic policies inspired by simulated annealing, including DynqLRU, a variant of qLRU. For such policies we prove asymptotic convergence to the optimum under the characteristic time approximation. In a real scenario, popularities vary over time and their estimation is very difficult. DynqLRU does not require popularity estimation, and our realistic, trace-driven evaluation shows that it significantly outperforms state-of-the-art policies, with up to 45% cost reduction. Giovanni Neglia, Damiano Carra, Pietro Michiardi |
INFOCOM | 2 |
| 2017 | HFSP: Bringing Size-Based Scheduling To HadoopabstractSize-based scheduling with aging has been recognized as an effective approach to guarantee fairness and near-optimal system response times. We present HFSP, a scheduler introducing this technique to a real, multi-server, complex, and widely used system such as Hadoop. Size-based scheduling requiresa priorijob size information, which is not available in Hadoop: HFSP builds such knowledge by estimating it on-line during job execution. Our experiments, which are based on realistic workloads generated via a standard benchmarking suite, pinpoint at a significant decrease in system response times with respect to the widely used Hadoop Fair scheduler, without impacting the fairness of the scheduler, and show that HFSP is largely tolerant to job size estimation errors. Mario Pastorelli, Damiano Carra, Matteo Dell'Amico, Pietro Michiardi |
IEEE Trans. Cloud Comput. | 2 |
| 2016 | Experimental Performance Evaluation of Cloud-Based Analytics-as-a-ServiceabstractAn increasing number of (AaaS) solutions has recently seen the light, in the landscape of cloud-based services. These services allow flexible composition of compute and storage components, that create powerful data ingestion and processing pipelines. This work is a first attempt at an experimental evaluation of analytic application performance executed using a wide range of storage service configurations. We present an intuitive notion of data locality, that we use as a proxy to rank different service compositions in terms of expected performance. Through an empirical analysis, we dissect the performance achieved by analytic workloads and unveil problems due to the impedance mismatch that arise in some configurations. Our work paves the way to a better understanding of modern cloud-based analytic services and their performance, both for its end-users and their providers. Francesco Pace, Marco Milanesio, Daniele Venzano, Damiano Carra, Pietro Michiardi |
CLOUD | 4 |
| 2016 | Google Dorks: Analysis, Creation, and New Defenses
Flavio Toffalini, Maurizio Abbà, Damiano Carra, Davide Balzarotti |
DIMVA | 3 |
| 2016 | PSBS: Practical Size-Based SchedulingabstractSize-based schedulers have very desirable performance properties: optimal or near-optimal response time can be coupled with strong fairness. Despite this, however, such systems are rarely implemented in practical settings, because they require knowinga priorithe amount of work needed to complete jobs: this assumption is difficult to satisfy in concrete systems. It is definitely more likely to inform the system with anestimateof the job sizes, but existing studies point to somewhat pessimistic results if size-based policies use imprecise job size estimations. We take the goal of designing scheduling policies thatexplicitly deal with inexact job sizes. First, we prove that, in the absence of errors, it is always possible to improve any scheduling policy by designing a size-based one thatdominatesit: in the new policy,no jobswill complete later than in the original one. Unfortunately, size-based schedulers can perform badly with inexact job size information when job sizes are heavily skewed; we show that this issue, and the pessimistic results shown in the literature, are due to problematic behavior when large jobs are underestimated. Once the problem is identified, it is possible to amend size-based schedulers to solve the issue. We generalize FSP—a fair and efficient size-based scheduling policy—to solve the problem highlighted above; in addition, our solution deals with different job weights (that can be assigned to a job independently from its size). We provide an efficient implementation of the resulting protocol, which we callPractical Size-Based Scheduler(PSBS). Through simulations evaluated on synthetic and real workloads, we show that PSBS has near-optimal performance in a large variety of cases with inaccurate size information, that it performs fairly and that it handles job weights correctly. We believe that this work shows that PSBS is indeed pratical, and we maintain that it could inspire the design of schedulers in a wide array of real-world use cases. Matteo Dell'Amico, Damiano Carra, Pietro Michiardi |
IEEE Trans. Computers | 2 |
| 2014 | Memory partitioning in Memcached: An experimental performance analysisabstractMemcached is a popular component of modern Web architectures, which allows fast response times - a fundamental performance index for measuring the Quality of Experience of end-users - for serving popular objects. In this work, we study how memory partitioning in Memcached works and how it affects system performance in terms of hit rate. Memcached divides the memory into different classes proportionally to the percentage of requests for objects of different sizes. Once all the available memory has been allocated, reallocation is not possible or limited, a problem called “calcification”. Calcification constitutes a symptom indicating that current memory partitioning mechanisms require a more careful design. Using an experimental approach, we show the negative impact of calcification on an important performance metric, the hit rate. We then proceed to design and implement a new memory partitioning scheme, called PSA, which replaces that of vanilla Memcached. With PSA, Memcached achieves a higher hit rate than what is obtained with the default memory partitioning mechanism, even in the absence of calcification. Moreover, we show that PSA is capable of “adapting” to the dynamics of clients' requests and object size distributions, thus defeating the calcification problem. Damiano Carra, Pietro Michiardi |
ICC | 1 |
| 2014 | Revisiting Size-Based Scheduling with Estimated Job SizesabstractWe study size-based schedulers, and focus on the impact of inaccurate job size information on response time and fairness. Our intent is to revisit previous results, which allude to performance degradation for even small errors on job size estimates, thus limiting the applicability of size-based schedulers. We show that scheduling performance is tightly connected to workload characteristics: in the absence of large skew in the job size distribution, even extremely imprecise estimates suffice to outperform size-oblivious disciplines. Instead, when job sizes are heavily skewed, known size-based disciplines suffer. In this context, we show - for the first time - the dichotomy of over-estimation versus under-estimation. The former is, in general, less problematic than the latter, as its effects are localized to individual jobs. Instead, under-estimation leads to severe problems that may affect a large number of jobs. We present an approach to mitigate these problems: our technique requires no complex modifications to original scheduling policies and performs very well. To support our claim, we proceed with a simulation-based evaluation that covers an unprecedented large parameter space, which takes into account a variety of synthetic and real workloads. As a consequence, we show that size-based scheduling is practical and outperforms alternatives in a wide array of use-cases, even in presence of inaccurate size information. Matteo Dell'Amico, Damiano Carra, Mario Pastorelli, Pietro Michiardi |
MASCOTS | 2 |
| 2013 | HFSP: Size-based scheduling for HadoopabstractSize-based scheduling with aging has, for long, been recognized as an effective approach to guarantee fairness and near-optimal system response times. We present HFSP, a scheduler introducing this technique to a real, multi-server, complex and widely used system such as Hadoop. Size-based scheduling requires a priori job size information, which is not available in Hadoop: HFSP builds such knowledge by estimating it on-line during job execution. Our experiments, which are based on realistic workloads generated via a standard benchmarking suite, pinpoint at a significant decrease in system response times with respect to the widely used Hadoop Fair scheduler, and show that HFSP is largely tolerant to job size estimation errors. Mario Pastorelli, Antonio Barbuzzi, Damiano Carra, Matteo Dell'Amico, Pietro Michiardi |
IEEE BigData | 3 |
| 2013 | Topic 7: Peer-to-Peer Computing - (Introduction)
Damiano Carra, Thorsten Strufe, György Dán, Marcel Karnstedt |
Euro-Par | 1 |
| 2013 | On the Impact of Incentives in eMule {Analysis and Measurements of a Popular File-Sharing Application}abstractMotivated by the popularity of content distribution and file sharing applications that nowadays dominate Internet traffic, we focus on the incentive mechanism of a very popular, yet not very well studied, peer-to-peer application, eMule. In our work, we recognize that the incentive scheme of eMule is more sophisticated than current alternatives (e.g., BitTorrent) as it uses a general, priority-based, time-dependent queuing discipline to differentiate service among cooperative users and free-riders. In this paper, we describe a general model of such an incentive mechanism and analyze its properties in terms of application performance. We validate our model using both numerical simulations (when analytical techniques become prohibitive) and with a measurement campaign of the live eMule system. Our results, in addition to validating our model, indicate that the incentive scheme of eMule suffers from starvation. Therefore, we present an alternative scheme that mitigates this problem, and validate it through numerical simulations and a second measurement campaign. Damiano Carra, Pietro Michiardi, Hani Salah, Thorsten Strufe |
IEEE J. Sel. Areas Commun. | 1 |
| 2013 | Characterization and Management of Popular Content in KADabstractThe endeavor of this work is to study the impact of content popularity in a large-scale Peer-to-Peer network, namely KAD. Based on an extensive measurement campaign, we pinpoint several deficiencies of KAD in handling popular content and provide a series of improvements to address such shortcomings. Our work reveals that keywords, which are associated with content, may become popular for two distinct reasons. First, we show that some keywords are intrinsically popular because they are common to many disparate contents: in such case we ameliorate KAD by introducing a simple mechanism that identifies stopwords. Then, we focus on keyword popularity that directly relates to popular content. We design and evaluate an adaptive load balancing mechanism that is backward compatible with the original implementation of KAD. Our scheme features the following properties: 1) it drives the process that selects the location of peers responsible to store references to objects, based on object popularity; 2) it solves problems related to saturated peers that would otherwise inflict a significant drop in the diversity of references to objects, and 3) if coupled with a load-aware content search procedure, it allows for a more fair and efficient usage of peer resources. Damiano Carra, Moritz Steiner, Pietro Michiardi, Ernst W. Biersack, Wolfgang Effelsberg, Taoufik En-Najjary |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2012 | Peer-assisted content distribution on a budget
Pietro Michiardi, Damiano Carra, Francesco Albanese, Azer Bestavros |
Comput. Networks | 2 |
| 2011 | Adaptive load balancing in KADabstractThe endeavor of this work is to study the impact of content popularity in a large-scale Peer-to-Peer network, namely KAD. Armed with the insights gained from an extensive measurement campaign, which pinpoints several deficiencies of the present KAD design in handling popular objects, we set off to design and evaluate an adaptive load balancing mechanism. Our mechanism is backward compatible with KAD, as it only modifies its inner algorithms, and presents several desirable properties: (i) it drives the process that selects the number and location of peers responsible to store references to objects, based on their popularity; (ii) it solves problems related to saturated peers, that entail a significant drop in the diversity of references to objects, and (iii) if coupled with an enhanced content search procedure, it allows a more fair and efficient usage of peer resources, at a reasonable cost. Our evaluation uses a trace-driven simulator that features realistic peer churn and a precise implementation of the inner components of KAD. Damiano Carra, Moritz Steiner, Pietro Michiardi |
Peer-to-Peer Computing | 1 |
| 2011 | On the Robustness of BitTorrent Swarms to Greedy PeersabstractThe success of BitTorrent has fostered the development of variants to its basic components. Some of the variants adopt greedy approaches aiming at exploiting the intrinsic altruism of the original version of BitTorrent in order to maximize the benefit of participating to a torrent. In this work, we study BitTyrant, a recently proposed strategic client. BitTyrant tries to determine the exact amount of contribution necessary to maximize its download rate by dynamically adapting and shaping the upload rate allocated to its neighbors. We evaluate in detail the various mechanisms used by BitTyrant to identify their contribution to the performance of the client. Our findings indicate that the performance gain is due to the increased number of connections established by a BitTyrant client, rather than to its subtle uplink allocation algorithm; surprisingly, BitTyrant reveals to be altruistic and particularly efficient in disseminating the content, especially during the initial phase of the distribution process. The possible gain of a single BitTyrant client, however, disappears in the case of a widespread adoption: our results indicate a severe loss of efficiency that we analyze in detail. Damiano Carra, Giovanni Neglia, Pietro Michiardi, Francesco Albanese |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2010 | Passive Online RTT Estimation for Flow-Aware Routers Using One-Way Traffic
Damiano Carra, Konstantin Avrachenkov, Sara Alouf, Alberto Blanc, Philippe Nain, Georg Post |
Networking | 1 |
| 2010 | Evaluating and improving the content access in KAD
Moritz Steiner, Damiano Carra, Ernst W. Biersack |
Peer-to-Peer Netw. Appl. | 2 |
| 2009 | Seed Scheduling for Peer-to-Peer NetworksabstractThe initial phase in a content distribution (file sharing) scenario is delicate due to the lack of global knowledge and the dynamics of the overlay. An unwise distribution of the pieces in this phase can cause delays in reaching steady state, thus increasing file download times. We devise a scheduling algorithm at the seed (source peer with full content), based on a proportional fair approach, and we implement it on a real file sharing client. In dynamic overlays, our solution improves by up to 25% the average downloading time of a standard protocol ala BitTorrent. Flavio Esposito, Abraham Matta, Pietro Michiardi, Nobuyuki Mitsutake, Damiano Carra |
NCA | 5 |
| 2008 | Uplink allocation beyond choke/unchoke: or how to divide and conquer bestabstractMotivated by emerging cooperative P2P applications we study new uplink allocation algorithms for substituting the rate-based choke/unchoke algorithm of BitTorrent which was developed for non-cooperative environments. Our goal is to shorten the download times by improving the uplink utilization of nodes. We develop a new family of uplink allocation algorithms which we call BitMax, to stress the fact that they allocate to each unchoked node the maximum rate it can sustain, instead of an 1/(k + 1) equal share as done in the existing BitTorrent. BitMax computes in each interval the number of nodes to be unchoked, and the corresponding allocations, and thus does not require any empirically preset parameters like k. We demonstrate experimentally that Bit-Max can reduce significantly the download times in a typical reference scenario involving mostly ADSL nodes. We also consider scenarios involving network bottlenecks caused by filtering of P2P traffic at ISP peering points and show that BitMax retains its gains also in these cases. Nikolaos Laoutaris, Damiano Carra, Pietro Michiardi |
CoNEXT | 2 |
| 2008 | On the Impact of Greedy Strategies in BitTorrent Networks: The Case of BitTyrantabstractThe success of BitTorrent has fostered the development of variants to its basic components. Some of the variants adopt greedy approaches aiming at exploiting the intrinsic altruism of the original version of BitTorrent in order to maximize the benefit of participating to a torrent. In this work we study BitTyrant, a recently proposed strategic client. BitTyrant tries to determine the exact amount of contribution necessary to maximize its download rate by dynamically adapting and shaping the upload rate allocated to its neighbors. We evaluate in detail the various mechanisms used by BitTyrant to identify their contribution to the performance of the client. Our findings indicate that the performance gain is due to the increased number of connections established by a BitTyrant client, rather than for its subtle uplink allocation algorithm; surprisingly, BitTyrant reveals to be altruistic and particularly efficient in disseminating the content, especially during the initial phase of the distribution process. The apparent gain of a single BitTyrant client, however, disappears in the case of a widespread adoption: our results indicate a severe loss of efficiency that we analyzed in detail. In contrast, a widespread adoption of the latest version of the mainline BitTorrent client would provide increased benefit for all peers. Damiano Carra, Giovanni Neglia, Pietro Michiardi |
Peer-to-Peer Computing | 1 |
| 2008 | Faster Content Access in KADabstractMany different distributed hash tables (DHTs) have been designed, but only few have been successfully deployed. The implementation of a DHT needs to deal with practical aspects (e.g. related to churn, or to the delay) that are often only marginally considered in the design. In this paper, we analyze in detail the content retrieval process in KAD, the implementation of the DHT Kademlia that is part of several popular peer-to-peer clients. In particular, we present a simple model to evaluate the impact of different design parameters on the overall lookup latency. We then perform extensive measurements on the lookup performance using an instrumented client. From the analysis of the results, we propose an improved scheme that is able to significantly decrease the overall lookup latency without increasing the overhead. Moritz Steiner, Damiano Carra, Ernst W. Biersack |
Peer-to-Peer Computing | 2 |
| 2008 | Stochastic Graph Processes for Performance Evaluation of Content Delivery Applications in Overlay NetworksabstractThis paper proposes a new methodology to model the distribution of finite size content to a group of users connected through an overlay network.Our methodology describes the distribution process as a constrained stochastic graph process (CSGP), where the constraints dictated by the content distribution protocol and the characteristics of the overlay network define the interaction among nodes. A CSGP is a semi-Markov process whose state is described by the graph itself. CSGPs offer a powerful description technique that can be exploited by Monte Carlo integration methods to compute in a very efficient way not only the mean but also the full distribution of metrics such as the file download times or number of hops from the source to the receiving nodes.We model several distribution architectures based on trees and meshes as CSGPs and solve them numerically. We are able to study scenarios with a very large number of nodes and we can precisely quantify the performance differences between the treeand mesh-based distribution architectures. Damiano Carra, Renato Lo Cigno, Ernst W. Biersack |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2007 | Building a reliable P2P system out of unreliable P2P clients: the case of KADabstractDistributed Hash Tables (DHT) provide a framework for managing information in a large distributed network of nodes. One of the main challenges DHT systems must face is node churn, i.e., nodes can arrive and depart at any time. To assure that information published in a DHT remains available despite node churn is equivalent to building a reliable system out of unreliable components. Damiano Carra, Ernst W. Biersack |
CoNEXT | 1 |
| 2007 | Graph Based Modeling of P2P Streaming Systems
Damiano Carra, Renato Lo Cigno, Ernst W. Biersack |
Networking | 1 |
| 2007 | Overlay architectures for file distribution: Fundamental performance analysis for homogeneous and heterogeneous cases
Ernst W. Biersack, Damiano Carra, Renato Lo Cigno, Pablo Rodriguez 0001, Pascal Felber |
Comput. Networks | 2 |
| 2007 | Graph Based Analysis of Mesh Overlay Streaming SystemsabstractThis paper studies fundamental properties of stream-based content distribution services. We assume the presence of an overlay network (such as those built by P2P systems) with limited degree of connectivity, and we develop a mathematical model that captures the essential features of overlay-based streaming protocols and systems. The methodology is based on stochastic graph theory, and models the streaming system as a stochastic process, whose characteristics are related to the streaming protocol. The model captures the elementary properties of the streaming system such as the number of active connections, the different play-out delay of nodes, and the probability of not receiving the stream due to node failures/misbehavior. Besides the static properties, the model is able to capture the transient behavior of the distribution graphs, i.e., the evolution of the structure over time, for instance in the initial phase of the distribution process. Contributions of this paper include a detailed definition of the methodology, its comparison with other analytical approaches and with simulative results, and a discussion of the additional insights enabled by this methodology. Results show that mesh based architectures are able to provide bounds on the receiving delay and maintain rate fluctuations due to system dynamics very low. Additionally, given the tight relationship between the stochastic process and the properties of the distribution protocol, this methodology gives basic guidelines for the design of such protocols and systems. Damiano Carra, Renato Lo Cigno, Ernst W. Biersack |
IEEE J. Sel. Areas Commun. | 1 |
| 2006 | Content Delivery in Overlay Networks: a Stochastic Graph Processes PerspectiveabstractWe consider the problem of distributing a content of finite size to a group of users connected through an overlay network that is built by a peer-to-peer application. The goal is the fastest possible diffusion of the content until it reaches all the peers. Applications like Bit-Torrent or SplitStream are examples where the problem we study is of great interest. In order to represent the content diffusion process, we model the system as a stochastic graph process and define the constraints the graph evolution is subject to. The evolution of the graph is a semi-Markov process where the sojourn times are the rewards of interest for the computation of the time needed to complete the file distribution. We discuss the general properties of the constrained stochastic graphs and we show preliminary results obtained with an ad-hoc Monte-Carlo technique. Damiano Carra, Renato Lo Cigno, Ernst W. Biersack |
GLOBECOM | 1 |
| 2006 | Fast Stochastic Analysis of P2P File Distribution ArchitecturesabstractIn this paper we investigate which is the most efficient architecture and protocol that can be used for file distribution. The focus of the analysis is to understand not only the parameters that influence the distribution process (constraints on the number of neighbors, bandwidth heterogeneity, etc.), but also the impact of the peer behavior, such as selfishness or neighbor selection strategies. The analysis also compares different tree- and mesh-based distribution architectures. We developed an ad-hoc Monte-Carlo technique that is able to analyze scenarios with millions of peers, a network size that traditional discrete- event simulators are not able to treat. The results give an accurate view of the fundamental protocol parameters and policies that impact on the final performance and allow designers to devise improved protocols. Damiano Carra, Renato Lo Cigno, Ernst W. Biersack |
GLOBECOM | 1 |
| 2006 | Fast Stochastic Exploration of Tree-Based File Distribution ArchitecturesabstractIn this work, generic systems for collaborative file distributionare considered. Tree-based structures are studied, that are important not only since they describe tree-based distribution architectures, but also for mesh-based architectures, since the initial diffusion phase usually follows a tree structure. Damiano Carra |
INFOCOM | 1 |