VLDB 2026 Research / reviewers in the wild / expert
Marco Canini
dblp:24/5715
· DBLP profile ↗
76ranked-venue papers
8as first author
31since 2021 · last 2026
0000-0002-5051-4283ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 38 · 4 first-author · 12 since 2021Systems, architecture and hardware · 20 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 8 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Security and privacy · 2Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reducing the GPU Memory Bottleneck with Lossless Compression for MLabstractMachine learning (ML) training and inference often process data sets far exceeding GPU memory capacity, forcing them to rely on PCIe for on-demand tensor transfers, causing critical transfer bottlenecks. Lossy compression has been proposed to relieve bottlenecks but introduces workload-dependent accuracy loss, making it complex or even prohibitive to use in existing ML deployments. Aditya K. Kamath, Arvind Krishnamurthy, Marco Canini, Simon Peter 0001 |
EuroSys | 3 |
| 2026 | Opportunistic Telemetry Transport in Hardware-Accelerated Observability Pipelines
Hadj Ahmed Chikh Dahmane, Alessandro Cornacchia, Marco Canini |
INFOCOM | 3 |
| 2026 | Observability Is Eating Your Cores: Fine-Grained Analysis of Microservice Metrics with IPU-Hosted Sketches
Alessandro Cornacchia, Theophilus Benson, Muhammad Bilal 0007, Marco Canini |
NSDI | 4 |
| 2026 | Cost-effective and Reliable Global Internet Peering with Programmable Switches
Congcong Miao, Zhiyi Yao, Jianchao Lv, Jinglin Wang, Shihan Lin, Xinyi Zhang 0004, Yunming Xiao, Jiwu Bu, Yachen Wang, Xianneng Zou, Yong Jiang 0001, Marco Canini, Gaogang Xie |
NSDI | 13 |
| 2026 | Label-free dataset profiling for federated client clusteringabstractAbstract Clustering clients into groups with relatively homogeneous data distributions is a key strategy for improving federated learning under non-independent and identically distributed data. However, most state-of-the-art clustering approaches require clients to possess labeled datasets and perform substantial local computation, limiting their applicability in real-world settings. To address these limitations, we introduce CoLEDS , a method for profiling unlabeled client datasets with minimal computational overhead. CoLEDS trains a model using a contrastive learning objective defined across multiple clients and optimized in a distributed fashion through joint client–server coordination. The resulting model embeds key properties of client datasets into low-dimensional vectors that are shared with the server for clustering. Extensive empirical evaluation shows that these profiles accurately capture latent dataset characteristics. By clustering clients based on these representations, CoLEDS yields federatively trained models that are better aligned with individual data distributions and enables appropriate model assignment even for clients that do not participate in federated training. Boris Radovic, Marco Canini, Veljko Pejovic |
Data Min. Knowl. Discov. | 2 |
| 2025 | DPFL: Decentralized Personalized Federated LearningabstractThis work addresses the challenges of data heterogeneity and communication constraints in decentralized federated learning (FL). We introduce decentralized personalized FL (DPFL), a bi-level optimization framework that enhances personalized FL by leveraging combinatorial relationships among clients, enabling fine-grained and targeted collaborations. By employing a constrained greedy algorithm, DPFL constructs a collaboration graph that guides clients in choosing suitable collaborators, enabling personalized model training tailored to local data while respecting a fixed and predefined communication and resource budget. Our theoretical analysis demonstrates that the proposed objective for constructing the collaboration graph yields superior or equivalent performance compared to any alternative collaboration structures, including pure local training. Extensive experiments across diverse datasets show that DPFL consistently outperforms existing methods, effectively handling non-IID data, reducing communication overhead, and improving resource efficiency in real-world decentralized FL scenarios. The code can be accessed at: \url{https://github.com/salmakh1/DPFL.} Salma Kharrat, Marco Canini, Samuel Horváth |
AISTATS | 2 |
| 2025 | Quantize Once, Train Fast: Allreduce-Compatible Compression with Provable GuaranteesabstractDistributed training enables large-scale deep learning, but suffers from high communication overhead, especially as models and datasets grow. Gradient compression, particularly quantization, is a promising approach to mitigate this bottleneck. However, existing quantization schemes are often incompatible with Allreduce, the dominant communication primitive in distributed deep learning, and many prior solutions rely on heuristics without theoretical guarantees. We introduce Global-QSGD, an Allreduce-compatible gradient quantization method that leverages global norm scaling to reduce communication overhead while preserving accuracy. Global-QSGD is backed by rigorous theoretical analysis, extending standard unbiased compressor frameworks to establish formal convergence guarantees. Additionally, we develop a performance model to evaluate its impact across different hardware configurations. Extensive experiments on NVLink, PCIe, and large-scale cloud environments show that Global-QSGD accelerates distributed training by up to 3.51× over baseline quantization methods, making it a practical and efficient solution for large-scale deep learning workloads. Jihao Xin, Marco Canini, Peter Richtárik, Samuel Horváth |
ECAI | 2 |
| 2025 | ACING: Actor-Critic for Instruction Learning in Black-Box LLMsabstractThe effectiveness of Large Language Models (LLMs) in solving tasks depends significantly on the quality of their instructions, which often require substantial human effort to craft.This underscores the need for automated instruction optimization.However, optimizing instructions is particularly challenging when working with black-box LLMs, where model parameters and gradients are inaccessible.We introduce ACING, an actor-critic reinforcement learning framework that formulates instruction optimization as a stateless, continuous-action problem, enabling exploration of infinite instruction spaces using only black-box feedback.ACING automatically discovers prompts that outperform human-written prompts in 76% of instruction-induction tasks, with gains of up to 33 points and a 10-point median improvement over the best automatic baseline in 33 tasks spanning instruction-induction, summarization, and chain-of-thought reasoning.Extensive ablations highlight its robustness and efficiency.An implementation of ACING is available at https://github.com/salmakh1/ACING. Salma Kharrat, Fares Fourati, Marco Canini |
EMNLP | 3 |
| 2025 | Query-based Knowledge Transfer for Heterogeneous Learning EnvironmentsabstractDecentralized collaborative learning under data heterogeneity and privacy constraints has rapidly advanced. However, existing solutions like federated learning, ensembles, and transfer learning, often fail to adequately serve the unique needs of clients, especially when local data representation is limited.
To address this issue, we propose a novel framework called Query-based Knowledge Transfer (QKT) that enables tailored knowledge acquisition to fulfill specific client needs without direct data exchange.
It employs a data-free masking strategy to facilitate the communication-efficient query-focused knowledge transformation while refining task-specific parameters to mitigate knowledge interference and forgetting. Our experiments, conducted on both standard and clinical benchmarks, show that QKT significantly outperforms existing collaborative learning methods by an average of 20.91% points in single-class query settings and an average of 14.32% points in multi-class query scenarios.
Further analysis and ablation studies reveal that QKT effectively balances the learning of new and existing knowledge, showing strong potential for its application in decentralized learning. Norah Alballa, Wenxuan Zhang 0003, Ziquan Liu, Ahmed M. Abdelmoniem, Mohamed Elhoseiny 0001, Marco Canini |
ICLR | 6 |
| 2025 | Assassyn: A Unified Abstraction for Architectural Simulation and ImplementationabstractThe continuous growth of on-chip transistors driven by technology scaling urges architecture developers to design and implement novel architectures to effectively utilize the excessive on-chip resources.Due to the challenges of programming in register-transfer level (RTL) languages, performance modeling based on simulation is typically developed alongside hardware implementation, allowing the exploration of high-level design decisions before dealing with the error-prone, low-level RTL details.However, this approach also introduces new challenges in coordinating across multiple teams to align implementation details separate codebases.In this paper, we address this issue by presenting Assassyn, a unified, high-level, and general-purpose programming framework for architectural simulation and implementation.By taking advantage of the concept of asynchronous event handling, a widely existing behavior in both hardware design and implementation and software engineering, a general-purpose, and high-level programming abstraction is proposed to mitigate the difficulties of RTL programming.Moreover, the unified programming interface naturally enables an accurate and faithful alignment between the simulation-based performance modeling and RTL implementation.Our evaluation demonstrates that Assassyn's high-level programming interface is sufficiently expressive to implement a wide range * Serve as both the first and correspondence author. Jian Weng 0002, Boyang Han, Derui Gao, Ruijie Gao, Wanning Zhang, An Zhong, Ceyu Xu, Jihao Xin, Yangzhixin Luo, Lisa Wu Wills, Marco Canini |
ISCA | 11 |
| 2025 | Scaling SCIERA: A Journey Through the Deployment of a Next-generation NetworkabstractThe SCION Next-Generation Network (NGN) architecture has expanded steadily since 2017, with today 20+ ISPs offering SCION connectivity. In production, IP-to-SCION-to-IP translation by SCION-IP-Gateways (SIGs) is used, such that applications are unaware of the NGN communication. To accelerate innovation and deployments, our aim is to increase the number of native SCION use cases, where the application is fully SCION-aware and optimizes communication across all path choices offered by the network. We set out to achieve two core objectives: (1) facilitating simple native connectivity for applications, and (2) enhancing the scalability of SCION deployment at academic sites. François Wirz, Marten Gartner, Jelte van Bommel, Elham Ehsani Moghadam, Grace H. Cimaszewski, Anxiao He, Yizhe Zhang 0006, Henry Birge-Lee, Felix Kottmann, Cyrill Krähenbühl, Jonghoon Kwon, Kyveli Mavromati, Liang Wang 0054, Daniel Bertolo, Marco Canini, Buseung Cho, Ronaldo A. Ferreira, Simon Peter Green, David Hausheer, Junbeom Hur, Xiaohua Jia, Heejo Lee, Prateek Mittal, Omo Oaiya, Chanjin Park, Adrian Perrig, Jerry Sobieski, Yixin Sun 0004, Cong Wang 0001, Klaas Wierenga |
SIGCOMM | 15 |
| 2024 | FilFL: Client Filtering for Optimized Client Participation in Federated LearningabstractFederated learning, an emerging machine learning paradigm, enables clients to collaboratively train a model without exchanging local data. Clients participating in the training process significantly impact the convergence rate, learning efficiency, and model generalization. We propose a novel approach, client filtering, to improve model generalization and optimize client participation and training. The proposed method periodically filters available clients to identify a subset that maximizes a combinatorial objective function with an efficient greedy filtering algorithm. Thus, the clients are assessed as a combination rather than individually. We theoretically analyze the convergence of federated learning with client filtering in heterogeneous settings and evaluate its performance across diverse vision and language tasks, including realistic scenarios with time-varying client availability. Our empirical results demonstrate several benefits of our approach, including improved learning efficiency, faster convergence, and up to 10% higher test accuracy than training without client filtering. Fares Fourati, Salma Kharrat, Vaneet Aggarwal, Mohamed-Slim Alouini, Marco Canini |
ECAI | 5 |
| 2024 | Continuous Exact Explanations of Neural NetworksabstractAccurate explanations of how a trained neural network (NN) behaves are desirable for a wide range of labor-intensive activities, including troubleshooting, validation, and understanding performance issues or identifying biases. We address the problem of explaining feedforward piecewise NNs (such as CNNs with Relu) by breaking them down into their linear components. Automatic encoding of the NN structure into linear program constraints is used to extract continuous exact explanations, which help interpret model behavior over continuous regions, provide model descriptions that convey the importance of input features and answer model queries that analyze the feature contributions under user-defined constraints. Our examples show that we can extract explanations for NNs with a moderate number of layers without relying on approximations. We demonstrate that the high-level explanations can help understand the outputs of NNs and the comparative importance of features by observing the linear models. We also show how explaining continuous inputs prevents certain attacks that have been proposed against existing explainers. Alice Dethise, Marco Canini |
ICDM | 2 |
| 2024 | Where is the Testbed for My Federated Learning Research?abstractProgressing beyond centralized AI is of paramount importance, yet, distributed AI solutions, in particular various federated learning (FL) algorithms, are often not comprehensively assessed, which prevents the research community from identifying the most promising approaches and practitioners from being convinced that a certain solution is deployment-ready. The largest hurdle towards FL algorithm evaluation is the difficulty of conducting real-world experiments over a variety of FL client devices and different platforms, with different datasets and data distribution, all while assessing various dimensions of algorithm performance, such as inference accuracy, energy consumption, and time to convergence, to name a few. In this paper, we present CoLExT, a real-world testbed for FL research. CoLExT is designed to streamline experimentation with custom FL algorithms in a rich testbed configuration space, with a large number of heterogeneous edge devices, ranging from single-board computers to smartphones, and provides real-time collection and visualization of a variety of metrics through automatic instrumentation. According to our evaluation, porting FL algorithms to CoLExT requires minimal involvement from the developer, and the instrumentation introduces minimal resource usage overhead. Furthermore, through an initial investigation involving popular FL algorithms running on CoLExT, we reveal previously unknown trade-offs, inefficiencies, and programming bugs. Janez Bozic, Amândio Faustino, Boris Radovic, Marco Canini, Veljko Pejovic |
SEC | 4 |
| 2023 | In-Network Aggregation with Transport Transparency for Distributed TrainingabstractRecent In-Network Aggregation (INA) solutions offload the all-reduce operation onto network switches to accelerate and scale distributed training (DT). On end hosts, these solutions build custom network stacks to replace the transport layer. The INA-oriented network stack cannot take advantage of the state-of-the-art performant transport layer implementation, and also causes complexity in system development and operation. Shuo Liu 0002, Qiaoling Wang, Junyi Zhang 0005, Wenfei Wu, Qinliang Lin, Yao Liu 0006, Marco Canini, Ray C. C. Cheung, Jianfei He |
ASPLOS (3) | 8 |
| 2023 | With Great Freedom Comes Great Opportunity: Rethinking Resource Allocation for Serverless FunctionsabstractCurrent serverless offerings give users limited flexibility for configuring the resources allocated to their function invocations. This simplifies the interface for users to deploy server-less computations but creates deployments that are resource inefficient. In this paper, we take a principled approach to the problem of resource allocation for serverless functions, analyzing the effects of automating this choice in a way that leads to the best combination of performance and cost. In particular, we systematically explore the opportunities that come with decoupling memory and CPU resource allocations and also enabling the use of different VM types, and we find a rich trade-off space between performance and cost. The provider can use this in a number of ways, e.g., exposing all these parameters to the user; eliding preferences for performance and cost from users and simply offer the same performance with lower cost; or exposing a small number of choices for users to trade performance for cost. Muhammad Bilal 0007, Marco Canini, Rodrigo Fonseca, Rodrigo Rodrigues 0001 |
EuroSys | 2 |
| 2023 | REFL: Resource-Efficient Federated LearningabstractFederated Learning (FL) enables distributed training by learners using local data, thereby enhancing privacy and reducing communication. However, it presents numerous challenges relating to the heterogeneity of the data distribution, device capabilities, and participant availability as deployments scale, which can impact both model convergence and bias. Existing FL schemes use random participant selection to improve the fairness of the selection process; however, this can result in inefficient use of resources and lower quality training. In this work, we systematically address the question of resource efficiency in FL, showing the benefits of intelligent participant selection, and incorporation of updates from straggling participants. We demonstrate how these factors enable resource efficiency while also improving trained model quality. Ahmed M. Abdelmoniem, Atal Narayan Sahu, Marco Canini, Suhaib A. Fahmy |
EuroSys | 3 |
| 2023 | TENSOR: Lightweight BGP Non-Stop RoutingabstractAs the solitary inter-domain protocol, BGP plays an important role in today's Internet. Its failures threaten network stability and will usually result in large-scale packet losses. Thus, the non-stop routing (NSR) capability that protects inter-domain connectivity from being disrupted by various failures, is critical to any Autonomous System (AS) operator. Replicating the BGP and underlying TCP connection status is key to realizing NSR. But existing NSR solutions, which heavily rely on OS kernel modifications, have become impractical due to providers' adoption of virtualized network gateways for better scalability and manageability. Congcong Miao, Yunming Xiao, Marco Canini, Ruiqiang Dai, Shengli Zheng, Jilong Wang 0001, Jiwu Bu, Aleksandar Kuzmanovic, Yachen Wang |
SIGCOMM | 3 |
| 2023 | SAGE: Software-based Attestation for GPU Execution
Andrei Ivanov, Benjamin Rothenberger, Alice Dethise, Marco Canini, Torsten Hoefler, Adrian Perrig |
USENIX ATC | 4 |
| 2023 | A Comprehensive Empirical Study of Heterogeneity in Federated LearningabstractFederated learning (FL) is becoming a popular paradigm for collaborative learning over distributed, private data sets owned by nontrusting entities. FL has seen successful deployment in production environments, and it has been adopted in services, such as virtual keyboards, auto-completion, item recommendation, and several IoT applications. However, FL comes with the challenge of performing training over largely heterogeneous data sets, devices, and networks that are out of the control of the centralized FL server. Motivated by this inherent challenge, we aim to empirically characterize the impact of device and behavioral heterogeneity on the trained model. We conduct an extensive empirical study spanning nearly 1.5K unique configurations on five popular FL benchmarks. Our analysis shows that these sources of heterogeneity have a major impact on both model quality and fairness, causing up to$4.6\times $and$2.2\times $degradation in the quality and fairness, respectively, thus shedding light on the importance of considering heterogeneity in FL system design. Ahmed M. Abdelmoniem, Chen-Yu Ho 0001, Pantelis Papageorgiou, Marco Canini |
IEEE Internet Things J. | 4 |
| 2022 | RDMA is Turing complete, we just did not know it yet!
Waleed Reda, Marco Canini, Dejan Kostic, Simon Peter 0001 |
NSDI | 2 |
| 2022 | Unlocking the Power of Inline Floating-Point Operations on Programmable Switches
Omar Alama, Jiawei Fei, Jacob Nelson 0001, Dan R. K. Ports, Amedeo Sapio, Marco Canini, Nam Sung Kim |
NSDI | 7 |
| 2022 | Renaissance: A self-stabilizing distributed SDN control plane using in-band communications
Marco Canini, Iosif Salem, Liron Schiff, Elad Michael Schiller, Stefan Schmid 0001 |
J. Comput. Syst. Sci. | 1 |
| 2021 | GRACE: A Compressed Communication Framework for Distributed Machine LearningabstractPowerful computer clusters are used nowadays to train complex deep neural networks (DNN) on large datasets. Distributed training increasingly becomes communication bound. For this reason, many lossy compression techniques have been proposed to reduce the volume of transferred data. Unfortunately, it is difficult to argue about the behavior of compression methods, because existing work relies on inconsistent evaluation testbeds and largely ignores the performance impact of practical system configurations. In this paper, we present a comprehensive survey of the most influential compressed communication methods for DNN training, together with an intuitive classification (i.e., quantization, sparsification, hybrid and low-rank). Next, we propose GRACE, a unified framework and API that allows for consistent and easy implementation of compressed communication on popular machine learning toolkits. We instantiate GRACE on TensorFlow and PyTorch, and implement 16 such methods. Finally, we present a thorough quantitative evaluation with a variety of DNNs (convolutional and recurrent), datasets and system configurations. We show that the DNN architecture affects the relative performance among methods. Interestingly, depending on the underlying communication library and computational cost of compression / decompression, we demonstrate that some methods may be impractical. GRACE and the entire benchmarking suite are available as open-source. Chen-Yu Ho 0001, Ahmed M. Abdelmoniem, Aritra Dutta, El Houcine Bergou, Konstantinos Karatsenidis, Marco Canini, Panos Kalnis |
ICDCS | 7 |
| 2021 | AutoLRS: Automatic Learning-Rate Schedule by Bayesian Optimization on the Fly
Tianyi Zhou 0001, Liangyu Zhao, Yibo Zhu 0001, Chuanxiong Guo, Marco Canini, Arvind Krishnamurthy |
ICLR | 6 |
| 2021 | DC2: Delay-aware Compression Control for Distributed Machine LearningabstractDistributed training performs data-parallel training of DNN models which is a necessity for increasingly complex models and large datasets. Recent works are identifying major communication bottlenecks in distributed training. These works seek possible opportunities to speed-up the training in systems supporting distributed ML workloads. As communication reduction, compression techniques are proposed to speed up this communication phase. However, compression comes at the cost of reduced model accuracy, especially when compression is applied arbitrarily. Instead, we advocate a more controlled use of compression and propose DC2, a delay-aware compression control mechanism. DC2 couples compression control and network delays in applying compression adaptively. DC2 not only compensates for network variations but can also strike a better trade-off between training speed and accuracy. DC2 is implemented as a drop-in module to the communication library used by the ML toolkit and can operate in a variety of network settings. We empirically evaluate DC2 in network environments exhibiting low and high delay variations. Our evaluation of different popular CNN models and datasets shows that DC2 improves training speed-ups of up to 41× and 5.3 × over baselines with no-compression and uniform compression, respectively. Ahmed M. Abdelmoniem, Marco Canini |
INFOCOM | 2 |
| 2021 | Analyzing Learning-Based Networked Systems with Formal VerificationabstractAs more applications of (deep) neural networks emerge in the computer networking domain, the correctness and predictability of a neural agent's behavior for corner case inputs are becoming crucial. Enabling the formal analysis of agents with nontrivial properties, we bridge between specifying intended high-level behavior and expressing low-level statements directly encoded into an efficient verification framework. Our results support that within minutes, one can establish the resilience of a neural network to adversarial attacks on its inputs, as well as formally prove properties that were previously relying on educated guesses. Finally, we also show how formal verification can help create an accurate visual representation of an agent behavior to perform visual inspection and improve its trustworthiness. Alice Dethise, Marco Canini, Nina Narodytska |
INFOCOM | 2 |
| 2021 | Rethinking gradient sparsification as total error minimizationabstractGradient compression is a widely-established remedy to tackle the communication bottleneck in distributed training of large deep neural networks (DNNs). Under the error-feedback framework, Top-$k$ sparsification, sometimes with $k$ as little as 0.1% of the gradient size, enables training to the same model quality as the uncompressed case for a similar iteration count. From the optimization perspective, we find that Top-$k$ is the communication-optimal sparsifier given a per-iteration $k$ element budget.We argue that to further the benefits of gradient sparsification, especially for DNNs, a different perspective is necessary — one that moves from per-iteration optimality to consider optimality for the entire training.We identify that the total error — the sum of the compression errors for all iterations — encapsulates sparsification throughout training. Then, we propose a communication complexity model that minimizes the total error under a communication budget for the entire training. We find that the hard-threshold sparsifier, a variant of the Top-$k$ sparsifier with $k$ determined by a constant hard-threshold, is the optimal sparsifier for this model. Motivated by this, we provide convex and non-convex convergence analyses for the hard-threshold sparsifier with error-feedback. We show that hard-threshold has the same asymptotic convergence and linear speedup property as SGD in both the case, and unlike with Top-$k$ sparsifier, has no impact due to data-heterogeneity. Our diverse experiments on various DNNs and a logistic regression model demonstrate that the hard-threshold sparsifier is more communication-efficient than Top-$k$. Atal Narayan Sahu, Aritra Dutta, Ahmed M. Abdelmoniem, Trambak Banerjee, Marco Canini, Panos Kalnis |
NeurIPS | 5 |
| 2021 | Scaling Distributed Machine Learning with In-Network Aggregation
Amedeo Sapio, Marco Canini, Chen-Yu Ho 0001, Jacob Nelson 0001, Panos Kalnis, Changhoon Kim, Arvind Krishnamurthy, Masoud Moshref, Dan R. K. Ports, Peter Richtárik |
NSDI | 2 |
| 2021 | Efficient sparse collective communication and its application to accelerate distributed deep learningabstractEfficient collective communication is crucial to parallel-computing applications such as distributed training of large-scale recommendation systems and natural language processing models. Existing collective communication libraries focus on optimizing operations for dense inputs, resulting in transmissions of many zeros when inputs are sparse. This counters current trends that see increasing data sparsity in large models. Jiawei Fei, Chen-Yu Ho 0001, Atal Narayan Sahu, Marco Canini, Amedeo Sapio |
SIGCOMM | 4 |
| 2021 | LineFS: Efficient SmartNIC Offload of a Distributed File System with Pipeline ParallelismabstractIn multi-tenant systems, the CPU overhead of distributed file systems (DFSes) is increasingly a burden to application performance. CPU and memory interference cause degraded and unstable application and storage performance, in particular for operation latency. Recent client-local DFSes for persistent memory (PM) accelerate this trend. DFS offload to SmartNICs is a promising solution to these problems, but it is challenging to fit the complex demands of a DFS onto simple SmartNIC processors located across PCIe. Jongyul Kim 0001, Insu Jang, Waleed Reda, Jaeseong Im, Marco Canini, Dejan Kostic, Youngjin Kwon, Simon Peter 0001, Emmett Witchel |
SOSP | 5 |
| 2020 | On the Discrepancy between the Theoretical Analysis and Practical Implementations of Compressed Communication for Distributed Deep LearningabstractCompressed communication, in the form of sparsification or quantization of stochastic gradients, is employed to reduce communication costs in distributed data-parallel training of deep neural networks. However, there exists a discrepancy between theory and practice: while theoretical analysis of most existing compression methods assumes compression is applied to the gradients of the entire model, many practical implementations operate individually on the gradients of each layer of the model.In this paper, we prove that layer-wise compression is, in theory, better, because the convergence rate is upper bounded by that of entire-model compression for a wide range of biased and unbiased compression methods. However, despite the theoretical bound, our experimental study of six well-known methods shows that convergence, in practice, may or may not be better, depending on the actual trained model and compression ratio. Our findings suggest that it would be advantageous for deep learning frameworks to include support for both layer-wise and entire-model compression. Aritra Dutta, El Houcine Bergou, Ahmed M. Abdelmoniem, Chen-Yu Ho 0001, Atal Narayan Sahu, Marco Canini, Panos Kalnis |
AAAI | 6 |
| 2020 | Finding the right cloud configuration for analytics clustersabstractFinding good cloud configurations for deploying a single distributed system is already a challenging task, and it becomes substantially harder when a data analytics cluster is formed by multiple distributed systems since the search space becomes exponentially larger. In particular, recent proposals for single system deployments rely on benchmarking runs that become prohibitively expensive as we shift to joint optimization of multiple systems, as users have to wait until the end of a long optimization run to start the production run of their job. Muhammad Bilal 0007, Marco Canini, Rodrigo Rodrigues 0001 |
SoCC | 2 |
| 2020 | Assise: Performance and Availability via Client-local NVM in a Distributed File System
Thomas E. Anderson, Marco Canini, Jongyul Kim 0001, Dejan Kostic, Youngjin Kwon, Simon Peter 0001, Waleed Reda, Henry Schuh, Emmett Witchel |
OSDI | 2 |
| 2020 | Do the Best Cloud Configurations Grow on Trees? An Experimental Evaluation of Black Box Algorithms for Optimizing Cloud Workloads Sub
Muhammad Bilal 0007, Marco Serafini, Marco Canini, Rodrigo Rodrigues 0001 |
Proc. VLDB Endow. | 3 |
| 2020 | Toward Consistent SDNs: A Case for Network State FuzzingabstractThe conventional wisdom is that a software-defined network (SDN) operates under the premise that the logically centralized control plane has an accurate representation of the actual data plane state. Unfortunately, bugs, misconfigurations, faults or attacks can introduce inconsistencies that undermine correct operation. Previous work in this area, however, lacks a holistic methodology to tackle this problem and thus, addresses only certain parts of the problem. Yet, the consistency of the overall system is only as good as its least consistent part. Motivated by an analogy of network consistency checking with program testing, we propose to add active probe-based network state fuzzing to our consistency check repertoire. Hereby, our system, Pazz, combines production traffic with active probes to periodically test if the actual forwarding path and decision elements (on the data plane) correspond to the expected ones (on the control plane). Our insight is that active traffic covers the inconsistency cases beyond the ones identified by passive traffic. Pazz prototype was built and evaluated on topologies of varying scale and complexity. Our results show that Pazz requires minimal network resources to detect persistent data plane faults through fuzzing and localize them quickly while outperforming baseline approaches. Apoorv Shukla, Said Jawad Saidi, Stefan Schmid 0001, Marco Canini, Thomas Zinner, Anja Feldmann |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2020 | P4xos: Consensus as a Network ServiceabstractIn this paper, we explore how a programmable forwarding plane offered by a new breed of network switches might naturally accelerate consensus protocols, specifically focusing on Paxos. The performance of consensus protocols has long been a concern. By implementing Paxos in the forwarding plane, we are able to significantly increase throughput and reduce latency. Our P4-based implementation running on an ASIC in isolation can process over 2.5 billion consensus messages per second, a four orders of magnitude improvement in throughput over a widely-used software implementation. This effectively removes consensus as a bottleneck for distributed applications in data centers. Beyond sheer performance, our approach offers several other important benefits: it readily lends itself to formal verification; it does not rely on any additional network hardware; and as a full Paxos implementation, it makes only very weak assumptions about the network. Huynh Tu Dang, Pietro Bressana, Han Wang 0009, Ki Suh Lee, Noa Zilberman, Hakim Weatherspoon, Marco Canini, Fernando Pedone, Robert Soulé |
IEEE/ACM Trans. Netw. | 7 |
| 2019 | Measurements As First-class ArtifactsabstractThe emergence of programmable switches has sparked a significant amount of work on new techniques to perform more powerful measurement tasks, for instance, to obtain fine-grained traffic and performance statistics. Previous work has focused on the efficiency of these measurements alone and has neglected flexibility, resulting in solutions that are hard to reuse or repurpose and that often overlap in functionality or goals. In this paper, we propose the use of a set of reusable primitive building blocks that can be composed to express measurement tasks in a concise and simple way. We describe the rationale for the design of our primitives, that we have named MAFIA (Measurements As FIrst-class Artifacts), and using several examples we illustrate how they can be combined to realize a comprehensive range of network measurement tasks. Writing MAFIA code does not require expert knowledge of low-level switch architecture details. Using a prototype implementation of MAFIA, we demonstrate the applicability of our approach and show that the use of our primitives results in compiled code that is comparable in size and resource usage with manually written specialized P4 code, and can be run in current hardware. Paolo Laffranchini, Luís E. T. Rodrigues, Marco Canini, Balachander Krishnamurthy |
INFOCOM | 3 |
| 2018 | Prelude: Ensuring Inter-Domain Loop-Freedom in SDN-Enabled NetworksabstractSoftware-Defined eXchanges (SDXes) promise to improve the interdomain routing ecosystem through SDN deployment. Yet, the naïve deployment of SDN on the Internet raises concerns about the correctness of the interdomain data-plane. By allowing operators to deflect traffic from default BGP routes, SDN policies can create permanent forwarding loops that are not visible to the control-plane. Alice Dethise, Marco Chiesa, Marco Canini |
APNet | 3 |
| 2018 | Fast and Accurate Load Balancing for Geo-Distributed Storage SystemsabstractThe increasing density of globally distributed datacenters reduces the network latency between neighboring datacenters and allows replicated services deployed across neighboring locations to share workload when necessary, without violating strict Service Level Objectives (SLOs). Kirill Bogdanov 0001, Waleed Reda, Gerald Q. Maguire Jr., Dejan Kostic, Marco Canini |
SoCC | 5 |
| 2018 | Dynam-IX: a dynamic interconnection eXchangeabstractAutonomous Systems (ASes) can reach hundreds of networks via Internet eXchange Points (IXPs), allowing improvements in traffic delivery performance and competitiveness. Despite the benefits, any pair of ASes needs first to agree on exchanging traffic. By surveying 100+ network operators, we discovered that most interconnection agreements are established through ad-hoc and lengthy processes heavily influenced by personal relationships and brand image. As such, ASes prefer long-term agreements at the expense of a potential mismatch between actual delivery performance and current traffic dynamics. ASes also miss interconnection opportunities due to trust reasons. To improve wide-area traffic delivery performance, we propose Dynam-IX, a framework that allows operators to build trust cooperatively and implement traffic engineering policies to exploit the rich interconnection opportunities at IXPs quickly. Dynam-IX offers a protocol to automate the interconnection process, an intent abstraction to express interconnection policies, a legal framework to digitally handle contracts, and a distributed tamper-proof ledger to create trust among ASes. We build and evaluate a Dynam-IX prototype and show that an AS can establish tens of agreements per minute with negligible overhead for ASes and IXPs. Pedro de B. Marcos, Marco Chiesa, Lucas F. Müller, Pradeeban Kathiravelu, Christoph Dietzel, Marco Canini, Marinho P. Barcellos |
CoNEXT | 6 |
| 2018 | Renaissance: A Self-Stabilizing Distributed SDN Control PlaneabstractBy introducing programmability, automated verification, and innovative debugging tools, Software-Defined Networks (SDNs) are poised to meet the increasingly stringent dependability requirements of today's communication networks. However, the design of fault-tolerant SDNs remains an open challenge. This paper considers the design of dependable SDNs through the lenses of self-stabilization - a very strong notion of fault-tolerance. In particular, we develop algorithms for an in-band and distributed control plane for SDNs, called Renaissance, which tolerates a wide range of (concurrent) controller, link, and communication failures. Our self-stabilizing algorithms ensure that after the occurrence of an arbitrary combination of failures, (i) every non-faulty SDN controller can eventually reach any switch in the network within a bounded communication delay (in the presence of a bounded number of concurrent failures) and (ii) every switch is managed by at least one non-faulty controller. We evaluate Renaissance through a rigorous worst-case analysis as well as a prototype implementation (based on OVS and Floodlight), and we report on our experiments using Mininet. Marco Canini, Iosif Salem, Liron Schiff, Elad Michael Schiller, Stefan Schmid 0001 |
ICDCS | 1 |
| 2018 | Sonata: query-driven streaming network telemetryabstractManaging and securing networks requires collecting and analyzing network traffic data in real time. Existing telemetry systems do not allow operators to express the range of queries needed to perform management or scale to large traffic volumes and rates. We present Sonata, an expressive and scalable telemetry system that coordinates joint collection and analysis of network traffic. Sonata provides a declarative interface to express queries for a wide range of common telemetry tasks; to enable real-time execution, Sonata partitions each query across the stream processor and the data plane, running as much of the query as it can on the network switch, at line rate. To optimize the use of limited switch memory, Sonata dynamically refines each query to ensure that available resources focus only on traffic that satisfies the query. Our evaluation shows that Sonata can support a wide range of telemetry tasks while reducing the workload for the stream processor by as much as seven orders of magnitude compared to existing telemetry systems. Arpit Gupta, Rob Harrison, Marco Canini, Nick Feamster, Jennifer Rexford, Walter Willinger |
SIGCOMM | 3 |
| 2018 | Methodology, measurement and analysis of flow table update characteristics in hardware openflow switches
Maciej Kuzniar, Peter Peresíni, Dejan Kostic, Marco Canini |
Comput. Networks | 4 |
| 2017 | Towards automatic parameter tuning of stream processing systemsabstractOptimizing the performance of big-data streaming applications has become a daunting and time-consuming task: parameters may be tuned from a space of hundreds or even thousands of possible configurations. In this paper, we present a framework for automating parameter tuning for stream-processing systems. Our framework supports standard black-box optimization algorithms as well as a novel gray-box optimization algorithm. We demonstrate the multiple benefits of automated parameter tuning in optimizing three benchmark applications in Apache Storm. Our results show that a hill-climbing algorithm that uses a new heuristic sampling approach based on Latin Hypercube provides the best results. Our gray-box algorithm provides comparable results while being two to five times faster. Muhammad Bilal 0007, Marco Canini |
SoCC | 2 |
| 2017 | DAIET: a system for data aggregation inside the networkabstractMany data center applications nowadays rely on distributed computation models like MapReduce and Bulk Synchronous Parallel (BSP) for data-intensive computation at scale [4]. These models scale by leveraging the partition/aggregate pattern where data and computations are distributed across many worker servers, each performing part of the computation. A communication phase is needed each time workers need to synchronize the computation and, at last, to produce the final output. In these applications, the network communication costs can be one of the dominant scalability bottlenecks especially in case of multi-stage or iterative computations [1]. Amedeo Sapio, Ibrahim Abdelaziz, Marco Canini, Panos Kalnis |
SoCC | 3 |
| 2017 | Distributed resource management across process boundariesabstractMulti-tenant distributed systems composed of small services, such as Service-oriented Architectures (SOAs) and Micro-services, raise new challenges in attaining high performance and efficient resource utilization. In these systems, a request execution spans tens to thousands of processes, and the execution paths and resource demands on different services are generally not known when a request first enters the system. In this paper, we highlight the fundamental challenges of regulating load and scheduling in SOAs while meeting end-to-end performance objectives on metrics of concern to both tenants and operators. We design Wisp, a framework for building SOAs that transparently adapts rate limiters and request schedulers system-wide according to operator policies to satisfy end-to-end goals while responding to changing system conditions. In evaluations against production as well as synthetic workloads, Wisp successfully enforces a range of end-to-end performance objectives, such as reducing average latencies, meeting deadlines, providing fairness and isolation, and avoiding system overload. Lalith Suresh 0001, Peter Bodík, Ishai Menache, Marco Canini, Florin Ciucu |
SoCC | 4 |
| 2017 | SIXPACK: Securing Internet eXchange Points Against Curious onlooKersabstractInternet eXchange Points (IXPs) play an ever-growing role in Internet inter-connection. To facilitate the exchange of routes amongst their members, IXPs provide Route Server (RS) services to dispatch the routes according to each member's peering policies. Nowadays, to make use of RSes, these policies must be disclosed to the IXP. This poses fundamental questions regarding the privacy guarantees of route-computation on confidential business information. Indeed, as evidenced by interaction with IXP administrators and a survey of network operators, this state of affairs raises privacy concerns among network administrators and even deters some networks from subscribing to RS services. We design Sixpack1, an RS service that leverages Secure Multi-Party Computation (SMPC) to keep peering policies confidential, while extending, the functionalities of today's RSes. As SMPC is notoriously heavy in terms of communication and computation, our design and implementation of Sixpack aims at moving computation outside of the SMPC without compromising the privacy guarantees. We assess the effectiveness and scalability of our system by evaluating a prototype implementation using traces of data from one of the largest IXPs in the world. Our evaluation results indicate that Sixpack can scale to support privacy-preserving route-computation, even at IXPs with many hundreds of member networks. Marco Chiesa, Daniel Demmler, Marco Canini, Michael Schapira, Thomas Schneider 0003 |
CoNEXT | 3 |
| 2017 | Rein: Taming Tail Latency in Key-Value Stores via Multiget SchedulingabstractWe tackle the problem of reducing tail latencies in distributed key-value stores, such as the popular Cassandra database. We focus on workloads of multiget requests, which batch together access to several data elements and parallelize read operations across the data store machines. We first analyze a production trace of a real system and quantify the skew due to multiget sizes, key popularity, and other factors. We then proceed to identify opportunities for reduction of tail latencies by recognizing the composition of aggregate requests and by carefully scheduling bottleneck operations that can otherwise create excessive queues. We design and implement a system called Rein, which reduces latency via inter-multiget scheduling using low overhead techniques. We extensively evaluate Rein via experiments in Amazon Web Services (AWS) and simulations. Our scheduling algorithms reduce the median, 95th, and 99th percentile latencies by factors of 1.5, 1.5, and 1.9, respectively. Waleed Reda, Marco Canini, Lalith Suresh 0001, Dejan Kostic, Sean Braithwaite |
EuroSys | 2 |
| 2017 | In-Network Computation is a Dumb Idea Whose Time Has ComeabstractProgrammable data plane hardware creates new opportunities for infusing intelligence into the network. This raises a fundamental question: what kinds of computation should be delegated to the network? Amedeo Sapio, Ibrahim Abdelaziz, Abdulla Aldilaijan, Marco Canini, Panos Kalnis |
HotNets | 4 |
| 2017 | A Self-Organizing Distributed and In-Band SDN Control PlaneabstractAdopting distributed control planes is critical towards ensuring high availability and fault-tolerance of dependable Software-Defined Networks (SDNs). However, designing and bootstrapping a distributed SDN control plane is a challenging task, especially if to be done in-band, without a dedicated control network, and without relying on legacy networking protocols. One of the most appealing and powerful notions of fault-tolerance is self-organization and this paper discusses the possibility of self-organizing algorithms for in-band control planes. Marco Canini, Iosif Salem, Liron Schiff, Elad Michael Schiller, Stefan Schmid 0001 |
ICDCS | 1 |
| 2017 | Correct by Construction Networks Using Stepwise Refinement
Leonid Ryzhyk, Nikolaj S. Bjørner, Marco Canini, Jean-Baptiste Jeannin, Cole Schlesinger, Douglas B. Terry, George Varghese |
NSDI | 3 |
| 2017 | ENDEAVOUR: A Scalable SDN Architecture For Real-World IXPsabstractInnovation in interdomain routing has remained stagnant for over a decade. Recently, Internet eXchange Points (IXPs) have emerged as economically-advantageous interconnection points for reducing path latencies and exchanging ever increasing traffic volumes among, possibly, hundreds of networks. Given their far-reaching implications on interdomain routing, IXPs are the ideal place to foster network innovation and extend the benefits of software defined networking (SDN) to the interdomain level. In this paper, we present, evaluate, and demonstrate ENDEAVOUR, an SDN platform for IXPs. ENDEAVOUR can be deployed on a multi-hop IXP fabric, supports a large number of use cases, and is highly scalable, while avoiding broadcast storms. Our evaluation with real data from one of the largest IXPs, demonstrates the benefits and scalability of our solution: ENDEAVOUR requires around 70% fewer rules than alternative SDN solutions thanks to our rule partitioning mechanism. In addition, by providing an open source solution, we invite everyone from the community to experiment (and improve) our implementation as well as adapt it to new use cases. Gianni Antichi, Ignacio Castro, Marco Chiesa, Eder Leão Fernandes, Remy Lapeyrade, Daniel Kopp, Jong Hun Han, Marc Bruyere, Christoph Dietzel, Mitchell Gusat, Andrew W. Moore 0002, Philippe Owezarski, Steve Uhlig, Marco Canini |
IEEE J. Sel. Areas Commun. | 14 |
| 2016 | Network Monitoring as a Streaming Analytics ProblemabstractProgrammable switches potentially make it easier to perform flexible network monitoring queries at line rate, and scalable stream processors make it possible to fuse data streams to answer more sophisticated queries about the network in real-time. However, processing such network monitoring queries at high traffic rates requires both the switches and the stream processors to filter the traffic iteratively and adaptively so as to extract only that traffic that is of interest to the query at hand. While the realization that network monitoring is a streaming analytics problem has been made earlier, our main contribution in this paper is the design and implementation of Sonata, a closed-loop system that enables network operators to perform streaming analytics for network monitoring applications at scale. To achieve this objective, Sonata allows operators to express a network monitoring query by considering each packet as a tuple. More importantly, Sonata allows them to partition the query across both the switches and the stream processor, and through iterative refinement, Sonata's runtime attempts to extract only the traffic that pertains to the query, thus ensuring that the stream processor can scale to satisfy a large number of queries for traffic at very high rates. We show with a simple example query involving DNS reflection attacks and traffic traces from one of the world's largest IXPs that Sonata can capture 95% of all traffic pertaining to the query, while reducing the overall data rate by a factor of about 400 and the number of required counters by four orders of magnitude. Arpit Gupta, Rüdiger Birkner, Marco Canini, Nick Feamster, Chris Mac-Stoker, Walter Willinger |
HotNets | 3 |
| 2016 | An Industrial-Scale Software Defined Internet Exchange Point
Arpit Gupta, Robert MacDavid, Rüdiger Birkner, Marco Canini, Nick Feamster, Jennifer Rexford, Laurent Vanbever |
NSDI | 4 |
| 2016 | An Industrial-Scale Software Defined Internet Exchange Point
Arpit Gupta, Robert MacDavid, Rüdiger Birkner, Marco Canini, Nick Feamster, Jennifer Rexford, Laurent Vanbever |
USENIX ATC | 4 |
| 2015 | A distributed and robust SDN control plane for transactional network updatesabstractSoftware-defined networking (SDN) is a novel paradigm that outsources the control of programmable network switches to a set of software controllers. The most fundamental task of these controllers is the correct implementation of the network policy, i.e., the intended network behavior. In essence, such a policy specifies the rules by which packets must be forwarded across the network. This paper studies a distributed SDN control plane that enables concurrent and robust policy implementation. We introduce a formal model describing the interaction between the data plane and a distributed control plane (consisting of a collection of fault-prone controllers). Then we formulate the problem of consistent composition of concurrent network policy updates (termed the CPC Problem). To anticipate scenarios in which some conflicting policy updates must be rejected, we enable the composition via a natural transactional interface with all-or-nothing semantics. We show that the ability of an f-resilient distributed control plane to process concurrent policy updates depends on the tag complexity, i.e., the number of policy labels (a.k.a. tags) available to the controllers, and describe a CPC protocol with optimal tag complexity f + 2. Marco Canini, Petr Kuznetsov, Dan Levin, Stefan Schmid 0001 |
INFOCOM | 1 |
| 2015 | C3: Cutting Tail Latency in Cloud Data Stores via Adaptive Replica Selection
Lalith Suresh 0001, Marco Canini, Stefan Schmid 0001, Anja Feldmann |
NSDI | 2 |
| 2015 | BRB: BetteR Batch Scheduling to Reduce Tail Latencies in Cloud Data StoresabstractA common pattern in the architectures of modern interactive web-services is that of large request fan-outs, where even a single end-user request (task) arriving at an application server triggers tens to thousands of data accesses (sub-tasks) to different stateful backend servers. The overall response time of each task is bottlenecked by the completion time of the slowest sub-task, making such workloads highly sensitive to the tail of latency distribution of the backend tier. The large number of decentralized application servers and skewed workload patterns exacerbate the challenge in addressing this problem. We address these challenges through BetteR Batch (BRB). By carefully scheduling requests in a decentralized and task-aware manner, BRB enables low-latency distributed storage systems to deliver predictable performance in the presence of large request fan-outs. Our preliminary simulation results based on production workloads show that our proposed design is at the 99th percentile latency within 38% of an ideal system model while offering latency improvements over the state-of-the-art by a factor of 2. Waleed Reda, Lalith Suresh 0001, Marco Canini, Sean Braithwaite |
SIGCOMM | 3 |
| 2015 | Systematically testing OpenFlow controller applications
Peter Peresíni, Maciej Kuzniar, Marco Canini, Daniele Venzano, Dejan Kostic, Jennifer Rexford |
Comput. Networks | 3 |
| 2014 | Panopticon: Reaping the Benefits of Incremental SDN Deployment in Enterprise Networks
Dan Levin, Marco Canini, Stefan Schmid 0001, Fabian Schaffert, Anja Feldmann |
USENIX ATC | 2 |
| 2013 | Incremental SDN deployment in enterprise networksabstractNo abstract available. Dan Levin, Marco Canini, Stefan Schmid 0001, Anja Feldmann |
SIGCOMM | 2 |
| 2012 | A SOFT way for openflow switch interoperability testingabstractThe increasing adoption of Software Defined Networking, and OpenFlow in particular, brings great hope for increasing extensibility and lowering costs of deploying new network functionality. A key component in these networks is the OpenFlow agent, a piece of software that a switch runs to enable remote programmatic access to its forwarding tables. While testing high-level network functionality, the correct behavior and interoperability of any OpenFlow agent are taken for granted. However, existing tools for testing agents are not exhaustive nor systematic, and only check that the agent's basic functionality works. In addition, the rapidly changing and sometimes vague OpenFlow specifications can result in multiple implementations that behave differently. Maciej Kuzniar, Peter Peresíni, Marco Canini, Daniele Venzano, Dejan Kostic |
CoNEXT | 3 |
| 2012 | A NICE Way to Test OpenFlow Applications
Marco Canini, Daniele Venzano, Peter Peresíni, Dejan Kostic, Jennifer Rexford |
NSDI | 1 |
| 2011 | Identifying and using energy-critical pathsabstractThe power consumption of the Internet and datacenter networks is already significant, and threatens to shortly hit the power delivery limits while the hardware is trying to sustain ever-increasing traffic requirements. Existing energy-reduction approaches in this domain advocate recomputing network configuration with each substantial change in demand. Unfortunately, computing the minimum network subset is computationally hard and does not scale. Thus, the network is forced to operate with diminished performance during the recomputation periods. In this paper, we propose REsPoNse, a framework which overcomes the optimality-scalability trade-off. The insight in REsPoNse is to identify a few energy-critical paths off-line, install them into network elements, and use a simple online element to redirect the traffic in a way that enables large parts of the network to enter a low-power state. We evaluate REsPoNse with real network data and demonstrate that it achieves the same energy savings as the existing approaches, with marginal impact on network scalability and application performance. Nedeljko Vasic, Prateek Bhurat, Dejan M. Novakovic, Marco Canini, Satyam Shekhar, Dejan Kostic |
CoNEXT | 4 |
| 2011 | Evaluation and design of cache replacement policies under flooding attacksabstractA flow cache is a fundamental building block for flow-based traffic processing. Its efficiency is critical for the overall performance of a number of networked devices and systems. However, if not properly managed, the flow cache can be easily filled up and rendered ineffective by traffic patterns such as flooding attacks and scanning activities which, unfortunately, commonly occur in the Internet. In this paper, we show that popular cache replacement policies such as LRU cause the flow caches to evict the so called heavy-hitter flows during flooding attacks. To address this shortcoming, we build upon our recent work and construct a replacement policy that is more resilient to floods and yet performs similarly to other policies under common network traffic conditions. Martin Zádník, Marco Canini |
IWCMC | 2 |
| 2011 | Evolution of Cache Replacement Policies to Track Heavy-Hitter Flows
Martin Zádník, Marco Canini |
PAM | 2 |
| 2011 | Online testing of federated and heterogeneous distributed systemsabstractDiCE is a system for online testing of federated and heterogeneous distributed systems. We have built a prototype of DiCE and integrated it with an open-source BGP router. DiCE quickly detects three important classes of faults, resulting from configuration mistakes, policy conflicts and programming errors. Marco Canini, Vojin Jovanovic, Daniele Venzano, Dejan M. Novakovic, Dejan Kostic |
SIGCOMM | 1 |
| 2011 | Insomnia in the access: or how to curb access network related energy consumptionabstractAccess networks include modems, home gateways, and DSL Access Multiplexers (DSLAMs), and are responsible for 70-80% of total network-based energy consumption. In this paper, we take an in-depth look at the problem of greening access networks, identify root problems, and propose practical solutions for their user- and ISP-parts. On the user side, the combination of continuous light traffic and lack of alternative paths condemns gateways to being powered most of the time despite having Sleep-on-Idle (SoI) capabilities. To address this, we introduce Broadband Hitch-Hiking (BH2), that takes advantage of the overlap of wireless networks to aggregate user traffic in as few gateways as possible. In current urban settings BH2 can power off 65-90% of gateways. Powering off gateways permits the remaining ones to synchronize at higher speeds due to reduced crosstalk from having fewer active lines. Our tests reveal speedup up to 25%. On the ISP side, we propose introducing simple inexpensive switches at the distribution frame for batching active lines to a subset of cards letting the remaining ones sleep. Overall, our results show an 80% energy savings margin in access networks. The combination of B2 and switching gets close to this margin, saving 66% on average. Eduard Goma Llairo, Marco Canini, Alberto López Toledo, Nikolaos Laoutaris, Dejan Kostic, Pablo Rodriguez 0001, Rade Stanojevic, Pablo Yagüe Valentin |
SIGCOMM | 2 |
| 2011 | Finding Almost-Invariants in Distributed SystemsabstractIt is notoriously hard to develop dependable distributed systems. This is partly due to the difficulties in foreseeing various corner cases and failure scenarios while implementing a system that will be deployed over an asynchronous network. In contrast, reasoning about the desired distributed system behavior and the corresponding invariants is easier than reasoning about the code itself. Further, the invariants can be used for testing, theorem proving, and runtime enforcement. In this paper, we propose an approach to observe the system behavior and automatically infer invariants which reveal implementation bugs. Using our tool, Avenger, we automatically generate a large number of potentially relevant properties, check them within the time and spatial domains using traces of system executions, and filter out all but a few properties before reporting them to the developer. Our key insight in filtering is that a good candidate for an invariant is the one that holds in all but a few cases, i.e., an "almost-invariant". Our experimental results with the XORP BGP implementation demonstrate Avenger's ability to identify the almost-invariants that lead the developer to programming errors. Maysam Yabandeh, Abhishek Anand, Marco Canini, Dejan Kostic |
SRDS | 3 |
| 2011 | Toward Online Testing of Federated and Heterogeneous Distributed Systems
Marco Canini, Vojin Jovanovic, Daniele Venzano, Boris Spasojevic, Olivier Crameri, Dejan Kostic |
USENIX ATC | 1 |
| 2010 | Evolution of cache replacement policies to track heavy-hitter flowsabstractFlow-based network traffic processing, that is, processing packets based on some state information associated to the flows to which the packets belong, is a key enabler for a variety of network services and applications. This form of stateful traffic processing is used in modern switches [1] and routers that contain flow tables to implement forwarding, firewalls, NAT, QoS, and collect measurements. Martin Zádník, Marco Canini |
ANCS | 2 |
| 2009 | Experience with high-speed automated application-identification for network-managementabstractAtoZ, an automatic traffic organizer, provides control of how network-resources are used by applications. It does this by combining the high-speed packet processing of the NetFPGA with an efficient method for application-behavior labeling. AtoZ can control network resources by prohibiting certain applications and controlling the resources available to others. We discuss deployment experience and use real traffic to illustrate how such an architecture enables several distinct features: high accuracy, high throughput, minimal delay, and efficient packet labeling --- all in a low-cost, robust configuration that works alongside the enterprise access-router. Marco Canini, Wei Li 0009, Martin Zádník, Andrew W. Moore 0002 |
ANCS | 1 |
| 2009 | Tracking elephant flows in internet backbone traffic with an FPGA-based cacheabstractThis paper presents an FPGA-friendly approach to tracking elephant flows in network traffic. Our approach, single step segmented least recently used (S3-LRU) policy, is a network traffic-friendly replacement policy for maintaining flow states in a Naiumlive hash table (NHT). We demonstrate that our S3-LRU approach preserves elephant flows: conservatively promoting potential elephants and evicting lowrate flows in LRU manner. Our approach keeps flow-state of any elephant since start-of-day and provides a significant improvement over filtering approaches proposed in previous work. Our FPGA-based implementation of the S3-LRU in combination with an NHT suites well the parallel access to block memories while capitalising on the retuning of parameters through dynamic-reprogramming. Martin Zádník, Marco Canini, Andrew W. Moore 0002, David J. Miller 0005, Wei Li 0009 |
FPL | 2 |
| 2009 | Efficient application identification and the temporal and spatial stability of classification schema
Wei Li 0009, Marco Canini, Andrew W. Moore 0002, Raffaele Bolla |
Comput. Networks | 2 |
| 2008 | On the Double-Faced Nature of P2P TrafficabstractOver the last few years, peer-to-peer (P2P) file sharing applications have evolved to become a major traffic source in the Internet. The ability to quantify their impact on the network, as a consequence of both signaling and download traffic, is fundamental to a number of network operations, including traffic engineering, capacity planning, quality of service, forecasting for long-term provisioning, etc. We present here a measurement study on the characteristics of the traffic associated with different P2P applications. Our aim is to offer useful insight into the nature of P2P traffic, which we consider a step toward building P2P traffic aggregates generators in simulative environments. We show that P2P traffic can be divided into two distinguished behavioral profiles, which, independently of the application protocol, present significant differences in the average and standard deviation of four measurements: arrival times, durations, volumes and average packet sizes ofP2P conversations. These profiles well represent the typical behavior of signaling and download traffic. Based on our findings, we argue that, if such distinction is not taken into account, the statistical measurements needed to model P2P traffic aggregates would result biased, and potentially bring to misleading results. Raffaele Bolla, Marco Canini, Riccardo Rapuzzi, Michele Sciuto |
PDP | 2 |