EDBT 2026 Demo / reviewers in the wild / expert
Marcelo Caggiani Luizelli
dblp:140/8276
· DBLP profile ↗
66ranked-venue papers
7as first author
44since 2021 · last 2026
0000-0003-0537-3052ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 18 · 5 first-author · 6 since 2021Systems, architecture and hardware · 13 · 10 since 2021Software engineering, systems software and programming languages · 10 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Why Large Language Models Struggle with Cloud Instance Selection
Matheus Machado, Matheus M. Costa, Marcelo Caggiani Luizelli, Fábio D. Rossi, Arthur Francisco Lorenzon |
CLOSER | 3 |
| 2026 | Mininet-AI: Emulating Network Agentic AI
Pedro da Silva Santiago, Diogo Mainart Monteiro, Victor Hugo Schneider Lopes, Francisco G. Vogt, Fábio D. Rossi, Christian Esteve Rothenberg, Marcelo Caggiani Luizelli |
NetSoft | 7 |
| 2026 | FlexUP: CU-UP Disaggregation on Programmable Data Planes
Francisco Germano Vogt, Victor Hugo Schneider Lopes, Fabricio Rodriguez, Marcelo Caggiani Luizelli, P. Gyanesh Patra, Christian Esteve Rothenberg, Gergely Pongrácz, Chrysa Papagianni |
NetSoft | 4 |
| 2026 | DEMO: QU4C: High-Throughput QUIC Traffic Reproduction on Programmable Switches
Filipo G. Costa, Francisco G. Vogt, Fabricio Rodriguez, Marcelo Caggiani Luizelli, Christian Esteve Rothenberg |
SIGCOMM | 4 |
| 2025 | Energy-Aware Node Selection for Cloud-Based Parallel Workloads with Machine Learning and Infrastructure as Code
Denis B. Citadin, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon |
CLOSER | 3 |
| 2025 | WFQ-Based SLA-Aware Edge Applications Provisioning
Pedro Henrique Sachete Garcia, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli, Paulo Silas Severo de Souza, Fábio D. Rossi |
CLOSER | 3 |
| 2025 | LLM-Based Adaptive Digital Twin Allocation for Microservice Workloads
Pedro Henrique Sachete Garcia, Ester S. Oribes, Ivan Mangini Lopes Júnior, Braulio Marques de Souza, Ângelo Vieira, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli, Paulo Silas Severo de Souza, Fábio D. Rossi |
CLOSER | 7 |
| 2025 | Towards Optimizing Cost and Performance for Parallel Workloads in Cloud Computing
William Maas, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon |
CLOSER | 3 |
| 2025 | OLEO: Optimizing LEO Satellites Offloading of Cloud-Edge ApplicationsabstractLow Earth Orbit (LEO) satellite constellations enable cloud-edge computing for latency-sensitive applications. However, frequent satellite mobility challenges resource allocation and service continuity, leading to disruptions and inefficient provisioning. Existing strategies often overlook temporal constraints, resulting in frequent migrations and degraded performance. We propose OLEO, a heuristic strategy that optimizes application offloading by prioritizing satellites with higher exposure time, reducing unnecessary migrations and improving resource utilization. Experimental results show that OLEO provisions up to 1.5 X more application requests while reducing migrations by up to 20% in comparison baselines. Gabriel P. Costa, Diogo Matos, Pedro Henrique Sachete Garcia, Arthur Francisco Lorenzon, Fábio D. Rossi, Marcelo Caggiani Luizelli |
ISCC | 6 |
| 2025 | Harnessing P4 for In-Network Unmanned Aerial Vehicle Collision AvoidanceabstractWith the advent of next-generation networks, new applications across multiple domains are gaining traction. This shift demands a redefined network paradigm, where ultrareliable, low-latency communication is key. In this work, we explore and extend the concept of in-network programmability in new directions. Unlike conventional approaches, we leverage P4 data plane programmability to implement an in-network collision avoidance algorithm in a UAV scenario. We evaluate our hardware-based implementation under different conditions, including latency and velocity, demonstrating that it efficiently detects and prevents collisions. Our results show the impact of end-to-end latency and highlight how in-network processing can be a valuable ally for time-sensitive tasks, paving the way for future advancements in hardware-based in-network applications. Fabricio Rodriguez, Francisco Germano Vogt, Marcelo Caggiani Luizelli, Christian Esteve Rothenberg, Géza Szabó |
NetSoft | 3 |
| 2025 | P4DMA: Unlocking High-Performance RDMA Traffic Generation on Programmable SwitchesabstractRemote Direct Memory Access (RDMA) is a key technology in modern data centers, enabling low-latency and high-throughput communication. However, evaluating RDMA performance and validating network designs often requires costly hardware setups or simulation tools with limited performance and realism. In this work, we present P4DMA, a system that leverages programmable switch ASICs to generate realistic high-performance RDMA traffic. By implementing RDMA traffic patterns using the P4 language, P4DMA enables researchers and practitioners to generate RDMA workloads at line rate, without relying on traditional RDMA NICs (RNICS). With Tofino's traffic generation capacity of up to Tbps, P4DMA offers a novel approach to stress and evaluate RDMA-capable infrastructures, and accelerate the prototyping of new RDMA-based applications. Filipo G. Costa, Francisco Germano Vogt, Fabricio Rodriguez, Suneet Kumar Singh, Marcelo Caggiani Luizelli, Christian Esteve Rothenberg |
NetSoft | 5 |
| 2025 | Bridging TSN and 5G: Synchronization and Flow Mapping for Smart ManufacturingabstractThe increasing demand for real-time industrial applications demands deterministic communication across networked devices. Time-sensitive networking (TSN) offers reliable performance to meet these requirements, compatible with the current Real-Time Ethernet (RTE) solutions applied in the industry. While TSN ensures determinism inside Ethernet networks, its integration with 5G/6G wireless technologies can overcome flexibility limitations. This paper explores key challenges in integrating TSN with mobile networks, focusing on device synchronization and flow mapping. We introduce a tool aligned with the 3GPP specifications that covers time synchronization across TSN domains via 5G and implements flow mapping based on application-specific quality of service (QoS) requirements. We evaluate these features under various scenarios, including a use case for an industrial production line use case. Our solution is developed as an open-source framework for the OMNeT++ simulator, supporting reproducibility and continuous updates and paving the way for customizable network solutions. Sergio Rossi Brito da Silva, Francisco Germano Vogt, Fabricio Rodriguez, Marcelo Caggiani Luizelli, Christian Esteve Rothenberg, P. Gyanesh Patra |
NetSoft | 4 |
| 2025 | P4Timely: Evaluating Time Synchronization Resilience in Programmable NetworksabstractAccurate time synchronization is essential for emerging networked applications that demand low latency, high precision, and coordinated operations, such as those found in data centers, time-sensitive networking (TSN), and distributed systems. In this demo, we introduce P4Timely, a flexible and programmable framework designed to evaluate and prototype time synchronization protocols in real time. Built using P4 and commodity programmable hardware, P4Timely enables users to easily configure synchronization settings through a simple interface, specify network conditions such as delay, jitter, and packet loss, and define whether nodes act as synchronizing or synchronized entities. P4Timely supports dynamic reconfiguration at runtime and provides built-in monitoring tools to assess synchronization accuracy under diverse network conditions. Sergio Rossi Brito da Silva, Francisco Germano Vogt, Marcelo Caggiani Luizelli, Fabricio Rodriguez, Flávio Geraldo Coelho Rocha, Christian Esteve Rothenberg |
NetSoft | 3 |
| 2025 | The Offloading Dilemma: Exploring the Boundaries of Programmable Data PlanesabstractIn-Network Computing (INC) is transforming how we design and evaluate modern networks by enabling programmable data planes to execute computational tasks at line rate. However, deciding which functions to offload, to which device and using which strategy is not a trivial task. This PhD project explores the opportunities, challenges and boundaries of offloading network functions and application logic to programmable switches and SmartNICs. Through a series of case studies, we demonstrate that INC-based solutions enhance both network efficiency and application performance. On the other hand, we also demonstrate the tradeoffs of these offloadings and how they can impact the infrastructure in terms of performance and resource utilization. Our findings provide actionable insights for the deployment of INC in real-world scenarios, highlighting where offloading delivers the most value and where traditional approaches remain preferable. Francisco Germano Vogt, Marcelo Caggiani Luizelli, Christian Esteve Rothenberg |
NetSoft | 2 |
| 2025 | Distributed Graph Neural Networks in Programmable Data PlanesabstractThe ability to redefine the data plane behavior with programmable network devices provides a plethora of novel possibilities for in-network computing. One of these possibilities is embedding Artificial Intelligence (AI) and Machine Learning (ML) techniques directly in the data plane. Motivations include reducing decision latency and closing the control loop-i.e., performing measurements, learning, decisions, and actions directly in the data plane. However, running entire AI/ML algorithms in a single device might be infeasible due to memory and computing constraints. This work addresses the research challenges of running a Graph Neural Network (GNN) in a set of devices of a programmable data plane. Our hypothesis is that by distributing the GNN processing across the devices, the GNN uses instantaneous snapshots of the global network state and can act more quickly. As a proof of concept, we trained and evaluated a distributed GNN to perform explicit congestion notifications based on Data Center Transmission Control Protocol (DCTCP). We verified the feasibility of GNN classification in the data plane through simulations and both software and hardware switch experiments with bmv2 and Intel Tofino. Ivan Peter Lamb, Pedro Arthur Pinheiro Rosa Duarte, Jonatas Adilson Marques, Marcelo Caggiani Luizelli, Luciano Paschoal Gaspary, Anderson Tavares, Ronaldo A. Ferreira, Ítalo S. Cunha, José Rodrigo Azambuja, Weverton Luis da Costa Cordeiro |
NOMS | 4 |
| 2025 | Toward real-time IoT multi-sensor data orchestration on wireless sensor networks
Pedro Henrique Sachete Garcia, Marcelo Caggiani Luizelli, Fábio D. Rossi |
J. Supercomput. | 2 |
| 2025 | MAPER: mobility-aware energy-efficient container registry migrations for edge computing infrastructures
Daniel Chaves Temp, Alexandre A. F. da Costa, Ângelo Vieira, Ester S. Oribes, Ivan M. Lopes, Paulo Silas Severo de Souza, Marcelo Caggiani Luizelli, Arthur Francisco Lorenzon, Fábio D. Rossi |
J. Supercomput. | 7 |
| 2024 | An ANN-Guided Multi-Objective Framework for Power-Performance Balancing in HPC SystemsabstractPower-performance efficiency has become one of the most critical issues in evolving High-Performance Computing systems (HPC) towards Exaflops. Thread-level parallelism (TLP) exploitation, dynamic voltage and frequency scaling (DVFS), and uncore frequency scaling (UFS) are methods widely applied to better balance the power consumption and performance improvements of parallel applications. However, selecting ideal combinations of these knobs for every application is challenging due to the massive number of possible solutions, as there is no unique combination that delivers at the same time the best performance and the lowest power consumption. Given that, we propose HPC-PPO (power-performance optimizer), a multi-objective optimization strategy driven by an artificial neural network that leverages hardware and software features of parallel applications to predict Pareto-efficient configurations of TLP degree, DVFS, and UFS that optimize the balance between power and performance. When validating HPC-PPO on three multicore processors with twenty-five applications, we show that HPC-PPO can predict combinations very close to the best ones found by an exhaustive search. We also show that the Pareto-efficient configurations predicted by HPC-PPO improve parallel applications' performance by 30.7% while spending 23.9% less power when compared to state-of-the-art strategies. William Maas, Paulo Silas Severo de Souza, Marcelo Caggiani Luizelli, Fábio D. Rossi, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon |
CF | 3 |
| 2024 | Multi-Tenant Programmable Switch Virtualization Leveraging Explicit Resource SharingabstractWith the migration of traditional computer networks to the Software-defined Networking paradigm, flexibility is a core feature that novel technologies must provide. In this context, virtualization is gaining traction in Programmable Data Planes (PDPs) as a means of achieving greater flexibility, with several solutions in the literature for instantiating virtual programmable switches on the same host device. Virtualization brings numerous advantages, enabling multi-tenancy in programmable data/research center networks and greater device resource utilization. Nevertheless, enabling a complete multitenant solution, in which the tenants have disjoint sets of virtual devices, requires management and security considerations not yet approached in previous investigations. Previous works focus mainly on the core underlying technology necessary to deploy multiple devices in the same physical host. This paper presents a PDP virtualization architecture based on program composition and access control for securely managing virtual switches from different tenants. Additionally, we define extensions to PDP programmability, allowing tenants to specify shared elements, such as tables, between their virtual devices. Our experiments highlight the ability to transparently manage multiple virtual switches hosted in the same physical device in networking scenarios with multiple tenants. Ivan Peter Lamb, Pedro Arthur Pinheiro Rosa Duarte, Marcelo Caggiani Luizelli, Luciano Paschoal Gaspary, José Rodrigo Azambuja, Weverton Luis sa Costa Cordeiro |
CNSM | 3 |
| 2024 | DigiNet: Scaling up Provisioning of Network Digital TwinabstractThe pursuit of self-driving networks is increasing pressure on adopting intelligent, edge-based networking services. However, deploying autonomous network models within operational and large-scale infrastructures entails substantial risks that require rigorous verification and validation procedures. In this context, the application of a Network Digital Twin (NDT) is emerging as a viable approach towards intelligent network decision-making based on high-fidelity models built upon digital representations of physical network devices (i.e., Digital Twins). In this paper, we take the first steps towards efficiently provisioning NDT models. To that end, we introduce the Digital Twin Network Provisioning Problem (DigiNet), which encompasses the optimal placement of NDT models and the efficient collection of telemetry data for synchronizing NDT models with their physical counterparts. We theoretically formalize DigiNet as a Mixed-Integer Linear Programming (MILP) model and present a polynomial-time heuristic. Our results show that DigiNet outperforms baseline approaches by up to 10x regarding the number of NDT models provisioned. Marcelo Caggiani Luizelli, Francisco Germano Vogt, Paulo Silas Severo de Souza, Arthur Francisco Lorenzon, Roberto Irajá Tavares da Costa Filho, Fábio D. Rossi, Rodrigo N. Calheiros, Christian Esteve Rothenberg |
NetSoft | 1 |
| 2024 | Spinner: Enabling In-network Flow Clustering Entirely in a Programmable Data PlaneabstractData plane programmability is redesigning the way we manage and operate forwarding devices. However, most of the algorithmic decisions performed by data planes are still deterministic and control-plane dependent. We argue that it is possible to break this dependency and make the data plane intelligent, so that it can learn the infrastructure state autonomously. Despite existing efforts to make data planes intelligent, little has been done to design unsupervised ML algorithms that fit the architectural constraints of programmable devices. Executing such approaches in the data plane has the potential to reduce the overall decision-making time, thus meeting packet processing deadlines (which are in the order of nanoseconds). In this paper, we propose Spinner, the first effort to deliver an unsupervised Machine Learning (ML) approach entirely in programmable devices. Spinner is a flow clustering algorithm designed to fit existing architectural constraints of SmartNICs, and that can reach line rate for most packet sizes with complexity O(k). To demonstrate the potential behind in-network clustering, we prototype and deploy Spinner in a programmable testbed and use it to enhance Explicit Congestion Notifications (ECN) at the server side. Spinner-enhanced TCP provides up to 2x higher throughput when comparing to de-facto TCP implementations. Luigi Cannarozzo, Thiago Bortoluzzi Morais, Paulo Silas Severo de Souza, Leonardo Gobatto, Ivan Peter Lamb, Pedro Arthur Pinheiro Rosa Duarte, José Rodrigo Azambuja, Arthur Francisco Lorenzon, Fábio D. Rossi, Weverton Luis da Costa Cordeiro, Marcelo Caggiani Luizelli |
NOMS | 11 |
| 2024 | PIPO-TG: Parameterizable High-Performance Traffic GenerationabstractIn recent years, the increasing demand for network resources due to real-time applications and data-intensive activities has posed challenges in managing and optimizing network performance. To assess network performance, security, and efficiency, traffic generation plays a crucial role. We introduce PIPO-TG, a Tofino-based traffic generation for high-performance experiments. The primary objective of PIPO-TG is to generate realistic and diverse traffic patterns, enabling researchers to evaluate network performance under varying conditions providing customizable packet forwarding with P4 programmable data planes. Our main contributions include user-defined packet header customization and open-source code for reproducibility. These efforts foster collaboration within the research community to advance traffic generation techniques. We show that PIPO-TG only requires a few lines of code to simulate heterogeneous network scenarios (e.g., traffic bursts and DDoS attacks) while maintaining hardware performance and flexibility. Filipo G. Costa, Francisco Germano Vogt, Fabricio Rodriguez, Ariel Góes de Castro, Marcelo Caggiani Luizelli, Christian Esteve Rothenberg |
NOMS | 5 |
| 2024 | A neural network framework for optimizing parallel computing in cloud servers
Everton Camargo de Lima, Fábio D. Rossi, Marcelo Caggiani Luizelli, Rodrigo N. Calheiros, Arthur Francisco Lorenzon |
J. Syst. Archit. | 3 |
| 2024 | Synergistically Rebalancing the EDP of Container-Based Parallel ApplicationsabstractThe use of containers has become standard in cloud environments. However, many parallel applications in containers will not present gains proportional to the extra available hardware. This inefficient use of hardware naturally leads to energy consumption waste. With that in mind, we proposeTT-Autoscaling. It works at two different levels: a) in the container, by automatically and transparently tuning the number of threads at runtime of the application, in a way to optimize the trade-off between energy and performance; b) in the cloud infrastructure, by smartly transferring the released resources to other containers that may run in parallel, making better use of the available resources. We compareTT-Autoscalingto the default execution of containers (serial execution with the maximum number of threads), showing 55.8% of performance improvements, 53.6% of energy reductions, and 79.5% of EDP improvements. We also show thatTT-Autoscalingoutperforms strategies that apply vertical autoscalers proposed by orchestrator tools. Vinicius S. da Silva, Everton Camargo de Lima, Janaina Schwarzrock, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2023 | Towards Optimizing the Edge-to-Cloud Continuum Resource Allocation
Igor Ferrazza Capeletti, Ariel Góes de Castro, Daniel Chaves Temp, Paulo Silas Severo de Souza, Arthur Francisco Lorenzon, Fábio D. Rossi, Marcelo Caggiani Luizelli |
CLOSER | 7 |
| 2023 | Latency-Aware Cost-Efficient Provisioning of Composite Applications in Multi-Provider Clouds
Daniel Chaves Temp, Igor Ferrazza Capeletti, Ariel Góes de Castro, Paulo Silas Severo de Souza, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli, Fábio D. Rossi |
CLOSER | 6 |
| 2023 | Taking Detours: An In-Network Fault-Tolerant Probing Planning for In-Band Network TelemetryabstractIn-band Network Telemetry (INT) is a novel network monitoring approach mainly fostered by programmable network devices. Despite existing efforts toward the orchestration of INT, little has yet been done to provide fault-tolerant mechanisms in the data plane (e.g., to address hardware failure). In this paper, we introduce InPatching - an in-network approach to fast recover INT-based monitoring from network link failures. InPatching is implemented in the data plane and allows the application of detours in an autonomous and coordinated manner without the control plane intervention. To provide efficient detours to INT solutions, we formalize the fault-tolerant probing planning for INT by means of a MILP (Mixed-Integer Linear Programming) model. We prototype InPatching in P4 and we show that it can recover from fault conditions much faster than control plane solutions (up to 18X), while not imposing substantial overhead. Ariel Góes de Castro, Igor Capelletti, Fábio D. Rossi, Arthur Francisco Lorenzon, Roberto Irajá Tavares da Costa Filho, Christian Esteve Rothenberg, Marcelo Caggiani Luizelli |
ICC | 7 |
| 2023 | QoEyes: Towards Virtual Reality Streaming QoE Estimation Entirely in the Data PlaneabstractIn recent years, advances in virtual reality (VR) technologies (e.g., high-quality VR headsets) have enabled a new perspective of experiences for users (e.g., gaming, online events). However, ensuring the user experience is still a challenge. Existing solutions are limited to measuring and estimating QoE at the user plane (e.g., VR player) or at the control plane, imposing unfeasible latency for different scenarios (5G networks and beyond). In this work, we propose QoEyes, an in-network QoE estimation based on the use of Inter-Packet-Gap (IPG) in programmable devices. Our results show that the IPG measured on the data plane is strongly linked to QoE, yielding an accurate data plane QoE estimate. Francisco Germano Vogt, Fabricio Rodriguez, Ariel Góes de Castro, Marcelo Caggiani Luizelli, Christian Esteve Rothenberg, Gergely Pongrácz |
NetSoft | 4 |
| 2023 | Demo of QoEyes: Towards Virtual Reality Streaming QoE Estimation Entirely in the Data PlaneabstractRecent advances in VR technology have created new user experiences (e.g., online events, gaming). However, ensuring the user experience is still a challenge. Mostly because Quality of Experience (QoE) measurement is limited to the user or control plane, causing high latencies for different scenarios (e.g., 5G networks and beyond). To address this challenge, we present QoEyes, an in-network QoE estimation technique based on Inter-Packet-Gap (IPG) measured in programmable devices. Our results show that a strong estimate of the user’s QoE can be provided by measuring the IPG on the data plane. Additionally, in this demonstration, we show this QoE estimate and other related metrics in real time, using a Grafana dashboard running in our monitoring server. Francisco Germano Vogt, Fabricio Rodriguez, Ariel Góes de Castro, Marcelo Caggiani Luizelli, Christian Esteve Rothenberg, Gergely Pongrácz |
NetSoft | 4 |
| 2023 | NeurOPar, A Neural Network-Driven EDP Optimization Strategy for Parallel WorkloadsabstractThe pursuit of energy efficiency has been driving the development of techniques to optimize hardware resource usage in high-performance computing (HPC) servers. On multicore architectures, thread-level parallelism (TLP) exploitation, dynamic voltage and frequency scaling (DVFS), and uncore frequency scaling (UFS) are three popular methods applied to improve the trade-off between performance and energy consumption, represented by the energy-delay product (EDP). However, the complexity of selecting the optimal configuration (TLP degree, DVFS, and UFS) for each application poses a challenge to software developers and end-users due to the massive number of possible configurations. To tackle this challenge, we propose NeurOpar, an optimization strategy for parallel workloads driven by an artificial neural network (ANN). It uses representative hardware and software metrics to build and train an ANN model that predicts combinations of thread count and core/uncore frequency levels that provide optimal EDP results. Through experiments on four multicore processors using twenty-five applications, we demonstrate that NeurOPar predicts combinations that yield EDP values close to the best ones achieved by an exhaustive search and improve the overall EDP by 42% compared to the default execution of HPC applications. We also show that NeurOPar can enhance the execution of parallel applications without incurring the performance and energy penalties associated with online methods by comparing it with two state-of-the-art strategies. Cristiano A. Künas, Fábio D. Rossi, Marcelo Caggiani Luizelli, Rodrigo N. Calheiros, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon |
SBAC-PAD | 3 |
| 2023 | Smart resource allocation of concurrent execution of parallel applicationsabstractAbstract Thread‐level parallelism (TLP) has been widely exploited to optimize computational resource usage in high‐performance systems. However, as many applications do not scale as the number of threads increase, resources will be wasted when the application executes with the maximum possible number of threads (i.e., the default execution) rather than fewer threads (thread throttling) that may use the resources more efficiently. Hence, instead of executing only one application with as many threads as possible, one can run more applications simultaneously by applying thread throttling to each one. The primary outcome of this strategy is a significant reduction in the total execution time and energy consumption when the system needs to execute a list of applications. Given that, we propose a smart resource allocation (SRA) for concurrent parallel application execution. It automatically finds the ideal degree of TLP for each application and guides the simultaneous parallel applications execution. When running 25 well‐known benchmarks on three multicore systems and comparing SRA to state‐of‐the‐art strategies (e.g., Batch, Equal policy, and Scalability), SRA improves the EDP by 87.4% over the Batch strategy; 75.5% over the Equal policy; and 38.8% over the scalability strategy. Vinicius S. da Silva, Angelo Gaspar Diniz Nogueira, Everton Camargo de Lima, Hiago Rocha, Matheus S. Serpa, Marcelo Caggiani Luizelli, Fábio D. Rossi, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
Concurr. Comput. Pract. Exp. | 6 |
| 2023 | Mobility-Aware Registry Migration for Containerized Applications on Edge Computing Infrastructures
Daniel Chaves Temp, Paulo Silas Severo de Souza, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli, Fábio D. Rossi |
J. Netw. Comput. Appl. | 4 |
| 2023 | HH-IPG: Leveraging Inter-Packet Gap Metrics in P4 Hardware for Heavy Hitter DetectionabstractThe research community has recently proposed several solutions based on modern programmable switches to detect entirely in the data plane the flows exceeding pre-determined threshold in a time window, i.e., Heavy Hitters (HH). This is commonly achieved by dividing the network stream into fixed time slots and identifying each separately without considering the traffic trends from previous intervals. In this work, we show that using specified time windows can lead to high inaccuracies. We make a case for rethinking how switches analyze the incoming packets and propose to leverage per-flow Inter Packet Gap (IPG) analytics instead of using flow counters for HH detection. We propose an algorithm and present a P4 pipeline design using this new metric in mind. We implement our solution on P4 hardware and experimentally evaluate it against real traffic traces. We show that our results are more accurate than related work by up to 20% while reducing the control channel overhead by up to two orders of magnitude. Finally, we showcase a QoS-oriented application of the proposed dataplane-only IPG-based HH detection in a mobile network scenario. Suneet Kumar Singh, Christian Esteve Rothenberg, Marcelo Caggiani Luizelli, Gianni Antichi, Pedro Henrique Gomes, Gergely Pongrácz |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2022 | Towards Efficient Selective In-Band Network Telemetry Report Using SmartNICs
Ronaldo Canofre, Ariel Góes de Castro, Arthur Francisco Lorenzon, Fábio D. Rossi, Marcelo Caggiani Luizelli |
AINA (1) | 5 |
| 2022 | Multivariate Interpolation at the Edge to Infer Faulty IoT Sensor Metrics
Marcos Paulo Konzen, Patric Lincoln Ramires Izolan, Fábio Júnior Griesang, Paulo Silas Severo de Souza, Tiago Ferreto, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli, Júlio C. B. de Mattos, Cinara Ewerling da Rosa, Fábio D. Rossi |
CLOSER | 7 |
| 2022 | DyPro: Dynamic Probing Planning for In-Band Network TelemetryabstractIn-band Network Telemetry (INT) is a novel net-work monitoring mechanism that improves fine-grained net-work visibility. Despite the increasing research efforts towards the orchestration of INT data acquisition, little has yet been done to efficiently collect telemetry data from the network considering monitoring applications requirements. In this paper, we introduce DyPro - a dynamic probing planning for INT. In particular, DyP ro ensures that telemetry dependencies are always satisfied by monitoring application requirements. We theoretically formalize it as a Mixed-Integer Linear Programming (MILP) optimization model and propose a heuristic procedure to efficiently solve it. Results show that DyP ro can outperform state-of-the-art solutions by up to 5x regarding the percentage of monitoring applications satisfied. Leandro M. Dallanora, Ariel Góes de Castro, Roberto Irajá Tavares da Costa Filho, Fábio D. Rossi, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli |
ISCC | 6 |
| 2022 | Intelligent Mapping of Virtualized Services on Multi-domain Networks
Vinicius Fulber-Garcia, Marcelo Caggiani Luizelli, Carlos Raniery Paula dos Santos, Eduardo Jaques Spinosa, Elias P. Duarte Jr. |
ISDA (1) | 2 |
| 2022 | Optimizing the EDP of OpenMP applications via concurrency throttling and frequency boosting
Sandro Matheus V. N. Marques, Matheus S. Serpa, Antoni Navarro Muñoz, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
J. Syst. Archit. | 5 |
| 2021 | The Actual Cost of Programmable SmartNICs: Diving into the Existing Limits
Pablo B. Viegas, Ariel Góes de Castro, Arthur Francisco Lorenzon, Fábio D. Rossi, Marcelo Caggiani Luizelli |
AINA (1) | 5 |
| 2021 | Synergically Rebalancing Parallel Execution via DCT and Turbo BoostingabstractThe increasing use of cloud and HPC systems put more pressure on the efficient utilization of hardware resources to keep costs low. Many dynamic concurrency throttling (DCT) techniques have successfully used to tune the number of executing threads to better balance a parallel application according to its available scalability. Similarly, boosting frequency strategies have been used to speed up the sequential parts’ execution. Given that, we propose Poseidon, the first transparent and automatic approach that cooperatively exploits both techniques to rebalance OpenMP applications without any preprocessing, with no code transformation, recompilation, or OS modification. Sandro Matheus V. N. Marques, Thiarles S. Medeiros, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
DAC | 4 |
| 2021 | Combining Thread Throttling and Mapping to Optimize the EDP of Parallel ApplicationsabstractThread-throttling and mapping strategies have been used together to make better use of hardware resources and improve the energy-delay product (EDP) of high-performance computing (HPC) systems. However, the design space exploration significantly grows with the increasing number of cores in those systems, making the task of finding the ideal number of active threads and allocating strategy a challenging task. On top of that, parallel applications present various patterns, such as irregularity, unbalanced computations, or high rates of communications. Given these considerations, we propose ETTM, an EDPaware thread-throttling and mapping optimization strategy that automatically finds an ideal combination of number of threads and thread mapping strategy. With the execution of eighteen well-known benchmarks on three multicore architectures, we show that EDP can be significantly improved when running applications with the solution found by EETM1. Gustavo Berned, Thiarles S. Medeiros, Matheus S. Serpa, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
PDP | 5 |
| 2021 | Optimizing Parallel Applications via Dynamic Concurrency Throttling and Turbo BoostingabstractWith the increasing number of cores in modern systems, dynamic concurrency throttling (DCT) and turbo-boosting techniques are becoming a solution to better use the hardware resources. While DCT techniques tune the number of running threads, boosting techniques speed up sequential phases or unbalanced threads. However, as each region of an application may behave differently, optimizing both knobs is not straightforward. Hence, we propose two strategies that apply DCT and turbo-boosting: DBF, which aims to find an ideal configuration for each parallel/sequential region, and DBC, which considers the combination of parallel/sequential regions during the optimization. We show that DBF and DBC improve the EDP by up to 19% and 27% compared to a DCT-only strategy and by up to 95% and 96% compared to a Boost-only technique. We also show that DBF is more suitable for applications with high variability in the CPU workload, while DBC is better when there is low workload variability. Sandro Matheus V. N. Marques, Thiarles S. Medeiros, Matheus S. Serpa, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
PDP | 5 |
| 2021 | Mitigating the processor aging through dynamic concurrency throttling
Thiarles S. Medeiros, Luan Pereira, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
J. Parallel Distributed Comput. | 4 |
| 2021 | Low learning-cost offline strategies for EDP optimization of parallel applications
Gustavo Berned, Fábio D. Rossi, Marcelo Caggiani Luizelli, Samuel Xavier de Souza, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
J. Syst. Archit. | 3 |
| 2020 | CUSCO: A Customizable Solution for NFV Composition
Vinicius Fulber-Garcia, Marcelo Caggiani Luizelli, Carlos Raniery Paula dos Santos, Elias P. Duarte Jr. |
AINA | 2 |
| 2020 | A Heuristic Approach for Large-Scale Orchestration of the In-band Data Plane Telemetry Problem
Rumenigue Hohemberger, Arthur Francisco Lorenzon, Fábio D. Rossi, Marcelo Caggiani Luizelli |
AINA | 4 |
| 2020 | Decreasing the Learning Cost of Offline Parallel Application Optimization StrategiesabstractMany parallel applications do not scale as the number of threads increases, which means that executing them with the maximum possible number of threads will not always deliver the best outcome in performance, energy consumption, or the tradeoff between both (represented by the energy-delay product- EDP). Given that, several strategies, online and offline, have already been proposed to rightly tune the number of threads according to the application. While the former can capture some behaviors that can only be known at runtime, the latter do not impose any execution overhead and can use more efficient and costly algorithms. However, these learning algorithms in static strategics may take several hours, precluding their use or a smooth migration across different systems. In this scenario, we propose a generic methodology for such offline strategies to significantly decrease the learning time by inferring the execution behavior of parallel applications using smaller input sets than the ones used by the target applications. Through the execution of eighteen well-known benchmarks on two multicore processors, we show that our methodology is capable of converging to results that are very close to those that use the regular input set, but converging 84.7% faster, on average. We also show that such a strategy delivers better results than a dynamic one, presenting an EDP 7.7% lower, on average, when executing the applications with the number of threads found during learning. Finally, we also compare our learning methodology with an exhaustive search. It has an average learning cost (i.e., the time spent by our search algorithm to find the best configuration) of only 3.1% to optimize the EDP of the entire benchmark set1. Gustavo Berned, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
PDP | 3 |
| 2020 | Modeling and Simulating Daily Power Budgets for Sustainable Data CentersabstractA novel energy-efficient scenario that makes possible to maintain sustainable data centers to feed part of resources through renewable energy sources has emerged. As renewable energies are accumulated in the form of power budgets, data centers must adapt a slice of the processing resources required to meet applications at those limits. This work is modeling and simulating the computing capacity of a data center according to daily power budgets from different sources of renewable energy. The results showed that based on the daily energy harvest of today's renewable energy sources, intelligent resource orchestration could use such energy so that up to 40% of what is needed to maintain a quality-of-service data center comes from non-polluting sources. Rumenigue Hohemberger, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli, Fábio D. Rossi |
PDP | 3 |
| 2020 | Dissecting the Performance of VR Video Streaming through the VR-EXP Experimentation PlatformabstractTo cope with the massive bandwidth demands of Virtual Reality (VR) video streaming, both the scientific community and the industry have been proposing optimization techniques such as viewport-aware streaming and tile-based adaptive bitrate heuristics. As most of the VR video traffic is expected to be delivered through mobile networks, a major problem arises: both the network performance and VR video optimization techniques have the potential to influence the video playout performance and the Quality of Experience (QoE). However, the interplay between them is neither trivial nor has it been properly investigated. To bridge this gap, in this article, we introduce VR-EXP, an open-source platform for carrying out VR video streaming performance evaluation. Furthermore, we consolidate a set of relevant VR video streaming techniques and evaluate them under variable network conditions, contributing to an in-depth understanding of what to expect when different combinations are employed. To the best of our knowledge, this is the first work to propose a systematic approach, accompanied by a software toolkit, which allows one to compare different optimization techniques under the same circumstances. Extensive evaluations carried out using realistic datasets demonstrate that VR-EXP is instrumental in providing valuable insights regarding the interplay between network performance and VR video streaming optimization techniques. Roberto Irajá Tavares da Costa Filho, Marcelo Caggiani Luizelli, Stefano Petrangeli, Maria Torres Vega, Jeroen van der Hooft, Tim Wauters, Filip De Turck, Luciano Paschoal Gaspary |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2019 | The Impact of Parallel Programming Interfaces on the Aging of a Multicore Embedded ProcessorabstractIn order to meet the increasing performance demand of applications, the amount of cores in a single chip package has been increasing. However, the heat has been rising at a higher scale, which accelerates the aging process in modern processors. Therefore, wisely balancing the use of resources is important to extend its longevity. Frequency performance stagnates after a certain amount of concurrent threads starts executing. In such cases, the only result is a temperature rise that directly influences the aging process, reducing the processor lifetime. This unbalance between threads can be originated from many factors, which includes the way threads communicate and synchronize. Considering that those characteristics are related to the Parallel Programming Interface (PPI) used to parallelize the application, this work proposes to evaluate three widely used PPIs executing on an embedded multicore. We show that, depending on the characteristic of the application, by only switching from one PPI to another, it is possible to reduce the effects of aging. For that, we have developed a model based on the Arrhenius equation. We show that OpenMP has a lower impact on the processor aging for memory-bound applications: up to 38% and 68% lower than PThreads and MPI, respectively. On the other hand, PThreads presents the lowest impact on the processor aging for CPU-bound applications. Ângelo Vieira, Paulo Silas Severo de Souza, Wagner dos Santos Marques, Marcelo Da Silva Conterato, Tiago Ferreto, Marcelo Caggiani Luizelli, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck, Fábio D. Rossi, Jorji Nonaka |
ISCAS | 6 |
| 2019 | Multilevel resource allocation for performance-aware energy-efficient cloud data centersabstractThe massive power consumption of data centers has been a recurring concern in current research. In cloud environments, lots of methods are being adopted that aim for energy efficiency. However, although such methods enable the decrease in power consumption, they regularly affect application performance. In this paper, we present a multilevel resource allocation approach towards dynamic network bandwidth at the physical substrate, managing different power-saving states and workload allocation at the cloud infrastructure at the same time employ virtual machine allocation and selection policies at the cloud platform. In order to evaluate our approach, tests were carried out in a simulated environment using scale-out application on a dynamic cloud infrastructure. Results showed that our proposal presents a better balance regarding a more energy-efficient data center with a smaller impact on application performance when compared with other works discussed in the literature. Fábio D. Rossi, Paulo Silas Severo de Souza, Wagner dos Santos Marques, Marcelo Da Silva Conterato, Tiago Ferreto, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli |
ISCC | 7 |
| 2019 | A Black-box Method for Accelerating Measurement Algorithms with Accuracy GuaranteesabstractNetwork Function Virtualization (NFV) enables software implementations of middleboxes such as load balancing, traffic engineering and quality of service. These often rely on network measurement such as per-flow frequency estimation, bandwidth estimation, counting distinct elements and estimating the traffic entropy. Keeping up with the line speed is an active challenge for NFV measurement techniques, and library algorithms are simply too slow. Sampling is a natural technique to increase the measurement throughput, but it requires a certain amount of traffic before accuracy is guaranteed. In this work, we introduce a throughput acceleration method that preserves accuracy from the very first packet. This technique works with a variety of existing measurement algorithms (e.g., the ones mentioned above), and improves their throughput while guaranteeing their correctness throughout the entire measurement. Our work includes a rigors analysis, an extensive evaluation with real network traces, and a real DPDK enabled Open vSwitch implementation. Ran Ben-Basat, Gil Einziger, Marcelo Caggiani Luizelli, Erez Waisbard |
Networking | 3 |
| 2019 | Transparent Aging-Aware Thread ThrottlingabstractTo satisfy the rising performance demands of modern applications, the number of cores in a single chip package has been increasing. However, the power dissipated and temperature have been growing at a higher rate, accelerating the aging process of new processors. Considering that a significant number of parallel applications are unbalanced, in many cases performance stagnates after a certain number of concurrent threads starts executing. In such cases, the only outcome is a temperature rise on the processor, which drastically accelerates aging. Given that, we propose an automatic and transparent approach to reduce the processor aging by automatically tuning the number of threads for OpenMP applications at run-time. Our tool, Geras, is entirely transparent to the end-user, so even already compiled binaries can be optimized. Through the execution of twelve well-known benchmarks on two multicore platforms, we show that Geras can improve the processor lifetime by up to 83% and 89% over the standard OpenMP execution and its built-in feature that dynamically adjusts the number of threads, respectively. We also show that Geras outperforms techniques that target performance or energy, which reinforces the need for a specific tool that optimizes aging1. Thiarles S. Medeiros, Luan Pereira, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
SBAC-PAD | 4 |
| 2019 | The Impact of Turbo Frequency on the Energy, Performance, and Aging of Parallel ApplicationsabstractTechnologies that improve the performance of parallel applications by increasing the nominal operating frequency of processors respecting a given TDP (Thermal Design Power) have been widely used. However, they may impact on other non-functional requirements in different ways (e.g. increasing energy consumption or aging). Therefore, considering the huge number of configurations available, represented by the range of all possible combinations among different parallel applications, amount of threads, dynamic voltage and frequency scaling (DVFS) governors, boosting technologies and simultaneous multithreading (SMT), selecting the one that offers the best tradeoff for a non-functional requirement is extremely challenging for software designers. Given that, in this work we assess the impact of changing these configurations on the energy consumption, performance, and aging of parallel applications on a turbo-compliant processor. Results show that there is no single configuration that would provide the best solution for all nonfunctional requirements at once. For instance, we demonstrate that the configuration that offers the best performance is the same one that has the worst impact on aging, accelerating it by up to 1.75 times. With our experiments, we provide guidelines for the developer when it comes to tuning performance using turbo boosting to save as much energy as possible and increase the lifespan of the hardware components. Sandro Matheus V. N. Marques, Thiarles S. Medeiros, Fábio D. Rossi, Marcelo Caggiani Luizelli, Alessandro Girardi, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
VLSI-SoC | 4 |
| 2018 | Scalable QoE-aware Path Selection in SDN-based Mobile NetworksabstractTo deal with the massive traffic produced by video applications, mobile operators rely on offloading technologies such as Small Cells, Content Delivery Networks and, shortly, Cloud Edge and 5G Device to Device communications. Although these techniques are fundamental for improving network efficiency, they produce a multitude of paths onto which the user traffic can be forwarded. Thus, a critical problem arises about how to handle the increasing video traffic while managing the interplay between infrastructure optimization and the user's Quality of Experience (QoE). Solving this problem is remarkably difficult, and recent investigations do not consider the large-scale context of mobile operator networks. To address this issue, we present a novel QoE-aware path deployment scheme for large-scale SDN-based mobile networks. The scheme relies on both a polynomial-time algorithm for composing multiple QoS metrics and a scalable QoS to QoE translation strategy. Considering real mobile operator network and video traffic traces, we show that the proposed algorithm outperformed state-of-the-art approaches by reducing impaired videos in aggregate MOS by at least 37% and lowering accumulated video stall length four times. Roberto Irajá Tavares da Costa Filho, William Lautenschlager, Nicolas Kagami, Marcelo Caggiani Luizelli, Valter Roesler, Luciano Paschoal Gaspary |
INFOCOM | 4 |
| 2018 | Optimizing NFV Chain Deployment through Minimizing the Cost of Virtual SwitchingabstractNetwork Function Virtualization (NFV) is a novel paradigm that enables flexible and scalable implementation of network services on cloud infrastructure. A key factor in the success of NFV is the ability to dynamically allocate physical resources according to the demand. This is particularly important when dealing with the data plane since additional resources are required in order to support the virtual switching of the packets between the Virtual Network Functions (VNFs). The exact amount of these resources depends on the way service chains are deployed and the amount of network traffic being handled. Thus, orchestrating service chains that require high traffic throughput is a very complex task and most existing solutions either concentrate on handcrafted tuning of the servers to achieve the needed performance level, or present theoretical placement functions that assume that the switching cost is part of the input. In this work, we bridge this gap by presenting a deployment algorithm for service chains that optimizes performance by minimizing the actual cost of virtual switching. The results are based on extensive measurements of the actual switching cost and the performance of service chains in a realistic NFV environment. Our evaluation indicates that this new algorithm significantly reduces virtual switching resource utilization when compared to the de-facto standard placement in OpenStack/Nova - allowing a much higher acceptance ratio of network services. Marcelo Caggiani Luizelli, Danny Raz, Yaniv Sa'ar |
INFOCOM | 1 |
| 2018 | Volumetric Hierarchical Heavy HittersabstractHierarchical heavy hitters (HHH) identification is useful for various network utilities such as anomaly detection, DDoS mitigation, and traffic analysis. However, the increasing support for jumbo frames enables attackers to overload the system with fewer packets, avoiding detection by packet counting techniques. This paper suggests an efficient algorithm for detecting HHH based on their traffic volume that asymptotically improves the runtime of previous works. We implement our algorithm in Open vSwitch (OVS) and incur a 4-6% overhead compared to a 42% throughput reduction experienced by the state-of-the-art. Ran Ben-Basat, Gil Einziger, Roy Friedman 0001, Marcelo Caggiani Luizelli, Erez Waisbard |
MASCOTS | 4 |
| 2018 | Predicting the performance of virtual reality video streaming in mobile networksabstractThe demand of Virtual Reality (VR) video streaming to mobile devices is booming, as VR becomes accessible to the general public. However, the variability of conditions of mobile networks affects the perception of this type of high-bandwidth-demanding services in unexpected ways. In this situation, there is a need for novel performance assessment models fit to the new VR applications. In this paper, we present PERCEIVE, a two-stage method for predicting the perceived quality of adaptive VR videos when streamed through mobile networks. By means of machine learning techniques, our approach is able to first predict adaptive VR video playout performance, using network Quality of Service (QoS) indicators as predictors. In a second stage, it employs the predicted VR video playout performance metrics to model and estimate end-user perceived quality. The evaluation of PERCEIVE has been performed considering a real-world environment, in which VR videos are streamed while subjected to LTE/4G network condition. The accuracy of PERCEIVE has been assessed by means of the residual error between predicted and measured values. Our approach predicts the different performance metrics of the VR playout with an average prediction error lower than 3.7% and estimates the perceived quality with a prediction error lower than 4% for over 90% of all the tested cases. Moreover, it allows us to pinpoint the QoS conditions that affect adaptive VR streaming services the most. Roberto Irajá Tavares da Costa Filho, Marcelo Caggiani Luizelli, Maria Torres Vega, Jeroen van der Hooft, Stefano Petrangeli, Tim Wauters, Filip De Turck, Luciano Paschoal Gaspary |
MMSys | 2 |
| 2017 | Constant Time Weighted Frequency Estimation for Virtual Network FunctionalitiesabstractMonitoring flow volumes is a fundamental capability in network measurement. Sampling is often used to cope with the line speed and the applied methods typically rely on uniform packet sampling. However, it is inaccurate when there is a large variance in packet sizes. In this work we introduce Byte Uniform Sampling (BUS), a sampling method for estimating flow volumes. We show that BUS can be combined with existing unweighted estimation algorithms and that the result is a weighted algorithm. BUS enables an asymptotic update time improvement as existing weighted algorithms are slower. We formally analyze BUS and evaluate it on five Internet traces. Finally, we extend the DPDK version of Open vSwitch to support BUS and demonstrate similar throughput when compared to uniform packet samples. Gil Einziger, Marcelo Caggiani Luizelli, Erez Waisbard |
ICCCN | 2 |
| 2017 | The actual cost of software switching for NFV chainingabstractNetwork Function Virtualization (NFV) is a novel paradigm that enables flexible and scalable implementation of network services on cloud infrastructure. An important enabler for the NFV paradigm is software switching, which should satisfy rigid network requirements such as high throughput and low latency. Despite recent research activities in the field of NFV, not much attention was given to understand the costs of software switching in NFV deployments. Existing approaches for traffic steering and orchestration of virtual network functions either neglect the cost of software switching or assume that it can be provided as an input, and therefore real NFV deployments of network services are often suboptimal. In this work, we conduct an extensive and in-depth evaluation that examines the impact of service chaining deployments on Open vSwitch - the de facto standard software switch for cloud environments. We provide insights on network performance metrics such as throughput, CPU utilization and packet processing, while considering different placement strategies of a service chain. We then use these insights to provide an abstract generalized cost function that accurately captures the CPU switching cost of deployed service chains. This cost is an essential building block for any practical optimized placement management and orchestration strategy for NFV service chaining. Marcelo Caggiani Luizelli, Danny Raz, Yaniv Sa'ar, Jose Yallouz |
IM | 1 |
| 2017 | Constant Time Updates in Hierarchical Heavy HittersabstractMonitoring tasks, such as anomaly and DDoS detection, require identifying frequent flow aggregates based on common IP prefixes. These are known as hierarchical heavy hitters (HHH), where the hierarchy is determined based on the type of prefixes of interest in a given application. The per packet complexity of existing HHH algorithms is proportional to the size of the hierarchy, imposing significant overheads. Ran Ben-Basat, Gil Einziger, Roy Friedman 0001, Marcelo Caggiani Luizelli, Erez Waisbard |
SIGCOMM | 4 |
| 2017 | A fix-and-optimize approach for efficient and large scale virtual network function placement and chaining
Marcelo Caggiani Luizelli, Weverton Luis da Costa Cordeiro, Luciana S. Buriol, Luciano Paschoal Gaspary |
Comput. Commun. | 1 |
| 2016 | How physical network topologies affect virtual network embedding quality: A characterization study based on ISP and datacenter networks
Marcelo Caggiani Luizelli, Leonardo Richter Bays, Luciana S. Buriol, Marinho P. Barcellos, Luciano Paschoal Gaspary |
J. Netw. Comput. Appl. | 1 |
| 2015 | Piecing together the NFV provisioning puzzle: Efficient placement and chaining of virtual network functionsabstractNetwork Function Virtualization (NFV) is a promising network architecture concept, in which virtualization technologies are employed to manage networking functions via software as opposed to having to rely on hardware to handle these functions. By shifting dedicated, hardware-based network function processing to software running on commoditized hardware, NFV has the potential to make the provisioning of network functions more flexible and cost-effective, to mention just a few anticipated benefits. Despite consistent initial efforts to make NFV a reality, little has been done towards efficiently placing virtual network functions and deploying service function chains (SFC). With respect to this particular research problem, it is important to make sure resource allocation is carefully performed and orchestrated, preventing over- or under-provisioning of resources and keeping end-to-end delays comparable to those observed in traditional middlebox-based networks. In this paper, we formalize the network function placement and chaining problem and propose an Integer Linear Programming (ILP) model to solve it. Additionally, in order to cope with large infrastructures, we propose a heuristic procedure for efficiently guiding the ILP solver towards feasible, near-optimal solutions. Results show that the proposed model leads to a reduction of up to 25% in end-to-end delays (in comparison to chainings observed in traditional infrastructures) and an acceptable resource over-provisioning limited to 4%. Further, we demonstrate that our heuristic approach is able to find solutions that are very close to optimality while delivering results in a timely manner. Marcelo Caggiani Luizelli, Leonardo Richter Bays, Luciana S. Buriol, Marinho P. Barcellos, Luciano Paschoal Gaspary |
IM | 1 |
| 2014 | Survivor: An enhanced controller placement strategy for improving SDN survivabilityabstractIn SDN, forwarding devices can only operate correctly while connected to a logically centralized controller. To avoid single-point-of-failure, controller architectures are usually implemented as distributed systems. In this context, recent literature identified fundamental issues, such as device isolation and controller overload, and proposed controller placement strategies to tackle them. However, current proposals have crucial limitations: (i) device-controller connectivity is modeled using single paths, yet in practice multiple concurrent connections may occur; (ii) peaks in the arrival of new flows are only handled on-demand, assuming that the network itself can sustain high request rates; and (iii) failover mechanisms require predefined information, which, in turn, has been overlooked. This paper proposes Survivor, a controller placement strategy that addresses these challenges. The strategy explicitly considers path diversity, capacity, and failover mechanisms at network design. Comparisons to the state-of-the-art on survivable controller placement show that Survivor is superior because (a) path diversity increases the survivability significantly; and (b) capacity-awareness is essential to handle overload during both normal and failover states. Lucas F. Müller, Rodrigo Ruas Oliveira, Marcelo Caggiani Luizelli, Luciano Paschoal Gaspary, Marinho P. Barcellos |
GLOBECOM | 3 |
| 2013 | Characterizing the impact of network substrate topologies on virtual network embeddingabstractNetwork virtualization is a mechanism that allows the coexistence of multiple virtual networks on top of a single physical substrate. One of the research challenges addressed recently in the literature is the efficient mapping of virtual resources on physical infrastructures. Although this challenge has received considerable attention, state-of-the-art approaches present, in general, a high rejection rate, i.e., the ratio between the number of denied virtual network requests and the total amount of requests is considerably high. In this work, we investigate the relationship between the quality of virtual network mappings and the topological structures of the underlying substrates. Exact solutions of an online embedding model are evaluated under different classes of network topologies. The obtained results demonstrate that the employment of physical topologies that contain regions with high connectivity significantly contributes to the reduction of rejection rates and, therefore, to improved resource usage. Marcelo Caggiani Luizelli, Leonardo Richter Bays, Luciana S. Buriol, Marinho P. Barcellos, Luciano Paschoal Gaspary |
CNSM | 1 |