Gianni Antichi

dblp:00/4948 · DBLP profile ↗
← Back
84ranked-venue papers
14as first author
39since 2021 · last 2026
0000-0002-6063-4975ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 65 · 11 first-author · 29 since 2021Software engineering, systems software and programming languages · 8 · 6 since 2021Systems, architecture and hardware · 7 · 1 first-author · 6 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Enabling Fast Networking in the Public Cloud
abstract
Despite a decade of research, most high-performance userspace network stacks remain impractical for public cloud tenants developing their applications atop Virtual Machines (VMs). We identify two root causes: (1) reliance on specialized NIC features (e.g., flow steering, deep buffers) absent in commodity cloud vNICs, and (2) rigid execution models ill-suited to diverse application needs. We present Machnet, a highperformance and flexible userspace network stack designed for public cloud VMs. Machnet uses only a minimal set of vNIC features that any major cloud provider supports. It also relies on a microkernel architecture to enable flexible application execution. We evaluate Machnet across three major public clouds and on production-grade applications, including a key-value store, an HTTP server, and a statemachine replication system. We release Machnet at https: //github.com/microsoft/machnet.
Alireza Sanaee, Vahab Jabrayilov, Ilias Marinos, Farbod Shahinfar, Divyanshu Saxena, Gianni Antichi, Kostis Kaffes
ASPLOS (2)6
2026 Spatiotemporal Sketch Disaggregation: Streaming Analytics with Heterogeneous Resources
abstract
Streaming analytics are essential in a large range of applications, including databases, networking, and machine learning. To optimize performance, practitioners are increasingly offloading such analytics to network nodes such as switches. However, resources such as fast SRAM memory available at switches are limited, not uniform, and may serve other functionalities as well (e.g., firewall). Moreover, resource availability changes over time due to the dynamic demands of in-network applications. In this paper, we propose a new approach to disaggregating data structures, leveraging any residual resources available at network nodes. We focus on sketches, which are fundamental for summarizing data for streaming analytics while providing beneficial space-accuracy tradeoffs. Our idea is to break sketches into multiple 'fragments' that are placed at different network nodes. The fragments cover different time periods and vary in size, and are combined to form a network-wide view of the underlying traffic. We apply our solution to three popular sketches (namely, Count Sketch, Count-Min Sketch, and UnivMon) and demonstrate that we can achieve approximately a 75% memory size reduction for the same error for many queries, or a near order-of-magnitude error reduction if memory is kept unchanged. Further, we demonstrate real-world feasibility through a hardware pipeline for high-speed commodity switches.
Jonatan Langlet, Peiqing Chen, Michael Mitzenmacher, Zaoxing Liu, Ran Ben-Basat, Gianni Antichi
ICDE6
2026 Defending against Traffic Analysis Attacks with Flexible In-Network Obfuscation
Guorui Xie, Qing Li 0006, Zhenning Shi, Gianni Antichi, Yijia Zhu, Changxing Weng, Sebastiano Miano, Yong Jiang 0001, Mingwei Xu 0001
NSDI4
2026 OSCAR: O(1)-Step Convergence and Readily-deployable Congestion Control
Zhaochen Zhang, Feiyang Xue, Rui Ning, Keqiang He, Gianni Antichi, Zhimeng Yin 0001, Rui Li 0020, Zhengqi Cui, Zhehao Lin, Peirui Cao, Guihai Chen, Chen Tian 0001
NSDI5
2026 POSTER: Beyond Probabilistic Data Structures for AI/ML Workload Monitoring
abstract
The rapid growth of AI models is placing unprecedented pressure on switch ASIC memory allocated for telemetry and flow monitoring. In this poster, we argue that the predictability of AI/ML traffic patterns can be exploited by perfect hashing techniques for accurate flow- and packet-tracking with minimal memory overhead.
Davide Palmiotti, Michele Ferrero, Gabriele Castellano, Massimo Gallo, Gianni Antichi
SIGCOMM5
2026 Don't Stall Me Now: Hiding Memory Latency in eBPF
abstract
eBPF has emerged as a popular platform for building high-performance I/O programs. However, eBPF's programming model limits the use of commonly used performance optimizations, leaving programs vulnerable to cache-miss-induced performance overheads. Specifically, the programming model restricts the choice of data structures used by a program, thus limiting the programmers' ability to improve data locality. Furthermore, it also makes it hard for programmers to overlap computation with data movement.
Farbod Shahinfar, Marco Molè, Aurojit Panda, Gianni Antichi
SIGCOMM4
2026 STORM: Enabling Traffic Scheduling for RDMA
abstract
Remote Direct Memory Access (RDMA) is increasingly used as a shared communication substrate across datacenter workloads with very different scheduling needs, from request-response services and storage fan-out to AI training collectives. Proper request scheduling can reduce communication time, but in practice, no RDMA flow scheduling is enabled in datacenters, leaving traffic to simple fair sharing. We present STORM, a NIC-level scheduler for all types of RDMA workloads using NIC-only information: the known RDMA request size, and per-queue-pair backlog. STORM converts these signals into a small number of extra priority levels on the wire and prioritizes requests that are either near completion or blocking queued dependent work. STORM requires no application hints and works with both in-order RoCEv2 and newer RDMA stacks that tolerate reordering. We prototype STORM on an FPGA NIC with negligible overhead. Across representative cloud and LLM training workloads, STORM reduces training iteration time by up to 12% and reduces average and P99 flow completion slowdown by up to 90% compared to fair scheduling.
Jichun Wu, Ran Shu 0001, Gianni Antichi, Yongqiang Xiong, Jon Crowcroft
SIGCOMM3
2026 Honey, I Shrunk the Headers With Flow.ZIP
abstract
Packet header overhead is a persistent source of inefficiency in packet-switched networks, reducing goodput and increasing network load. Trends like tunneling further increase this overhead, significantly impacting flow completion times. While, in principle, it is possible to compress these headers, existing methods require specialized hardware on every hop to compress/decompress the packet to/from custom header formats.
Yinda Zhang 0002, Liangcheng Yu, Gianni Antichi, Ran Ben-Basat, Vincent Liu 0001
SIGCOMM3
2026 Elastic Scaling of Real-Time Communication Services
abstract
Real-time Communications (RTC) services, including multiparty conferencing, live streaming, and cloud-gaming, rely on a large-scale media plane infrastructure that provides real-time audio/video processing to clients. Unfortunately, offthe- shelf RTC services are not elastically scalable. As a result, operators must provision media servers to meet peak demand, resulting in resource under-utilization and high cost. Given that today microservice orchestrators like Kubernetes allow web-services to scale transparently and econimically, this paper looks at applying the same approach to scale large-scale RTC services. We find that this is challenging for two reasons: (a) the default network dataplane underlying Kubernetes does not meet the compelling traffic management, performance and real-time requirements of RTC; and (b) current autoscaling policies are ill-suited to RTC. We address these challenges by designing a RTC-specific service mesh that pushes media traffic processing into the OS kernel and designing new RTC-specific Kubernetes autoscaling policies. Our evaluation on a functional VoIP test-bed shows that this combination allows to deploy elatically scalable RTC services with 100× lower-jitter and 700× lower RTT than the current state-of-the art.
Máté Nagy 0002, Tamás Lévai, Felician Németh, Aurojit Panda, Gianni Antichi, Gábor Rétvári
IEEE Trans. Netw. Serv. Manag.5
2025 Aqua: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
abstract
Inference on large-language models (LLMs) is constrained by GPU memory capacity. A sudden increase in the number of inference requests to a cloud-hosted LLM can deplete GPU memory, leading to contention between multiple prompts for limited resources. Modern LLM serving engines deal with the challenge of limited GPU memory using admission control, which causes them to be unresponsive during request bursts. We propose that preemptive scheduling of prompts in time slices is essential for ensuring responsive LLM inference, especially under conditions of high load and limited GPU memory. However, preempting prompt inference incurs a high paging overhead, which reduces inference throughput. We present Aqua, a GPU memory management framework that significantly reduces the overhead of paging inference state; achieving both responsive and high throughput inference even under bursty request patterns. We evaluate Aqua by hosting several state-of-the-art large generative ML models of different modalities on servers with 8 Nvidia H100 80G GPUs. Aqua improves the responsiveness of LLM inference by 20X compared to the state-of-the-art. It improves LLM inference throughput over a single long prompt by 4X.
Abhishek Vijaya Kumar, Gianni Antichi, Rachee Singh
ASPLOS (2)2
2025 Gigaflow: Pipeline-Aware Sub-Traversal Caching for Modern SmartNICs
abstract
The success of modern public/edge clouds hinges heavily on the performance of their end-host network stacks if they are to support the emerging and diverse tenants' workloads (e.g., distributed training in the cloud to fast inference at the edge). Virtual Switches (vSwitches) are vital components of this stack, providing a unified interface to enforce high-level policies on incoming packets and route them to physical interfaces, containers, or virtual machines. As performance demands escalate, there has been a shift toward offloading vSwitch processing to SmartNICs to alleviate CPU load and improve efficiency. However, existing solutions struggle to handle the growing flow rule space within the NIC, leading to high miss rates and poor scalability.
Annus Zulfiqar, Ali Imran 0005, Venkat Kunaparaju, Ben Pfaff, Gianni Antichi, Muhammad Shahbaz 0001
ASPLOS (2)5
2025 Enabling Virtual Priority in Data Center Congestion Control
abstract
In data center networks, various types of traffic with strict performance requirements operate simultaneously, necessitating effective isolation and scheduling through priority queues. However, most switches support only around ten priority queues. Virtual priority can address this limitation by emulating multi-priority queues on a single physical queue, but existing solutions often require complex switch-level scheduling and hardware changes. Our key insight is that virtual priority can be achieved by carefully managing bandwidth contention in a physical queue, which is traditionally handled by congestion control (CC) algorithms. Hence, the virtual priority mechanism needs to be tightly coupled with CC. In this paper, we propose PrioPlus, a CC enhancement algorithm that can be integrated with existing congestion control schemes to enable virtual priority transmission. PrioPlus assigns specific delay ranges to different priority levels, ensuring that flows transmit only when the delay is within the assigned range, effectively meeting virtual priority requirements. Compared to Swift CC with physical priority queues, PrioPlus provides strict priority for high-priority flows without impacting performance sensibly. Meanwhile, it benefits low-priority flows from 25% to 41% as its priority-aware design enhances CC's ability to fully utilize available bandwidth once higher-priority traffic completes. As a result, in coflow and model training scenarios, PrioPlus improves job completion times by 21% and 33%, respectively, compared to Swift with physical priority queues.
Zhaochen Zhang, Feiyang Xue, Keqiang He, Zhimeng Yin 0001, Gianni Antichi, Yizhi Wang 0004, Rui Ning, Haixin Nan, Xu Zhang 0006, Peirui Cao, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
EuroSys5
2025 Enabling Silent Telemetry Data Transmission with InvisiFlow
Yinda Zhang 0002, Liangcheng Yu, Gianni Antichi, Ran Ben-Basat, Vincent Liu 0001
NSDI3
2025 State-Compute Replication: Parallelizing High-Speed Stateful Packet Processing
Qiongwen Xu, Sebastiano Miano, Tao Wang 0088, Adithya Murugadass, Songyuan Zhang, Anirudh Sivaraman, Gianni Antichi, Srinivas Narayana
NSDI8
2025 Astral: A Datacenter Infrastructure for Large Language Model Training at Scale
abstract
The flourishing of Large Language Models (LLMs) calls for increasingly ultra-scale training. In this paper, we share our experience in designing, deploying, and operating our novel Astral datacenter infrastructure, along with operational lessons and evolutionary insights gained from its production use. Astral has three important innovations: (i) a same-rail interconnection network architecture on tier-2, which enables the scaling of LLM training. To physically deploy this high-density infrastructure, we introduce a distributed high-voltage direct current power system and a new air-liquid integrated cooling system. (ii) a full-stack monitoring system featuring cross-host and hierarchical logging correlation, which diagnoses failures at scale and precisely localizes root causes. (iii) an operator-granular forecasting component Seer that efficiently generates operator execution timelines with acceptable accuracy, aiding in fault diagnosis, model tuning, and network architecture upgrading. Astral infrastructure has been gradually deployed over 18 months, supporting LLM training and inference for multiple customers.
Qingkai Meng 0001, Zhenhui Zhang, ChonLam Lao, Chengyuan Huang, Baojia Li 0002, Weizhen Dang, Zitong Lin, Yuanyuan Gong, Chunzhi He, Xiaoyuan Hu, Yinben Xia, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Kun Yang 0001, Gianni Antichi, Guihai Chen, Chen Tian 0001
SIGCOMM22
2025 Switch bypass: End-host cloud networking revisited
Antonio Le Caldare, Luigi Leonardi, Sebastiano Miano, Gregorio Procissi, Gianni Antichi, Giuseppe Lettieri
Comput. Networks5
2025 Troubleshooting Programmable Data Planes via Real-Time Table Information Recording
abstract
While the flexibility of programmable switches brings opportunities, it also introduces security risks. Hence, it is vital to conduct effective troubleshooting in the programmable switch to mitigate frequent network failures. However, troubleshooting programmable switch failures is challenging due to their enhanced flexibility and functionality compared to regular switches, posing increased difficulty in debugging, particularly with limited debugging tools and information. To address this problem, we propose an efficient troubleshooting method that records real-time information about packets in the data plane, including the tables involved in packet processing. Unfortunately, due to hardware limitations, it is infeasible to record all tables’ information in the data plane. Thus, the key is to find the table set reflecting the execution path a packet goes through while minimizing the resource overhead. We first represent P4 programs as a probabilistic transition directed acyclic graph (DAG) and employ information entropy to quantify the information within a set of tracked tables. Then, we adopt a two-step approach and design algorithms to find both optimal and approximately optimal table record plans. The evaluation results show the efficacy of the proposed method, including achieving the same path recovery rate as the related works with less than one-third of the resource consumption.
Chengyuan Huang, Yibo Xiao, Tianfan Zhang, Bingheng Yan, Ahmed M. Abdelmoniem, Gianni Antichi, Xiaoliang Wang 0001, Fu Xiao 0001, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
IEEE Trans. Netw.8
2024 A Smart Cache for a SmartNIC! Scaling End-Host Networking to 400Gbps and Beyond
abstract
•Virtual switches optimize performance by caching multi-table lookup traversals to single-table Megaflow cache, which SmartNICs offload directly to hardware •We present Gigaflow: a multi-table sub-traversal cache for SmartNICs, designed to capture a much larger rule space using the same cache size •Open vSwitch caches traversals into Megaflow and can't share sub-traversals among traffic, making the captured rule space proportional to cache size •By caching sub-traversals into a multi-table cache, we can capture 3 orders of magnitude more rule space, attain 51% higher cache hit rate, and 31% lower end-to-end packet latency, with manageable processing overhead
Annus Zulfiqar, Ali Imran 0005, Venkat Kunaparaju, Ben Pfaff, Gianni Antichi, Muhammad Shahbaz 0001
HCS5
2024 Rethinking the Switch Architecture for Stateful In-network Computing
abstract
Programmable switches are a disruptive technology that has seen increasing adoption in the past decade. Since their inception, however, there has been tension regarding how to design these switches. Classic programmable switches operate at line rate but impose significant limitations on the expressiveness of their programming models. In contrast, alternative designs relax the strict line rate requirement but are more easily programmable. The common belief is that a switch's performance and its programmability are at odds.
Alberto Lerner, Davide Zoni, Paolo Costa, Gianni Antichi
HotNets4
2024 Incremental Specialization of Network Programs
abstract
Programmable network devices process packets using limited time and space. Consequently, much effort has been spent making network programs run as efficiently as possible. One promising line of work focuses on specializing the implementation of a network program to a particular---presumed constant---control-plane configuration. However, while some parts of the control plane configurations are constant for long periods of time, others change frequently, and in bursts (e.g., due to routing table updates).
Fabian Ruffy, Zhanghan Wang, Gianni Antichi, Aurojit Panda, Anirudh Sivaraman
HotNets3
2024 Rethinking Cloud Network Stacks with Switch Bypass
abstract
Virtual switches are one of the most important building blocks in public cloud network stacks as they apply high-level policies to traffic enabling communication between virtual machines (VMs) and the rest of the world. The problem is that virtual switches need CPU cores to process packets and the more cores assigned to them, the less are available to VMs that are rented to customers and hence generate revenue. With this paper, we show that it is potentially possible to find a sweet-spot between performance and costs. The insight is that applications running on VMs are not always using 100% of their CPU processing power: we use this to design switch bypass, a new technique that allow virtual switches to opportunistically offload part of their processing to the virtual NIC drivers associated with guest VMs. Using packet classification as use-case, we show that with switch bypass we obtain a performance boost up to 40% without the need of additional core processing power.
Antonio Le Caldare, Luigi Leonardi, Sebastiano Miano, Gregorio Procissi, Gianni Antichi, Giuseppe Lettieri
HPSR5
2024 Accelerating network analytics with an on-NIC streaming engine
Sebastiano Miano, Giuseppe Lettieri, Gianni Antichi, Gregorio Procissi
Comput. Networks3
2024 Morpheus: A Run Time Compiler and Optimizer for Software Data Planes
abstract
State-of-the-art approaches to design, develop and optimize software packet-processing programs are based on static compilation: the compiler’s input is a description of the forwarding plane semantics and the output is a binary that can accommodate any control plane configuration or input traffic. In this paper, we demonstrate that tracking control plane actions and packet-level traffic dynamics at run time opens up new opportunities for code specialization. We present Morpheus, a system working alongside static compilers that continuously optimizes the targeted networking code. We introduce a number of new techniques, from static code analysis to adaptive code instrumentation, and we implement a toolbox of domain specific optimizations that are not restricted to a specific data plane framework or programming language. We apply Morpheus to several systems, from eBPF and DPDK programs including Katran, Meta’s production-grade load balancer to container orchestration solutions such a Kubernets. We compare Morpheus to state-of-the-art optimization frameworks and show that it can bring up to 2x throughput improvement, while halving the 99th percentile latency.
Sebastiano Miano, Alireza Sanaee, Fulvio Risso, Gábor Rétvári, Gianni Antichi
IEEE/ACM Trans. Netw.5
2023 Automatic Kernel Offload Using BPF
abstract
BPF support in Linux has made kernel extensions easier. Recent efforts have shown that using BPF to offload portions of server applications, e.g., memcached and service proxies, can improve application performance and efficiency. However, thus far, the community has not looked at the question of what parts of an application should be offloaded? This paper first shows that blindly offloading application functionality to the kernel is neither beneficial nor desirable, and care must be taken when deciding what to offload. Furthermore, when deciding what to offload, developers must consider not just the application, but also the workload being handled, and the kernel being targetted, Therefore, we advocate automating this decision process in a compiler, that can analyze application code, and produce two executables, a kernel offload and a userspace program, that jointly implement the application's functionality. This paper discusses the challenges that must be addressed to build such a compiler, and why they can be feasibly addressed.
Farbod Shahinfar, Sebastiano Miano, Giuseppe Siracusano, Roberto Bifulco, Aurojit Panda, Gianni Antichi
HotOS6
2023 Efficient Attack Detection with Multi-Latency Neural Models on Heterogeneous Network Devices
abstract
To achieve fast and accurate attack detection, some works manually tailor neural networks (NNs) for deployment on CPUs of gateways, routers, or even programmable switches. However, with such solutions, NNs must be custom-tailored across different devices to meet the heterogeneous settings (e.g., OS and CPU types). Even worse, a model may require frequent adjustments to adapt to the same device's varying traffic rates. In this paper, we present Soteria, an automated multi-latency NN generation and scheduling system for fast and accurate detection against fluctuating traffic rates across heterogeneous hardware. Soteria first uses an evolutionary training algorithm to evolve the Pareto front, i.e., the set of NNs with a good spread on accuracy and model size. Then, for each device, Soteria filters the optimal multi-latency NNs by non-dominating sorting on the NNs' test latency on the device. Finally, to cope with the dynamic traffic rate, we design a heuristic scheduling scheme that adaptively selects NN s to maintain a balance between the detection accuracy and latency.
Guorui Xie, Qing Li 0006, Haolin Yan, Dan Zhao 0003, Gianni Antichi, Yong Jiang 0001
ICNP5
2023 Dryad: Deploying Adaptive Trees on Programmable Switches for Networking Classification
abstract
Decision trees (DT) have been used for high-speed networking classification on programmable switches. Most DT solutions, however, are static and cannot be deployed once the switch resource changes. In this paper, we propose Dryad to fast reprogram tree models when resource budgets change. In Dryad, we first develop a large and accurate “one-training-for-all“ DT (ODT) that can be quickly resized without computational retraining. ODTs are deployed in switches using a progressive search algorithm that searches the adaptations according to their resources. To achieve high accuracy and low packet latency, the adaptation leverages 1) innovative hard and soft pruning methods to compress the ODT rapidly with minimal performance loss; and 2) P4 scaling operations of match-action table arrangement and joint range-ternary match, which allow the switch to accommodate a larger (i.e., more accurate) ODT. Finally, an ODTCompiler is proposed to automatically convert the adapted ODT into a P4 program and then install it. Experimental results on three commodity switches under different resource scenarios show that Dryad achieves a higher classification F1-score (3.78 % higher), and completes the adaptation 161 × faster than other solutions.
Guorui Xie, Qing Li 0006, Jiaye Lin, Gianni Antichi, Dan Zhao 0003, Zhenhui Yuan, Ruoyu Li 0003, Yong Jiang 0001
ICNP4
2023 Poster: Continual Network Learning
abstract
We make a case for in-network Continual Learning as a solution for seamless adaptation to evolving network conditions without forgetting past experiences. We propose implementing Active Learning-based selective data filtering in the data plane, allowing for data-efficient continual updates. We explore relevant challenges and propose future research directions.
Nicola Di Cicco, Amir Al Sadi, Chiara Grasselli, Andrea Melis 0001, Gianni Antichi, Massimo Tornatore
SIGCOMM5
2023 Direct Telemetry Access
abstract
Fine-grained network telemetry is becoming a modern datacenter standard and is the basis of essential applications such as congestion control, load balancing, and advanced troubleshooting. As network size increases and telemetry gets more fine-grained, there is a tremendous growth in the amount of data needed to be reported from switches to collectors to enable network-wide view. As a consequence, it is progressively hard to scale data collection systems.
Jonatan Langlet, Ran Ben-Basat, Gabriele Oliaro, Michael Mitzenmacher, Minlan Yu, Gianni Antichi
SIGCOMM6
2023 HH-IPG: Leveraging Inter-Packet Gap Metrics in P4 Hardware for Heavy Hitter Detection
abstract
The research community has recently proposed several solutions based on modern programmable switches to detect entirely in the data plane the flows exceeding pre-determined threshold in a time window, i.e., Heavy Hitters (HH). This is commonly achieved by dividing the network stream into fixed time slots and identifying each separately without considering the traffic trends from previous intervals. In this work, we show that using specified time windows can lead to high inaccuracies. We make a case for rethinking how switches analyze the incoming packets and propose to leverage per-flow Inter Packet Gap (IPG) analytics instead of using flow counters for HH detection. We propose an algorithm and present a P4 pipeline design using this new metric in mind. We implement our solution on P4 hardware and experimentally evaluate it against real traffic traces. We show that our results are more accurate than related work by up to 20% while reducing the control channel overhead by up to two orders of magnitude. Finally, we showcase a QoS-oriented application of the proposed dataplane-only IPG-based HH detection in a mobile network scenario.
Suneet Kumar Singh, Christian Esteve Rothenberg, Marcelo Caggiani Luizelli, Gianni Antichi, Pedro Henrique Gomes, Gergely Pongrácz
IEEE Trans. Netw. Serv. Manag.4
2022 Domain specific run time optimization for software data planes
abstract
State-of-the-art approaches to design, develop and optimize software packet-processing programs are based on static compilation: the compiler's input is a description of the forwarding plane semantics and the output is a binary that can accommodate any control plane configuration or input traffic.
Sebastiano Miano, Alireza Sanaee, Fulvio Risso, Gábor Rétvári, Gianni Antichi
ASPLOS5
2022 Backdraft: a Lossless Virtual Switch that Prevents the Slow Receiver Problem
Alireza Sanaee, Farbod Shahinfar, Gianni Antichi, Brent E. Stephens
NSDI3
2022 Re-architecting Traffic Analysis with Neural Network Interface Cards
Giuseppe Siracusano, Salvator Galea, Davide Sanvito, Mohammad Malekzadeh, Gianni Antichi, Paolo Costa, Hamed Haddadi 0001, Roberto Bifulco
NSDI5
2022 Isolation Mechanisms for High-Speed Packet-Processing Pipelines
Tao Wang 0088, Xiangrui Yang 0002, Gianni Antichi, Anirudh Sivaraman, Aurojit Panda
NSDI3
2022 ClassBench-ng: Benchmarking Packet Classification Algorithms in the OpenFlow Era
abstract
Packet classification, i.e., the process of categorizing packets into flows, is a first-class citizen in any networking device. Every time a new packet has to be processed, one or more header fields need to be compared against a set of pre-installed rules. This is done for basic forwarding operations, to apply security policies, application-specific processing, or quality-of-service guarantees. A lot of research efforts have identified better lookup techniques, i.e., finding the best match between packet headers and rules, by capitalizing on the rule sets characteristics. Here, ClassBench has greatly served the community by enabling the generation of IPv4 rule sets. In this paper, we present a new tool, ClassBench-ng, that creates synthetic IPv4, IPv6, and OpenFlow rules. We start from an analysis of classification rules deployed in-the-wild and we use the findings to craft our solution. ClassBench-ng can generate a user-defined number of rules as well as an associated header trace matching them. Compared to state-of-the-art solutions, the rule set generation process is usually more accurate and it is able to produce rules matching a number of different use cases, i.e., from an IPv4 router to an OpenFlow switch, which is unique among current rule set generation tools.
Jirí Matousek 0002, Adam Lucanský, David Janecek, Jozef Sabo, Jan Korenek, Gianni Antichi
IEEE/ACM Trans. Netw.6
2021 The case for network functions decomposition
abstract
This paper makes a case for writing unrestricted eBPF network functions which then get automatically decomposed between kernel and user-space.
Farbod Shahinfar, Sebastiano Miano, Alireza Sanaee, Giuseppe Siracusano, Roberto Bifulco, Gianni Antichi
CoNEXT6
2021 Zero-CPU Collection with Direct Telemetry Access
abstract
Programmable switches are driving a massive increase in fine-grained measurements. This puts significant pressure on telemetry collectors that have to process reports from many switches. Past research acknowledged this problem by either improving collectors' stack performance or by limiting the amount of data sent from switches. In this paper, we take a different and radical approach: switches are responsible for directly inserting queryable telemetry data into the collectors' memory, bypassing their CPU, and thereby improving their collection scalability. We propose to use a method we call direct telemetry access, where switches jointly write telemetry reports directly into the same collector's memory region, without coordination. Our solution, DART, is probabilistic, trading memory redundancy and query success probability for CPU resources at collectors. We prototype DART using commodity hardware such as P4 switches and RDMA NICs and show that we get high query success rates with a reasonable memory overhead. For example, we can collect INT path tracing information on a fat tree topology without a collector's CPU involvement while achieving 99.9% query success probability and using just 300 bytes per flow.
Jonatan Langlet, Ran Ben-Basat, Sivaramakrishnan Ramanathan, Gabriele Oliaro, Michael Mitzenmacher, Minlan Yu, Gianni Antichi
HotNets7
2021 Providing In-network Support to Coflow Scheduling
abstract
Emerging distributed applications, such as big data analytics, generate a large number of flows that concurrently transport data across data center networks. To improve their performance, it is required to account for the behavior of such a collection of flows, i.e., coflows, rather than individual ones. State-of-the-art solutions achieve near-optimal completion time by continuously reordering unfinished coflows at the end-host and using network priorities.This paper shows that dynamically changing flow priorities at the end-host, without considering in-flight packets, can cause high degrees of packet reordering, thus imposing pressure on the congestion control and potentially harming network performance in the presence of switches with shallow buffers. We present pCoflow, a new solution that integrates end-host based coflow ordering with in-network scheduling based on packet history. Our evaluation shows that pCoflow improves in coflow completion time upon state-of-the-art solutions by up to 34% for varying loads.
Cristian Hernandez Benet, Andreas Kassler, Gianni Antichi, Theophilus Benson, Gergely Pongrácz
NetSoft3
2021 revisiting the open vSwitch dataplane ten years later
abstract
This paper shares our experience in supporting and running the Open vSwitch (OVS) software switch, as part of the NSX product for enterprise data center virtualization used by thousands of VMware customers. Starting in 2009, the OVS design split its code between tightly coupled kernel and userspace components. This split was necessary at the time for performance, but it caused maintainability problems that persist today. In addition, in-kernel packet processing is now much slower than newer options.
William Tu, Yi-Hung Wei, Gianni Antichi, Ben Pfaff
SIGCOMM3
2021 Fast ReRoute on Programmable Switches
abstract
Highly dependable communication networks usually rely on some kind of Fast Re-Route (FRR) mechanism which allows to quickly re-route traffic upon failures, entirely in the data plane. This paper studies the design of FRR mechanisms for emerging reconfigurable switches. Our main contribution is an FRR primitive for programmable data planes, PURR, which provides low failover latency and high switch throughput, by avoiding packet recirculation. PURR tolerates multiple concurrent failures and comes with minimal memory requirements, ensuring compact forwarding tables, by unveiling an intriguing connection to classic “string theory” (i.e., stringology), and in particular, the shortest common supersequence problem. PURR is well-suited for high-speed match-action forwarding architectures (e.g., PISA) and supports the implementation of a broad variety of FRR mechanisms. Our simulations and prototype implementation (on an FPGA and a Tofino switch) show that PURR improves TCAM memory occupancy by a factor of 1.5 ×- 10.8 × compared to a naïve encoding when implementing state-of-the-art FRR mechanisms. PURR also improves the latency and throughput of datacenter traffic up to a factor of 2.8 ×- 5.5 × and 1.2 ×- 2 ×, respectively, compared to approaches based on recirculating packets.
Marco Chiesa, Roshan Sedar, Gianni Antichi, Michael Borokhovich, Andrzej Kamisinski, Georgios Nikolaidis, Stefan Schmid 0001
IEEE/ACM Trans. Netw.3
2020 Detecting routing loops in the data plane
abstract
Routing loops can harm network operation. Existing loop detection mechanisms, including mirroring packets, storing state on switches, or encoding the path onto packets, impose significant overheads on either the switches or the network.
Jan Kucera 0004, Ran Ben-Basat, Mário Kuka, Gianni Antichi, Minlan Yu, Michael Mitzenmacher
CoNEXT4
2020 DISCOvering the heavy hitters with disaggregated sketches
abstract
We propose DISCO - a lightweight approach to flow monitoring in the data plane. The idea is to disaggregate the computation of a single (logically) centralized sketch into multiple small "sketch fragments" that are distributed across the flows' paths. This allows use less resources at switches without trading on telemetry capabilities.
Valerio Bruschi, Ran Ben-Basat, Zaoxing Liu, Gianni Antichi, Giuseppe Bianchi 0001, Michael Mitzenmacher
CoNEXT4
2020 PINT: Probabilistic In-band Network Telemetry
abstract
Commodity network devices support adding in-band telemetry measurements into data packets, enabling a wide range of applications, including network troubleshooting, congestion control, and path tracing. However, including such information on packets adds significant overhead that impacts both flow completion times and application-level performance.
Ran Ben-Basat, Sivaramakrishnan Ramanathan, Gianni Antichi, Minlan Yu, Michael Mitzenmacher
SIGCOMM4
2020 An Incrementally-Deployable P4-Enabled Architecture for Network-Wide Heavy-Hitter Detection
abstract
The advent of Software-Defined Networking with OpenFlow first, and subsequently the emergence of programmable data planes, has boosted lots of research around many networking aspects: monitoring, security, traffic engineering. In the context of monitoring, most of the proposed solutions show the benefits of data plane programmability by simplifying the network complexity with a one big-switch abstraction. Only few papers look at network-wide solutions, but consider the network only composed by programmable devices. In this paper, we argue that the primary challenge for a successful adoption of those solutions is the deployment problem: how to compose and monitor a network consisting of both legacy and programmable switches? We propose an approach for incrementally deploy programmable devices in an ISP network with the goal of monitoring as many distinct network flows as possible. While assessing the benefits of our solution, we realized that proposed network-wide monitoring algorithms might not be optimized for a partial deployment scenario. We then also developed and implemented in P4 a novel strategy capable of detecting network-wide heavy flows: results show that it can achieve better accuracy than state-of-the-art solutions while relying on less information from the data plane and leading to only marginal additional packet processing time.
Damu Ding, Marco Savi, Gianni Antichi, Domenico Siracusa
IEEE Trans. Netw. Serv. Manag.3
2019 Towards Cheap Scalable Browser Multiplayer
abstract
The barrier to entry for the development of independent, browser-based multiplayer games is high for two reasons: complexity and cost. In this work, we introduce and evaluate a method and library that aims to make this barrier as small as possible, by utilising appropriate development abstractions and peer-to-peer communication between players. Our preliminary evaluation shows that we can lower both the technical development overhead, as well as minimise server costs, at no loss to performance.
Yousef Amar, Gareth Tyson, Gianni Antichi, Lucio Marcenaro
CoG3
2019 PURR: a primitive for reconfigurable fast reroute: hope for the best and program for the worst
abstract
Highly dependable communication networks usually rely on some kind of Fast Re-Route (FRR) mechanism which allows to quickly re-route traffic upon failures, entirely in the data plane. This paper studies the design of FRR mechanisms for emerging reconfigurable switches.
Marco Chiesa, Roshan Sedar, Gianni Antichi, Michael Borokhovich, Andrzej Kamisinski, Georgios Nikolaidis, Stefan Schmid 0001
CoNEXT3
2019 Event-Driven Packet Processing
abstract
The rise of programmable network devices and the P4 programming language has sparked an interest in developing new applications for packet processing data planes. Current data-plane programming models allow developers to express packet processing on a synchronous packet-by-packet basis, motivated by the goal of line rate processing in feed-forward pipelines. But some important data-plane operations do not naturally fit into this programming model. Sometimes we want to perform periodic tasks, or update the same state variables multiple times, or base a decision on state sitting at a different pipeline stage. While a P4-programmable device might contain special features to handle these tasks, such as packet generators and recirculation paths, there is currently no clean and consistent way to expose them to P4 programmers. We therefore propose a common, general way to express event processing using the P4 language, beyond just processing packet arrival and departure events. We believe that this more general notion of event processing can be supported without sacrificing line rate packet processing and we have developed a prototype event-driven architecture on the NetFPGA SUME platform to serve as an initial proof of concept.
Stephen Ibanez, Gianni Antichi, Gordon J. Brebner, Nick McKeown
HotNets2
2019 An Empirical Study of the Cost of DNS-over-HTTPS
abstract
DNS is a vital component for almost every networked application. Originally it was designed as an unencrypted protocol, making user security a concern. DNS-over-HTTPS (DoH) is the latest proposal to make name resolution more secure.
Timm Böttger, Félix Cuadrado, Gianni Antichi, Eder Leão Fernandes, Gareth Tyson, Ignacio Castro, Steve Uhlig
Internet Measurement Conference3
2019 FAST: enabling fast software/hardware prototype for network experimentation
abstract
The evolution of new technologies in network community is getting ever faster. Yet it remains the case that prototyping those novel mechanisms on a real-world system (i.e. CPU-FPGA platforms) is both time and labor consuming, which has a serious impact on the research timeliness. In order to bring researchers out of trivial process in prototype development, this paper proposed FAST, a software hardware co-design framework for fast network prototyping. With the programming abstraction of FAST, researchers are able to prototype (using C, verilog or both) a wide spectrum of network boxes rapidly based on all kinds of CPU-FPGA platforms. FAST framework takes care of managing DMA, PCIe and Linux Kernel while providing a unified API for researchers so they can focus only on the packet processing functions. We demonstrate FAST framework's easy to use features with a number of prototypes and show we can get over 10x gains in performance or 1000x better accuracy in clock synchronization compared with their software versions.
Xiangrui Yang 0002, Zhigang Sun 0002, Junnan Li 0002, Jinli Yan, Tao Li 0008, Wei Quan 0004, Donglai Xu, Gianni Antichi
IWQoS8
2019 Incremental Deployment of Programmable Switches for Network-wide Heavy-hitter Detection
abstract
The advent of Software-Defined Networking with OpenFlow first, and subsequently the emergence of programmable data planes, has boosted lot of research around many networking aspects: monitoring, security, traffic engineering. In the context of network monitoring, most of the proposed solutions show the benefits of data plane programmability by simplifying the complexity of the network with a one big-switch abstraction. Only few papers look at network-wide solutions, but consider the network as non heterogeneous: only composed by programmable devices. In this paper, we argue that the primary challenge for a successful adoption of those solutions is the deployment problem: how to compose and monitor a network consisting of both legacy and programmable switches? We propose an approach for incrementally deploy programmable devices in an ISP network with the goal of monitoring as many distinct network flows as possible. While assessing the benefits of our solution, we realized that proposed network-wide monitoring algorithms might not be optimized for a partial deployment scenario. We then also developed a novel strategy capable of detecting network-wide heavy flows with the same accuracy of state-of-the-art solutions but by relying on less information from the data plane.
Damu Ding, Marco Savi, Gianni Antichi, Domenico Siracusa
NetSoft3
2018 An SDN-inspired Model for Faster Network Experimentation
abstract
Assessing the impact of changes in a production network (e.g., new routing protocols or topologies) requires simulation or emulation tools capable of providing results as close as possible to those from a real-world experiment. Large traffic loads and complex control-data plane interactions constitute significant challenges to these tools. To meet these challengeswe propose a model for the fast and convenient evaluation of SDN as well as legacy networks. Our approach emulates the network's control plane and simulates the data plane, to achieve high fidelity necessary for control plane behavior, while being capable of handling large traffic loads. We design and implement a proof of concept from the proposed model. The initial results of the prototype, compared to a state-of-the-art solution, shows it can increase the speed of network experiments by nearly 95% in the largest tested network scenario.
Eder Leão Fernandes, Gianni Antichi, Ignacio Castro, Steve Uhlig
SIGSIM-PADS2
2018 Understanding PCIe performance for end host networking
abstract
In recent years, spurred on by the development and availability of programmable NICs, end hosts have increasingly become the enforcement point for core network functions such as load balancing, congestion control, and application specific network offloads. However, implementing custom designs on programmable NICs is not easy: many potential bottlenecks can impact performance.
Rolf Neugebauer, Gianni Antichi, Jose Fernando Zazo, Yury Audzevich, Sergio López-Buedo, Andrew W. Moore 0002
SIGCOMM2
2018 Rethinking IXPs' Architecture in the Age of SDN
abstract
Software-defined Internet eXchange points (SDXs) are a promising solution to the long-standing limitations and problems of interdomain routing. While the proposed SDX architectures have improved the scalability of the control plane, these solutions have ignored the underlying fabric upon which they should be deployed. In this paper, we present Umbrella, a software-defined interconnection fabric that complements and enhances those architectures. Umbrella is a switching fabric architecture and management approach that improves the overall robustness, limiting control plane dependence, and suitable for the topology of any existing Internet eXchange Point (IXP). We validate Umbrella through a real-world deployment on two production IXPs, TouSIX and NSPIXP-3, and demonstrate its use in practice, sharing our experience of the challenges faced.
Marc Bruyere, Gianni Antichi, Eder Leão Fernandes, Remy Lapeyrade, Steve Uhlig, Philippe Owezarski, Andrew W. Moore 0002, Ignacio Castro
IEEE J. Sel. Areas Commun.2
2018 OFLOPS-SUME and the Art of Switch Characterization
abstract
The philosophy of software-defined networking (SDN) has introduced new challenges in network system management. In contrast to traditional network devices that contained both the control and the data plane functionality in a tightly coupled manner, SDN technologies separate the two network planes and define a remote API for low-level device configuration. Nonetheless, the enhanced flexibility of the SDN paradigm is prone to create novel performance and scalability bottlenecks in the network. To help network managers and application developers better understand the actual behavior of SDN implementations, we present a hardware/software co-design that enables switch characterization at 40 Gbps and beyond. We conduct an evaluation of both software and hardware switches. We expose the unwanted effects of the OpenFlow barrier primitive, potential misbehaviors when adding or modifying a batch of rules, and how simple operations, such as packet modification, can impact the switch forwarding performance. We release the code publicly as open source to promote experiments reproducibility as well as encourage the network community to evolve our solution.
Rémi Oudin, Gianni Antichi, Charalampos Rotsos, Andrew W. Moore 0002, Steve Uhlig
IEEE J. Sel. Areas Commun.2
2017 Mind the Gap - A Comparison of Software Packet Generators
abstract
Network research relies on packet generators to assess performance and correctness of new ideas. Software-based generators in particular are widely used by academic researchers because of their flexibility, affordability, and open-source nature. The rise of new frameworks for fast IO on commodity hardware is making them even more attractive. Longstanding performance differences of software generation versus hardware in terms of throughput are no longer as big of a concern as they used to be few years ago. This paper investigates the properties of several high-per-formance software packet generators and the implications on their precision when a given traffic pattern needs to be generated. We believe that the evaluation strategy presented in this paper helps understanding the actual limitations in high-performance software packet generation, thus helping the research community to build better tools.
Paul Emmerich, Sebastian Gallenmüller, Gianni Antichi, Andrew W. Moore 0002, Georg Carle
ANCS3
2017 ClassBench-ng: Recasting ClassBench after a Decade of Network Evolution
abstract
Internet evolution is driven by a continuous stream of new applications and users driving the demand for services. To keep up with this, a never-stopping research has been transforming the Internet ecosystem over the time. Technological changes on both protocols (the uptake of IPv6) and network architectures (the adoption of Software Defined Networking) introduced new challenges for ASIC designers. In particular, IPv6 and OpenFlow increased the complexity of the rule matching problem, pushing researchers to build new packet classiffication algorithms capable to keep pace with a steady growth of link speed. A lot of research effort identifies better lookup techniques capitalizing on the characteristics of rule sets. So far, the availability of small numbers of real rule sets and synthetic ones, generated with tools such as ClassBench, has boosted research in the IPv4 world. Starting from an analysis of rule sets taken from operational environments, we present ClassBench-ng, a new open source tool for the generation of synthetic IPv4, IPv6, and OpenFlow 1.0 rule sets exposing the same properties of real ones. We feel this tool can meet the requirements of nowadays researchers, boosting the rule matching research as ClassBench has done since ten years ago.
Jirí Matousek 0002, Gianni Antichi, Adam Lucanský, Andrew W. Moore 0002, Jan Korenek
ANCS2
2017 Where Has My Time Gone?
Noa Zilberman, Matthew P. Grosvenor, Diana Andreea Popescu, Neelakandan Manihatty Bojan, Gianni Antichi, Marcin Wójcik, Andrew W. Moore 0002
PAM5
2017 Re-architecting datacenter networks and stacks for low latency and high performance
abstract
Modern datacenter networks provide very high capacity via redundant Clos topologies and low switch latency, but transport protocols rarely deliver matching performance. We present NDP, a novel data-center transport architecture that achieves near-optimal completion times for short transfers and high flow throughput in a wide range of scenarios, including incast. NDP switch buffers are very shallow and when they fill the switches trim packets to headers and priority forward the headers. This gives receivers a full view of instantaneous demand from all senders, and is the basis for our novel, high-performance, multipath-aware transport protocol that can deal gracefully with massive incast events and prioritize traffic from different senders on RTT timescales. We implemented NDP in Linux hosts with DPDK, in a software switch, in a NetFPGA-based hardware switch, and in P4. We evaluate NDP's performance in our implementations and in large-scale simulations, simultaneously demonstrating support for very low-latency and high throughput.
Mark Handley, Costin Raiciu, Alexandru Agache, Andrei Voinescu, Andrew W. Moore 0002, Gianni Antichi, Marcin Wójcik
SIGCOMM6
2017 ENDEAVOUR: A Scalable SDN Architecture For Real-World IXPs
abstract
Innovation in interdomain routing has remained stagnant for over a decade. Recently, Internet eXchange Points (IXPs) have emerged as economically-advantageous interconnection points for reducing path latencies and exchanging ever increasing traffic volumes among, possibly, hundreds of networks. Given their far-reaching implications on interdomain routing, IXPs are the ideal place to foster network innovation and extend the benefits of software defined networking (SDN) to the interdomain level. In this paper, we present, evaluate, and demonstrate ENDEAVOUR, an SDN platform for IXPs. ENDEAVOUR can be deployed on a multi-hop IXP fabric, supports a large number of use cases, and is highly scalable, while avoiding broadcast storms. Our evaluation with real data from one of the largest IXPs, demonstrates the benefits and scalability of our solution: ENDEAVOUR requires around 70% fewer rules than alternative SDN solutions thanks to our rule partitioning mechanism. In addition, by providing an open source solution, we invite everyone from the community to experiment (and improve) our implementation as well as adapt it to new use cases.
Gianni Antichi, Ignacio Castro, Marco Chiesa, Eder Leão Fernandes, Remy Lapeyrade, Daniel Kopp, Jong Hun Han, Marc Bruyere, Christoph Dietzel, Mitchell Gusat, Andrew W. Moore 0002, Philippe Owezarski, Steve Uhlig, Marco Canini
IEEE J. Sel. Areas Commun.1
2016 Horse: towards an SDN traffic dynamics simulator for large scale networks
abstract
The Software Defined Networking (SDN) paradigm can be successfully applied to the inter-domain ecosystem to empower network fabrics with finer grained policies and traffic engineering capabilities. However, introducing SDN at the inter-domain level might also lead to misconfigurations with potential to negatively impact on the Internet. Simulators are a popular approach to verify network behaviour and test applications before going into production. In the case of SDN, the available options do not scale for large scale networks nor high traffic loads. In this paper we propose a new simulator to foster SDN research and improve our understanding on the impact of the new use cases over the traffic flow. A simulation tool capable of efficiently reproducing large scale networks, high traffic loads, and policies, by abstracting the interactions between switches and controllers of the SDN network.
Eder Leão Fernandes, Gianni Antichi, Ignacio Castro, Steve Uhlig
SIGCOMM2
2015 Blueswitch: Enabling Provably Consistent Configuration of Network Switches
abstract
Previous research on consistent updates for distributed network configurations has focused on solutions for centralized networkconfiguration controllers. However, such work does not address the complexity of modern switch datapaths. Modern commodity switches expose opaque configuration mechanisms, with minimal guarantees for datapath consistency and with unclear configuration semantics. Furthermore, would-be solutions for distributed consistent updates must take into account the configuration guarantees provided by each individual switch - plus the compositional problems of distributed control and multi-switch configurations that considerably transcend the single-switch problems. In this paper, we focus on the behavior of individual switches, and demonstrate that even simple rule updates result in inconsistent packet switching in multi-table datapaths. We demonstrate that consistent configuration updates require guarantees of strong switch-level atomicity from both hardware and software layers of switches - even in a single switch. In short, the multiple-switch problems cannot be reasonably approached until single-switch consistency can be resolved. We present a hardware design that supports a transactional configuration mechanism, and provides packet-consistent configuration: all packets traversing the datapath will encounter either the old configuration or the new one, and never an inconsistent mix of the two. Unlike previous work, our design does not require modifications to network packets. We precisely specify the hardwaresoftware protocol for switch configuration; this enables us to prove the correctness of the design, and to provide well-specified invariants that the software driver must maintain for correctness. We implement our prototype switch design using the NetFPGA-10G hardware platform, and evaluate our prototype against commercial off-the-shelf switches.
Jong Hun Han, Prashanth Mundkur, Charalampos Rotsos, Gianni Antichi, Nirav Dave, Andrew W. Moore 0002, Peter G. Neumann
ANCS4
2015 Towards an SDN network control application for differentiated traffic routing
abstract
In the last years, Software Defined Networking has emerged as a promising paradigm to foster network innovation and address the issues coming from the ossification of the TCP/IP architecture. The clean separation between control and data plane, the definition of northbound and southbound interfaces are key features of the Software Defined Networking paradigm. Moreover, a centralised control plane allows network operators to deploy advanced control and management strategies. Effective traffic engineering and resources management policies allow to achieve a better utilisation of network resources and improve end-to-end service performance. This paper deals with the architectural design and experimental validation of a control application that enables differentiated routing for traffic flows belonging to different service classes. The new control application makes routing decisions leveraging on OpenFlow network statistics, i.e., taking advantage of real-time network status information. Moreover, a Deep Packet Inspection module has been developed and integrated in the control application to detect VoIP traffic with Session Initiation Protocol signalling, enforcing this way policies for a differentiated treatment of VoIP traffic. Finally, a functional validation is performed in emulated environment.
Davide Adami, Gianni Antichi, Rosario Giuseppe Garroppo, Stefano Giordano, Andrew W. Moore 0002
ICC2
2015 OFLOPS-Turbo: Testing the next-generation OpenFlow switch
abstract
The heterogeneity barrier breakthrough achieved by the OpenFlow protocol is currently paced by the variability in performance semantics among network devices, which reduces the ability of applications to take complete advantage of programmable control. As a result, control applications remain conservative on performance requirements in order to be generalizable and trade performance for explicit state consistency in order to support varying performance behaviours. In this paper we argue that network control must be optimized towards network device capabilities and network managers and application developers must perform informed design decision using accurate switch performance profiles. This becomes highly critical for modern OpenFlow-enabled 10 GbE optical switches which significantly elevate switch performance requirements. We present OFLOPS-Turbo, the integration of the OFLOPS switch evaluation platform, with the OSNT platform, a hardware-accelerated traffic generation and capture system supporting lossless 10 GbE functionality. Using OFLOPS-Turbo, we conduct an evaluation of flow table manipulation capabilities in a representative collection of 10 GbE production OpenFlow switch devices and interpret the evolution of OpenFlow support by comparison with historical data.
Charalampos Rotsos, Gianni Antichi, Marc Bruyere, Philippe Owezarski, Andrew W. Moore 0002
ICC2
2015 An integrated environment for open-source network softwarization
abstract
Network softwarization drives innovation both in software and hardware. This demo introduces a highly integrated environment that enables open source solutions for software defined network (SDN) in both hardware and software. This environment is built upon the NetFPGA platform for rapid prototyping of networking devices. It showcases tools (OSNT and OFLOPS) for evaluating the performance of networking devices, and demonstrates them using a pipelined multi-table OpenFlow enabled switch application. An open-source environment integrating both software and hardware that fully inter-operate, as demonstrated here, is essential for high-quality software defined networking solutions.
Jong Hun Han, Gianni Antichi, Noa Zilberman, Charalampos Rotsos, Andrew W. Moore 0002
NetSoft2
2015 Enabling Performance Evaluation Beyond 10 Gbps
abstract
Despite network monitoring and testing being critical for computer networks, current solutions are both extremely expensive and inflexible. This demo presents OSNT (www.osnt.org), a community-driven, high-performance, open-source traffic generator and capture system built on top of the NetFPGA-10G board which enables flexible network testing. The platform supports full line-rate traffic generation regardless of packet size across the four card ports, packet capture filtering and packet thinning in hardware and sub-msec time precision in traffic generation and capture, corrected using an external GPS device. Furthermore, it provides a software APIs to test the dataplane performance of multi-10G switches, providing a starting point for a number of different test cases. OSNT flexibility is further demonstrated through the OFLOPS-turbo platform: an integration of OSNT with the OFLOPS OpenFlow switch performance evaluation platform, enabling control and data plane evaluation of 10G switches. This demo showcases the applicability of the OSNT platform to evaluate the performance of legacy and OpenFlow-enabled networking devices, and demonstrates it using commercial switches.
Gianni Antichi, Charalampos Rotsos, Andrew W. Moore 0002
SIGCOMM1
2015 Extreme Data-rate Scheduling for the Data Center
abstract
Designing scalable and cost-effective data center interconnect architectures based on electrical packet switches is challenging. To overcome this challenge, researchers have tried to harness the advantages of optics in data center environment. This has resulted in exploration of hybrid switching architectures that contains an optical circuit switch to serve long bursts of traffic along with an electrical packet switch serving short bursts of traffic. The performance of such hybrid switching architectures in data center is dependent on the schedulers. Building hybrid schedulers is challenging because of varying properties of data center traffic, increasing network demands, requirements imposed by hybrid network architecture etc. Slow schedulers can negatively impact the performance of the data center network because of poor resource utilization. With future demands, this problem is going to escalate motivating the need for faster schedulers. One approach to do this would be to use a hardware based scheduler. In this paper we propose a framework that can be used to explore and evaluate hardware based hybrid schedulers.
Neelakandan Manihatty Bojan, Noa Zilberman, Gianni Antichi, Andrew W. Moore 0002
SIGCOMM3
2014 JA-trie: Entropy-based packet classification
abstract
Any improvement in packet classification performance is crucial to ensure Internet functions continue to track the ever-increasing link capacities. Packet classification is the foundation of many Internet functions: from fundamental packet-forwarding to advanced features such as Quality of Service en-forcement, monitoring and security functions. This work proposes a novel trie-based classification algorithm, named Jump-Ahead Trie (JA-trie), utilizing an entropy-based pre-processing phase and a novel approach to wildcard matching. Through extensive experimental tests, we demonstrate that our proposed algorithm is able to outperform a range of state-of-the-art classification algorithms.
Gianni Antichi, Christian Callegari, Andrew W. Moore 0002, Stefano Giordano, Enrico Anastasi
HPSR1
2014 On virtualization-aware traffic engineering in OpenFlow Data Centers networks
abstract
Oversubscription of intra-Data Center network links and high volatility of VM deployments require a flexible and agile control of Data Center network infrastructures, also integrated with computing and storage resources. In this scenario, the Software-Defined Network paradigm and, specifically, the OpenFlow protocol, opens up new opportunities for the design of innovative resource management platforms that enable dynamic and fine-grain control of DC networks through traffic engineering algorithms. This paper investigates the performance of two different sets of cloud-fluent traffic engineering algorithms. Conceived to work during cloud service deployments, the main target of such algorithms is to achieve a better utilization of network resources by exploiting OpenFlow capabilities for traffic-aware deployments of Virtual Machines. The effectiveness of the proposed solutions is evaluated in terms of network link utilization against VM requests acceptance ratio through simulations and experimental tests carried out by using an ad-hoc emulator.
Molka Gharbaoui, Barbara Martini, Davide Adami, Gianni Antichi, Stefano Giordano, Piero Castoldi
NOMS4
2013 Architecture for an open source network tester
abstract
To make networks more reliable, enormous resources are poured into all phases of the network-equipment lifecycle. The process starts early in the design phase when simulation is used to verify the correctness of a design, and continues through manufacturing and perhaps months of rigorously trials. With over 7,000 Internet RFCs and hundreds of IEEE standards, a typical piece of networking equipment undergoes hundreds of conformance tests before being deployed. Finally, when deployed in a production network, the equipment is tested regularly. Throughout the process, a relentless battery of tests and measurement help ensure the correct operation of the equipment.
Muhammad Shahbaz 0001, Gianni Antichi, Yilong Geng, Noa Zilberman, G. Adam Covington, Marc Bruyere, Nick Feamster, Nick McKeown, Bob Felderman, Michaela Blott, Andrew W. Moore 0002, Philippe Owezarski
ANCS2
2013 Effective resource control strategies using OpenFlow in cloud data center
Davide Adami, Barbara Martini, Molka Gharbaoui, Piero Castoldi, Gianni Antichi, Stefano Giordano
IM5
2012 An open hardware implementation of CUSUM based network anomaly detection
abstract
The detection of anomalies in backbone networks is posing serious performance issues, not only in terms of accuracy, but also in terms of detection speed. Indeed current software solutions to the problem, even promising from the point of view of detection and false alarm rates, suffer from the inability of performing the required operations in real time, when working in high speed backbone networks. On the other hand, hardware solutions are based on costly and inflexible niche systems.
Gianni Antichi, Christian Callegari, Stefano Giordano
GLOBECOM1
2012 Enabling open-source high speed network monitoring on NetFPGA
abstract
Network measurement both as diagnostic and within measurement-based techniques of traffic engineering and management, alongside network measurement for security has maintained the needs of researchers and network operators for the ongoing development of measurement tools for traffic monitoring/characterisation and to support Intrusion Detection Systems (IDSs). Many such tools capitalise on the pricing of commodity hardware by operating on general purpose architectures. Many are based on the well known libpcap API, a de facto standard in this area. Despite the many improvements that have been applied to packet capturing, packet-monitoring implementations still suffer from either: performance flaws on commodity hardware due mainly to unresolvable hardware bottlenecks, or costly and inflexible niche systems. To address such issues, the paper proposes a system architecture based on the cooperation of NetFPGA and a general purpose host PC. The NetFPGA is an open networking platform accelerator that enables rapid development of hardware-accelerated packet processing applications. The objective is to combine the high performance of a hardware-oriented solution with the flexibility of general purpose PCs.
Gianni Antichi, Stefano Giordano, David J. Miller 0005, Andrew W. Moore 0002
NOMS1
2011 Design and Development of an OpenFlow Compliant Smart Gigabit Switch
abstract
In this paper we propose a novel hardware-software co-design vision that aims at enhancing flexibility and reusability of hardware based packet forwarding engines. In particular, we move on the path of the well-known OpenFlow architecture that allows the user to decide the action to be performed over the packet (drop, forward through a given port etc.) upon interaction with a software control plane. Although such an approach is certainly powerful and is gaining more and more attention in both academia and industry, it is biased towards routing application: its main goal is to allow the software control plane to arbitrarily route a packet flow. However, we think that a similar paradigm, encompassing high performance packet forwarding hardware driven by a flexible software control plane, may be beneficial even to other kinds of applications, like monitoring and measurements. However, the primitives that the OpenFlow protocol provides are not flexible enough for such purposes. For this reason, we propose a flexible packet forwarding architecture based on regular expression that, besides enabling standard-compliant OpenFlow switching, can be easily reconfigured through its control plane to support other kinds of applications.
Gianni Antichi, Andrea Di Pietro, Stefano Giordano, Gregorio Procissi, Domenico Ficara
GLOBECOM1
2011 Scaling Regular Expression Matching Performance in Parallel Systems through Sampling Techniques
abstract
Modern network devices need to perform deep packet inspection at high speed for security and application- specific services. For this purpose, regular expressions are used, due to their high expressive power, and Deterministic Finite Automata (DFAs) are adopted to match them. Many works have been proposed to improve DFAs, especially in terms of memory consumption and speed. Instead, we address another issue: the scalability of DFAs to parallel systems and their buffer requirements. To our knowledge, a single attempt to parallelize DFA walk on regular multicore systems (which ex- ploits speculation with limited efficiency) has been proposed in literature. We propose a solution in which a number of processing units are committed to walk in parallel a DFA for the same packet; at this aim, sampling techniques on both text and regular expressions are adopted. This scheme is the first in literature that proposes effective parallelization of DFA walk, hence allowing for packet processing time reduction and less memory for reordering buffers. The result is that speed scales as the number of processing units.
Domenico Ficara, Gianni Antichi, Fabio Vitucci, Nicola Bonelli, Andrea Di Pietro, Stefano Giordano, Gregorio Procissi
GLOBECOM2
2011 Differential encoding of DFAs for fast regular expression matching
abstract
Deep packet inspection is a fundamental task to improve network security and provide application-specific services. State-of-the-art systems adopt regular expressions due to their high expressive power. They are typically matched through deterministic finite automata (DFAs), but large rule sets need a memory amount that turns out to be too large for practical implementation. Many recent works have proposed improvements to address this issue, but they increase the number of transitions (and then of memory accesses) per character. This paper presents a new representation for DFAs, orthogonal to most of the previous solutions, called delta finite automata ($\delta$FA), which considerably reduces states and transitions while preserving a transition per character only, thus allowing fast matching. A further optimization exploits$N$th order relationships within the DFA by adopting the concept of “temporary transitions.”
Domenico Ficara, Andrea Di Pietro, Stefano Giordano, Gregorio Procissi, Fabio Vitucci, Gianni Antichi
IEEE/ACM Trans. Netw.6
2010 A Randomized Scheme for IP Lookup at Wire Speed on NetFPGA
abstract
Because of the rapid growth of both traffic and links capacity, the time budget to perform IP address lookup on a packet continues to decrease and lookup tables of routers unceasingly grow. Therefore, new lookup algorithms and new hardware platform are required to perform fast IP lookup. This paper presents a new scheme on top of the NetFPGA board which takes advantage of parallel queries made on perfect hash functions. Such functions are built by using a very compact and fast data structure called Blooming Trees, thus allowing the vast majority of memory accesses to involve small and fast on-chip memories only.
Gianni Antichi, Andrea Di Pietro, Domenico Ficara, Stefano Giordano, Gregorio Procissi, Fabio Vitucci
ICC1
2010 Achieving Perfect Hashing through an Improved Construction of Bloom Filters
abstract
A Bloom Filter is an efficient randomized data structure for membership queries on a set with a certain known false positive probability. Bloom Filters (BFs) are very attractive for their limited memory requirements and their easy construction which make them a popular choice for many tasks in network devices. However, in a number of network applications, more than simple probabilistic membership queries is required, and BFs can be adopted as a coarse filtering stage, leaving the ultimate filtering and classification process to other techniques, such as hash tables or tree-like structures. In this paper we propose a scheme to extend BFs with "indexing" features so that when an element x is queried, an univocal index of that element is returned, which in turn can be used as an address for a table, just as a perfect hashing scheme. This extension, called indexed Bloom Filter (iBF), comes at the cost of a small increment of false positive probability and simply fits in existing BF-based applications.
Gianni Antichi, Andrea Di Pietro, Domenico Ficara, Stefano Giordano, Franco Russo, Fabio Vitucci
ICC1
2010 Sampling Techniques to Accelerate Pattern Matching in Network Intrusion Detection Systems
abstract
Modern network devices need to perform deep packet inspection at high speed for security and application-specific services. Instead of standard strings to represent the dataset to be matched, state-of-the-art systems adopt regular expressions, due to their high expressive power. The current trend is to use Deterministic Finite Automata (DFAs) to match regular expressions. However, while the problem of the large memory consumption of DFAs has been solved in many different ways, only a few works have focused on increasing the lookup speed. This paper introduces a novel yet simple idea to accelerate DFAs for security applications: payload sampling. Our approach allows to skip a large portion of the text, thus processing less bytes. The price to pay is a slight number of false alarms which require a confirmation stage. Therefore, we propose a double-stage matching scheme providing two new different automata. Results show a significant speed-up in regular traffic processing, thus confirming the effectiveness of the approach.
Domenico Ficara, Gianni Antichi, Andrea Di Pietro, Stefano Giordano, Gregorio Procissi, Fabio Vitucci
ICC2
2009 A Prefix-Distribution Adaptive Scheme for Routing Lookup Acceleration
abstract
IP address lookup is a fundamental task for Internet routers, due to the rapid growth of both traffic and links capacity. Many algorithms have been proposed to improve lookup performance in terms of memory consumption, search speed and update complexity. Due to the presence of wildcards and netmasks, such algorithms adopt several techniques to deal with longest prefix matching. However, the analysis of lookup tables reveals that the first 16 bits of forwarding rules are almost always specified. Therefore, more powerful exact-matching schemes can be applied to the first half of addresses. This paper presents a routing lookup accelerator (RLA) which allows the lookup of the first 16 bits to be sped up. The target is an efficient scheme to be implemented in small and fast memories of recent hardware platforms. Specifically, since in several forwarding tables the distribution of the first 16 bits is characterized by empty gaps as well as pronounced peaks, we propose to divide the address space in different ranges and to encode each address only as a difference with respect to a given address chosen as reference for that range. Then, a hybrid direct-addressing / multibit trie scheme is used for each range. As RLA is orthogonal to all other schemes, any other lookup algorithm can be used to perform longest prefix matching on the remaining bits.
Gianni Antichi, Andrea Di Pietro, Domenico Ficara, Stefano Giordano, Gregorio Procissi, Fabio Vitucci
GLOBECOM1
2009 End-to-End Inference of Link Level Queueing Delay Statistics
abstract
Characterizing delay distribution over the links of a network provides a remarkable amount of information which can be useful for troubleshooting, traffic engineering, adaptive multimedia flow coding, overlay network design, etc. Since querying each and every node of a path in order to retrieve this kind of information can be unfeasible or just too resource demanding, the recent research trend is to infer the internal state of a network by means of end-to-end measurements. Many algorithms in literature require active measurements and are based on a single-sender multiple-receivers scheme, thus relying on the cooperation of a possibly wide number of nodes, which is a quite strong assumption. Moreover, many previous works adopt Expectation-Maximization algorithms to cope with large and under-determined equation systems, thus increasing the uncertainty of the final delay estimation. This paper, instead, proposes a technique to infer the cumulants of the delay distribution over each link of a given network path, based on two-points measurements only. The cumulants, in turn, can be used to approximate the distribution function through the Edgeworth series. The results of our approach are assessed through a wide series of model-based and ns2 based simulations and show fairly good performance under different network load conditions.
Gianni Antichi, Andrea Di Pietro, Domenico Ficara, Stefano Giordano, Gregorio Procissi, Fabio Vitucci
GLOBECOM1
2009 Network Topology Discovery through Self-Constrained Decisions
abstract
Network Topology Discovery is crucial to a number of network management tasks. Traditional topology discovery techniques require internal nodes to take actions on measurement packets, which makes them unpractical in many cases. For these reasons, tomographic techniques have been introduced, which allow for the reconstruction of network topologies with no need for cooperation from internal routers. The usual approach to tomographic topology discovery is based on clustering nodes into tree structures according to soft similarity metrics. We recently proposed a novel technique based on decision theoretic considerations that help the topology reconstruction by limiting the set of hypotheses to a finite and well-defined set, thus determining hard metrics. In the scheme, probe traffic is sent to all couples of end-nodes and a metric is assigned to each measurement. In this paper, we extend the technique by ordering the topology reconstruction procedure according to metrics reliability defined in terms of their variances. The algorithms presented in the paper are validated through extensive simulations in several network scenarios. The results show that such a methodology allows to retrieve a complete picture of the network that includes the detection of all the internal nodes along with the values of capacities of the interconnecting links.
Gianni Antichi, Andrea Di Pietro, Domenico Ficara, Stefano Giordano, Gregorio Procissi, Fabio Vitucci
GLOBECOM1
2009 Second-Order Differential Encoding of Deterministic Finite Automata
abstract
Deep packet inspection is required in an increasing number of network devices, in order to improve network security and provide application-specific services. Instead of standard strings to represent the data set to be matched, state-of-the-art systems adopt regular expressions, due to their high expressive power and flexibility. Typically regular expressions are matched through deterministic finite automata (DFAs), but large rule sets need a memory amount which turns out to be too large for practical implementation. Many recent works have proposed improvements to address this issue, but they increase the number of transitions (and then of memory accesses) per character. In a previous work, we have presented a smart representation for DFA which, while preserving fast matching (i.e., a transition per character only), considerably reduces states and transitions. In this paper we introduce a novel optimized automaton, which exploits second order relationships within the DFA and is based on the key concept of "temporary transitions". Results for real data sets show that it allows for a further memory saving.
Gianni Antichi, Andrea Di Pietro, Domenico Ficara, Stefano Giordano, Gregorio Procissi, Fabio Vitucci
GLOBECOM1
2009 Faster DFAs through Simple and Efficient Inverse Homomorphisms
abstract
Performing deep packet inspection at high speed is a fundamental task for network security and application-specific services. In state-of-the-art systems, sets of signatures to be searched are described by regular expressions, and finite automata (FAs) are employed for the search. In particular, deterministic FAs (DFAs) need a large amount of memory to represent current sets, therefore the target of many works has been the reduction of memory footprint of DFAs. This paper, instead, focuses on speed multiplication by enlarging the amount of bytes observed in the text (i.e., searching for k-bytes per state-traversal). For this purpose, an interesting yet simple inverse homomorphism is employed to reduce the amount of transitions in the modified DFA. The algorithm results to be remarkably faster than standard DFAs, and provides also a good compression scheme that is orthogonal to other schemes.
Domenico Ficara, Stefano Giordano, Gregorio Procissi, Fabio Vitucci, Gianni Antichi, Andrea Di Pietro
INFOCOM5
2008 Design of a High Performance Traffic Generator on Network Processor
abstract
Evaluating the performance of high-speed networks is a critical task due to the lack of reliable tools to generate traffic workloads at high rates. The current open-source software tools are not suitable to deal with high-speed networks as they present poor performance in terms of generated frames per second and scarce timing/rate accuracy in traffic generation. These issues are due to the intrinsic limitations of the PC architecture, for which these tools are designed. This paper proposes a different approach based on the Intel Network Processor IXP2400. The design aims to maintain the high flexibility of PC solutions while outperforming them in terms of throughput and packet rate. This is obtained by combining a general-purpose PC with the processing units of a network processor.
Gianni Antichi, Andrea Di Pietro, Domenico Ficara, Stefano Giordano, Gregorio Procissi, Fabio Vitucci
DSD1
2008 Blooming Trees for Minimal Perfect Hashing
abstract
Hash tables are used in many networking applications, such as lookup and packet classification. But the issue of collisions resolution makes their use slow and not suitable for fast operations. Therefore, perfect hash functions have been introduced to make the hashing mechanism more efficient. In particular, a minimal perfect hash function is a function that maps a set of n keys into a set of n integer numbers without collisions. In literature, there are many schemes to construct a minimal perfect hash function, either based on mathematical properties of polynomials or on graph theory. This paper proposes a new scheme which shows remarkable results in terms of space consumption and processing speed. It is based on an alternative to Bloom Filters and requires about 4 bits per key and 12.8 seconds to construct a MPHF with 3.8times109elements.
Gianni Antichi, Domenico Ficara, Stefano Giordano, Gregorio Procissi, Fabio Vitucci
GLOBECOM1