Tom Barbette

dblp:162/5287 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0003-1269-2190ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 14 · 6 first-author · 9 since 2021Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 xPUBench: Scalable and Energy-Efficient GPU and DPU-Accelerated Network Functions
Maxime Vanliefde, Romain Van Hauwaert, Nikita Tyunyayev, Clément Delzotti, Elena Agostini, Tom Barbette
PAM6
2025 OpenDesc: From Static NIC Descriptors to Evolvable Metadata Interfaces
abstract
Modern NICs offer rich functionalities, but host-side software lacks a unified way to express or adapt to their capabilities. Instead, developers rely on device-specific code and ad-hoc glue layers, leading to duplication, reduction to the lowest common denominator, inefficiency, and poor reuse. As NICs grow more flexible, the lack of a shared interface description becomes a core architectural bottleneck.
Seyyidahmed Lahmer, Nikita Tyunyayev, Tom Barbette
HotNets3
2025 PAMO: Pattern Matching Offload for Intrusion Detection Systems
abstract
Intrusion Detection Systems (IDS) play a crucial role in network security. An IDS recognizes malicious activity in network traffic by matching it against patterns defined in a set of rules. The complexity and size of rule sets lead to substantial computational load. In a state-of-the-art IDS, such as Suricata, a single CPU core processes a few hundred MB to a few GB of network traffic per second, and rule evaluation accounts for over 60% of CPU consumption. Scaling IDS to today's high-speed networks is, therefore, a significant challenge.
Lukás Sismis, Colin Evrard, Etienne Rivière, Tom Barbette
Middleware4
2025 Sloth: A Kernel-Bypass Scheduler Maximizing Energy Efficiency under Latency Constraints
Clément Delzotti, Pol Maistriaux, Tom Barbette
Networking3
2024 Poster: Enhancing the Performance of a Single Connection Using Multipath Quic
abstract
The QUIC protocol, designed to reduce latency and improve internet security, faces goodput performance challenges in high-speed networks, particularly with single-thread implementations. This poster extends an existing userspace Multipath QUIC (MPQUIC) implementation to enhance the goodput of a single connection by pinning different network path logics to different cores. Our solution, mcMPQUIC, achieves a goodput of up to 20 Gbps with ten paths/cores, surpassing the baseline MPQUIC performance by more than five times.
Vany V. Ingenzi, Tom Barbette, Olivier Bonaventure
ICNP2
2024 A High-Speed Robust Tunnel Using Forward Erasure Correction in Segment Routing
abstract
Low-latency applications drive an increasing number of modern applications. Latency depends on factors such as link layer technologies and how higher-layer protocols cope with transmission errors and packet losses. Most transport protocols rely exclusively on retransmissions to cope with losses with minimal overhead but potentially large tail latency. This paper leverages network coding to propose the high-speed robust tunnel (HIRT), providing timely packet delivery to any application independently of the transport protocol. A network code recovers lost data without requiring time-consuming retransmissions by adding redundancy packets, thereby slightly increasing the bandwidth usage to reduce the tail latency. Our algorithm dynamically adapts the rate of redundancy packets by measuring the network loss patterns. We implement HIRT using IPv6 Segment Routing (SRv6). We suggest an efficient software implementation and demonstrate on CloudLab that our solution can protect traffic at high speeds ($>50 \text{Gbps}$) on standard servers even when facing severe packet losses in the network. We evaluate HIRT with HTTP over TCP/QUIC, and file system benchmarks over a real network with losses, Starlink. HIRT reduces the tail latency of short HTTP requests by$2 \times$and the mean request completion time of longer requests by up to$20 \%$. HIRT also decreases the tail latency of NFS requests by up to$20 \%$.
Louis Navarre, François Michel, Tom Barbette
ICNP3
2023 A High-Speed Stateful Packet Processing Approach for Tbps Programmable Switches
Mariano Scazzariello, Tommaso Caiazzi, Hamid Ghasemirahni, Tom Barbette, Dejan Kostic, Marco Chiesa
NSDI4
2022 Packet Order Matters! Improving Application Performance by Deliberately Delaying Packets
Hamid Ghasemirahni, Tom Barbette, George P. Katsikas, Alireza Farshin, Amir Roozbeh, Massimo Girondi, Marco Chiesa, Gerald Q. Maguire Jr., Dejan Kostic
NSDI2
2022 Retina: analyzing 100GbE traffic on commodity hardware
abstract
As network speeds have increased to over 100 Gbps, operators and researchers have lost the ability to easily ask complex questions of reassembled and parsed network traffic. In this paper, we introduce Retina, a software framework that lets users analyze over 100 Gbps of real-world traffic on a single server with no specialized hardware. Retina supports running arbitrary user-defined analysis functions on a wide variety of extensible data representations ranging from raw packets to parsed application-layer handshakes. We introduce a novel filtering mechanism and subscription interface to safely and efficiently process high-speed traffic. Under the hood, Retina implements an efficient data pipeline that strategically discards unneeded traffic and defers expensive processing operations to preserve computation for complex analyses. We present the framework architecture, evaluate its performance on production traffic, and explore several applications. Our experiments show that Retina is capable of running sophisticated analyses at over 100 Gbps on a single commodity server and can support 5--100× higher traffic rates than existing solutions, dramatically reducing the effort to complete investigations on real-world networks.
Gerry Wan, Fengchen Gong, Tom Barbette, Zakir Durumeric
SIGCOMM3
2022 Cheetah: A High-Speed Programmable Load-Balancer Framework With Guaranteed Per-Connection-Consistency
abstract
Large service providers use load balancers to dispatch millions of incoming connections per second towards thousands of servers. There are two basic yet critical requirements for a load balancer:uniform load distributionof the incoming connections across the servers, which requires to support advanced load balancing mechanisms, andper-connection-consistency(PCC), i.e, the ability to map packets belonging to the same connection to the same server even in the presence of changes in the number of active servers and load balancers. Yet, simultaneously meeting these requirements has been an elusive goal. Today’s load balancers minimize PCC violations at the price of non-uniform load distribution. This paper presents Cheetah, a load balancer that supports advanced load balancing mechanismsandPCC while being scalable, memory efficient, fast at processing packets, and offers comparable resilience to clogging attacks as with today’s load balancers. The Cheetah LB design guarantees PCC foranyrealizable server selection load balancing mechanism and can be deployed in both stateless and stateful manners, depending on operational needs. We implemented Cheetah on both a software and a Tofino-based hardware switch. Our evaluation shows that a stateless version of Cheetah guarantees PCC, has negligible packet processing overheads, and can support load balancing mechanisms that reduce the flow completion time by a factor of$2-3 \times $.
Tom Barbette, Erfan Wu, Dejan Kostic, Gerald Q. Maguire Jr., Panagiotis Papadimitratos, Marco Chiesa
IEEE/ACM Trans. Netw.1
2021 PacketMill: toward per-Core 100-Gbps networking
abstract
We present PacketMill, a system for optimizing software packet processing, which (i) introduces a new model to efficiently manage packet metadata and (ii) employs code-optimization techniques to better utilize commodity hardware. PacketMill grinds the whole packet processing stack, from the high-level network function configuration file to the low-level userspace network (specifically DPDK) drivers, to mitigate inefficiencies and produce a customized binary for a given network function. Our evaluation results show that PacketMill increases throughput (up to 36.4 Gbps -- 70%) & reduces latency (up to 101 us -- 28%) and enables nontrivial packet processing (e.g., router) at ~100 Gbps, when new packets arrive >10× faster than main memory access times, while using only one processing core.
Alireza Farshin, Tom Barbette, Amir Roozbeh, Gerald Q. Maguire Jr., Dejan Kostic
ASPLOS2
2021 High-speed Connection Tracking in Modern Servers
abstract
The rise of commodity servers equipped with high-speed network interface cards poses increasing demands on the efficient implementation of connection tracking, i.e., the task of associating the connection identifier of an incoming packet to the state stored for that connection. In this work, we thoroughly investigate and compare the performance obtainable by different implementations of connection tracking using high-speed real traffic traces. Based on a load balancer use case, our results show that connection tracking is an expensive operation, achieving at most 24 Gbps on a single core. Core-sharding and lock-free hash tables emerge as the only suitable multi-thread approaches for enabling 100 Gbps packet processing. In contrast to recent beliefs, we observe that newly proposed techniques to "lazily" delete connection states are not more effective than properly tuned traditional deletion techniques based on timer wheels.
Massimo Girondi, Marco Chiesa, Tom Barbette
HPSR3
2021 What You Need to Know About (Smart) Network Interface Cards
George P. Katsikas, Tom Barbette, Marco Chiesa, Dejan Kostic, Gerald Q. Maguire Jr.
PAM2
2021 Combined Stateful Classification and Session Splicing for High-Speed NFV Service Chaining
abstract
Network functionssuch as firewalls, NAT, DPI, content-aware optimizers, and load-balancers are increasingly realized as software to reduce costs and enable outsourcing. To meet performance requirements thesevirtualnetwork functions (VNFs) often bypass the kernel and use their own user-space networking stack. A naïve realization of a chain of VNFs will exchange raw packets, leading to many redundant operations, wasting resources. In this work, we design a system to execute a pipeline of VNFs. We provide the user facilities to define (i) a traffic class of interest for the VNF, (ii) a session to group the packets (such as the TCP 4-tuple), and (iii) the amount of space per session. The system synthesizes a classifier and builds an efficient flow table that when possible will automatically be partially offloaded and accelerated by the network interface. We utilize an abstract view of flows to support seamless inspection and modification of the content of any flow (such as TCP or HTTP). By applying only surgical modifications to the protocol headers, we avoid the need for a complex, hard-to-maintain user-space TCP stack and can chain multiple VNFswithout re-constructing the stream multiple times, allowing up to 5x improvement over standard approaches.
Tom Barbette, Cyril Soldani, Laurent Mathy
IEEE/ACM Trans. Netw.1
2020 Stateless CPU-aware datacenter load-balancing
abstract
Today, datacenter operators deploy Load-balancers (LBs) to efficiently utilize server resources, but must over-provision server resources (by up to 30%) because of load imbalances and the desire to bound tail service latency. We posit one of the reasons for these imbalances is the lack of per-core load statistics in existing LBs. As a first step, we designed CrossRSS, a CPU core-aware LB that dynamically assigns incoming connections to the least loaded cores in the server pool. CrossRSS leverages knowledge of the dispatching by each server's Network Interface Card (NIC) to specific cores to reduce imbalances by more than an order of magnitude compared to existing LBs in a proof-of-concept datacenter environment, processing 12% more packets with the same number of cores.
Tom Barbette, Marco Chiesa, Gerald Q. Maguire Jr., Dejan Kostic
CoNEXT1
2020 A High-Speed Load-Balancer Design with Guaranteed Per-Connection-Consistency
Tom Barbette, Haoran Yao, Dejan Kostic, Gerald Q. Maguire Jr., Panagiotis Papadimitratos, Marco Chiesa
NSDI1
2020 Metron: High-performance NFV Service Chaining Even in the Presence of Blackboxes
abstract
Deployment of 100Gigabit Ethernet (GbE) links challenges the packet processing limits of commodity hardware used for Network Functions Virtualization (NFV). Moreover, realizing chained network functions (i.e., service chains) necessitates the use of multiple CPU cores, or even multiple servers, to process packets from such high speed links. Our system Metron jointly exploits the underlying network and commodity servers’ resources: ( i ) to offload part of the packet processing logic to the network, ( ii ) by using smart tagging to setup and exploit the affinity of traffic classes, and ( iii ) by using tag-based hardware dispatching to carry out the remaining packet processing at the speed of the servers’ cores, with zero inter-core communication. Moreover, Metron transparently integrates, manages, and load balances proprietary “blackboxes” together with Metron service chains. Metron realizes stateful network functions at the speed of 100GbE network cards on a single server, while elastically and rapidly adapting to changing workload volumes. Our experiments demonstrate that Metron service chains can coexist with heterogeneous blackboxes, while still leveraging Metron’s accurate dispatching and load balancing. In summary, Metron has ( i ) 2.75–8× better efficiency, up to ( ii ) 4.7× lower latency, and ( iii ) 7.8× higher throughput than OpenBox, a state-of-the-art NFV system.
George P. Katsikas, Tom Barbette, Dejan Kostic, Gerald Q. Maguire Jr., Rebecca Steinert
ACM Trans. Comput. Syst.2
2019 RSS++: load and state-aware receive side scaling
abstract
While the current literature typically focuses on load-balancing among multiple servers, in this paper, we demonstrate the importance of load-balancing within a single machine (potentially with hundreds of CPU cores). In this context, we propose a new load-balancing technique (RSS++) that dynamically modifies the receive side scaling (RSS) indirection table to spread the load across the CPU cores in a more optimal way. RSS++ incurs up to 14x lower 95th percentile tail latency and orders of magnitude fewer packet drops compared to RSS under high CPU utilization. RSS++ allows higher CPU utilization and dynamic scaling of the number of allocated CPU cores to accommodate the input load while avoiding the typical 25% over-provisioning.
Tom Barbette, George P. Katsikas, Gerald Q. Maguire Jr., Dejan Kostic
CoNEXT1
2018 Building a chain of high-speed VNFs in no time: Invited Paper
abstract
To cope with the growing performance needs of appliances in datacenters or the network edge, current middle-box functionalities such as firewalls, NAT, DPI, content-aware optimizers or load-balancers are often implemented on multiple (perhaps virtual) machines. In this work, we design a system able to run a pipeline of VNFs with a high level of parallelism to handle many flows. We provide the user facilities to define the traffic class of interest for the VNF, a definition of session to group the packets such as the TCP 4-tuples, and the amount of space per sessions. The system will then synthesize the classification and build a unique, efficient flow table. We build an abstract view of flows and use it to implement support for seamless inspection and modification of the content of any flow (such as TCP or HTTP), automatically reflecting a consistent view, across layers, of flows modified on-the-fly. Our prototype gives rise to a user-space software NFV data-plane enabling easy implementation of middlebox functionalities, as well as the deployment of complex scenarios. Our prototype implementation is able to handle our testbed limit of -34 Gbps of HTTP requests (for 8-KB files) through a service chain of multiples stateful VNFs, on a single Xeon core.
Tom Barbette, Cyril Soldani, Romain Gaillard, Laurent Mathy
HPSR1
2018 Metron: NFV Service Chains at the True Speed of the Underlying Hardware
George P. Katsikas, Tom Barbette, Dejan Kostic, Rebecca Steinert, Gerald Q. Maguire Jr.
NSDI2
2015 Fast Userspace Packet Processing
abstract
In recent years, we have witnessed the emergence of high speed packet I/O frameworks, bringing unprecedented network performance to userspace. Using the Click modular router, we rst review and quantitatively compare several such packet I/O frameworks, showing their superiority to kernel-based forwarding. We then reconsider the issue of software packet processing, in the context of modern commodity hardware with hardware multi-queues, multi-core processors and non-uniform memory access. Through a combination of existing techniques and improvements of our own, we derive modern general principles for the design of software packet processors. Our implementation of a fast packet processor framework, integrating a faster Click with both Netmap and DPDK, ex-hibits up-to about 2.3x speed-up compared to other software implementations, when used as an IP router.
Tom Barbette, Cyril Soldani, Laurent Mathy
ANCS1