Elie F. Kfoury

dblp:223/8072 · DBLP profile ↗
← Back
35ranked-venue papers
6as first author
31since 2021 · last 2026
0000-0003-1236-6168ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 24 · 6 first-author · 21 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Accelerating Anomaly Detection in Industrial Control Systems Using SmartNICs and DPDK
abstract
Industrial Control Systems (ICS) are critical infrastructures that integrate physical processes with programmable logic controllers (PLCs), human–machine interfaces (HMIs), and network communication. Given their dual exposure to cyber and physical threats, continuous monitoring is essential to ensure reliability and safety. A key requirement of ICS is maintaining low latency, as delays can desynchronize control loops and compromise system stability.This paper presents a hybrid detection framework that combines sensor-level time-series fault analysis with flow-based network anomaly detection for Modbus/TCP-based ICS. The framework leverages an NVIDIA BlueField-3 SmartNIC to offload machine-learning inference and sequential signal processing directly to the network interface using DPDK. The proposed system employs cumulative sum (CUSUM) and exponentially weighted moving average (EWMA) techniques to extract features for a multilayer perceptron (MLP) classifier, which identifies normal and faulty sensor behavior with high per-device accuracy. Network flow statistics are also analyzed to detect cyberattacks targeting the ICS infrastructure. Experimental results show sub–1 μs average inference latency for binary classification and an average of 5 μs per 32-packet burst on the BlueField-3 SmartNIC.
Sergio Elizalde, Samia Choueiri, Ali Mazloum, Elie F. Kfoury, Jorge Crichigno
CCNC5
2026 A Testbed to Evaluate Next-Generation Security Solutions in Cyber-Physical Systems using Hardware Acceleration
abstract
At the core of modern manufacturing systems lie Cyber-Physical Systems (CPS) that prioritize operational continuity over security, resulting in a rising number of cyberattacks targeting critical infrastructures. This paper presents a work-in-progress testbed that modernizes Smart Manufacturing Systems (SMS) by integrating Domain-Specific Accelerators (DSAs)—Data Processing Units (DPUs) and Programmable Data Plane (PDP) switches—to strengthen Operational Technology (OT) security without compromising availability or reliability. These accelerators provide fine-grained visibility, real-time anomaly detection, and efficient policy enforcement at line rate. Preliminary results show that accelerator-based applications outperform CPU-based implementations by several orders of magnitude. Demonstrated use cases include a DPU that performs memory inspection via Direct Memory Access (DMA) to detect injected anomalies and a PDP that implements inline detection using pre-trained Machine Learning (ML) models. With low processing overhead, the system also enables continuous telemetry collection for digital-twin generation without disrupting critical operations. The testbed, deployed on the South Carolina Cloud (SC Cloud), offers remote access for developing and evaluating next-generation CPS and OT security applications.
Ali AlSabeh, Ali Mazloum, Elie F. Kfoury, Ramy F. Harik, Thorsten Wuest, Jorge Crichigno
CCNC4
2026 Design and Deployment of a Testbed for SmartNIC and Programmable Data Plane Experimentation
abstract
This paper presents the design and deployment of a virtualized testbed that facilitates experimentation and instruction in programmable network systems. The platform integrates Smart Network Interface Cards (SmartNICs), Programmable Data Plane (PDP) switches, and the Data Plane Development Kit (DPDK), within a cloud-based orchestration framework to reproduce high-performance, real-world networking scenarios. It supports line-rate processing, enabling real-time applications such as telemetry, encrypted traffic inspection, and malware detection. Through a series of use cases, we demonstrate how the testbed enables advanced experimentation by offloading infrastructure functions to the data plane, achieving low latency, high throughput, and high scalability. The system also provides users with guided labs for learners and is accessible via NETLAB+ for remote use. Future work includes federation with national-scale infrastructures such as FABRIC to broaden access and support multi-institutional collaboration.
Samia Choueiri, Ali Mazloum, Sergio Elizalde, Amith GSPN, Ali AlSabeh, Elie F. Kfoury, Jorge Crichigno
CCNC7
2026 Real-Time Encrypted Traffic Classification with P4-DPDK
Amith Gorthi Srinivasa Prabhakara Narasimha, Ali Mazloum, Samia Choueiri, Sergio Elizalde, Elie F. Kfoury, Jorge Crichigno
ICC5
2025 Detection and Mitigation of Volumetric DDoS Attacks using Adaptive Rate-Limiting in P4-DPDK
abstract
Distributed Denial of Service (DDoS) attacks are increasingly targeting network environments. This paper presents a high-performance, adaptive system for detecting and mitigating volumetric DDoS attacks using P4-DPDK. The proposed system operates entirely in the user space, leveraging multicore CPUs and SmartNICs to process traffic at line rate while enabling flexible and efficient control plane operations. DDoS detection is implemented in a linear prediction model that forecasts traffic based on historical observations and dynamically adjusts ratelimiting thresholds. The hyperparameters for the model are tuned using an optimization algorithm. The system is evaluated on the FABRIC testbed and tested using real traffic traces. Experimental results demonstrate robust mitigation against diverse attack types, effective adaptation to real-world traffic, and reduced packet loss under high-throughput conditions approaching 100 Gbps, compared to Suricata-DPDK implementations.
Samia Choueiri, Ali Mazloum, Sergio Elizalde, Elie F. Kfoury, Jorge Crichigno
GLOBECOM4
2025 Toward Fingerprinting Encrypted C2 Traffic in the Data Plane
abstract
• Transport Layer Security (TLS) is the dominant protocol that enables users to securely interact with the Internet. • Threat actors are using TLS to bypass traditional cybersecurity defenses like firewalls and intrusion detection systems. • Modern malware attacks are hiding behind TLS secure channels. • Many malware families that infect users receive malicious instructions from the command and control (C2) server. • As the communication between malware and the C2 server is encrypted, it can easily bypass modern security appliances that rely on deep packet inspection (DPI). • In response to this threat, this project aims at utilizing ML to identify encrypted C2 communication. • The project implements a distributed ML model over two hardware accelerators. • The system achieves 99.3% detection accuracy with a microsecond-level processing latency.
Ali Mazloum, Elie F. Kfoury, Ali AlSabeh, Jorge Crichigno
GLOBECOM2
2025 Real-Time Flow Statistics Collection Using RDMA and P4 Programmable Data Planes
abstract
Measuring network traffic in real time is essential for applications such as traffic profiling, anomaly detection, resource allocation, and network performance improvement. As network speeds and traffic volumes increase, traditional solutions (e.g., NetFlow, sFlow, Zeek) face challenges in processing and summarizing traffic efficiently, often leading to incomplete measurements. This paper introduces a system that summarizes network traffic and provides per-flow measurements in real time by leveraging P4 Programmable Data Planes (PDPs). The system computes per-flow traffic statistics directly in the data plane at line rate. The statistics are then transmitted to a server using the low-latency, high-throughput RDMA over Converged Ethernet (RoCEv2) protocol. On the server, worker threads process the received reports and update a global data structure that maintain the flows. The system was implemented and tested using an Intel Tofino-based PDP and an RDMA-capable SmartNIC (NVIDIA BlueField-2). Experiments on real packet traces show that the system is capable of analyzing traffic at scale without compromising the accuracy of the measurements, outperforming traditional Network Security Monitors (NSMs).
Elie F. Kfoury, Ali Mazloum, Ali AlSabeh, Jorge Crichigno
ICC1
2025 Domain Name Security Inspection at Line Rate: Tls Sni Extraction in the Data Plane Using P4 and Dpdk
abstract
A widely adopted approach to monitor HTTPS traffic leverages the Server Name Identification (SNI) extension of TLS. Generally, the hostname is transferred in plain text over the SNI field and Deep Packet Inspection (DPI) is used to parse the TLS header and extract the hostname. However, DPI is often performed on general-purpose processors and utilizes the kernel of the operating system, which results in an overhead to the network, especially under high traffic loads. To this end, this paper proposes offloading the identification of SNI hostnames to the data plane using P4 and the Data Plane Development Kit (DPDK). In the proposed system, a P4 Programmable Data Plane (PDP) switch is the first line of defense where most of the TLS traffic is processed. DPDK is the second line of defense which processes all TLS packets that require processing capabilities beyond what the P4 PDP switch provides. To support line rate pattern matching on the hostname, the DPDK application is offloaded to a SmartNIC, leveraging its Regex engine. Experiments on various recent and public datasets from different regions and platforms reveal that the P4 switch is capable of parsing 85%99 % of hostnames. Furthermore, performance analysis shows that the P4 switch and the DPDK application, respectively, inspect a hostname in around 1 microsecond ($\mu \mathrm{s}$) and$7 \mu ~\mathrm{s}$, achieving an order of magnitude improvement over solution running on general-purpose processors.
Ali Mazloum, Ali AlSabeh, Elie F. Kfoury, Jorge Crichigno
ICC3
2025 Enabling Line-Rate TLS SNI Inspection in P4 Programmable Data Planes
abstract
With the increasing adoption of the HyperText Transfer Protocol Secure (HTTPS), organizations face new challenges in monitoring traffic to defend against attacks and enforce security policies, such as filtering malicious websites. One widely used technique to monitor HTTPS is by scrutinizing the hostname in the Server Name Identification (SNI) extension during the Transport Layer Security (TLS) handshake. Parsing the SNI typically involves Deep Packet Inspection (DPI), often performed on general-purpose processors, which can create bottlenecks and significantly impact network throughput. In response, this paper introduces a novel framework for parsing and identifying SNI hostnames in the data plane at line-rate using P4. Evaluation results on recent publicly available datasets from various regions and platforms demonstrate that our framework can successfully parse 85%-99% of hostnames in P4. Furthermore, performance analysis reveals that the proposed data plane solution can inspect the hostname in approximately 1 microsecond (μ s), representing orders of magnitude improvement over solutions running on Central Processing Units (CPUs).
Ali AlSabeh, Ali Mazloum, Elie F. Kfoury, Jorge Crichigno, Hala Strohmier Berry
NOMS3
2025 Performance Evaluation of Stateless Firewalling: Host-Based, SmartNIC, and P4 Switch
abstract
In modern data centers, traditional software-based firewalls often struggle to keep up with the growing demands for robust security. One adopted approach to improve the packet processing efficiency is through software acceleration techniques like the Data Plane Development Kit (DPDK), which significantly enhances the performance of software-based firewalls. Another approach is to offload the firewall functionalities to the hardware either using the new generation of Network Interface Cards (SmartNICs) or by using P4 Programmable Data Plane (PDP) switches. SmartNICs integrate dedicated processing units optimized to efficiently handle networking tasks, including security. P4 PDP switches enable custom packet processing in the data plane, allowing security applications to run at line rates within the network. This study compares the performance of stateless firewalls implemented using nftables (host-based), DPDK (host-based), DOCA Flow and OvS hardware (SmartNIC-based), and P4 (PDP-based). It evaluates the achievable throughput and processing latency for the five implementations under different testing scenarios. It also compares the CPU utilization of the host-based implementations. The results demonstrate that while the software acceleration technique significantly enhances host-side performance, SmartNIC-based and PDP-based firewalls provide a superior performance over all software-based implementations.
Sergio Elizalde, Ali Mazloum, Samia Choueiri, Elie F. Kfoury, Jorge Crichigno
NOMS4
2025 Real-Time Congestion Control Algorithm Identification with P4 Programmable Switches
abstract
The classification of Congestion Control Algorithms (CCAs) is vital in current networks, where increasing traffic demands and dynamic conditions challenge their stability and performance. CCAs play a fundamental role in managing congestion, balancing throughput, minimizing latency, and reducing packet loss to ensure reliable data transmission across diverse scenarios. Despite their importance, accurately identifying the CCA in use remains a challenging task; critical for optimizing resource allocation and enhancing Quality of Service (QoS). This paper presents a framework that leverages P4-programmable switches for real-time extraction of key traffic metrics, including queuing delay, interarrival time, queue depth, RTT, and sending rate. These metrics are analyzed using a Random Forest classifier to predict the CCA in use with high accuracy. Extensive experiments in a controlled network environment, featuring a bottleneck link and flows utilizing CCAs such as Cubic, Reno, BBR, and Vegas, validate the effectiveness of our approach.
Andrés García-López, Elie F. Kfoury, Jorge Crichigno, Jaime Galán-Jiménez
NOMS2
2025 Improving flow fairness in non-programmable networks using P4-programmable Data Planes
Elie F. Kfoury, Ali Mazloum, Jorge Crichigno
Comput. Networks2
2025 Security applications in P4: Implementation and lessons learned
Ali Mazloum, Ali AlSabeh, Elie F. Kfoury, Jorge Crichigno
Comput. Networks3
2025 A survey on security applications with SmartNICs: Taxonomy, implementations, challenges, and future trends
abstract
Over the last decade, network applications have grown exponentially, demanding high-speed interconnects. Unfortunately, chip manufacturers are approaching the upper limits of silicon-based computing with slow improvements in computational performance and energy efficiency. This trend has forced the industry to shift paradigms, moving from monolithic architectures to heterogeneous, domain-specific designs. Moreover, the ever-evolving threats compromise digital services and demand more scalable and flexible solutions to ensure service continuity in production networks. Smart Network Interface Cards (SmartNICs) are a product of this new paradigm, integrating domain-specific engines and general-purpose cores to offload various network infrastructure tasks, including those related to security. This paper provides a comprehensive overview of SmartNICs, with a particular focus on their role in strengthening network defenses. It introduces SmartNIC technology and presents a taxonomy of security applications offloaded to SmartNICs, categorized into Intrusion Detection and Prevention Systems (IDS/IPS), defenses against volumetric attacks, and data confidentiality mechanisms. Additionally, the paper explores vulnerabilities associated with adopting SmartNICs in the cloud, examining the threat model and reviewing proposed remediations in the literature. Finally, it discusses challenges and future trends in SmartNIC security applications, highlighting current initiatives and open research areas.
Sergio Elizalde, Ali AlSabeh, Ali Mazloum, Samia Choueiri, Elie F. Kfoury, Jorge Crichigno
J. Netw. Comput. Appl.5
2025 Enhancing visibility on a science DMZ with P4-perfSONAR
abstract
The Science Demilitarized Zone (Science DMZ) is a specialized network designed to facilitate the transfer of large-scale scientific data. One of the key elements of the Science DMZ is perfSONAR, an active performance measurement device that monitors end-to-end paths over multiple domains. Although versatile, perfSONAR faces limitations such as restricted visibility of events and coarse-grained measurements. This paper proposes a scheme that integrates P4 programmable data plane (PDP) switches with perfSONAR. P4 PDP switches are passively installed and operate on real-time traffic copies, providing flexibility to collect fine-grained custom measurements and report events in the data plane. This integration enables perfSONAR to collect per-flow granular statistics of actual traffic, identify a broader range of networking issues, and enhance visibility while reducing the overhead of active tests. Additionally, the scheme uses an adaptive linear prediction (LP) model that dynamically adjusts the rate of reports sent from the P4 PDP switch to perfSONAR, minimizing the storage and processing needed for the latter. Experimental results show that the system reduces the number of reports by a factor of five while maintaining a small and configurable relative mean error (RME).
Ali Mazloum, Elie F. Kfoury, Ali AlSabeh, Jorge Crichigno
J. Netw. Comput. Appl.2
2024 Scalable Heavy Hitter Detection: A DPDK-based Software Approach with P4 Integration
abstract
Identifying heavy hitters is vital for applications like Denial of Service (DoS) detection and traffic engineering. Current solutions fall into hardware or software categories. Hardware solutions (e.g., P4 programmable data plane switches) offer high performance but require adding hardware, which may not be ideal for virtualized environments (e.g., cloud). Software solutions are cost-effective and flexible but suffer from performance issues due to the packet processing overhead in the Operating System (OS) kernel. This paper presents a scalable heavy hitter detection algorithm in the software, bypassing the kernel using the Data Plane Development Kit (DPDK). The Count-min Sketch (CMS) data structure is used to estimate the frequency of packets per flow. The system is implemented in P4 and deployed on the P4-DPDK target running on CPU cores. The experiments analyzed the impact of various parameters such as the packet size distribution, the number of CPU cores, and the number of hash functions, on the performance and the accuracy of the detection. The system's performance is further evaluated through comparison with another DPDK-based approach for heavy hitter detection. The results show accurate identification of heavy hitters and improved performance, even at a high traffic rate approaching 100Gbps.
Samia Choueiri, Ali Mazloum, Elie F. Kfoury, Jorge Crichigno
GLOBECOM3
2024 Enabling Fairness in Flow Allocation using P4-programmable Data Planes
abstract
This paper presents a system designed to enhance Transmission Control Protocol (TCP) fairness by rebalancing router queues and reducing the impact of Round-Trip Time (RTT) unfairness. The proposed system utilizes a P4-programmable Data Plane (PDP) to process a copy of the traffic from the link between two non-programmable routers. The PDP measures the throughput and calculates the RTT of competing flows in the data plane. Then, the control plane generates the rules to be implemented in a non-programmable router that will allocate flows in different queues to isolate their dynamics. The limits for each queue result from the Jenks optimization algorithm. This approach ensures that flows with similar characteristics share the same queue.The results demonstrate that the system efficiently identifies and segregates flows into multiple queues, thereby enforcing fairness among competing flows and enhancing the Flow Completion Time (FCT). The experiments were executed on traffic provided by Measurement and Analysis on the WIDE Internet (MAWI). The system effectively rebalances queues and dynamically redistributes underutilized bandwidth independently of the design principles of the transport protocol. Furthermore, the results show that the system effectively mitigates the effects of bufferbloat and successfully detects and reduces the impact of protocol abuses at the network layer.
Elie F. Kfoury, Ali Mazloum, Jorge Crichigno
GLOBECOM2
2024 Reducing the Impact of RTT Unfairness using P4-Programmable Data Planes
abstract
This paper presents a system that mitigates the Round-trip Time (RTT) unfairness issue in non-programmable networks using P4-programmable data planes. In traditional loss-based congestion control algorithms (CCAs), RTT unfairness occurs when the flows with shorter RTTs obtain higher bandwidth shares with respect to the flows with longer RTTs. This behavior occurs due to the faster recovery period that flows with shorter RTTs experience after a loss event. On the other hand, more recent CCAs, such as the Bottleneck Bandwidth and Round-trip Time (BBR), present the opposite behavior, where the flows with longer RTTs achieve higher throughput than the ones with shorter RTTs. In this paper, the proposed system employs a P4-programmable data plane to monitor the RTT of flows traversing a non-programmable router at line rate using passive taps. The P4-programmable data plane analyzes the RTT of each flow, sub-sequently segregating them into different queues. This separation is aimed at minimizing the interaction between flows with varying RTTs. Results show that implementing flow separation improves the fairness of long flows, reduces the RTT of individual flows allocated in different queues, and improves the Flow Completion Times (FCTs) of short flows. P4, RTT unfairness, Transmission Control Protocol (TCP), Congestion Control Algorithm (CCA), Bottleneck Bandwidth and Round-trip Time (BBR).
Elie F. Kfoury, Jorge Crichigno, Gautam Srivastava 0001
ICC2
2024 perfSONAR: Enhancing Data Collection through Adaptive Sampling
abstract
perfSONAR IS a tool used to monitor and troubleshoot problems in high-speed networks such as Science Demilitarized Zones (DMZs). It is essential to validate that data transfers are performing as expected. However, perfSONAR suffers from the trade-off between the measurement accuracy and the overhead induced by its active testsThis paper presents a scheme that offloads the traffic monitoring to a programmable data plane (PDP) switch. The scheme integrates a PDP switch with perfSONAR, where the switch continuously collects network measurements (e.g., latency, throughput, packet loss rate) and periodically reports the measurements to the perfSONAR archiver. This integration significantly enhances the granularity, visibility, and troubleshooting capabilities of perfSONAR. Additionally, the scheme automates the reporting period according to the variability of the monitored measurements, which eliminates the need of human intervention observed in today’s networks. In contrast to traditional schemes that report all measurements, the proposed approach uses the Linear Prediction (LP) method to only report the samples that reveal a variation on the measurements. Experimental results show that the system reduces the number of reports by five times under stable network conditions and sustains a relative mean error (RME) below 0.06.
Ali Mazloum, Ali AlSabeh, Elie F. Kfoury, Jorge Crichigno
NOMS3
2024 Machine learning controller for data rate management in science DMZ networks
Christian Vega Caicedo, Elie F. Kfoury, Jorge E. Pezoa, Miguel E. Figueroa, Jorge Crichigno
Comput. Networks2
2024 Evaluating TCP BBRv3 performance in wired broadband networks
Elie F. Kfoury, Jorge Crichigno, Gautam Srivastava 0001
Comput. Commun.2
2024 On DGA Detection and Classification Using P4 Programmable Switches
Ali AlSabeh, Kurt Friday, Elie F. Kfoury, Jorge Crichigno, Elias Bou-Harb
Comput. Secur.3
2024 P4BS: Leveraging Passive Measurements From P4 Switches to Dynamically Modify a Router's Buffer Size
abstract
The performance of networked applications can be dramatically impacted by the size of the buffer at the bottleneck router. Shallow buffers may increase packet losses and decrease link utilization, while deep buffers may increase the queueing delays for latency-sensitive flows. Operators nowadays configure large buffers statically without considering the characteristics of flows or dynamic traffic patterns. This paper presents P4BS, a system that dynamically modifies the buffer size of a legacy router. P4BS leverages programmable switches as passive instruments to measure various metrics that are vital when deciding on buffer size. The measured metrics include the number of long-lived flows and their round-trip times, the packet loss rates, and the queueing delays. Using these measurements, the programmable switch sequentially searches for a buffer size that minimizes the queueing delays and the packet loss rates. The system was implemented on a Tofino hardware switch and the system was tested on a wide range of network scenarios. The results show improvements in the quality of service of various applications including Web browsing, video streaming, and voice over IP.
Elie F. Kfoury, Jorge Crichigno, Elias Bou-Harb
IEEE Trans. Netw. Serv. Manag.1
2023 P4CCI: P4-Based Online TCP Congestion Control Algorithm Identification for Traffic Separation
abstract
Congestion Control Algorithms (CCAs) regulate the sending rates of hosts to avoid congestion in the network. Studies have shown that when flows belonging to different CCAs coexist on the same link, their shares on that link are significantly different. If the CCAs of active flows can be determined on live traffic, then flows belonging to the same CCA can be allocated into a dedicated queue. Unfortunately, identifying the CCA at line rate is not straightforward since the CCA is not advertised in the header fields of a packet. Moreover, with Gigabits per second (Gbps) traffic crossing a network, analyzing each packet to infer the CCA is not possible, especially with general-purpose CPUs. This paper proposes P4CCI, a system that detects the CCA of a flow at line rate by leveraging Programmable Data Planes (PDP). The PDP computes and extracts the flow's bytes-in-flight and sends them to a Deep Learning model for classification. Once classified, the flows are allocated into dedicated queues based on their CCA type. The system was implemented and tested on real hardware that uses Intel's Tofino ASIC. The experiments were executed on traffic provided by CAIDA. Results show that P4CCI can detect the CCAs with high accuracy. Furthermore, the performance of the network is greatly improved when the flows are separated by their CCAs.
Elie F. Kfoury, Jorge Crichigno, Elias Bou-Harb
ICC1
2023 A survey on network simulators, emulators, and testbeds used for research and education
Elie F. Kfoury, Jorge Crichigno, Gautam Srivastava 0001
Comput. Networks2
2023 A Survey on Rerouting Techniques with P4 Programmable Data Plane Switches
Ali Mazloum, Elie F. Kfoury, Jorge Crichigno
Comput. Networks2
2022 Enabling P4 Hands-on Training in an Academic Cloud
abstract
This paper describes a cloud infrastructure and virtual laboratories on P4 programmable data plane switches. P4 programmable data planes emerged as a technology that enables innovation in networking. P4 is a programming language used to describe how network packets are processed. This paper explains an entry-level training library on P4. The virtual laboratories introduce the learner to P4 and data plane concepts by providing step-by-step guides and exercises. The virtual laboratories are hosted in the Academic Cloud, a distributed platform that manages and orchestrates computing resources. Additionally, the paper describes a work in progress of P4 virtual laboratories that uses Intel Tofino switches. Lastly, the paper discusses the use of the Academic Cloud as a network testbed.
Elie F. Kfoury, Jorge Crichigno
DCOSS2
2022 INC: In-Network Classification of Botnet Propagation at Line Rate
Kurt Friday, Elie F. Kfoury, Elias Bou-Harb, Jorge Crichigno
ESORICS (1)2
2022 A survey on security applications of P4 programmable switches and a STRIDE-based vulnerability assessment
Ali AlSabeh, Joseph Khoury, Elie F. Kfoury, Jorge Crichigno, Elias Bou-Harb
Comput. Networks3
2022 A survey on TCP enhancements using P4-programmable devices
Elie F. Kfoury, Jorge Crichigno, Gautam Srivastava 0001
Comput. Networks2
2021 Dynamic Router's Buffer Sizing using Passive Measurements and P4 Programmable Switches
abstract
The router's buffer size imposes significant impli-cations on the performance of the network. Network operators nowadays configure the router's buffer size manually and stati-cally. They typically configure large buffers that fill up and never go empty, increasing the Round-trip Time (RTT) of packets significantly and decreasing the application performance. Few works in the literature dynamically adjust the buffer size, but are implemented only in simulators, and therefore cannot be tested and deployed in production networks with real traffic. Previous work suggested setting the buffer size to the Bandwidth-delay Product (BDP) divided by the square root of the number of long flows. Such formula is adequate when the RTT and the number of long flows are known in advance. This paper proposes a system that leverages programmable switches as passive instruments to measure the RTT and count the number of flows traversing a legacy router. Based on the measurements, the programmable switch dynamically adjusts the buffer size of the legacy router in order to mitigate the unnecessary large queuing delays. Results show that when the buffer is adjusted dynamically, the RTT, the loss rate, and the fairness among long flows are enhanced. Additionally, the Flow Completion Time (FCT) of short flows sharing the queue is greatly improved. The system can be adopted in campus, enterprise, and service provider networks, without the need to replace legacy routers.
Elie F. Kfoury, Jorge Crichigno, Elias Bou-Harb, Gautam Srivastava 0001
GLOBECOM1
2020 Offloading Media Traffic to Programmable Data Plane Switches
abstract
According to estimations, approximately 80% of Internet traffic represents media traffic. Much of it is generated by end users communicating with each other (e.g., voice, video sessions). A key element that permits the communication of users that may be behind Network Address Translation (NAT) is the relay server. This paper presents a scheme for offloading media traffic from relay servers to programmable switches. The proposed scheme relies on the capability of a P4 switch with a customized parser to de-encapsulate and process packets carrying media traffic. The switch then applies multiple switch actions over the packets. As these actions are simple and collectively emulate a relay server, the scheme is capable of moving relay functionality to the data plane operating at terabits per second. Performance evaluations show that the proposed scheme not only produces optimal results regarding Quality of Service (QoS) parameters (no packet loss, minimum delay, negligible delay variation, high Mean Opinion Score) but also scales much better than current solutions. Evaluations conducted with up to 35Gbps of media traffic or its equivalent of 400,000 simultaneous G.711 media sessions (limited only by the traffic generator rather than by the switch) show an ideal operation of the switch-based solution (using$\sim \text{l}$% of the switching capacity). In contrast, a relay server with a modern CPU model used for evaluations can process up to 900 simultaneous G.711 media sessions per core.
Elie F. Kfoury, Jorge Crichigno, Elias Bou-Harb
ICC1
2020 Towards a Unified In-Network DDoS Detection and Mitigation Strategy
abstract
Distributed Denial of Service (DDoS) attacks have terrorized our networks for decades, and with attacks now reaching 1.7 Tbps, even the slightest latency in detection and subsequent remediation is enough to bring an entire network down. Though strides have been made to address such maliciousness within the context of Software Defined Networking (SDN), they have ultimately proven ineffective. Fortunately, P4 has recently emerged as a platform-agnostic language for programming the data plane and in turn allowing for customized protocols and packet processing. To this end, we propose a first-of-a-kind P4-based detection and mitigation scheme that will not only function as intended regardless of the size of the attack, but will also overcome the vulnerabilities of SDN that have characteristically been exploited by DDoS. Moreover, it successfully defends against the broad spectrum of currently relevant attacks while concurrently emphasizing the Quality of Service (QoS) of legitimate end-users and overall SDN functionality. We demonstrate the effectiveness of the proposed scheme using a software programmable P4-switch, namely, the Behavorial Model version 2 (BMv2), showing its ability to withstand a variety of DDoS attacks in real-time via three use cases that can be generalized to most contemporary attack vectors. Specifically, the results substantiate that the mechanism herein is orders of magnitude faster than traditional polling techniques (e.g., NetFlow or sFlow) while minimizing the impact on benign traffic. We concur that the approach's design particularities facilitate seamless and scalable deployments in high-speed networks requiring line-rate functionality, in addition to being generic enough to be integrated into viable network topologies.
Kurt Friday, Elie F. Kfoury, Elias Bou-Harb, Jorge Crichigno
NetSoft2
2020 An emulation-based evaluation of TCP BBRv2 Alpha for wired broadband
Elie F. Kfoury, Jorge Crichigno, Elias Bou-Harb
Comput. Commun.1
2019 A Flow-Based Entropy Characterization of a NATed Network and Its Application on Intrusion Detection
abstract
This paper presents a flow-based entropy characterization of a small/medium-sized campus network that uses network address translation (NAT). Although most networks follow this configuration, their entropy characterization has not been previously studied. Measurements from a production network show that the entropies of flow elements (external IP address, external port, campus IP address, campus port) and tuples have particular characteristics. Findings include: i) entropies may widely vary in the course of a day. For example, in a typical weekday, the entropies of the campus and external ports may vary from below 0.2 to above 0.8 (in a normalized entropy scale 0-1). A similar observation applies to the entropy of the campus IP address; ii) building a granular entropy characterization of the individual flow elements can help detect anomalies. Data shows that certain attacks produce entropies that deviate from the expected patterns; iii) the entropy of the 3-tuple {external IP, campus IP, campus port} is high and consistent over time, resembling the entropy of a uniform distribution's variable. A deviation from this pattern is an encouraging anomaly indicator; iv) strong negative and positive correlations exist between some entropy time-series of flow elements.
Jorge Crichigno, Elie F. Kfoury, Elias Bou-Harb, Nasir Ghani, Yasmany Prieto, Christian Vega Caicedo, Jorge E. Pezoa, David Torres
ICC2