Matty Kadosh

dblp:190/1358 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
8since 2021 · last 2025
0009-0000-2585-9581ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2025 OptiReduce: Resilient and Tail-Optimal AllReduce for Distributed Deep Learning in the Cloud
Ertza Warraich, Omer Shabtai, Khalid Manaa, Shay Vargaftik, Yonatan Piasetzky, Matty Kadosh, Lalith Suresh 0001, Muhammad Shahbaz 0001
NSDI6
2024 Towards Accelerating the Network Performance on DPUs by optimising the P4 runtime
abstract
Data Processing Units (DPUs) are becoming increasingly popular, especially for use in conjunction with Warehouse-Scale Computers (WSCs) due to their ability to handle networking functions and data-centric workloads. Cost-performance, energy efficiency, network 1/0, and batch processing workloads are important design factors for WSCs. Recent developments in AI and the never-ending increase in demand for data processing, cloud computing, and HPC set the optimisation of all those design factors as a high priority. DPUs can be utilised to achieve significant improvements in all those areas. This includes in-line network processing and upcoming enhanced security paradigms such as post-quantum cryptography (PQC) for quantum re-silient communications or software-defined perimeters (SDP) for confidential computing implementations. Being P4-enabled and dRMT-based, DPUs allow for the reconfigurability of the network traffic without the need to change the hardware. However, the network performance on such devices is only sometimes deter-ministic since the actual traffic and the rules, both of which have to do with packet processing, are not known during compile time. In this paper, we envision how the network performance on DPUs can be accelerated. We describe the challenges that negatively impact the bandwidth and latency: the complex steering pipeline and the massive runtime needed to optimise. These challenges arise from the lack of information during compile time that is only known during runtime. Thus, we envision optimising during runtime by leveraging DPUs' reconfigurability on the network 110. For this, we discuss the significant factors that must be considered to accelerate the network performance on such devices and we propose a solution for them.
Dimosthenis Iliadis-Apostolidis, Khalid Manaa, Matty Kadosh, Iacovos Ioannou, Vasos Vassiliou, Sokol Kosta, Juan Jose Vegas Olmos
PDP3
2023 Unleashing SmartNIC Packet Processing Performance in P4
abstract
SmartNICs are on the rise as a packet processing platform, with the trend towards a uniform P4 programming model. However, unleashing SmartNIC packet processing performance in P4 is a formidable task. Traditional SmartNIC optimizations rely on low-level program tuning, but P4 abstractions operate at one level above. At the same time, today's P4 optimizations primarily focus on resource packing rather than performance tuning. We develop Pipeleon, an automated performance optimization framework for P4 programmable SmartNICs. We introduce techniques that are tailored to the performance characteristics of SmartNICs, and further leverage dynamic workload patterns for profile-guided optimization. Pipeleon pinpoints program hotspots at the P4 level and computes runtime optimization plans to specialize the program layout based on the latest profile. We have prototyped Pipeleon and applied it to optimize two popular P4 SmartNICs---Nvidia BlueField2 and Netronome Agilio CX---as well as a software SmartNIC emulator extended based on BMv2. Our results show that Pipeleon significantly improves SmartNIC packet processing performance in realistic scenarios.
Jiarong Xing, Yiming Qiu 0001, Kuo-Feng Hsu, Songyuan Sui, Khalid Manaa, Omer Shabtai, Yonatan Piasetzky, Matty Kadosh, Arvind Krishnamurthy, T. S. Eugene Ng, Ang Chen 0001
SIGCOMM8
2023 On the Protection of a High Performance Load Balancer Against SYN Attacks**This is an extended journal version of [2]
abstract
SYN flooding is a simple and effective denial-of-service attack. In this attack, many TCP SYN requests are sent to the targeted server, in an attempt to consume its resources and make it unresponsive to legitimate traffic. While SYN attacks have traditionally targeted web servers, they are also known to be very harmful to intermediate cloud devices, and in particular to stateful load balancers (LBs). Fighting against a SYN attack without negatively affecting legitimate connections is not easy, especially if the LB needs to perform frequent server pool updates during the attack, which is very likely since attacks can often last for many hours or even days. This paper is the first to propose LB schemes that guarantee high throughput of one million connections per second, while supporting a high pool update rate without breaking connections and fighting against a high rate SYN attack. Using an analysis and a proof of concept, we show that the LB can handle up to 10 million fake SYNs per second when the RTT is 10ms, and up to 5 million fake SYNs per second when the RTT is 20ms.
Reuven Cohen, Matty Kadosh, Alan Lo, Qasem Sayah
IEEE Trans. Cloud Comput.2
2022 Runtime Programmable Switches
Jiarong Xing, Kuo-Feng Hsu, Matty Kadosh, Alan Lo, Yonatan Piasetzky, Arvind Krishnamurthy, Ang Chen 0001
NSDI3
2022 LB Scalability: Achieving the Right Balance Between Being Stateful and Stateless
abstract
A high performance Layer-4 load balancer (LB) is one of the most important components of a cloud service infrastructure. Such an LB uses network and transport layer information for deciding how to distribute client requests across a group of servers. A crucial requirement for a stateful LB is per connection consistency (PCC); namely, that all the packets of the same connection will be forwarded to the same server, as long as the server is alive, even if the pool of servers or the assignment function changes. The challenge is in designing a high throughput, low latency solution that is also scalable. This paper proposes a highly scalable LB, called Prism, implemented using a programmable switch ASIC. As far as we know, Prism is the first reported stateful LB that can process millions of connections per second and hundreds of millions connections in total, while ensuring PCC. This is due to the fact that Prism forwards all the packets in hardware, even during server pool changes,while avoiding the need to maintain a hardware state per every active connection. We implemented a prototype of the proposed architecture and showed that Prism can scale to 100 million simultaneous connections, and can accommodate more than one pool update per second.
Reuven Cohen, Matty Kadosh, Alan Lo, Qasem Sayah
IEEE/ACM Trans. Netw.2
2021 Hardware SYN Attack Protection For High Performance Load Balancers
abstract
SYN flooding is a simple and effective denial-of-service attack, in which an attacker sends many SYN requests to a target's server in an attempt to consume server resources and make it unresponsive to legitimate traffic. While SYN attacks have traditionally targeted web servers, they are also known to be very harmful to intermediate cloud devices, and in particular to stateful load balancers (LBs). We propose LB schemes that guarantee high throughput of one million connections per second, while supporting a high pool update rate without breaking connections, and fighting against a high rate SYN attack, of up to 10 million fake SYNs per second.
Reuven Cohen, Matty Kadosh, Alan Lo, Qasem Sayah
HOTI2
2021 A Vision for Runtime Programmable Networks
abstract
Our community has made significant progress in developing programmable network infrastructure, starting from the control plane and expanding to the data plane. As a latest trend, network devices are becoming runtime programmable while serving live traffic. This allows for reprogramming of individual device programs at fine-grained timescales to add or remove network functions. Many applications and services, however, need control over a combination of devices, including end host stacks, NICs, and switches, to accomplish their goals. We lay out our vision for runtime programmable networks, building upon device-level features to provide live, network-wide, runtime reprogramming. A whole-stack approach is needed with new programming models, compiler support, and network management abstractions. We outline a research agenda as a call to arms to the community.
Jiarong Xing, Yiming Qiu 0001, Kuo-Feng Hsu, Matty Kadosh, Alan Lo, Aditya Akella, Thomas E. Anderson, Arvind Krishnamurthy, T. S. Eugene Ng, Ang Chen 0001
HotNets5
2018 Switch ASIC Programmability in Hybrid Mode
abstract
Programmable ASIC technology enables the switching data plane to rapidly support emergent technologies such as VNF offloading, custom tunneling and in-band telemetry. We propose a new approach for a "hybrid mode" of ASIC programmability, which maintains a discrete legacy hardware pipeline and control functions (e.g. routing, bridging) while providing a way to extend it. This places requirements on the switching hardware, programming language, data plane APIs and the network OS in order to achieve this goal. In this paper we present two hardware agnostic hybrid mode applications using a Mellanox programmable switch ASIC, P4-16 programming language, SAI flexible APIs and the SONIC Open Network OS and Linux TC. Also applications based on the Onyx OS and Spectrum SDK is discussed as a hardware specific example.
Yonatan Piasetzky, Matty Kadosh, Marian Pritsak, Omer Shabtai, Alan Lo, Guohan Lu
ICNP2
2016 Unlocking Credit Loop Deadlocks
abstract
The recently emerging Converged Enhanced Ethernet (CEE) data center networks rely on layer-2 flow control in order to support packet loss sensitive transport protocols, such as RDMA and FCoE. Although lossless networks were proven to improve end-to-end network performance, without careful design and operation, they might suffer from in-network deadlocks, caused by cyclic buffer dependencies. These dependencies are called credit loops. Although existing credit loops rarely deadlock, when they do they can block large parts of the network. Naive solutions recover from credit loop deadlock by draining buffers and dropping packets. Previous works suggested credit-loop avoidance by central routing algorithms, but these assume specific topologies and are slow to react to failures.
Alexander Shpiner, Eitan Zahavi, Vladimir Zdornov, Tal Anker, Matty Kadosh
HotNets5