Ralf Kundel

dblp:242/0398 · DBLP profile ↗
← Back
21ranked-venue papers
10as first author
15since 2021 · last 2026
0000-0003-1711-5990ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 3 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 User Plane Performance in Beyond 5G Networks: Comprehensive Analysis and Evaluation
abstract
Emerging applications such as autonomous driving, virtual reality, and smart factories place greater demands on the Quality of Service of existing network infrastructure, particularly radio networks. The current 5th and new 6th generation of cellular networks aim to meet these requirements and provide ubiquitous connectivity to devices with diverse demands. These networks comprise a control plane and a user plane. While the control plane is responsible for managing the network and its devices, the user plane forwards data and directly influences the experienced Quality of Service. A key network function in the user plane is the User Plane Function (UPF), which forwards packets between cellular network devices and the data network, such as the Internet or an edge data center. However, the extent to which existing UPF implementations can provide sufficient Quality of Service for emerging applications remains largely unexplored. In this work, we analyze and compare various UPF implementations from both theoretical and practical perspectives. We consider both software-based and hardware-accelerated implementations and compare them in terms of performance and latency under load. The setup enables up to 10,000 subscriber sessions while enforcing QoS mechanisms such as rate limiting. The evaluation demonstrates that three of the four investigated UPFs provide QoS enforcement, while their latency behavior differs by orders of magnitude depending on the employed technology.
Fridolin Siegmund, Ralf Kundel, Tobias Meuser, Ralf Steinmetz
Comput. Commun.2
2025 Massive QoS-Aware Packet Queueing and Traffic Shaping at the Access Edge Using FPGAs
abstract
Large-scale packet queueing and scheduling is the basis for today's Quality of Service (QoS) in computer access networks, especially to achieve guaranteed high throughput and low latency. The throughput of a single network function, implementing the QoS functionality for a large number of customers, is in the range of hundreds of gigabits or even more. Therefore, a good performance of the underlying hardware is mandatory. While highly-performant fixed-function ASICs offer sufficient functionality for most data center use cases, as of today, they cannot support all functionality required for access networks, including QoS-aware packet queuing. In this paper, we first analyze mobile and residential Internet access requirements from an Internet service provider perspective, focusing on the QoS-aware packet queueing needs. Considering this analysis, we present a universal and generic FPGA design for high-performance packet queueing and scheduling. Our evaluation results show that FPGAs can be used to implement a deterministic QoS packet queueing system with high performance. This concept can extend today's programmable networking ASICs with the desired functionality as an offloaded “sidecar”.
Ralf Kundel, Lisa Wernet, Fridolin Siegmund, Leonhard Nobach, Hans-Jörg Kolbe, Tobias Meuser
NOMS1
2025 Instant P4STA: Beyond Tbit/s Network Function Evaluation with P4 Programmable Hardware
abstract
Cloud data center, backbone, and access networks constantly push the boundaries towards lower latencies, jitter, and scalable throughput. Evaluating data plane devices, i.e., switches, routers, and complex network functions, by developers and service operators under demanding settings is imperative to ensure service resilience in real-world deployments. Our proposed prototype, Instant P4STA, extends a packet timestamping framework for programmable hardware by combining a hardware packet generator with a uniform browser-based packet editor for dynamic packet generation. The user can specify the packet template bit-by-bit, utilizing the Python library Scapy with a vast variety of packet templates. This way, our prototype combines the best features of software and hardware-based packet generators. We demonstrate packet generation up to 3.2 Tbit/s on eight egress ports with up to four packet types in parallel. More packet generation throughput is possible with more egress ports, capped only by the number of physical ports in the programmable hardware.
Fridolin Siegmund, Matthias Hollick, Ralf Kundel
NOMS3
2024 RDA: Residence Delay Aggregation for Time-Sensitive Networking
abstract
Time-Sensitive Networking (TSN) enables deterministic and low-latency communication for real-time applications over Ethernet. That is accomplished by leveraging scheduling and shaping techniques configured for each egress port within the network switches. Although Time Aware Shaper (TAS) is a promising solution for TSN, its adoption often involves substantial complexity. In this work, we propose Residence Delay Aggregation (RDA), a novel asynchronous TSN mechanism that offers dynamic traffic scheduling adapted to the traffic load. Specifically, the proposed RDA mechanism provides upper bound delays similar to other asynchronous TSN mechanisms while improving the flexibility of traffic scheduling and reducing the deployment complexity.
Chengbo Zhou, Christoph Gärtner, Amr Rizk, Boris Koldehofe, Björn Scheuermann 0001, Ralf Kundel
NOMS6
2023 CML-IDS: Enhancing Intrusion Detection in SDN Through Collaborative Machine Learning
abstract
The centralized control plane in Software-Defined Networking (SDN) offers significant advancements in network management capabilities. However, SDN is also susceptible to cybersecurity risks and vulnerabilities. Deploying the Machine Learning (ML) approach in an Intrusion Detection System (IDS) can facilitate early detection of potential vulnerabilities. However, deploying an ML-based IDS solely in either the SDN control plane or the data plane has its benefits and drawbacks. For instance, a high-capacity ML model deployed in the control plane can enhance the detection performance but may increase network latency and the risk of overwhelming the control plane. In contrast, lightweight ML models deployed in the data plane could accelerate intrusion detection with lower detection performance. However, a functional IDS should provide a good detection performance at a line rate. To accomplish these objectives, we introduce a novel method called Collaborative ML-based IDS (CML-IDS), which involves deploying ML models in both the control and data planes to detect network attacks collaboratively. To facilitate this collaboration, we assess the confidence of the classification model, which is flexibly deployed within the programmable data plane. Our evaluation results demonstrate that the CML-IDS enhances the average intrusion detection performance to 93.46% and reduces the misclassification rate by 54.66% when compared to an IDS that solely relies on the ML model deployed in the data plane. Furthermore, CML-IDS effectively reduces network latency caused by forwarding flows to the control plane.
Pegah Golchin, Chengbo Zhou, Pratyush Agnihotri, Mehrdad Hajizadeh, Ralf Kundel, Ralf Steinmetz
CNSM5
2023 RPM: Reverse Path Congestion Marking on P4 Programmable Switches
abstract
Transport layer congestion control relies on feedback signals that travel from the congested link to the receiver and back to the sender. This forward congestion control loop, first, requires at least one Round-Trip Time (RTT) to react to congestion and secondly, it depends on the downstream path after the bottleneck. The former property leads to a reaction time in the order of RTT+ bottleneck queue delay, while the second may amplify the unfairness due to heterogeneous RTT. In this paper, we present Reverse Path Congestion Marking (RPM) to accelerate the reaction to network congestion events without changing the end-host stack. RPM decouples the congestion signal from the downstream path after the bottleneck while maintaining the stability of the congestion control loop. We show that RPM improves throughput fairness for RTT-heterogeneous TCP flows as well as the flow completion time, especially for small Data Center TCP (DCTCP) flows. Finally, we show RPM evaluation results in a testbed built around P4 programmable ASIC switches.
Nehal Baganal Krishna, Tuan-Dat Tran, Ralf Kundel, Amr Rizk
LCN3
2023 Demo: Flexibility-aware Network Management of Time-Sensitive Flows
abstract
We investigate the application of a recently published metric for flexibility in the context of combined port queue schedules of network paths in Time-Sensitive Networks (TSN). TSN comprises a set of specifications for deterministic networking, including support for scheduled traffic with guaranteed deterministic end-to-end delays. Typically, scheduler resource allocation in TSN disregards flexibility of scheduler configurations. Essentially, the notion of flexibility of paths comprising multiple concatenated ports having each a TSN configuration is based on the number of possible embeddings, i.e., resource allocations, for a new flow of a given specification (size and delay deadline) along that path. This demonstration allows the user to define TSN schedules along network paths and, hence, illustrates the behavior and benefit of performing flexibility-aware TSN configuration.
Christoph Gärtner, Amr Rizk, Boris Koldehofe, René Guillaume, Ralf Kundel, Ralf Steinmetz
SIGCOMM5
2023 Fast incremental reconfiguration of dynamic time-sensitive networks at runtime
Christoph Gärtner, Amr Rizk, Boris Koldehofe, René Guillaume, Ralf Kundel, Ralf Steinmetz
Comput. Networks5
2022 FPGA-assisted Massive Packet Queueing and Traffic Shaping at the Network Edge
abstract
Large-scale packet queueing and scheduling is the basis for today’s Quality of Service (QoS) in computer access networks, especially to achieve guaranteed high throughput and low latency. While highly performant fixed-function ASICs offer sufficient functionality for most data center use-cases, as of today, they cannot support all functionality required for access networks, e.g., QoS-aware packet queueing.In this poster, we present an FPGA-based architecture for network packet queueing optimized for residential and mobile Internet access networks.
Ralf Kundel, Leonhard Nobach, Hans-Jörg Kolbe, Tobias Meuser, Ralf Steinmetz
FCCM1
2022 Host Bypassing: Let your GPU speak Ethernet
abstract
Hardware acceleration of network functions is essential to meet the challenging Quality of Service requirements in nowadays computer networks. Graphical Processing Units (GPU) are a widely deployed technology that can also be used for computing tasks, including acceleration of network functions. In this work, we demonstrate how commodity GPUs, which do not provide any network interfaces, can be used to accelerate network functions. Our approach leverages PCIe peer-to-peer capabilities and allows the GPU to control the network interface card directly, without any assistance from the operating system or control application. The presented evaluation results demonstrate the feasibility of our approach and its performance of up to 10 Gbit/s, even for small packets.
Ralf Kundel, Leonard Anderweit, Jonas Markussen, Carsten Griwodz, Osama Abboud, Benjamin Becker, Tobias Meuser
NetSoft1
2022 Improving DDoS Attack Detection Leveraging a Multi-aspect Ensemble Feature Selection
abstract
DDoS attack detection is crucial in computer networks to meet the reliability and accessibility requirements of online services. The ability of machine learning to discriminate between DDoS attacks and benign flows makes it a promising candidate for DDoS detection. Correctly classifying the flows with high performance in near real-time is a critical issue for an ML-based DDoS detector to reduce the damages of DDoS attacks. In order to improve the performance of classification and reduce the prediction time, we propose a multi-aspect Ensemble Feature Selection (EFS) for DDoS attack detection in this work. The presented EFS selects the most relevant features of each attack separately, leveraging a combination of statistical filtering approaches and machine learning methods. We evaluate our method on two different datasets to demonstrate the EFS robustness toward model-specific biases. Last, we demonstrate that the prediction time is reduced leveraging the proposed EFS.
Pegah Golchin, Ralf Kundel, Tim Steuer, Rhaban Hark, Ralf Steinmetz
NOMS2
2021 P4-CoDel: Experiences on Programmable Data Plane Hardware
abstract
Fixed buffer sizing in computer networks, especially the Internet, is a compromise between latency and bandwidth. A decision in favor of high bandwidth, implying larger buffers, subordinates the latency as a consequence of constantly filled buffers. This phenomenon is called Bufferbloat. Active Queue Management (AQM) algorithms such as CoDel or PIE, designed for the use on software based hosts, offer a flow agnostic remedy to Bufferbloat by controlling the queue filling and hence the latency through subtle packet drops.In previous work, we have shown that the data plane programming language P4 is powerful enough to implement the CoDel algorithm. While legacy software algorithms can be easily compiled onto almost any processing architecture, this is not generally true for AQM on programmable data plane hardware, i.e., programmable packet processors. In this work, we highlight corresponding challenges, demonstrate how to tackle them, and provide techniques enabling the implementation of such AQM algorithms on different high speed P4-programmable data plane hardware targets. In addition, we provide measurement results created on different P4-programmable data plane targets. The resulting latency measurements reveal the feasibility and the constraints to be considered to perform Active Queue Management within these devices. Finally, we release the source code and instructions to reproduce the results in this paper as open source to the research community.
Ralf Kundel, Amr Rizk, Jeremias Blendin, Boris Koldehofe, Rhaban Hark, Ralf Steinmetz
ICC1
2021 Poster: Reverse-Path Congestion Notification: Accelerating the Congestion Control Feedback Loop
abstract
Congestion control mechanisms in computer networks rely mainly on a feedback loop having a reaction time equal to the flow RTT. Reducing this feedback time helps the sender to react faster to changing network conditions such as congestion. In this work, we propose reverse-path congestion notification on top of programmable networking switches. Our approach can significantly lower the reaction time, such that the congestion control implementation can adapt much faster to changing network conditions. The proposed approach aims to work with current TCP implementations with no required changes to the communication endpoints. Last, we show how the presented approach could be realized by utilizing off-the-shelf programmable switches.
Ralf Kundel, Nehal Baganal Krishna, Christoph Gärtner, Tobias Meuser, Amr Rizk
ICNP1
2021 POSTER: Leveraging PIFO Queues for Scheduling in Time-Sensitive Networks
abstract
Time-Sensitive Networking emerged as a convergent Ethernet-based real-time networking standard for industrial applications. To support real-time, jitter-free isochronous traffic the corresponding TSN mechanism denoted Time Aware Shaper requires special hardware support. In this work, we propose a path to building TSN networks on top of programmable switches. Specifically, we show here how to leverage a data structure amenable to programmable data planes known as Push-in First-out (PIFO) queue to support TSN traffic scheduling for isochronous real-time, as well as, best effort traffic.
Christoph Gärtner, Amr Rizk, Boris Koldehofe, Rhaban Hark, René Guillaume, Ralf Kundel, Ralf Steinmetz
LANMAN6
2021 Monitoring Flows with Per-Application Granularity using Programmable Data Planes
abstract
The accurate and timely knowledge of a network's internal state is essential for various network management operations like routing, resource allocation, or even intrusion detection. This especially holds true for highly flexible, programmable networks that quickly react to dynamic conditions. However, current approaches of state monitoring in such networks rely on per-rule counter information. Due to limited rule space, their granularity is strongly limited. This generally yields an aggregated and therefore altered representation of the network state. Utilizing the programmability of today's data planes, we tackle this problem and present a novel approach to increase the measurement granularity up to per-application statistics. For demonstration purposes, we show how our approach greatly improves the estimation of the Flow Size Distribution.
Rhaban Hark, Mohamed Ghanmi, Ralf Kundel, Patrick Lieser, Ralf Steinmetz
LANMAN3
2020 Flexible Content-based Publish/Subscribe over Programmable Data Planes
abstract
Publish/subscribe systems have to react fast on changes in their environment while handling many events with low end-to-end latency and high throughput. Moving the broker functionality of publish/subscribe systems to the underlying network layer reduces the path length of events and, in addition, forwarding benefits from powerful and programmable hardware. So far attempts of underlay publish/subscribe depend on a specific API of the network devices, e. g., the OpenFlow protocol, which have restrictions in dealing with dynamic devices and corresponding changes in the introduced attribute names for matching and filtering events.In this work, we focus on the next generation of network devices, which are envisioned to provide reconfigurable hardware components, specified by the open P4 description language. We introduce two new approaches that enable a flexible and generic attribute/value encoding, understandable by P4-capable packet processors, to benefit from the performance properties of hardware. Furthermore, the proposed approaches reduce the effort in encoding and decoding event messages.
Ralf Kundel, Christoph Gärtner, Manisha Luthra, Sukanya Bhowmik, Boris Koldehofe
NOMS1
2020 Microbursts in Software and Hardware-based Traffic Load Generation
abstract
Many software based traffic load generators suffer from packet rate variation which is known as rate jitter. In this Demo, we show how this varying rate burstiness can affect the device under test even if the generated average data rate seems constant. To this end, we compare a hardware rate shaping, which is implemented using a programmable P4-switch, and a conventional software load generator and show their impact on a software device under test. The results show, that microbursts within the test load significantly impact the experiment results. Our recommendation is to benchmark the traffic load generator before conducting measurement experiments especially when the device under test is sensitive to microbursts.
Ralf Kundel, Amr Rizk, Boris Koldehofe
NOMS1
2020 P4STA: High Performance Packet Timestamping with Programmable Packet Processors
abstract
QoS requirements of current network control and management applications require the ability to conduct precise measurements of network elements, including switches, routers and Virtual Network Functions (VNFs). State-of-the-art network switches have a forwarding delay of 1µs and below and offer high bandwidths of hundreds Gigabits per second. This imposes high time accuracy and loss-detection requirements on measurement equipment that are not met by existing, software-based measurement tools. The use of specialized tools, meeting these requirements, is restricted by limited flexibility and high cost.In this work, we introduce P4STA, an open source frame-work that combines the flexibility of software-based traffic load generation with the accuracy of hardware packet timestamping. Our evaluation results, obtained using an off-the-shelf P4-programmable switch, show that a time resolution up to 1ns can be achieved on these programmable data plane platforms. Moreover we show how to combine the traffic load of multiple software-based load generators to achieve a measurement load of up to 100Gbit/s per port. Experiments on further programmable platforms, specifically on P4-SmartNICs and FPGAs, show similar results. With this work, we make P4STA available for the research community to advance high performance experiment measurements at nanosecond accuracy.
Ralf Kundel, Fridolin Siegmund, Jeremias Blendin, Amr Rizk, Boris Koldehofe
NOMS1
2019 How to measure the speed of light with programmable data plane hardware?
abstract
Driven by real-time applications such as IIoT, TSN and vehicular networks, the optimization of networks and its elements regarding latency and throughput becomes more and more important. With this demo we show how latencies of network components can be identified within nanosecond accuracy by use of commodity P4 hardware. We show a measured propagation speed of$5ns/m$in fiber optical cables. Besides that, our approach scales up to$100Gbit/s$link speed by the aggregation of many low-cost load generators to a flexible software-based load generation.
Ralf Kundel, Fridolin Siegmund, Boris Koldehofe
ANCS1
2019 INetCEP: In-Network Complex Event Processing for Information-Centric Networking
abstract
Emerging network architectures like Information-Centric Networking (ICN)offer simplicity in the data plane by addressing named data. Such flexibility opens up the possibility to move data processing inside network elements for high-performance computation, known as in-network processing. However, existing ICN architectures are limited in terms of (i)in-network processing and (ii)data plane programming abstractions. Such architectures can benefit from Complex Event Processing (CEP), an in-network processing paradigm to efficiently process data inside the data plane. Yet, it is extremely challenging to integrate CEP because the current communication model of ICN is limited to consumer-initiated interaction that comes with significant overhead in number of requests to process continuous data streams. In contrast, a change to producer-initiated interaction, as favored by CEP, imposes severe limitations for request-reply interactions. In this paper, we propose an in-network CEP architecture, INETCEP that supports unified interaction patterns (consumer- and producer-initiated). In addition, we provide a CEP query language and facilitate CEP operations while increasing the range of applications that can be supported by ICN. We provide an open source implementation and evaluation of INETCEP over an ICN architecture, Named Function Networking, and two applications: energy forecasting in smart homes and a disaster scenario.
Manisha Luthra, Boris Koldehofe, Jonas Höchst, Patrick Lampe, Ali Haider Rizvi, Ralf Kundel, Bernd Freisleben
ANCS6
2019 P4-BNG: Central Office Network Functions on Programmable Packet Pipelines
abstract
Large-scale telecommunications providers have to continuously challenge and evolve their network infrastructure to efficiently serve growing markets demands. They must increase performance, lower time-to-market, provide new services, and lower the cost of the infrastructure and its operation. Network Functions Virtualization (NFV) on commodity hardware offers an attractive, low-cost platform to establish innovations much faster than with purpose-built hardware products. Unfortunately, implementing NFV on commodity processors does not match the performance requirements of the high-throughput data plane components in large carrier access networks. In this article, we propose a way to offer residential network access with programmable packet processing architectures. Based on the highly flexible P4 programming language, we present a design and open source implementation of a BNG data plane that meets the challenging demands of Broadband Network Gateways in carrier-grade environments. The proposed evaluation results show the desired performance characteristics and our proposed design together with upcoming P4 hardware can offer a giant leap towards highest performance NFV network access.
Ralf Kundel, Leonhard Nobach, Jeremias Blendin, Hans-Jörg Kolbe, Georg Schyguda, Vladimir Gurevich, Boris Koldehofe, Ralf Steinmetz
CNSM1