EDBT 2026 Demo / reviewers in the wild / expert
Martin Spinler
dblp:222/7730
· DBLP profile ↗
6ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DMA Calypte: Open-Source Ultra-Low Latency DMA Engine for FPGAsabstractAchieving the lowest possible communication latency between an FPGA accelerator card and host CPU software is critical for many applications, such as high-frequency trading, in-memory storage systems, and 5G processing. However, no freely available PCIe DMA engine (DMAE) is currently optimized for low latency. To address this, we present DMA Calypte, an open-source, platform-independent, ultra-low-latency DMAE. We implemented numerous optimizations at the DMAE, software, and device driver levels to minimize communication latency over PCIe. DMA Calypte's functionality was validated on accelerator cards featuring AMD Kintex UltraScale+ and Agilex 7 F-Series chips with a PCIe Gen3 x8 interface. We achieved a round-trip time between software and FPGA of just 790 ns and a throughput of up to 35 Gbps. Vladislav Válek, Martin Spinler, Jakub Cabal, Tomás Martínek |
FPL | 2 |
| 2022 | FPL Demo: 400G FPGA Packet Capture Based on Network Development KitabstractCESNET, the Czech NREN (National Research and Education Network), has a long research history in the area of high-speed network monitoring using FPGA accelerated cards. Now, we are ready to present our open-source Network Development Kit for FPGAs11https://github.com/CESNET/ndk-app-minimal/ which is ready for 400 Gbps data transfers via Ethernet and PCI Express. The demo aims to show the possibilities of NDK, which allows users to quickly and easily develop new network applications for FPGA-based acceleration cards. Even high-speed DMA Module fully supported in NDK is available free of charge for academic purposes. It can thus significantly contribute to the spread of 400G technology in the academic community and also among other users. The accelerator card equipped with the Intel Agilex I-Series FPGA will transmit and receive back 400G Ethernet (400GBASE) traffic via external loopback. The received packets will be forwarded via very fast packet DMA transfers directly to the RAM of the host computer. Jakub Cabal, Jiri Sikora, Stepan Friedl, Martin Spinler, Jan Korenek |
FPL | 4 |
| 2021 | DMA Medusa: A Vendor-Independent FPGA-Based Architecture for 400 Gbps DMA TransfersabstractFPGA accelerator cards are used for packet capture and monitoring in high-speed networks. With the 400G Ethernet technology, there is a need for an ability to transfer data to and from the host memory at the speed of 400Gbps. Currently available architectures (for example [1],[2],[3]) are limited to throughput up to 100Gbps and are therefore not suitable for this use case.This paper presents a vendor-independent DMA architecture that is capable of scaling up to 400Gbps throughput in a single FPGA using two PCIe Gen4 ×16 slots bifurcated into four ×8 interfaces. This architecture is designed to support hundreds of independent DMA channels and supports one or more PCIe endpoints with different configurations. We also demonstrate the performance of the proposed DMA architecture using results measured on an accelerator card with Intel Stratix 10 DX FPGA. Jan Kubálek, Jakub Cabal, Martin Spinler, Radek Isa |
FCCM | 3 |
| 2018 | Accelerated Wire-Speed Packet Capture at 200 GbpsabstractWe present our latest FPGA acceleration card NFB-200G2QL that is specifically designed to enable traffic processing at 200 Gbps. Unique high-speed DMA engines in the FPGA together with highly optimized Linux drivers enable data transfer through PCIe interfaces with minimal CPU overhead. Captured traffic can be independently distributed between individual cores of two physical CPUs (NUMA nodes) without utilization of QPI. As a result, wire-speed packet capture to the host memory from two fully saturated 100 Gbps Ethernet interfaces (QSFP28+ cages) is achieved and various network monitoring applications can utilize the power of the latest FPGAs and CPUs for data processing. This is especially useful when both directions of a single 100GbE link are monitored. The live demonstration shows how the packets are received from two 100 Gbps Ethernet links at wire-speed and captured to the host memory at 200 Gbps without a loss. The opposite direction of communication is also shown, i.e. how the packets are transmitted from the host memory and fully saturate the two 100GbE network interfaces. Achieved speeds are demonstrated by counters and gauges showing generated, received/transmitted and captured packets. We also show statistics of CPU load during the packet capture/transmission for different packet lengths. Lukas Kekely, Martin Spinler, Stepan Friedl, Jiri Sikora, Jan Korenek |
FPL | 2 |
| 2018 | Demonstration of Full-Duplex Packet Transfers Over PCI Express with Sustained 200 Gbps ThroughputabstractCESNET (Czech NREN) and Netcope Technologies have a long research history in the area of high-speed network monitoring using FPGA accelerated cards (i.e. SmartNICs). Now, we are ready to demonstrate a new NFB-200G2QL accelerator specifically designed to push the achievable traffic processing throughput to 200 Gbps in a single card. The card is equipped with two 100 Gbps Ethernet interfaces (QSFP28+ standard), powerful Virtex UltraScale+ FPGA, and two PCIe Gen3 x16 interfaces. Unique high-speed DMA engines in the FPGA together with highly optimized Linux drivers enable to achieve 200 Gbps data transfer throughput through the PCIe interfaces with minimal CPU overhead. Captured network traffic can be independently distributed among individual cores of two physical CPUs (NUMA nodes) without utilization of QPI. As a result, wire-speed packet capture to the host memory from both fully saturated 100 Gbps Ethernet interfaces is achieved and various network monitoring applications can utilize the power of the latest FPGAs and CPUs for data processing. This is especially useful when traffic of both directions of a single 100GbE link needs to be processed. The proposed demonstration shows how packets of arbitrary length can be received from two 100 Gbps Ethernet links at wire-speed and captured to the host memory at sustained 200 Gbps without any loss. The opposite direction of communication is also shown, i.e. how packets can be transmitted from the host memory and fully saturate two 100 Gbps Ethernet network interfaces. The reception and the transmission of data can be even shown operating simultaneously (full-duplex) without any degradation of performance in either direction. Achieved throughputs are demonstrated by counters and graphs showing live statistics of generated, received/transmitted and captured packets. We can also show detailed statistics of CPU load during the transfers of data for different packet lengths. Lukas Kekely, Martin Spinler, Stepan Friedl, Jiri Sikora, Jan Korenek, Viktor Pus |
FPT | 2 |
| 2018 | Live demonstration of FPGA based networking accelerator for 200 Gbps data transfersabstractCESNET (Czech NREN) is ready to demonstrate a new NFB-200G2QL accelerator with Virtex UltraScale+ FPGA specifically designed to push the achievable traffic processing throughput to 200 Gbps in a single card. Unique high-speed DMA engines in the FPGA together with highly optimized Linux drivers enable to achieve 200 Gbps data transfer through two PCIe Gen3 χ 16 interfaces with minimal CPU overhead. Captured network traffic can be independently distributed among individual cores of two physical CPUs (NUMA nodes) without utilization of QPI. As a result, wire-speed packet capture to the host memory from two fully saturated 100 Gbps Ethernet interfaces (QSFP28+) is achieved and various network monitoring applications can utilize the power of the latest FPGAs and CPUs for data processing. This is especially useful when traffic of both directions of a single 100GbE link needs to be processed. The proposed demonstration will show how the packets can be received from two 100 Gbps Ethernet links at full speed and captured to the host memory at 200 Gbps without any loss. The opposite direction of communication will also be shown, i.e. how the packets can be transmitted from the host memory towards the two 100GbE network interfaces. Achieved speeds will be demonstrated by counters and graphs showing generated, received/transmitted and captured packets. We will also show detailed statistics of CPU load during the packet capture/transmission for different packet lengths. Lukas Kekely, Martin Spinler, Stepan Friedl, Jiri Sikora, Jan Korenek |
NOMS | 2 |