EDBT 2026 Demo / reviewers in the wild / expert
Jakub Cabal
dblp:214/9809
· DBLP profile ↗
7ranked-venue papers
4as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DMA Calypte: Open-Source Ultra-Low Latency DMA Engine for FPGAsabstractAchieving the lowest possible communication latency between an FPGA accelerator card and host CPU software is critical for many applications, such as high-frequency trading, in-memory storage systems, and 5G processing. However, no freely available PCIe DMA engine (DMAE) is currently optimized for low latency. To address this, we present DMA Calypte, an open-source, platform-independent, ultra-low-latency DMAE. We implemented numerous optimizations at the DMAE, software, and device driver levels to minimize communication latency over PCIe. DMA Calypte's functionality was validated on accelerator cards featuring AMD Kintex UltraScale+ and Agilex 7 F-Series chips with a PCIe Gen3 x8 interface. We achieved a round-trip time between software and FPGA of just 790 ns and a throughput of up to 35 Gbps. Vladislav Válek, Martin Spinler, Jakub Cabal, Tomás Martínek |
FPL | 3 |
| 2022 | FPL Demo: 400G FPGA Packet Capture Based on Network Development KitabstractCESNET, the Czech NREN (National Research and Education Network), has a long research history in the area of high-speed network monitoring using FPGA accelerated cards. Now, we are ready to present our open-source Network Development Kit for FPGAs11https://github.com/CESNET/ndk-app-minimal/ which is ready for 400 Gbps data transfers via Ethernet and PCI Express. The demo aims to show the possibilities of NDK, which allows users to quickly and easily develop new network applications for FPGA-based acceleration cards. Even high-speed DMA Module fully supported in NDK is available free of charge for academic purposes. It can thus significantly contribute to the spread of 400G technology in the academic community and also among other users. The accelerator card equipped with the Intel Agilex I-Series FPGA will transmit and receive back 400G Ethernet (400GBASE) traffic via external loopback. The received packets will be forwarded via very fast packet DMA transfers directly to the RAM of the host computer. Jakub Cabal, Jiri Sikora, Stepan Friedl, Martin Spinler, Jan Korenek |
FPL | 1 |
| 2021 | DMA Medusa: A Vendor-Independent FPGA-Based Architecture for 400 Gbps DMA TransfersabstractFPGA accelerator cards are used for packet capture and monitoring in high-speed networks. With the 400G Ethernet technology, there is a need for an ability to transfer data to and from the host memory at the speed of 400Gbps. Currently available architectures (for example [1],[2],[3]) are limited to throughput up to 100Gbps and are therefore not suitable for this use case.This paper presents a vendor-independent DMA architecture that is capable of scaling up to 400Gbps throughput in a single FPGA using two PCIe Gen4 ×16 slots bifurcated into four ×8 interfaces. This architecture is designed to support hundreds of independent DMA channels and supports one or more PCIe endpoints with different configurations. We also demonstrate the performance of the proposed DMA architecture using results measured on an accelerator card with Intel Stratix 10 DX FPGA. Jan Kubálek, Jakub Cabal, Martin Spinler, Radek Isa |
FCCM | 2 |
| 2020 | Multi Buses: Theory and Practical Considerations of Data Bus Width Scaling in FPGAsabstractAs the throughput of computer networks and other peripheral interfaces is rising, developers are forced to use ever-wider data buses in FPGA designs. However, utilization of wide buses poses a serious threat of performance degradation, especially for the shortest data transactions (packets), as aliasing and alignment overheads on the bus can be extremely increased. In this paper, we propose a novel design method for the description of very wide data buses that we call Multi Buses. The key idea is to enable the processing of multiple transactions per clock cycle with very high and predictable effective throughput even in the worst-case. The feasibility of the proposed method is shown via analysis of achievable performance by both theoretical means and selected proof of concept implementations. Thanks to the proposed method, we were able to design FPGA cores for key operations in networking (e.g. parser, match table, CRC, deparser) with sufficient throughputs for wire-speed packet processing of 400Gbps, lTbps and even 2 Tbps Ethernet links. Lukas Kekely, Jakub Cabal, Viktor Pus, Jan Korenek |
DSD | 2 |
| 2019 | Scalable P4 Deparser for Speeds Over 100 GbpsabstractThe P4 language is a language suitable for the description of packet processing inside a network device. The typical P4 device consists of three main building blocks: Parser, Match+Action Tables and Deparser. The deparsing is the most challenging block because the main task of this block is to assemble the output packet based on changes in Match+Action Tables. This operation can be quite complicated in the case of high-speed networks. In this work, we present the scalable architecture (in term of the throughput) of a deparsing circuit which is suitable for implementation in FPGAs. Jakub Cabal, Pavel Benácek, Jana Foltova, Juraj Holub |
FCCM | 1 |
| 2018 | Configurable FPGA Packet Parser for Terabit Networks with Guaranteed Wire-Speed ThroughputabstractAs throughput of computer networks is on a constant rise, there is a need for ever-faster packet parsing modules at all points of the networking infrastructure. Parsing is a crucial operation which has an influence on the final throughput of a network device. Moreover, this operation must precede any kind of further traffic processing like filtering/classification, deep packet inspection, and so on. This paper presents a parser architecture which is capable to currently scale up to a terabit throughput in a single FPGA, while the overall processing speed is sustained even on the shortest frame lengths and for an arbitrary number of supported protocols. The architecture of our parser can be also automatically generated from a high-level description of a protocol stack in the P4 language which makes the rapid deployment of new protocols considerably easier. The results presented in the paper confirm that our automatically generated parsers are capable of reaching an effective throughput of over 1 Tbps (or more than 2000 Mpps) on the Xilinx UltraScale+ FPGAs and around 800 Gbps (or more than 1200 Mpps) on their previous generation Virtex-7 FPGAs. Jakub Cabal, Pavel Benácek, Lukas Kekely, Michal Kekely, Viktor Pus, Jan Korenek |
FPGA | 1 |
| 2018 | High-Speed Computation of CRC Codes for FPGAsabstractAs the throughput of networks and memory interfaces is on a constant rise, there is a need for ever-faster error-detecting codes. Cyclic redundancy checks (CRC) are a common and widely used to ensure consistency or detect accidental changes of data. We propose a novel FPGA architecture for the computation of the CRC designed for general high-speed data transfers. Its key feature is allowing a processing of multiple independent data packets (transactions) in each clock cycle, what is a necessity for achieving high overall throughput on very wide data buses. Experimental results confirm that the proposed architecture reaches an effective throughput sufficient for utilization in multi-terabit Ethernet networks (over 2 Tbps or over 3000 Mpps) on a single Xilinx UltraScale+ FPGA. Jakub Cabal, Lukas Kekely, Jan Korenek |
FPT | 1 |