Jan Korenek

dblp:19/640 · DBLP profile ↗
← Back
68ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-4662-7349ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 48 · 2 first-author · 6 since 2021Computer networks · 16 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 FRAPPE: Feasibility Report on Accelerating Payload Pattern-Matching Engines in Intrusion Detection Systems with FPGAs
abstract
Modern Intrusion Detection Systems (IDS) struggle to scale to$100+$Gbps throughput, as typically, the MultiPattern Matching (MPM) stage overwhelms CPU resources. While FPGA-based acceleration offers a theoretical solution, its adoption in production environments remains negligible despite decades of research. This disconnect stems not from a lack of raw hardware performance, but from unmatched IDS requirements and seemingly invasive integration. In this paper, we propose a standardized, stateless offload architecture based on the DPDK rte_flow API that decouples hardware acceleration from complex IDS logic. By using the FPGA as a smart tagger that annotates packets with matched pattern IDs via a compact metadata interface, we enable inline, scalable, integration with existing IDS pipelines like Suricata. We validate this approach through a trace-driven co-design study, demonstrating that a small metadata budget of three pattern IDs per packet is sufficient to offload the vast majority of traffic, resulting in up to 57 % throughput increase. Finally, we survey state-of-the-art 100+ Gbps FPGA engines against our derived integration criteria to highlight the critical features that future designs must implement to enable practical deployment.
Lukás Sismis, Jan Korenek
DDECS2
2024 LGBM2VHDL: Mapping of LightGBM Models to FPGA
abstract
Gradient boosting (GB) is an effective and widely used type of ensemble machine-learning method. The opportunity to transform the trained GB models to the hardware level represents the potential for significant acceleration of many applications and their availability as embedded systems. In this work, we have therefore developed the LGBM2VHDL tool for the automated mapping of models trained by the LightGBM library to circuits described by VHDL. Compared to existing tools, we have used an architecture that is better suited for large-scale GB models involving up to thousands of decision trees. We have further optimized the architecture using two newly proposed techniques. By applying these techniques to the tested models, the amount of memory required was significantly reduced to almost half of the original resources, and the amount of basic configurable blocks was reduced by up to 4 times on average. The developed tool is available as open-source.
Tomás Martínek, Jan Korenek, Tomás Cejka
FCCM2
2023 Optimizing Packet Classification on FPGA
abstract
Packet classification is a crucial time-critical operation for many different networking tasks ranging from switching or routing to monitoring and security devices like firewalls or IDS. Accelerated architectures implementing packet classification must satisfy the ever-growing demand for current high-speed networks. However, packet classification is generally used together with other packet processing algorithms, which decreases the available hardware resources on the FPGA chip. The introduction of the P4 language requires the packet classification to be even more flexible while maintaining a high throughput with limited resources. Thus, we need flexible and high-performance architectures to balance processing speed and hardware resources for specific types of rules. DCFL algorithm provides high performance and flexibility. Therefore, we propose optimizations to the DCFL algorithm and overall packet processing hardware architecture. The goal is to maximize the throughput and minimize the resource strain. The main idea of the approach is to analyze the ruleset, identify some conflicting rules and offload these rules to other hardware modules. This approach allows us to process packets faster, even in the worst-case scenarios. Moreover, we can fit more packet processing into the FPGA and fine-tune the packet processing architecture to meet a specific network application’s throughput and resource demands. With the proposed optimizations we can achieve up to a 76 % increase in the throughput of the packet classification. Alternatively, we can achieve up to a 37 % decrease in resources needed.
Michal Kekely, Jan Korenek
DDECS2
2023 Accelerating IDS Using TLS Pre-Filter in FPGA
abstract
Intrusion Detection Systems (IDSes) are a widely used network security tool. However, achieving sufficient throughput is challenging as network link speeds increase to 100 or 400 Gbps. Despite the large number of papers focusing on the hardware acceleration of IDSes, the approaches are mostly limited to the acceleration of pattern matching or do not support all types of IDS rules. Therefore, we propose hardware acceleration that significantly increases the throughput of IDSes without limiting the functionality or the types of rules supported. As the IDSes cannot match signatures in encrypted network traffic, we propose a hardware TLS pre-filter that removes encrypted TLS traffic from IDS processing and doubles the average processing speed. Implemented on an acceleration card with an Intel Agilex FPGA, the pre-filter supports 100 and 400 Gbps throughput. The hardware design is optimized to achieve a high frequency and to utilize only a few hardware resources.
Vlastimil Kosar, Lukás Sismis, Jirí Matousek 0002, Jan Korenek
ISCC4
2023 Analysis of TLS Prefiltering for IDS Acceleration
Lukás Sismis, Jan Korenek
PAM2
2022 FPL Demo: 400G FPGA Packet Capture Based on Network Development Kit
abstract
CESNET, the Czech NREN (National Research and Education Network), has a long research history in the area of high-speed network monitoring using FPGA accelerated cards. Now, we are ready to present our open-source Network Development Kit for FPGAs11https://github.com/CESNET/ndk-app-minimal/ which is ready for 400 Gbps data transfers via Ethernet and PCI Express. The demo aims to show the possibilities of NDK, which allows users to quickly and easily develop new network applications for FPGA-based acceleration cards. Even high-speed DMA Module fully supported in NDK is available free of charge for academic purposes. It can thus significantly contribute to the spread of 400G technology in the academic community and also among other users. The accelerator card equipped with the Intel Agilex I-Series FPGA will transmit and receive back 400G Ethernet (400GBASE) traffic via external loopback. The received packets will be forwarded via very fast packet DMA transfers directly to the RAM of the host computer.
Jakub Cabal, Jiri Sikora, Stepan Friedl, Martin Spinler, Jan Korenek
FPL5
2022 ClassBench-ng: Benchmarking Packet Classification Algorithms in the OpenFlow Era
abstract
Packet classification, i.e., the process of categorizing packets into flows, is a first-class citizen in any networking device. Every time a new packet has to be processed, one or more header fields need to be compared against a set of pre-installed rules. This is done for basic forwarding operations, to apply security policies, application-specific processing, or quality-of-service guarantees. A lot of research efforts have identified better lookup techniques, i.e., finding the best match between packet headers and rules, by capitalizing on the rule sets characteristics. Here, ClassBench has greatly served the community by enabling the generation of IPv4 rule sets. In this paper, we present a new tool, ClassBench-ng, that creates synthetic IPv4, IPv6, and OpenFlow rules. We start from an analysis of classification rules deployed in-the-wild and we use the findings to craft our solution. ClassBench-ng can generate a user-defined number of rules as well as an associated header trace matching them. Compared to state-of-the-art solutions, the rule set generation process is usually more accurate and it is able to produce rules matching a number of different use cases, i.e., from an IPv4 router to an OpenFlow switch, which is unique among current rule set generation tools.
Jirí Matousek 0002, Adam Lucanský, David Janecek, Jozef Sabo, Jan Korenek, Gianni Antichi
IEEE/ACM Trans. Netw.5
2021 Efficient Acceleration of Decision Tree Algorithms for Encrypted Network Traffic Analysis
abstract
Network traffic analysis and deep packet inspection are time-consuming tasks, which current processors can not handle at 100 Gbps speed. Therefore security systems need fast packet processing with hardware acceleration. With the growing of encrypted network traffic, it is necessary to extend Intrusion Detection Systems (IDSes) and other security tools by new detection methods. Security tools started to use classifiers trained by machine learning techniques based on decision trees. Random Forest, Compact Random Forest and AdaBoost provide excellent result in network traffic analysis. Unfortunately, hardware architectures for these machine learning techniques need high utilisation of on-chip memory and logic resources. Therefore we propose several optimisations of highly pipelined architecture for acceleration of machine learning techniques based on decision trees. The optimisations use the various encoding of a feature vector to reduce hardware resources. Due to the proposed optimisations, it was possible to reduce LUTs by 70.5 % for HTTP brute force attack detection and BRAMs by 50 % for application protocol identification. Both with only negligible impact on classifiers' accuracy. Moreover, proposed optimisations reduce wires and multiplexors in the processing pipeline, positively affecting the proposed architecture's maximal achievable frequency.
Roman Vrana, Jan Korenek
DDECS2
2021 Increasing Memory Efficiency of Hash-Based Pattern Matching for High-Speed Networks
abstract
Increasing speed of network links continuously pushes up requirements on the performance of network security and monitoring systems, including their typical representative and its core function: an intrusion detection system (IDS) and pattern matching. To allow the operation of IDS applications like Snort and Suricata in networks supporting throughput of 100Gbps or even more, a recently proposed pre-filtering architecture approximates exact pattern matching using hash-based matching of short strings that represent a given set of patterns. This architecture can scale supported throughput by adjusting the number of parallel hash functions and on-chip memory blocks utilized in the implementation of a hash table. Since each hash function can address every memory block, scaling throughput also increases the total capacity of the hash table. Nevertheless, the original architecture utilizes the available capacity of the hash table inefficiently. We therefore propose three optimization techniques that either reduce the amount of information stored in the hash table or increase its achievable occupancy. Moreover, we also design modifications of the architecture that enable resource-efficient utilization of all three optimization techniques together in synergy. Compared to the original pre-filtering architecture, combined use of the proposed optimizations in the 100Gbps scenario increases the achievable capacity for short strings by three orders of magnitude. It also reduces the utilization of FPGA logic resources to only a third.
Tomas Fukac, Jirí Matousek 0002, Jan Korenek, Lukas Kekely
FPT3
2021 Scalability of Hash-Based Pattern Matching for High-Speed Network Security and Monitoring
abstract
Gradually increasing throughput of high-speed networks puts continuous pressure on the performance of operations over a stream of network data. Probably the most affected operation in the area of network security and monitoring is pattern matching, which is in the core of widely deployed intrusion detection systems (IDSes) like Snort, Suricata and Bro. This paper therefore proposes several optimizations of a hash-based pattern matching architecture that together allow to increase its throughput to 100 Gbps and beyond. Proposed optimizations target an interconnection network between parallel hash function engines and independent memory blocks addressed by hashes computed over short strings of input data. Specifically, the optimizations reduce resource utilization by sharing parts of the full interconnection network among its several outputs and lower collision rate in these shared parts by aggregation and distributed buffering of memory access requests. The optimized pattern matching architecture is therefore able to utilize a higher number of parallel hash functions, each of which can use the interconnection network to access any memory block. This allows not only to increase the throughput of a key component within IDSes to more than 100 Gbps, but also to support a larger set of network threat patterns and to update this set dynamically.
Tomas Fukac, Jan Korenek, Jirí Matousek 0002
ISCC2
2020 Multi Buses: Theory and Practical Considerations of Data Bus Width Scaling in FPGAs
abstract
As the throughput of computer networks and other peripheral interfaces is rising, developers are forced to use ever-wider data buses in FPGA designs. However, utilization of wide buses poses a serious threat of performance degradation, especially for the shortest data transactions (packets), as aliasing and alignment overheads on the bus can be extremely increased. In this paper, we propose a novel design method for the description of very wide data buses that we call Multi Buses. The key idea is to enable the processing of multiple transactions per clock cycle with very high and predictable effective throughput even in the worst-case. The feasibility of the proposed method is shown via analysis of achievable performance by both theoretical means and selected proof of concept implementations. Thanks to the proposed method, we were able to design FPGA cores for key operations in networking (e.g. parser, match table, CRC, deparser) with sufficient throughputs for wire-speed packet processing of 400Gbps, lTbps and even 2 Tbps Ethernet links.
Lukas Kekely, Jakub Cabal, Viktor Pus, Jan Korenek
DSD4
2020 Increasing Throughput of Intrusion Detection Systems by Hash-Based Short String Pre-filter
abstract
With an increasing speed of network links, it is also necessary to increase the throughput of network security systems. An intrusion detection system (IDS) is one of the key components in the protection of network infrastructure. Unfortunately, the IDS has to match a large set of regular expressions (REs) in network streams, which has a negative impact on its throughput. A fast pre-filtration of network traffic can allow to achieve a higher overall throughput. Therefore, we have designed a new algorithm, which is able to select short strings that represent an RE set utilized in the IDS. Compared to previous methods, strings are selected in less than a second for an RE and can reduce network traffic up to 3.3 times better. As all selected strings have the same length, they can be used in a hash-based pre-filter, which is able to process more 100 Gbps of network traffic.
Tomas Fukac, Vlastimil Kosar, Jan Korenek, Jirí Matousek 0002
LCN3
2019 Hash-based Pattern Matching for High Speed Networks
abstract
Regular expression matching is a complex task which is widely used in network security monitoring applications. With the growing speed of network links and the number of regular expressions, pattern matching architectures have to be improved to retain wire-speed processing. Multi-striding is a well-known technique to increase processing speed but it requires a lot of FPGA resources. Therefore, we focus on the design of new hardware architecture for fast pre-filtering of network traffic. The proposed pre-filter performs fast hash-based matching of short strings, which are specific for matched regular expressions. As the proposed pre-filter significantly reduces input traffic, exact pattern matching can operate on significantly lower speeds. Then the exact pattern match can be done by CPU or by a slow automaton with a few hardware resources. The paper provides analyses of false-positive detection of the pre-filter with respect to the length of matching strings. The number of false-positives is low, even if the length of the selected strings is short. Therefore input traffic can be significantly reduced. For 100 Gb links, the pre-filter reduced the input data to 1.83 Gbps using four-symbol strings.
Tomas Fukac, Jan Korenek
DDECS2
2019 Acceleration of Feature Extraction for Real-Time Analysis of Encrypted Network Traffic
abstract
With the growing amount of encrypted network traffic, it is important to have tools for the analysis and classification of encrypted network data. Encrypted network traffic is usually analysed by statistical methods because Deep Packet Inspection or pattern matching is not applicable. However, the statistical methods are usually designed to work offline on already captured network traffic. For real-time analysis, hardware acceleration is needed to achieve wire-speed 10 Gbps throughput. Therefore, we focus on real-time monitoring of encrypted network traffic and propose a new acceleration method to extract features from encrypted network data. Approximate computing is used to speed up the computation of entropy for the input data stream and to reduce FPGA logic utilization. As can be seen in the results, the precision of classification has decreased only by 0.1 to 0.2. Moreover, proposed hardware architecture has very low FPGA logic utilization and can operate on high frequency.
Roman Vrana, Jan Korenek, David Novak
DDECS2
2019 Deep Packet Inspection in FPGAs via Approximate Nondeterministic Automata
abstract
Deep packet inspection via regular expression (RE) matching is a crucial task of network intrusion detection systems (IDSes), which secure Internet connection against attacks and suspicious network traffic. Monitoring high-speed computer networks (100 Gbps and faster) in a single-box solution demands that the RE matching, traditionally based on finite automata (FAs), is accelerated in hardware. In this paper, we describe a novel FPGA architecture for RE matching that is able to process network traffic beyond 100 Gbps. The key idea is to reduce the required FPGA resources by leveraging approximate nondeterministic FAs (NFAs). The NFAs are compiled into a multi-stage architecture starting with the least precise stage with a high throughput and ending with the most precise stage with a low throughput. To obtain the reduced NFAs, we propose new approximate reduction techniques that take into account the profile of the network traffic. Our experiments showed that using our approach, we were able to perform matching of large sets of REs from SNORT, a popular IDS, on unprecedented network speeds.
Milan Ceska 0002, Vojtech Havlena, Lukás Holík, Jan Korenek, Ondrej Lengál, Denis Matousek, Jirí Matousek 0002, Jakub Semric, Tomás Vojnar
FCCM4
2018 Regular expression matching with pipelined delayed input DFAs for high-speed networks
abstract
Regular expression matching (RE matching) is a widely used operation in network security monitoring applications. With the speed of network links increasing to 100 Gbps and 400 Gbps, it is necessary to speed up packet processing and provide RE matching at such high speeds. Although many RE matching algorithms and architectures have been designed, none of them supports 100 Gbps throughput together with fast updates of an RE set. Therefore, this paper focuses on the design of a new hardware architecture that addresses both these requirements. The proposed architecture uses multiple highly memory-efficient Delayed Input DFAs (D2FAs), which are organized to a processing pipeline. As all D2FAs in the pipeline have only local communication, the proposed architecture is able to operate at high frequency even for a large number of parallel engines, which allows scaling throughput to hundreds of gigabits per second. The paper also analyses how to scale the number of engines and the capacity of buffers to achieve desired throughput. Using the parameters obtained while matching a sample RE set represented by a D2FA in a real network traffic, the architecture can be tuned for wire-speed throughput of 400 Gbps.
Denis Matousek, Juraj Kubis, Jirí Matousek 0002, Jan Korenek
ANCS4
2018 Memory Aware Packet Matching Architecture for High-Speed Networks
abstract
Packet classification is a crucial operation for many different networking tasks ranging from switching or routing to monitoring and security devices like firewall or IDS. Generally, accelerated architectures implementing packet classification must be used to satisfy ever-growing demands of current high-speed networks. Furthermore, to keep up with the rising network throughputs, the accelerated architectures for FPGAs must be able to classify more than one packet in each clock cycle. This can be mainly achieved by utilization of multiple processing pipelines in parallel, what brings replication of FPGA logic and more importantly scarce on-chip memory resources. Therefore in this paper, we propose a novel parallel hardware architecture for hash-based exact match classification of multiple packets per clock cycle with reduced memory replication requirements. The basic idea is to leverage the fact that modern FPGAs offer hundreds of BlockRAM tiles that can be accessed (addressed) independently to maintain high throughput of matching even without fully replicated memory architecture. Our results show that the proposed approach can use memory very efficiently and scales exceptionally well with increased record capacities. For example, the designed architecture is able to achieve throughput of more than 2 Tbps (over 3 000 Mpps) with an effective capacity of more than 40 000 IPv4 flow records for the cost of only 366 BlockRAM tiles and around 57 000 LUTs.
Michal Kekely, Lukas Kekely, Jan Korenek
DSD3
2018 High-Speed Regular Expression Matching with Pipelined Memory-Based Automata
abstract
The paper proposes an architecture of a high-speed regular expression (RE) matching system with fast updates of an RE set. The architecture uses highly memory-efficient Delayed Input DFAs (D2FAs), which are organized to a processing pipeline. The architecture is designed so that it communicates only locally among its components in order to achieve high frequency even for a large number of parallel matching engines (MEs), which allows scaling throughput to hundreds of gigabits per second (Gbps). The architecture is able to achieve processing throughput of up to 400 Gbps on current FPGA chips.
Denis Matousek, Jirí Matousek 0002, Jan Korenek
FCCM3
2018 Configurable FPGA Packet Parser for Terabit Networks with Guaranteed Wire-Speed Throughput
abstract
As throughput of computer networks is on a constant rise, there is a need for ever-faster packet parsing modules at all points of the networking infrastructure. Parsing is a crucial operation which has an influence on the final throughput of a network device. Moreover, this operation must precede any kind of further traffic processing like filtering/classification, deep packet inspection, and so on. This paper presents a parser architecture which is capable to currently scale up to a terabit throughput in a single FPGA, while the overall processing speed is sustained even on the shortest frame lengths and for an arbitrary number of supported protocols. The architecture of our parser can be also automatically generated from a high-level description of a protocol stack in the P4 language which makes the rapid deployment of new protocols considerably easier. The results presented in the paper confirm that our automatically generated parsers are capable of reaching an effective throughput of over 1 Tbps (or more than 2000 Mpps) on the Xilinx UltraScale+ FPGAs and around 800 Gbps (or more than 1200 Mpps) on their previous generation Virtex-7 FPGAs.
Jakub Cabal, Pavel Benácek, Lukas Kekely, Michal Kekely, Viktor Pus, Jan Korenek
FPGA6
2018 Accelerated Wire-Speed Packet Capture at 200 Gbps
abstract
We present our latest FPGA acceleration card NFB-200G2QL that is specifically designed to enable traffic processing at 200 Gbps. Unique high-speed DMA engines in the FPGA together with highly optimized Linux drivers enable data transfer through PCIe interfaces with minimal CPU overhead. Captured traffic can be independently distributed between individual cores of two physical CPUs (NUMA nodes) without utilization of QPI. As a result, wire-speed packet capture to the host memory from two fully saturated 100 Gbps Ethernet interfaces (QSFP28+ cages) is achieved and various network monitoring applications can utilize the power of the latest FPGAs and CPUs for data processing. This is especially useful when both directions of a single 100GbE link are monitored. The live demonstration shows how the packets are received from two 100 Gbps Ethernet links at wire-speed and captured to the host memory at 200 Gbps without a loss. The opposite direction of communication is also shown, i.e. how the packets are transmitted from the host memory and fully saturate the two 100GbE network interfaces. Achieved speeds are demonstrated by counters and gauges showing generated, received/transmitted and captured packets. We also show statistics of CPU load during the packet capture/transmission for different packet lengths.
Lukas Kekely, Martin Spinler, Stepan Friedl, Jiri Sikora, Jan Korenek
FPL5
2018 High-Speed Computation of CRC Codes for FPGAs
abstract
As the throughput of networks and memory interfaces is on a constant rise, there is a need for ever-faster error-detecting codes. Cyclic redundancy checks (CRC) are a common and widely used to ensure consistency or detect accidental changes of data. We propose a novel FPGA architecture for the computation of the CRC designed for general high-speed data transfers. Its key feature is allowing a processing of multiple independent data packets (transactions) in each clock cycle, what is a necessity for achieving high overall throughput on very wide data buses. Experimental results confirm that the proposed architecture reaches an effective throughput sufficient for utilization in multi-terabit Ethernet networks (over 2 Tbps or over 3000 Mpps) on a single Xilinx UltraScale+ FPGA.
Jakub Cabal, Lukas Kekely, Jan Korenek
FPT3
2018 Demonstration of Full-Duplex Packet Transfers Over PCI Express with Sustained 200 Gbps Throughput
abstract
CESNET (Czech NREN) and Netcope Technologies have a long research history in the area of high-speed network monitoring using FPGA accelerated cards (i.e. SmartNICs). Now, we are ready to demonstrate a new NFB-200G2QL accelerator specifically designed to push the achievable traffic processing throughput to 200 Gbps in a single card. The card is equipped with two 100 Gbps Ethernet interfaces (QSFP28+ standard), powerful Virtex UltraScale+ FPGA, and two PCIe Gen3 x16 interfaces. Unique high-speed DMA engines in the FPGA together with highly optimized Linux drivers enable to achieve 200 Gbps data transfer throughput through the PCIe interfaces with minimal CPU overhead. Captured network traffic can be independently distributed among individual cores of two physical CPUs (NUMA nodes) without utilization of QPI. As a result, wire-speed packet capture to the host memory from both fully saturated 100 Gbps Ethernet interfaces is achieved and various network monitoring applications can utilize the power of the latest FPGAs and CPUs for data processing. This is especially useful when traffic of both directions of a single 100GbE link needs to be processed. The proposed demonstration shows how packets of arbitrary length can be received from two 100 Gbps Ethernet links at wire-speed and captured to the host memory at sustained 200 Gbps without any loss. The opposite direction of communication is also shown, i.e. how packets can be transmitted from the host memory and fully saturate two 100 Gbps Ethernet network interfaces. The reception and the transmission of data can be even shown operating simultaneously (full-duplex) without any degradation of performance in either direction. Achieved throughputs are demonstrated by counters and graphs showing live statistics of generated, received/transmitted and captured packets. We can also show detailed statistics of CPU load during the transfers of data for different packet lengths.
Lukas Kekely, Martin Spinler, Stepan Friedl, Jiri Sikora, Jan Korenek, Viktor Pus
FPT5
2018 General IDS Acceleration for High-Speed Networks
abstract
Network Intrusion Detection Systems have gained popularity as one of the key technologies to secure communication infrastructures. However, their high computational complexity poses performance challenges for practical deployment in modern high-speed networks. To achieve the highest quality of detection, IDS should process as much relevant data as it can without becoming the bottleneck of a network connection. At the same time, IDS implementation should be flexible enough to accommodate detection methods of ever emerging new security threats. This paper aims at an acceleration of IDS by means of informed packet discarding, effectively focusing the available resources of overloaded IDS to the most relevant parts of analyzed traffic. Unlike previous works, the proposed scheme does not move the IDS nor any specific portion of it into the hardware accelerator. Rather it uses smart software based or hardware accelerated offload (bypass) of the traffic parts that are not likely to represent a security threat. The flexible nature of software-based IDS is therefore fully maintained, while the quality of threat detection remains sufficiently high even when processing high-speed traffic. We show that controlled (informed) discarding of well-defined portions of input traffic yields better detection rates, compared to the default uncontrolled (blind) buffer overflow discarding in high throughput scenarios. Our results show that it is entirely possible to run an IDS on a high-speed network link using single CPU with an FPGA accelerated packet pre-filtering.
Jan Kucera 0004, Lukas Kekely, Adam Piecek, Jan Korenek
ICCD4
2018 Live demonstration of FPGA based networking accelerator for 200 Gbps data transfers
abstract
CESNET (Czech NREN) is ready to demonstrate a new NFB-200G2QL accelerator with Virtex UltraScale+ FPGA specifically designed to push the achievable traffic processing throughput to 200 Gbps in a single card. Unique high-speed DMA engines in the FPGA together with highly optimized Linux drivers enable to achieve 200 Gbps data transfer through two PCIe Gen3 χ 16 interfaces with minimal CPU overhead. Cap­tured network traffic can be independently distributed among individual cores of two physical CPUs (NUMA nodes) without utilization of QPI. As a result, wire-speed packet capture to the host memory from two fully saturated 100 Gbps Ethernet interfaces (QSFP28+) is achieved and various network monitoring applications can utilize the power of the latest FPGAs and CPUs for data processing. This is especially useful when traffic of both directions of a single 100GbE link needs to be processed. The proposed demonstration will show how the packets can be received from two 100 Gbps Ethernet links at full speed and captured to the host memory at 200 Gbps without any loss. The opposite direction of communication will also be shown, i.e. how the packets can be transmitted from the host memory towards the two 100GbE network interfaces. Achieved speeds will be demonstrated by counters and graphs showing generated, received/transmitted and captured packets. We will also show detailed statistics of CPU load during the packet capture/transmission for different packet lengths.
Lukas Kekely, Martin Spinler, Stepan Friedl, Jiri Sikora, Jan Korenek
NOMS5
2017 ClassBench-ng: Recasting ClassBench after a Decade of Network Evolution
abstract
Internet evolution is driven by a continuous stream of new applications and users driving the demand for services. To keep up with this, a never-stopping research has been transforming the Internet ecosystem over the time. Technological changes on both protocols (the uptake of IPv6) and network architectures (the adoption of Software Defined Networking) introduced new challenges for ASIC designers. In particular, IPv6 and OpenFlow increased the complexity of the rule matching problem, pushing researchers to build new packet classiffication algorithms capable to keep pace with a steady growth of link speed. A lot of research effort identifies better lookup techniques capitalizing on the characteristics of rule sets. So far, the availability of small numbers of real rule sets and synthetic ones, generated with tools such as ClassBench, has boosted research in the IPv4 world. Starting from an analysis of rule sets taken from operational environments, we present ClassBench-ng, a new open source tool for the generation of synthetic IPv4, IPv6, and OpenFlow 1.0 rule sets exposing the same properties of real ones. We feel this tool can meet the requirements of nowadays researchers, boosting the rule matching research as ClassBench has done since ten years ago.
Jirí Matousek 0002, Gianni Antichi, Adam Lucanský, Andrew W. Moore 0002, Jan Korenek
ANCS5
2017 Packet Classification with Limited Memory Resources
abstract
Network security and monitoring devices use packet classification to match packet header fields in a set of rules. Many hardware architectures have been designed to accelerate packet classification and achieve wire-speed throughput for 100 Gbps networks. The architectures are designed for high throughput even for the shortest packets. However, FPGA SoC and Intel Xeon with FPGA have limited resources for multiple accelerators. Usually, it is necessary to balance between available resources and the level of acceleration. Therefore, we have designed new hardware architecture for packet classification, which can balance between the processing speed and hardware resources. To achieve 10 Gbps average throughput the architecture need only 20 BlockRAMs for 5500 rules. Moreover, the architecture can scale the processing speed to wire-speed throughput on 100 Gbps line at the cost of additional memory resources.
Michal Kekely, Jan Korenek
DSD2
2017 Line rate programmable packet processing in 100Gb networks
abstract
The P4 language provides a way to describe a custom network packet processing behavior that involves header parsing, matching and assembling modified packets. Such abstraction represents a significant step towards removing the limitation of fixed-function networking devices. Our live demonstration shows a straightforward usage of an algorithm and tool that maps a P4 program to a general architecture of FPGA-based networking device. Network traffic is received, parsed, filtered and modified by the generated circuit at the full line rate of 100 Gbps Ethernet. The results of our ongoing joint research project NFV200 show that the FPGA technology can be used to improve network flexibility without the usual burden of tedious and error-prone HDL coding.
Pavel Benácek, Viktor Pus, Jan Korenek, Michal Kekely
FPL3
2017 Mapping of P4 match action tables to FPGA
abstract
Current networks are changing very fast. Network administrators need more flexible and powerful tools to be able to support new protocols or services very fast. The P4 language provides new level of abstraction for flexible packet processing. Therefore, we have designed new architecture for memory efficient mapping of P4 match/action tables to FPGA. The architecture is based on DCFL algorithm and is able to balance the processing speed and available memory resources.
Michal Kekely, Jan Korenek
FPL2
2016 Dynamically Reconfigurable Architecture with Atomic Configuration Updates for Flexible Regular Expressions Matching in FPGA
abstract
Regular expressions matching is commonly used in network security devices in order to detect malicious network traffic. New network attacks and other threats are emerging frequently. Therefore, the security device must be able to update the set of used regular expressions as soon as possible. The update operation must not disrupt normal operations of the security device. Therefore, the update must be done atomically. Current reconfigurable architectures are not suitable for highly integrated embedded network security devices because they require either additional external memory, ASICs or partial reconfiguration of the FPGA. Also, architectures based on deterministic finite automaton have an exponential time complexity even for real-word sets of regular expressions. Therefore, in this paper, we introduce a reconfigurable architecture with atomic updates suitable for real-world sets of regular expressions. Inspired by previous designs for both ASICs and FPGAs, we propose regular expressions matching architecture with significantly lower consumption of FPGA resources than previous dynamically reconfigurable FPGA design. The proposed architecture uses an interconnection matrix with a linear space complexity, while the previous one uses an interconnection matrix with a quadratic space complexity. The proposed architecture consumes from 6.9 to 48.9 times less LUTs than previous dynamically reconfigurable FPGA design. Single matched symbol utilizes between 4.35 and 32.2 LUTs.
Vlastimil Kosar, Jan Korenek
DSD2
2016 Packet processing on FPGA SoC with DPDK
abstract
One of the most important topics of today is a packet processing in data centers with respect to the power consumption and efficient utilization of computational resources. The ARM architecture has proved to be an energy efficient computational system. Together with an integrated FPGA on a single die, it offers potentially a high performance with respect to the power consumption. DPDK - a set of libraries and drivers intended primarily for fast packet processing - is becoming to be a standard approach for packet processing, especially in data centers. In this paper, we exploit the potential of packet processing based on DPDK and FPGA SoC architectures. Especially, we aim at the potential of utilizing the ARM Cortex-A9 and Cortex-A53 CPUs.
Jan Viktorin, Jan Korenek
FPL2
2016 High-speed regular expression matching with pipelined automata
abstract
Pattern matching is a complex task which is widely used in network security monitoring applications. With the growing speed of network links, pattern matching architectures have to be improved in order to retain wire-speed processing. Multi-striding is a well-known technique on how to increase throughput of pattern matching architectures. In the paper we provide an analysis of scalability of multi-striding and show that it does not scale well and cannot be used for 100Gbps throughput because utilization of FPGA resources grows exponentially. Therefore, we have designed a new hardware architecture for high-speed pattern matching that combines the multi-striding technique and parallel processing using pipelined finite state machines (FSMs). The architecture shares a single packet buffer for all parallel FSMs. Efficient implementation of the packet buffer reduces the number of BlockRAMs to 18% when compared to simple parallel implementation. Instead of multiplexing input data, the architecture pipelines the states of FSMs. Such pipelined processing with only local communication has a direct positive impact on frequency and throughput and allows us to scale the architecture to hundreds of Gbps.
Denis Matousek, Jan Korenek, Viktor Pus
FPT2
2016 Software Defined Monitoring of Application Protocols
abstract
With the ongoing shift of network services to the application layer also the monitoring systems focus more on the data from the application layer. The increasing speed of the network links, together with the increased complexity of application protocol processing, require a new way of hardware acceleration. We propose a new concept of hardware acceleration for flexible flow-based application level traffic monitoring which we call Software Defined Monitoring. Application layer processing is performed by monitoring tasks implemented in the software in conjunction with a configurable hardware accelerator. The accelerator is a high-speed application-specific processor tailored to stateful flow processing. The software monitoring tasks control the level of detail retained by the hardware for each flow in such a way that the usable information is always retained, while the remaining data is processed by simpler methods. Flexibility of the concept is provided by a plugin-based design of both hardware and software, which ensures adaptability in the evolving world of network monitoring. Our high-speed implementation using FPGA acceleration board in a commodity server is able to perform a 100 Gb/s flow traffic measurement augmented by a selected application-level protocol analysis.
Lukas Kekely, Jan Kucera 0004, Viktor Pus, Jan Korenek, Athanasios V. Vasilakos
IEEE Trans. Computers4
2015 Towards Efficient Field Programmable Pattern Matching Array
abstract
Pattern matching is used in most of the network security devices in order to detect attacks, threats and malicious network traffic. Many hardware architectures have been designed to accelerate this time-critical operation in order to increase processing speed and achieve multi-gigabit throughput. Recently introduced automata processor is an powerful architecture which represents a new class of field programmable circuits called Field Programmable Pattern Matching Array (FPPMA). In the paper, we propose to improve the FPPMA architecture by Deterministic Units (DU), which have been originally introduced in NFA Split and significantly reduce the amount of required resources for mapping NFA to hardware. Dual position automaton is used to create data structure for the new FPPMA architecture from set of regular expressions. Moreover, we investigate efficiency of DU based on a dual position automation. Since the DU can implement a finite automaton without structural restrictions, we propose to use unrestricted finite automaton for deterministic unit and the dual position automaton for parts mapped to basic FPPMA elements. The results show that the DU based on the unrestricted finite automaton utilizes up to 42.43% less State-Transition elements than the DU based on the dual position automaton. Utilization of DUs provides significant reduction of FPPMA resources. For example, State-transition elements were reduced by more than 71% for the Snort module spyware-put.
Vlastimil Kosar, Jan Korenek
DSD2
2015 A Fast FPGA-Based Classification of Application Protocols Optimized Using Cartesian GP
David Grochol, Lukás Sekanina, Martin Zádník, Jan Korenek
EvoApplications4
2015 Hardware accelerated flow measurement of 100 Gb ethernet
abstract
This demo demonstrates results of a joint research project of CESNET and INVEA-TECH focused on 100 GbE network flow monitoring using FPGA. It shows, to the best of our knowledge, the first flow monitoring setup capable of handling fully saturated 100 G Ethernet line. We present COMBO-CG card that provides accurate timestamps for high-resolution traffic monitoring. The card is complemented by fast DMA engine and optimized Linux drivers which were designed and implemented to achieve 100 Gbps data transfers through PCIe bus with low CPU utilization. Network traffic can be distributed among multiple CPU cores based on configurable hash functions. Our flow exporter is able to fully utilize available CPU cores to provide wire-speed performance for processing of the 100 Gbps traffic. The demo will show complete 100 G flow monitoring setup - from packet generator to flow collector.
Viktor Pus, Petr Velan, Lukas Kekely, Jan Korenek, Pavel Minarík
IM4
2014 Network monitoring probe based on Xilinx Zynq
abstract
To provide reliable network and cloud services, it is necessary to perform precise monitoring and security analysis of cloud, ISP and local networks.Current SOHO (Small Office Home Office) devices have very limited resources and can not provide precise network security monitoring in local networks. Therefore we have designed small and low-power network probe which is able to analyse the network traffic at the application layer. The Xilinx Zynq enables to divide the task between hardware and software efficiently. The FPGA logic provides preprocessing (filtering) of data and the processor performs deep packet inspection to analyse application protocols. Moreover, the probe is ready to offload any time consuming operation (eg. regular expression matching) to the FPGA logic to increase processing speed.
Jan Viktorin, Pavol Korcek, Tomas Fukac, Jan Korenek
ANCS4
2014 Low latency book handling in FPGA for high frequency trading
abstract
Recent growth in algorithmic trading has caused a demand for lowering the latency of systems for electronic trading. FPGA cards are widely used to reduce latency and accelerate market data processing. To create a low latency trading system, it is crucial to effectively build a representation of the market state (book) in hardware. Thus, we have designed a new hardware architecture, which updates the book with the best bid/offer prices based on the incoming messages from the exchange. For each message a corresponding financial instrument needs to be looked up and its record needs to be updated. Proposed architecture is utilizing cuckoo hashing for the book handling, which enables low latency symbol lookup and high memory utilization. In this paper we discuss a trade-off between lookup latency and memory utilization. With average latency of 253 ns the proposed architecture is able to handle 119 275 instruments while using only 144 Mbit QDR SRAM.
Milan Dvorak, Jan Korenek
DDECS2
2014 Fast lookup for dynamic packet filtering in FPGA
abstract
Rapidly growing speed and complexity of computer networks impose new requirements on fast lookup structures which are utilized in many networking applications (SDN, firewalls, NATs, etc.). We propose a novel lookup concept based on the well-known cuckoo hashing, which can achieve good memory utilization, supplemented by a binary search tree for offloading the colliding keys and supporting LPM lookup. We also propose a hardware architecture implementing this lookup concept in the FPGA. Our solution is suitable for lookup of the variable-length keys in 100+ Gbps networks. Memory utilization of the proposed concept is thoroughly evaluated and it is shown that the concept is scalable to external memory components.
Lukas Kekely, Martin Zádník, Jirí Matousek 0002, Jan Korenek
DDECS4
2014 On NFA-split architecture optimizations
abstract
Fast regular expression matching is widely used in many network devices. The NFA-Split hardware architecture is an efficient approach to match large set of regular expression at multigigabit speed with very low FPGA logic utilization. We propose optimizations of NFA-Split architecture, which further reduce FPGA logic utilization and significantly reduce memory utilization. The amount of utilized BlockRAMs was reduced by 97% for a Snort web-cgi module and FPGA logic utilization was reduced by 34% for a Snort backdoor module. Moreover, we propose new NFA-Split construction algorithm which decrease overall construction time up to 39 times.
Vlastimil Kosar, Jan Korenek
DDECS2
2014 Design methodology of configurable high performance packet parser for FPGA
abstract
Packet parsing is among basic operations that are performed at all points of a network infrastructure. Modern networks impose challenging requirements on the performance and configurability of packet parsing modules. However, high-speed parsers often use a significant amount of hardware resources. We propose a novel architecture of a pipelined packet parser for FPGA, which offers low latency in addition to high throughput (over 100 Gb/s). Moreover, the latency, throughput and chip area can be finely tuned to fit the needs of a particular application. The parser is hand-optimized thanks to a direct implementation in VHDL, yet the structure is uniform and easily extensible for new protocols.
Viktor Pus, Lukas Kekely, Jan Korenek
DDECS3
2014 Trade-offs and progressive adoption of FPGA acceleration in network traffic monitoring
abstract
Current hardware acceleration cores for network traffic processing are often well optimized for one particular task and therefore provide high level of hardware acceleration. But for many applications, such as network traffic monitoring and security, it is also necessary to achieve rapid development cycle to provide fast response to security threats.We propose and evaluate a new concept of hardware acceleration for flexible flow-based network traffic monitoring with support of application protocol analysis. The concept is called Software Defined Monitoring (SDM) and it relies on a configurable hardware accelerator implemented in FPGA, coupled with smart monitoring tasks running as software on general CPU. The monitoring tasks in the software control the level of detail and type of information retained during the hardware processing. This arrangement allows rapid application prototyping in the software, followed by further shifting of the timing critical parts of the processing to the hardware accelerator. The concept is proposed with the scalability in mind, therefore it is suitable for different FPGA based platforms ranging from embedded single-chip solutions (such as Zynq or CycloneV) to high-speed backbone network monitoring boxes. Our pilot high-speed implementation using FPGA acceleration board in a commodity server performs a 100Gb/s flow traffic measurement augmented by a selected application protocol analysis.
Lukas Kekely, Viktor Pus, Pavel Benácek, Jan Korenek
FPL4
2014 Software Defined Monitoring of application protocols
abstract
Current high-speed network monitoring systems focus more and more on the data from the application layers. Flow data is usually enriched by the information from HTTP, DNS and other protocols. The increasing speed of the network links, together with the time consuming application protocol parsing, require a new way of hardware acceleration. Therefore we propose a new concept of hardware acceleration for flexible flow-based application level monitoring which we call Software Defined Monitoring (SDM). The concept relies on smart monitoring tasks implemented in the software in conjunction with a configurable hardware accelerator. The hardware accelerator is an application-specific processor tailored to stateful flow processing. The monitoring tasks reside in the software and can easily control the level of detail retained by the hardware for each flow. This way the measurement of bulk/uninteresting traffic is offloaded to the hardware while the advanced monitoring over the interesting traffic is performed in the software. The proposed concept allows one to create flexible monitoring systems capable of deep packet inspection at high throughput. Our pilot implementation in FPGA is able to perform a 100 Gb/s flow traffic measurement augmented by a selected application-level protocol parsing.
Lukas Kekely, Viktor Pus, Jan Korenek
INFOCOM3
2013 Hardware architecture for the fast pattern matching
abstract
As the speed of current computer networks increases, it is necessary to protect networks by security systems such as firewalls and Intrusion Detection Systems (IDS) operating at multigigabit speeds. As attacks on modern networks became more and more complex, it is necessity to detect attack placed not only in single packet but at the level of network flows. Pattern matching in the network flows is the time-critical operation of many modern IDS. Most of the regularly used patterns are described by the regular expression. This work describes advanced hardware architecture for the fast regular expression matching based on the perfect hashing. The proposed architecture is scalable and can achieve multigigabit throughput per network flow.
Jan Kastil, Vlastimil Kosar, Jan Korenek
DDECS3
2013 Hardware acceleration in computer networks
abstract
Summary form only given. Network traffic processing speed is crucial in most of network devices, because any packet drop can lead to lower quality of network services, affect precise monitoring or disallow detection of security threats. General purpose processors are not able to process all data on high-speed network links. For 100 Gb lines, packet can arrive every 5 ns. Therefore network devices use hardware acceleration to speed up time-critical operations. In this tutorial, we will introduce these time-critical operations together with hardware architectures, which are able to achieve 10, 40 or even 100 Gbps throughput. In particular, the tutorial deals with parsing of packet headers, longest prefix matching (IP look-up), packet classification and regular expression matching. All these operations are widely used in network security and monitoring devices and have to be accelerated to achieve 100 Gb throughput. We will present that deep pipelines, perfect hashing and efficient utilization of on-chip memory can help to achieve high throughput or decrease hardware resources. The end of the presentation will be devoted to the rapid development of hardware accelerated network applications for 100 Gbps networks. In summary, tutorial participants will become familiar with the state of the art algorithms and hardware architectures for high speed packet processing. They will learn how to utilize these architectures and accelerate network applications.
Jan Korenek
DDECS1
2013 Towards hardware architecture for memory efficient IPv4/IPv6 Lookup in 100 Gbps networks
abstract
With growing speed of computer networks, core routers have to increase performance of longest prefix match (LPM) operation on IP addresses. While existing LPM algorithms are able to achieve high throughput for IPv4 addresses, an IPv6 processing speed is limited. To achieve 100 Gbps throughput, LPM operation has to be processed in dedicated hardware and a forwarding table has to fit into an on-chip memory. Current LPM algorithms need large memory to store IPv6 forwarding tables or use compression with dynamic data structres, which can not be simply implemented in hardware. Therefore, we provide analysis of available forwarding tables of core routers and propose a new representation of prefix sets. The proposed representation has very low memory demands and is suitable for high-speed pipelined processing, which is shown on a new highly pipelined hardware architecture with 100 Gbps throughput.
Jirí Matousek 0002, Martin Skacan, Jan Korenek
DDECS3
2013 Memory efficient IP lookup in 100 GBPS networks
abstract
The increasing number of devices connected to the Internet together with video on demand have a direct impact to the speed of network links and performance of core routers. To achieve 100 Gbps throughput, core routers have to implement IP lookup in dedicated hardware and represent a forwarding table using a data structure, which fits into the on-chip memory. Current IP lookup algorithms have high memory demands when representing IPv6 prefix sets or introduce very high pre-processing overhead. Therefore, we performed analysis of IPv4 and IPv6 prefixes in forwarding tables and propose a novel memory representation of IP prefix sets, which has very low memory demands. The proposed representation has better memory utilization in comparison to the highly optimized Shape Shifting Trie (SST) algorithm and it is also suitable for IP lookup in 100 Gbps networks, which is shown on a new pipelined hardware architecture with 170 Gbps throughput.
Jirí Matousek 0002, Martin Skacan, Jan Korenek
FPL3
2013 NFA reduction for regular expressions matching using FPGA
abstract
Many algorithms have been proposed to accelerate regular expression matching via mapping of a nondeterministic finite automaton into a circuit implemented in an FPGA. These algorithms exploit unique features of the FPGA to achieve high throughput. On the other hand the FPGA poses a limit on the number of regular expressions by its limited resources. In this paper, we investigate applicability of NFA reduction techniques - a formal aparatus to reduce the number of states and transitions in NFA prior to its mapping into FPGA. The paper presents several NFA reduction techniques, each with a different reduction power and time complexity. The evaluation utilizes regular expressions from Snort and L7 decoder. The best NFA reduction algorithms achieve more than 66% reduction in the number of states for a Snort ftp module. Such a reduction translates directly into 66% LUT-FF pairs saving in the FPGA.
Vlastimil Kosar, Martin Zádník, Jan Korenek
FPT3
2012 A new embedded platform for rapid development of network applications
abstract
No abstract available.
Jan Korenek, Pavol Korcek, Vlastimil Kosar, Martin Zádník, Jan Viktorin
ANCS1
2012 Low-latency modular packet header parser for FPGA
abstract
Packet parsing is the basic operation performed at all points of the network infrastructure. Modern networks impose challenging requirements on the performance and configurability of packet parsing modules, however the high-speed parsers often use very large chip area. We propose novel architecture of pipelined packet parser, which in addition to high throughput (over 100 Gb/s) offers also low latency. Moreover, the latency to throughput ratio can be finely tuned to fit the particular application.
Viktor Pus, Lukas Kekely, Jan Korenek
ANCS3
2012 Reducing memory in high-speed packet classification
abstract
Many packet classification algorithms were proposed to deal with the rapidly growing speed of computer networks. Unfortunately all of these algorithms are able to achieve high throughput only at the cost of excessively large memory and can be used only for small sets of rules. We propose new algorithm that uses four techniques to lower the memory requirements: division of rule set into subsets, removal of critical rules, prefix coloring and perfect hashing. The algorithm is designed for pipelined hardware implementation, can achieve the throughput of 266 million packets per second, which corresponds to 178Gb/s for the shortest 64B packets, and outperforms older approaches in terms of memory requirements by 66% in average for the rule sets available to us.
Viktor Pus, Jan Korenek
IWCMC2
2011 Netbench: Framework for Evaluation of Packet Processing Algorithms
abstract
Many algorithms and hardware architectures are proposed to increase processing speed of time-critical operations in the field of longest prefix matching, packet classification and regular expression matching. Despite this fact, there is still no free and easily extensible platform for evaluation, comparison and experiments with existing approaches. We propose the Net bench Framework which aims to serve as an independent platform for researchers seeking the easiest way to implement their algorithms, as well as the comparison of their algorithms with reference implementations of other approaches. The framework is provided as an open source and can be easily extended to support new algorithms or new comparison methodology. Net bench is publicly available at http://www.fit.vutbr.cz/netbench.
Viktor Pus, Jiri Tobola, Vlastimil Kosar, Jan Kastil, Jan Korenek
ANCS5
2011 Reduction of FPGA resources for regular expression matching by relation similarity
abstract
Intrusion Detection Systems have to match large sets of regular expressions to detect malicious traffic on multi-gigabit networks. Many algorithms and architectures have been proposed to accelerate pattern matching, but formal methods for reduction of Nondeterministic finite automata have not been used yet. We propose to use reduction of automata by similarity to match larger set of regular expressions in FPGA. Proposed reduction is able to decrease the number of states by more than 32% and the amount of transitions by more than 31%. The amount of look-up tables is reduced by more than 15% and the amount of flip-flops by more than 34%.
Vlastimil Kosar, Jan Korenek
DDECS2
2011 Hardware architecture for packet classification with prefix coloring
abstract
Packet classification is a widely used operation in network security devices. As network speeds are increasing, the demand for hardware acceleration of packet classification in FPGAs or ASICs is growing. Nowadays algorithms implemented in hardware can achieve multigigabit speeds, but suffer with great memory overhead. We propose a new algorithm and hardware architecture which reduces memory requirements of decomposition based methods for packet classification. The algorithm uses prefix coloring to reduce large amount of Cartesian product rules at the cost of an additional pipelined processing and a few bits added into results of the longest prefix match operation. The proposed hardware architecture is designed as a processing pipeline with the throughput of 266 million packets per second using commodity FPGA and one external memory. The greatest strength of the algorithm is the constant time complexity of the search operation, which makes the solution resistant to various classes of network security attacks.
Viktor Pus, Michal Kajan, Jan Korenek
DDECS3
2011 Effective hash-based IPv6 longest prefix match
abstract
With the growing speed of computer networks, the core routers have to increase performance of longest prefix match (LPM) operation on IP address. While existing LPM algorithms are able to achieve high throughput for IPv4 addresses, the IPv6 processing speed is limited. In this paper we propose a new Hast-Tree Bitmap algorithm for fast longest prefix match for both IPv4 and IPv6 networks. The algorithm is able to achieve high throughput for a long IPv6 addresses by fast hash function which is used to jump over the sparse part of the IP prefix tree. The proposed algorithm was mapped to the highly pipelined hardware architecture, which offers well balanced resource requirements for IPv6 look-up and is able to achieve a wire-speed throughput for 100 Gbps networks.
Jiri Tobola, Jan Korenek
DDECS2
2010 Efficient packet classification algorithm based on entropy
abstract
This paper deals with packet classification in high-speed networks. It introduces a novel method for packet classification based on the amount of information stored in the ruleset. Basic principles of the algorithm based on the effort to reduce the amount of the necessary memory space and number of computational steps are presented together with analysis of the input rulesets.
Michal Kajan, Jan Korenek
ANCS2
2010 High speed pattern matching algorithm based on deterministic finite automata with faulty transition table
abstract
Regular expression matching is the time-critical operation of many modern intrusion detection systems (IDS). This paper proposes pattern matching algorithm to match regular expression against multigigabit data stream. As usually used regular expressions are only subjectively tested and often generates many false positives/ negatives, proposed algorithm support the possibility to reduce memory requirements by introducing small amount of faults into the pattern matching. Algorithm is based on the perfect hashing and is suitable for hardware implementation.
Jan Kastil, Jan Korenek
ANCS2
2010 NFA split architecture for fast regular expression matching
abstract
Many hardware architectures have been designed to accelerate regular expression matching in network security devices, but most of them can achieve high throughput only for strings or small sets of regular expressions. We propose new NFA Split architecture which reduces the amount of consumed FPGA resources in order to match larger set of regular expressions. New algorithm is introduced to find non-collision sets of states and determine part of nondeter-ministic automaton which can be mapped to the memory based architecture. For all analysed sets of regular expressions, the algorithm was able to find non-collision sets with 67.8% of states in average and reduces the amount of consumed flip-flops to 37.6% and look-up tables to 63.9% in average.
Jan Korenek, Vlastimil Kosar
ANCS1
2010 Hardware accelerated pattern matching based on Deterministic Finite Automata with perfect hashing
abstract
With the increased amount of data transferred by computer networks, the amount of the malicious traffic also increases and therefore it is necessary to protect networks by security systems such as firewalls and Intrusion Detection Systems (IDS) operating at multigigabit speeds. Pattern matching is the time critical operation of current IDS. This paper deals with the analysis of regular expressions used by modern IDS to describe malicious traffic. According to our analysis, more than 64 percent of regular expressions create Deterministic Finite Automaton (DFA) with less than 20 percent of saturation of the transition table which allows efficient implementation of pattern matching into FPGA platform. We propose architecture for fast pattern matching using perfect hashing suitable for implementation into FPGA platform. The memory requirements of presented architecture is closed to the theoretical minimum for sparse transition tables.
Jan Kastil, Jan Korenek
DDECS2
2010 Efficient mapping of nondeterministic automata to FPGA for fast regular expression matching
abstract
With the growing number of viruses and network attacks, Intrusion Detection Systems have to match a large set of regular expressions at multi-gigabit speed to detect malicious activities on the network. Many algorithms and architectures have been designed to accelerate pattern matching, but most of them can be used only for strings or a small set of regular expressions. We propose new NFA-Split architecture, which reduces the amount of consumed FPGA resources in order to match larger set of regular expressions at multi-gigabit speed. The proposed reduction uses model of nondeterministic and deterministic automaton for effective mapping of regular expressions to FPGA. A new algorithm is designed to split the nondeterministic automaton transition table in order to map a part of the table into memory. The algorithm can place more than 49% of transition table to memory, which reduces the amount of look-up tables by more than 43% and flip-flops by more than 38% for all selected sets of regular expressions. Moreover, a sparse transition table is mapped to memory with overlapped rows, which enables to store the table in a highly compact form.
Jan Korenek, Vlastimil Kosar
DDECS1
2010 Memory optimizations for packet classification algorithms in FPGA
abstract
Packet classification algorithms are widely used in network security devices. As network speeds are increasing, the demand for hardware acceleration of packet classification in FPGAs or ASICs is growing. Nowadays hardware architectures can achieve multigigabit speeds only at the cost of large data structures, which can not fit into the on-chip memory. We propose novel method how to reduce data structure size for the family of decomposition architectures at the cost of additional pipelined processing with only small amount of logic resources. The reduction significantly decreases overhead given by the Cartesian product nature of classification rules. Therefore the data structure can be compressed to 10% on average. As high compression ratio is achieved, fast on-chip memory can be used to store data structures and hardware architectures can process network traffic at significantly higher speed.
Viktor Pus, Juraj Blaho, Jan Korenek
DDECS3
2009 Memory optimization for packet classification algorithms
abstract
We propose novel method how to reduce data structure size for the family of packet classification algorithms at the cost of additional pipelined processing with only small amount of logic resources. The reduction significantly decreases overhead given by the crossproduct nature of classification rules. Therefore the data structure can be compressed to 10% on average. As high compression ratio is achieved, fast on-chip memory can be used to store data structures and hardware architectures can process network traffic at significantly higher speed.
Juraj Blaho, Jan Korenek, Viktor Pus
ANCS2
2009 Packet header analysis and field extraction for multigigabit networks
abstract
Packet header analysis and extraction of header fields needs to be performed in all network devices. As network speed is increasing quickly, high speed packet header processing is required. We propose a new architecture of packet header analysis and fields extraction intended for high-speed FPGA-based network applications. The architecture is able to process 20 Gbps network links with less than 12 percent of available resources of Virtex 5 110 FPGA. Moreover, the presented solution can balance between network throughput and consumed hardware resources to fit application needs. The architecture for packet header processing is generated from standard XML protocol scheme and is strongly optimised for resource consumption and speed by an automatic HDL code generator. Our solution also enables to change the set of extracted header fields on-line without FPGA reconfiguration.
Petr Kobierský, Jan Korenek, Libor Polcak
DDECS2
2009 Methodology for Fast Pattern Matching by Deterministic Finite Automaton with Perfect Hashing
abstract
As the speed of current computer networks increases, it is necessary to protect networks by security systems such as firewalls and intrusion detection systems operating at multigigabit speeds. Pattern matching is the time-critical operation of current IDS on multigigabit networks. Regular expressions are often used to describe malicious network patterns. This paper deals with fast regular expression matching using the deterministic finite automaton (DFA) with perfect hash function. We introduce decomposition of the problem on two parts: transformation of the input alphabet and usage of a fast DFA, and usage of perfect hashing to reduce space/speed tradeoff for DFA transition table.
Jan Kastil, Jan Korenek, Ondrej Lengál
DSD2
2009 Fast and scalable packet classification using perfect hash functions
abstract
Packet classification is an important operation for applications such as routers, firewalls or intrusion detection systems. Many algorithms and hardware architectures for packet classification have been created, but none of them can compete with the speed of TCAMs in the worst case. We propose new hardware-based algorithm for packet classification. The solution is based on problem decomposition and is aimed at the highest network speeds. A unique property of the algorithm is the constant time complexity in terms of external memory accesses. The algorithm performs exactly two external memory accesses to classify a packet. Using FPGA and one commodity SRAM chip, a throughput of 150 million packets per second can be achieved. This makes throughput of 100 Gbps for the shortest packets. Further performance scaling is possible with more or faster SRAM chips.
Viktor Pus, Jan Korenek
FPGA2
2008 GICS: Generic interconnection system
abstract
The division of an application between a conventional processor and an acceleration card with FPGA chips has been proved as a suitable way for an acceleration of computationally intensive tasks. In such applications, the designer usually has to implement an interconnection between components placed in FPGA and the host system bus. This task is often complicated by different requirements of user components for throughput, latency of reading operations, need for DMA transfers etc. The objective of this work is to show a new approach for implementation of interconnection systems and to enable the designer to focus on the development of the target application. The proposed interconnection system is based on tree topology. The system eliminates the sensitivity of wide buses to the distance, supports the connection of components with different requirements for throughput, supports split transaction model and many other features. The proposed system is implemented and evaluated on chips with Virtex 5 technology.
Tamas Malek, Tomás Martínek, Jan Korenek
FPL3
2007 Online Protocol Testing for FPGA Based Fault Tolerant Systems
abstract
In this paper, the methodology for automated design of checker for communication protocol testing is presented. Based on the level of checking, different design strategies can be performed - in the paper the lowest level is presented. The definition of dedicated language for the description of possible communication faults is presented. The core generator is used to produce VHDL code describing the behaviour of the checker.
Jiri Tobola, Zdenek Kotásek, Jan Korenek, Tomás Martínek, Martin Straka
DSD3
2007 FlowContext: Flexible Platform for Multigigabit Stateful Packet Processing
abstract
Network security systems become an essential part of many network structures in both company and university domains. These systems however require a higher semantic level of network traffic analysis like stateful filtration or TCP stream reassembling. This paper deals with an architecture of flexible FlowContext platform capable of stateful processing at multigigabit speeds. It allows to analyze and process incoming network traffic with a flow-based approach rather than packet-based one. The proposed architecture is flexible in supporting wide range of applications, allows performance scalability and state information consistency checking. The advantages and flexibility of proposed platform is demonstrated on several network security applications.
Martin Kosek, Jan Korenek
FPL2
2005 NetFlow Probe Intended for High-Speed Networks
abstract
With growing speed of communication over the Internet there is a need for a reliable monitoring devices which are able to provide information about spectrum of traffic mix, attacks, applications, etc. This paper proposes architecture of network flow monitoring adapter based on hardware platform COMBO6. With use of field programmable gate arrays (FPGA) placed on these cards it is possible to monitor flows in high-speed environment. Component parts of the architecture and implementation platform are described. Several different models have been created to analyze and prove important characteristics of the architecture and results are derived. The probe is able to monitor 1 million simultaneous flows on an 2Gbps network link.
Martin Zádník, Tomas Pecenka, Jan Korenek
FPL3