Karel Hynek

dblp:247/0217 · DBLP profile ↗
← Back
17ranked-venue papers
1as first author
14since 2021 · last 2026
0000-0002-8281-618XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 5 since 2021Security and privacy · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Advanced similarity metrics for IP flow data analytics
abstract
Machine learning techniques provide powerful tools for analysis of encrypted network traffic. We present a novel set of features and a distance measure suitable for a wide variety of distance-based machine learning techniques useful in classification, clustering, and novelty detection in encrypted traffic flow data. The proposed features are given by distances between probability distributions of such characteristics as packet sizes or inter-arrival times. The proposed distance measure incorporates those features in a fractional l p -metric with p close to 0.1. The effectiveness of the distance measure is evaluated using an extensive labeled dataset containing 154 web services. The dataset was captured on the ISP backbone, and we present it as supplementary material. Application of k -nearest neighbors (kNN) algorithm in combination with the proposed distance measure and feature selection gives the classification accuracy of 91.0%. Comparison with a deep learning model shows that the kNN is competitive with state-of-the-art models. On a different publicly available dataset with a low number of classes, our approach reaches accuracy 99.6%, outperforming models presented in the literature. The benefits of novel features are further demonstrated using the LightGBM model, i.e. without relying on distance-based techniques. Besides direct classification, we have used the local outlier factor and rank-based detection algorithms to detect novel traffic flows. We show that their performance improves when using the proposed distance measure. Finally, to speed up traffic classification we compared the performance of kNN and Approximate Nearest Neighbors (ANN) algorithm, and applied the Affinity propagation clustering algorithm to select a suitable subset of kNN/ANN training points.
Ivo Petr, Jitka Hrabáková, Magda Friedjungová, Daniel Vasata, Karel Hynek
Comput. Networks5
2026 Universal Embedding Function for Traffic Classification via QUIC Domain Recognition Pretraining: A Transfer Learning Success
Jan Luxemburk, Karel Hynek, Richard Plný, Tomás Cejka
IEEE Trans. Netw. Serv. Manag.2
2025 Taming Volatility: Stable and Private QUIC Classification with Federated Learning
abstract
Federated Learning (FL) is a promising approach for privacy-preserving network traffic analysis, but its practical deployment is challenged by the non-IID nature of real-world data. While prior work has addressed statistical heterogeneity, the impact of temporal traffic volatility-the natural daily ebb and flow of network activity-on model stability remains largely unexplored. This volatility can lead to inconsistent data availability at clients, destabilizing the entire training process. In this paper, we systematically address the problem of temporal volatility in federated QUIC classification. We first demonstrate the instability of standard FL in this dynamic setting. We then propose and evaluate a client-side data buffer as a practical mechanism to ensure stable and consistent local training, decoupling it from real-time traffic fluctuations. Using the real-world CESNETQUIC22 dataset partitioned into 14 autonomous clients, we then demonstrate that this approach enables robust convergence. Our results show that a stable federated system achieves a 95.2% F1 score, a mere 2.3 percentage points below a non-private centralized model. This work establishes a blueprint for building operationally stable FL systems for network management, proving that the challenges of dynamic network environments can be overcome with targeted architectural choices.
Richard Jozsa, Karel Hynek, Adrián Pekár
CNSM2
2025 CESNET TS-Zoo: A Library for Reproducible Analysis of Network Traffic Time Series
abstract
Time Series Analysis (TSA) is an essential tool in computer networking, supporting tasks such as traffic forecasting, capacity planning, load balancing, quality of service monitoring, behavior profiling, and anomaly detection. Despite its widespread use, the community was limited by the lack of sufficient datasets. Our recent dataset, CESNET-TimeSeries24, finally fills this gap. However, its substantial size presents significant challenges for practical use in research. Therefore, inspired by the other machine learning communities that often develop supportive tools and benchmarks to accelerate research, we introduced a CESNET TS-Zoo library. It is designed to streamline dataset management, experiment setting, and reproducibility in the TSA of network traffic. TS-Zoo provides a standardized API for accessing the CESNET-TimeSeries24 dataset and includes methods for time series preprocessing, multiple dataset partitioning, and data loading for experiments. Furthermore, the preprocessing steps can be exported and imported, enabling reproducible experiments. Therefore, the TS-Zoo library simplifies TSA experiments and enables reproducibility of TSA research applied in computer networking.
Milan Kures, Josef Koumar, Karel Hynek
CNSM3
2025 On Evaluation of Data Fusion Methods for Network Traffic Classification
abstract
Data fusion plays a critical role in modern network security, enhancing explainability and trust in ML/AI-based detectors. Yet, the effect of standard fusion methods has been largely overlooked. This paper evaluates seven common lateand hard-fusion techniques-Majority Voting, Weighted Majority Voting, Recall Combiner, Naive Bayes, Behavior-Knowledge Space, Decision Tree, and Logistic Regression-in the context of traffic classification and attack detection. Using both synthetic data from controlled environments and real-world datasets, we analyze their performance across diverse scenarios. The results highlight the strengths and limitations of each method and offer practical guidance for selecting effective fusion strategies in network security applications.
Richard Plný, Karel Hynek
CNSM2
2024 WIF: Efficient Library for Network Traffic Analysis
abstract
Network traffic classification and analysis are crucial for maintaining computer security. Nevertheless, the rise of encrypted traffic has made reliable threat detection increasingly challenging, requiring more complex algorithms such as heterogeneous ensembles. These types of algorithms proved to be effective in complex threat detection while maintaining high accuracy and explainability. However, their complexity and time-consuming development process limit their widespread adoption. Therefore, we created a new library called Weak Indication Framework (WIF) for the faster development of heterogeneous ensembles, which minimizes the time between attack discovery and detection capability. Moreover, WIF-based detectors are efficient enough to operate on large Internet Service Provider networks—a single detector can protect millions of users. We demonstrate the effectiveness of the WIF library through four different detectors (TOR, Cryptomining, IoT Malware, and Tunnel detector), each achieving outstanding performance and quick deployment times.
Richard Plný, Karel Hynek, Pavel Siska
CNSM2
2024 Comparative analysis of DNS over HTTPS detectors
abstract
DNS over HTTPS (DoH) is a protocol that encrypts DNS traffic to improve user privacy and security. However, its use also poses challenges for network operators and security analysts who need to detect and monitor network traffic for security purposes. Therefore, there are multiple DoH detection proposals that leverage machine learning to identify DoH connections; however, these proposals were often tested on different datasets, and their evaluation methodologies were not consistent enough to allow direct performance comparison. In this study, seven DoH detection proposals were recreated and evaluated with six different experiments to answer research questions that targeted specific deployment scenarios concerning ML-model transferability, usability, and longevity. For thorough testing, a large Collection of DoH datasets along with a novel 5-week dataset was used, which enabled the evaluation of models’ longevity. This study provides insights into the current state of DoH detection techniques and evaluates the models in scenarios that have not been previously tested. Therefore, this paper goes beyond classical replication studies and shows previously unknown properties of seven published DoH detectors.
Kamil Jerábek, Karel Hynek, Ondrej Rysavý
Comput. Networks2
2024 NetTiSA: Extended IP flow with time-series features for universal bandwidth-constrained high-speed network traffic classification
abstract
Network traffic monitoring based on IP Flows is a standard monitoring approach that can be deployed to various network infrastructures, even the large ISP networks connecting millions of people. Since flow records traditionally contain only limited information (addresses, transport ports, and amount of exchanged data), they are also commonly extended by additional features that enable network traffic analysis with high accuracy. These flow extensions are, however, often too large or hard to compute, which then allows only offline analysis or limits their deployment only to smaller-sized networks. This paper proposes a novel extended IP flow called NetTiSA (Network Time Series Analysed) flow, based on analysing the time series of packet sizes. By thoroughly testing 25 different network traffic classification tasks, we show the broad applicability and high usability of NetTiSA flow. For practical deployment, we also consider the sizes of flows extended by NetTiSA features and evaluate the performance impacts of their computation in the flow exporter. The novel features proved to be computationally inexpensive and showed excellent discriminatory performance. The trained machine learning classifiers with proposed features mostly outperformed the state-of-the-art methods. NetTiSA finally bridges the gap and brings universal, small-sized, and computationally inexpensive features for traffic classification that can be scaled up to extensive monitoring infrastructures, bringing the machine learning traffic classification even to 100 Gbps backbone lines.
Josef Koumar, Karel Hynek, Jaroslav Pesek, Tomás Cejka
Comput. Networks2
2023 Network Traffic Classification Based on Single Flow Time Series Analysis
abstract
Network traffic monitoring using IP flows is used to handle the current challenge of analyzing encrypted network communication. Nevertheless, the packet aggregation into flow records naturally causes information loss; therefore, this paper proposes a novel flow extension for traffic features based on the time series analysis of the Single Flow Time series, i.e., a time series created by the number of bytes in each packet and its timestamp. We propose 69 universal features based on the statistical analysis of data points, time domain analysis, packet distribution within the flow timespan, time series behavior, and frequency domain analysis. We have demonstrated the usability and universality of the proposed feature vector for various network traffic classification tasks using 15 well-known publicly available datasets. Our evaluation shows that the novel feature vector achieves classification performance similar or better than related works on both binary and multiclass classification tasks. In more than half of the evaluated tasks, the classification performance increased by up to 5 %.
Josef Koumar, Karel Hynek, Tomás Cejka
CNSM2
2023 BOTA: Explainable IoT Malware Detection in Large Networks
abstract
Explainability and alert reasoning are essential but often neglected properties of intrusion detection systems. The lack of explainability reduces security personnel’s trust, limiting the overall impact of alerts. This article proposes the botnet analysis (BOTA) system, which uses the concepts of weak indicators and heterogeneous meta-classifiers to maintain accuracy compared with state-of-the-art systems while also providing explainable results that are easy to understand. To evaluate the proposed system, we have implemented a demonstration of intrusion weak-indication detectors, each working on a different principle to ensure robustness. We tested the architecture with various real-world and lab-created data sets, and it correctly identified 94.3% of infected Internet of Things (IoT) devices without false positives. Furthermore, the implementation is designed to work on top of extended bidirectional flow data, making it deployable on large 100-Gb/s large-scale networks at the level of Internet Service Providers. Thus, a single instance of BOTA can protect millions of devices connected to end-users’ local networks and significantly reduce the threat arising from powerful IoT botnets.
Daniel Uhrícek, Karel Hynek, Tomás Cejka, Dusan Kolár
IEEE Internet Things J.2
2022 Tunneling through DNS over TLS providers
abstract
DNS over TLS (DoT) is one of the approaches for private DNS resolution, which has already gained support by open resolvers. Moreover, DoT is used by default in Android operating systems. This study investigates the possibility of creating DNS covert channels using DoT, which is a security threat that benefits from the increased privacy of encrypted communication. We evaluated the performance and usability of DoT tunnels created via commonly used resolvers. Our results show that the performance characteristics of DoT tunnels differ vastly depending on the used DoT resolver; however, the creation of a DoT tunnel is possible, reaching speeds up to 232 Kbps. Moreover, we successfully transferred data via DoT servers claiming Anti-Virus protection and family-friendly content.
Lukás Melcher, Karel Hynek, Tomás Cejka
CNSM2
2022 Large Scale Analysis of DoH Deployment on the Internet
Sebastián García, Joaquín Bogado, Karel Hynek, Dmitrii Vekshin, Tomás Cejka, Armin Wasicek
ESORICS (3)3
2021 Towards Evaluating Quality of Datasets for Network Traffic Domain
abstract
This paper deals with the quality of network traffic datasets created to train and validate machine learning classification and detection methods. Naturally, there is a long epoch of research targeted at data quality; however, it is focused mainly on data consistency, validity, precision, and other metrics, which are insufficient for network traffic use-cases. The rise of Machine learning usage in network monitoring applications requires a new methodology for evaluation datasets. There is a need to evaluate and compare traffic samples captured at different conditions and decide the usability of the already captured and annotated data. This paper aims to explain a use case of dataset creation, propose definitions regarding the quality of the network traffic datasets, and finally, describe a framework for datasets analysis.
Dominik Soukup, Peter Tisovcík, Karel Hynek, Tomás Cejka
CNSM3
2021 Novel HTTPS classifier driven by packet bursts, flows, and machine learning
abstract
Encryption of network traffic recently starts to cover remaining readable information, which is heavily used by current monitoring systems; thus, it is time to focus on novel methods of encrypted traffic analysis and classification. The aim of this paper is to define a new network traffic characteristic called Sequence of packet Burst Length and Time (SBLT), which was inspired by existing approaches and definitions. Contrary to other works, SBLT is feasible even for high-speed backbone networks as a part of IP flow data. The advantage of SBLT features is shown using a machine learning classification model for HTTPS traffic types as an example. This paper presents the definition of SBLT, proposes a new annotated public dataset of HTTPS traffic with 5 categories, and evaluates the developed classifier reaching accuracy over 99 %. This classifier can help analysts to deal with a huge amount of encrypted traffic and maintain situational awareness.
Zdena Tropková, Karel Hynek, Tomás Cejka
CNSM2
2020 DoH Insight: detecting DNS over HTTPS by machine learning
abstract
Over the past few years, a new protocol DNS over HTTPS (DoH) has been created to improve users' privacy on the internet. DoH can be used instead of traditional DNS for domain name translation with encryption as a benefit. This new feature also brings some threats because various security tools depend on readable information from DNS to identify, e.g., malware, botnet communication, and data exfiltration. Therefore, this paper focuses on the possibilities of encrypted traffic analysis, especially on the accurate recognition of DoH. The aim is to evaluate what information (if any) can be gained from HTTPS extended IP flow data using machine learning. We evaluated five popular ML methods to find the best DoH classifiers. The experiments show that the accuracy of DoH recognition is over 99.9 %. Additionally, it is also possible to identify the application that was used for DoH communication, since we have discovered (using created datasets) significant differences in the behavior of Firefox, Chrome, and cloudflared. Our trained classifier can distinguish between DoH clients with the 99.9 % accuracy.
Dmitrii Vekshin, Karel Hynek, Tomás Cejka
ARES2
2020 Pipelined ALU for effective external memory access in FPGA
abstract
The external memories in digital design are closely related to high response time. The most common approach to mitigate latency is adding a caching mechanism into the memory subsystem. This solution might be sufficient in CPU architecture, where we can reschedule operations when a cache miss occurs. However, the FPGA architectures are usually accelerators with simple functionality, where it is not possible to postpone work. The cache miss often leads to whole pipeline stall or even to data loss. The architecture we present in this paper reduces this problem by aggregating arithmetic operations into the memory subsystem itself. Fast data processing is achieved because arithmetic operations working with external data are offloaded. Our architecture reaches a speed of 200 Mp/s (operations carried out). It is designed to be used in systems with link speeds of 100 Gb/s. It outperforms other implementations by a factor of at least 3. The additional benefit of our architecture is reducing the number of memory transactions by a factor of two on real-world datasets.
Tomás Benes, Michal Kekely, Karel Hynek, Tomás Cejka
DSD3
2020 Refined Detection of SSH Brute-Force Attackers Using Machine Learning
Karel Hynek, Tomás Benes, Tomás Cejka, Hana Kubátová
SEC1