Marcin Mrukowicz

dblp:273/7405 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
9since 2021 · last 2024
0000-0002-5348-8703ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 9 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2024 An Ensemble Classifier Based on kNN with an Interval Threshold Strategy
Urszula Bentkowska, Marcin Mrukowicz, Wojciech Galka, Karol Lech
ACIIDS (2)2
2024 Self-tuning framework to reduce the number of false positive instances using aggregation functions in ensemble classifier
abstract
In this contribution, the model which is dedicated to reducing the number of false positive instances is proposed. This is a self-tuning model using aggregation functions and time-series data periods. As a case study, the proposed model is tested in the context of phishing link detection. In the proposed model, well-known aggregation functions are applied to combine the confidence values of multiple Classification models for email phishing. The division of the dataset into multiple segments and subsets facilitates the implementation of incremental learning strategies. This approach enables the iterative enhancement of model performance through the training of new data while leveraging previously acquired knowledge. In our research, two datasets are considered, namely the existing PhiUSIIL phishing URL dataset as well as the dataset provided by the FreshMail company are applied. The proposed algorithm achieves a small number of expected false positives. This reduces the costs associated with manual analysis of such cases by domain experts (in the case of incorrect prediction as phishing mail).
Wojciech Galka, Jan G. Bazan, Urszula Bentkowska, Marcin Mrukowicz, Pawel Drygas, Marcin Ochab, Piotr Suszalski, Sebastian Obara
KES4
2024 Real-time anonymisation of DNS network traffic
abstract
The article covers the architecture developed to perform the anonymisation of active DNS traffic measurement. Data is anonymised on the fly using a proposed solution based on iptables. The proposed system was tested in the computer network of the University of Rzeszów. Furthermore, this solution could be applied to the bigger networks. The active DNS traffic based datasets are currently still rare and simultaneously they are valuable to researchers and IT admins. Collecting DNS data could be used in network monitoring tools, protecting the network from exfiltration and infiltration, detecting other malicious activities and finally performing machine learning and other research. Unfortunately, active DNS collected data contains information, which could affect the privacy of the computer network’s user. The anonymisation of the data is therefore necessary. Since the amount of DNS queries could be enormous the anonymisation system needs to be flexible and scallable. In this contribution, a system, which meets these requirements is proposed. Finally it is worth to note, that its complexity is moderate and based on open-source software.
Marcin Ochab, Marcin Mrukowicz, Jaromir Sarzynski, Piotr Molenda
KES2
2024 A scalable data acquisition system for the efficient processing of DNS network traffic
abstract
The article covers the architecture developed to efficiently collect large-volume DNS traffic. The resulting collected dataset can be utilised in a multitude of scenarios, including anomaly detection, machine learning, and network monitoring. The system enables the automatic enrichment of this data with information from external reputation databases and additional data, such as location and AS number. Data is anonymised on the fly using a proposed solution based on iptables. The pre-processed data is sent via the Kafka broker to a dedicated Clickhouse database, which allows for efficient analysis of this type of large data sets. The important aspect of the proposed DNS network acquisition system is that it is focused on active DNS measurement, which remains a relatively uncommon practice. It is essential that the detailed description of the proposed architecture is readily reproducible. The learning dataset obtained at the final stage allows for convenient querying and filtering using SQL. Furthermore, it is easily adaptable to work with a multitude of machine learning environments. The proposed modular system could be easily extended to other purposes, such as DNS traffic monitoring or generally collecting other network protocol data.
Marcin Ochab, Marcin Mrukowicz, Jaromir Sarzynski, Wojciech Rzasa
KES2
2024 Binary ensemble kNN based classifier for microarray datasets
abstract
In this contribution, the binary ensemble classifier based on kNN in the case of microarrays is discussed. There are considered diverse values of hyperparameters, such as the number of neighbors, the distance metric discussed, and other related types of hyperparameters. Moreover, interval modelling is used to improve the performance of the Classification. The algorithm is based on the interval-valued aggregation. To obtain several intervals, we use column-wise partitioning. Several kNN classifiers with different numbers of neighbors are used to compute certainty coefficients on each partition. Based on these values, intervals are determined. The interval-valued aggregation is treated as a hyperparameter of the model and is applied to combine obtained intervals. The performance of the proposed classifier is compared to the well-known ensemble classifier, which is a Bagging classifier. The results of the statistical tests prove that the proposed ensemble model may be effectively applied to the high-dimensional datasets, especially microarrays.
Aleksander Wojtowicz, Marcin Mrukowicz, Wojciech Galka, Krzysztof Balicki, Wojciech Rzasa, Urszula Bentkowska
KES2
2024 The effectiveness of aggregation functions used in fuzzy local contrast constructions
Barbara Pekala, Urszula Bentkowska, Michal Kepski, Marcin Mrukowicz
Fuzzy Sets Syst.4
2022 Comparison of aggregation classes in ensemble classifiers for high dimensional datasets
abstract
In the paper, we consider a combination of classifiers for the microarray datasets - examples of high dimensional datasets. Aggregation functions, in this case, are used to combine the output values of the constituent classifiers. Some known families of aggregation functions with several examples are studied and compared concerning their usefulness in the given classification method. Based on the proposed ranking of aggregations, the top aggregation functions considering their properties were explored to find which predispose to yield better classification results.
Jan G. Bazan, Stanislawa Bazan-Socha, Urszula Bentkowska, Wojciech Galka, Marcin Mrukowicz, Lech Zareba
FUZZ-IEEE5
2022 Interval modelling in optimization of k- N N classifiers for large number of attributes in data sets on an example of DNA microarrays
abstract
In this contribution there are considered interval-valued methods of improving the quality of classification by the k- N N binary classifiers in the case of large number of attributes in data sets. One of the possible applications of the introduced method are the microarray data so the presented results may be applied in medical diagnosis support, for example, in identification of marker genes. The proposed algorithm involving reduction methods of time complexity and interval modeling of data using interval-valued aggregation functions has a significantly higher classification quality than the other considered algorithm with the aggregation method involving the numerical arithmetic mean. The former algorithm is based on the interval-valued aggregation of uncertainty intervals determined on the basis of certainty coefficients of individual k- N N classifiers while the latter is based on the aggregation of certainty coefficients of individual k- N N classifiers with the use of the numerical arithmetic mean. Moreover, the performance of the new algorithm based on interval-valued methods was compared with the known from the literature classifiers in cancer diagnosis based on gene expression microarrays. The obtained results prove that in practical applications the newly proposed algorithm can be successfully used.
Urszula Bentkowska, Jan G. Bazan, Lech Zareba, Jerzy Socha, Stanislawa Bazan-Socha, Marcin Mrukowicz
Int. J. Intell. Syst.6
2021 Human- and Machine-Generated Traffic Distinction by DNS Protocol Analysis
abstract
In this contribution we analyze a real DNS traffic collected at the University of Rzeszów campus. All DNS queries and responses observed in the entire network were gathered. Data include traffic generated by students, scholars, and other staff members as well as servers, IoT and all other devices connected to network. Data was collected using the Tshark network protocol analyzer and stored in a ClickHouse columnar-oriented database dedicated for high volume data analyses. Fuzzy C-means clustering was applied to analyze DNS traffic and to distinguish between human- and machine generated traffic. Analysis was performed on a representative sample containing 3 516 094 records and 33 proposed features.
Marcin Ochab, Marcin Mrukowicz, Jaromir Sarzynski, Urszula Bentkowska
FUZZ-IEEE2
2020 Multi-class classification problems for the k-NN algorithm in the case of missing values
abstract
In this contribution methods for improving the quality of multi-class classification by the k nearest neighborhood classifiers in the case of large number of missing values in data sets are considered. Two versions of classifiers are compared. In the first case the aggregation of certainty coefficients of the individual classifiers with the use of the arithmetic mean is applied. In the second case interval modelling and interval-valued aggregation functions are involved. It is proved that the classifier which uses interval methods entails a much slower decrease in classification quality.
Urszula Bentkowska, Jan G. Bazan, Marcin Mrukowicz, Lech Zareba, Piotr Molenda
FUZZ-IEEE3
2020 New fuzzy local contrast measures: definitions, evaluation and comparison
abstract
In this contribution the concept of a local contrast of a fuzzy relation with the use of a consensus measure is introduced. A construction method of such local contrast using aggregation functions and fuzzy implications is considered. Other construction methods using similarity measure are also pointed out. Several examples of local contrasts are provided. The usability of introduced local contrast measures is evaluated by applying them in image processing for salient region detection.
Urszula Bentkowska, Michal Kepski, Marcin Mrukowicz, Barbara Pekala
FUZZ-IEEE3
2020 Application of similarity measures with uncertainty in classification methods
abstract
In this paper, the problem of measuring the degree of inclusion and similarity measure for interval-valued fuzzy sets is considered. We recall inclusion and similarity measures with uncertainty by using the partial or linear order on intervalvalued fuzzy sets. Moreover, we discuss an influence of inclusion and similarity measures with uncertainty to decision-making algorithm that uses those new measures.
Barbara Pekala, Ewa Rak, Dawid Kosior, Marcin Mrukowicz, Jan G. Bazan
FUZZ-IEEE4