Pawel Karczmarek

dblp:86/8717 · DBLP profile ↗
← Back
36ranked-venue papers
12as first author
23since 2021 · last 2025
0000-0002-6215-297XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 12 first-author · 19 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021
YearPublicationVenuePosition
2025 Smooth Ordered Weighted Averaging operators
Alicja Rachwal, Pawel Karczmarek, Albert Rachwal
Inf. Sci.2
2025 Rough set-inspired isolation forest
Albert Rachwal, Pawel Karczmarek, Alicja Rachwal
Inf. Sci.2
2025 Money Cannot Buy Happiness: Emotions in the IT Industry
abstract
The COVID-19 pandemic triggered a sudden shift toward remote and hybrid work, placing IT technologies at the forefront of organizational practices and revealing a spectrum of emotional responses among employees. While conventional wisdom suggests that high salaries in the IT industry safeguard well-being, this study challenges the notion that “money can't buy happiness” by demonstrating that technostress, social isolation, and “Zoom fatigue” persist regardless of income level. Drawing on a longitudinal dataset collected at three one-year intervals, the research employs fuzzy semantics to translate qualitative survey data into quantitative descriptors, basket analysis to identify consistent behavioral patterns, and a Minimal Spanning Tree-Based Isolation Forest enhanced by Takagi–Sugeno fuzzy rules to detect anomalies. The findings indicate a nuanced interplay of negative and positive emotions, with fear, anxiety, and fatigue frequently coexisting alongside pride, energy, and satisfaction. Some anomalies in responses reveal issues such as random guessing or minimal IT usage, underscoring the importance of filtering out low-quality data. Crucially, the emotional outcomes are influenced not only by pandemic-related disruption but also by psychosocial and organizational factors – such as role complexity, skill requirements, and managerial support – providing a holistic view of how and why IT-based work can foster both strain and fulfillment. These insights hold practical implications for employers, suggesting that comprehensive well-being strategies, rather than monetary incentives alone, are pivotal in promoting a healthier, more resilient IT workforce.
Adam Kiersztyn, Lukasz Galka, Krystyna Wojciechowska, Krystyna Kiersztyn, Agnieszka Rzepka, Kamil Jonak, Pawel Karczmarek
IEEE Trans. Fuzzy Syst.7
2024 Analysis of smooth and enhanced smooth quadrature-inspired generalized Choquet integral
Pawel Karczmarek, Adam Gregosiewicz, Zbigniew A. Lagodowski, Michal Dolecki, Lukasz Galka, Pawel Powroznik, Witold Pedrycz, Kamil Jonak
Fuzzy Sets Syst.1
2024 Deterministic attribute selection for isolation forest
Lukasz Galka, Pawel Karczmarek
Pattern Recognit.2
2023 Effective enhancement of isolation Forest method based on Minimal Spanning tree clustering
Lukasz Galka, Pawel Karczmarek, Mikhail Tokovarov
Inf. Sci.2
2023 Choquet Integral-Based Aggregation for the Analysis of Anomalies Occurrence in Sustainable Transportation Systems
abstract
Anomaly detection is one of the most important problems of modern data science due to the threat to the security of information systems as well as their users. This applies in particular to logistic data, which is used to predict costs, times, and organization of travel routes. Data anomalies may endanger the welfare and safety of transport users, goods, handling companies, and consumers. Moreover, they contribute to the overexploitation of the natural environment. Therefore, it is extremely important to find methods that are responsible for their effective detection. The desired approach may be the Choquet integral and its extensions, which in various applications have proven that with their help it is possible to efficiently increase the quality of the classification measured, for example, with the help of the accuracy. Due to the fact that the Choquet integral is resistant to data fluctuations and takes into account the quality (significance) of the information source, it appears to be an effective proposition for the final determination of what data, or more precisely, which records can be considered anomalous. The innovative approach to analyze transport data has not been used before. This article considers four publicly available databases covering different fields of application of transport systems. In a series of comprehensive numerical experiments, the Choquet integral-based approach has proven high efficiency for each of them. Moreover, we made a comparative analysis of the solutions before applying the Choquet integral and the results after its application.
Pawel Karczmarek, Lukasz Galka, Adam Kiersztyn, Michal Dolecki, Krystyna Kiersztyn, Witold Pedrycz
IEEE Trans. Fuzzy Syst.1
2022 On the Understanding of Anomalies in the Oculography Data and Their Classification with an Application of Fuzzy Aggregators
abstract
Modern medicine has been increasingly using information technologies and computer systems to improve decision-making processes. Important examples of such activities are the collection, processing, and analysis of oculographic data. The correct interpretation of the measurement results plays a very important role here. Medical examinations based on eye-tracking allow for early detection of many diseases or disfunctions of human brain. Moreover, it can be seen that implementation of such methods improves human-computer interfaces. In this study, we propose an innovative solution based on ten approaches of anomaly detection with the use of fuzzy aggregation of their results. Also, our experiment is supported by the analysis of obtained anomaly scores by the specialists in the field of medicine. The results of the automatic and expert evaluation show high potential of our method. There are also some differences in the perception of the anomaly by machine learning techniques and experts’ judgement. Hence, in this paper we try to understand and explain them.
Michal Dolecki, Pawel Karczmarek, Lukasz Galka, Malgorzata Plechawska-Wójcik, Monika Kaczorowska, Mikhail Tokovarov, Dariusz Czerwinski
FUZZ-IEEE2
2022 On the Detection of Anomalies with the Use of Choquet Integral and Their Interpretability in Motion Capture Data
abstract
Modern information technologies allow for the collection, processing, and data analysis. A very important role of these systems can be observed in the analysis of medical records, particularly in the analysis of motion capture data. Detection and interpretation of data collected from movement recording devices enables for a fast diagnosis, e.g. disease or misfunction. Moreover, it can be assistive in the introduction of appropriate treatment or rehabilitation. Anomaly detection methods play a key role in the evaluation of this type of medical research results. In this study, we introduce an innovative approach based on the aggregation of the results of eleven anomaly detection classifiers outcomes with the fuzzy Choquet integral. Furthermore, the results of numerical experiments are confronted with the assessments of the experts in medical field. The results show the great potential of our method in supporting the decision-making process based on the motion capture data analysis. Moreover, we have caught the differences between the expert's understanding of anomaly and the anomalies in data found by the modern machine learning methods.
Michal Dolecki, Pawel Karczmarek, Lukasz Galka, Magdalena Zawadka, Jakub Smolka, Maria Skublewska-Paszkowska, Edyta Lukasik, Pawel Powroznik, Piotr Gawda, Dariusz Czerwinski
FUZZ-IEEE2
2022 Quadrature-Inspired Generalized Choquet Integral
abstract
In this study, we present an innovative approach to deriving an aggregate classification score based on multiple classifiers based on generalizations of the Choquet integral. These generalizations are inspired by the quadratures known from numerical analysis, used to calculate integrals, e.g. the Newton-Cotes formula. The previous formulas for calculating generalizations of the Choquet integral used two (e.g. the case of the difference of these values), or three adjacent values related to the degrees of belonging of a given element to individual classes, related to one density of the fuzzy measure. In this article, we offer an interesting generalization. The novel enhancement is based on the replacement of typical product or t-norm appearing under the integral sign by forms related to mathematical quadratures. The formulas become more precise and better reflect the idea of integration. Moreover, a series of numerical experiments confirmed the advantage of the new approach over the existing ones.
Pawel Karczmarek, Michal Dolecki, Pawel Powroznik, Lukasz Galka, Witold Pedrycz, Dariusz Czerwinski
FUZZ-IEEE1
2022 Enhanced Tree-Based Anomaly Detection
abstract
Anomaly detection in data sets is one of the most important challenges for modern analysts and data administrators. It is usually based on algorithms that use raw data. In this study, we analyze the possibilities of improving the well-known Isolation Forest algorithm based on binary search trees for data preprocessing using the grouping of both attributes first, and then records within attribute groups. Attribute clustering is based on hierarchical grouping, while record grouping uses K-Means and Fuzzy C-Means. To describe the relationships between records, data membership functions are also used, built on the basis of record distances from centroids. This approach gives a new look at the possibilities of the Isolation Forest method and leads to a significant improvement in the results for selected public databases.
Pawel Karczmarek, Lukasz Galka, Michal Dolecki, Witold Pedrycz, Dariusz Czerwinski, Adam Kiersztyn, Rafal Stegierski
FUZZ-IEEE1
2022 Analysis of Sub-Integral Functions in the Aggregation of Classification Results Using Generalizations of the Choquet Integral on the Example of Emotion Classification
abstract
Speech emotion recognition is a complicated and challenging task in the human-computer interaction. Its performance depends on the relevance of the considered features in addition to the extent by which the speakers express their emotions. The well-known methods for determining the characteristics of a speech signal for the classification of emotional states are: The Mel Frequency Cepstral Coefficients, spectral methods like Short Time Fourier Transform and wavelet transform. However, the conventional approach to classifying emotions encounters problems, such as lack of sharp boundaries between separate emotional states, features correlation or expressing the same emotions differently depending on gender, culture orage. To prevent these issues, the fuzzy techniques such as an aggregation on a basis of the Choquet integral can be efficiently applied. Therefore, the main goal of this study is to determine the most efficient integrand of the popular Choquet integral generalizations. To find such a kind of functions, we run six separate tests for non-fuzzy classifiers: Multilayer Perceptron, Naive Bayes Network, Decision Trees, Probabilistic Neural Network, Random Forest and fuzzy ones: Fuzzy Multilayer Perceptron, Fuzzy Rule Classifier and Fuzzy Decision Trees. Additionally, the aggregation of the classification processes using 25 families of t-norms serving as integrands was performed. The results obtained in this extensive set of numerical experiments helped to establish the most effective operators which can be applied successfully in the multi-modal emotion recognition tasks. The used aggregation increased the emotions recognition accuracy for non-fuzzy classifiers from 11.10% to 15.50%, and for fuzzy classifiers from 3.31% to 6.00% depending on tested dataset.
Pawel Karczmarek, Pawel Powroznik, Maria Skublewska-Paszkowska, Slawomir Przylucki, Edyta Lukasik
FUZZ-IEEE1
2022 Aggregation of Tennis Groundstrokes on the Basis of the Choquet Integral and Its Generalizations
abstract
This paper focuses on the aggregation of two tennis groundstrokes: Forehand and backhand, which are performed in every match and training session. Recognition of tennis movements is very challenging task. Therefore it is important to correctly classify the proper patterns of various tennis players. This kind of study may be a part of training or development of a player’s skills. In this study, Support Vector Machines, Multilayer Perceptron, and Spatial-Temporal Graph Convolutional Neural Networks are applied to find a match and put the move trajectories into the proper classes. The images obtained from three dimensional data recorded using the Vicon optical motion capture system are the input data for the classifiers. The images containing forehand and backhand shots are divided into two phases: The preparation and the shot together with the racket swinging as the finishing element of the move. Due to the problems with strict classification of these data, the input is also fuzzy. In order to improve the accuracy of the forehand and backhand recognition, the non-typical Choquet-like integral aggregation functions are applied as well as traditional aggregation operators like median, or voting. In particular, the so-called pre-aggregation operators with specific t-norms, overlap functions and modification of the shape of the function under the integral sign give the best results, reaching accuracy at the level of 98.88%, which is better than the above mentioned individual classifiers.
Maria Skublewska-Paszkowska, Pawel Powroznik, Pawel Karczmarek, Edyta Lukasik
FUZZ-IEEE3
2022 A probabilistic generalization of isolation forest
abstract
The problem of finding anomalies and outliers in datasets is one of the most important challenges of modern data analysis. Among the commonly dedicated tools to solve this task one can find Isolation Forest (IF) that is an efficient, conceptually simple, and fast method. In this study, we propose the Probabilistic Generalization of Isolation Forest (PGIF) that is an intuitively appealing and efficient enhancement of the original approach. The proposed generalization is based on nonlinear dependence of segment-cumulated probability from the length of segment. Introduction of the generalization allows to achieve more effective splits that are rather performed between the clusters, i.e. regions where datapoints constitute dense formations and not through them. In a comprehensive series of experiments, we show that the proposed method allows us to detect anomalies hidden between clusters more effectively. Moreover, it is demonstrated that our approach favorably affects the quality of anomaly detection in both artificial and real datasets. In terms of time complexity our method is close to the original one since the generalization is related only to the building of the trees while the scoring procedure (which takes the main time) is kept unchanged.
Mikhail Tokovarov, Pawel Karczmarek
Inf. Sci.2
2022 Detection and Classification of Anomalies in Large Datasets on the Basis of Information Granules
abstract
Anomaly (outlier) detection is one of the most important problems of modern data analysis. The sources of anomalies are varying. They can be the results of database users’ mistakes, operational errors, or just missing values. The problem is very important because of the fast growth of large datasets. Therefore, in this article, we present detailed results of work on the concept of granular computing-based approach to anomaly detection, classification, and gradation. The aim of the study is to introduce an innovative solution that allows the use of information granules to identify and classify anomalies. The novelty of the proposed solution consists in the use of fuzzy semantics implied by the statistical properties of the data considered. Moreover, instead of the classic approach to detecting anomalies in the data, it is proposed to determine the degree of anomaly for the data transformed to the new resulting state space. Thanks to the use of an innovative approach using the universal descriptor space, it is possible to determine the degree of anomaly, and by using various aggregation methods one can also specify its type.
Adam Kiersztyn, Pawel Karczmarek, Krystyna Kiersztyn, Witold Pedrycz
IEEE Trans. Fuzzy Syst.2
2021 Fuzzy Extensions of Isolation Forests for Road Anomaly Detection
abstract
In the presented paper the authors are showing the usage of fuzzy extensions of isolations forests for detecting road anomalies like potholes. Using the data acquired by the accelerometer in the smartphone and the proper smartphone application, the vibrations while driving over road were analyzed using multiple variants of extended isolation forests - n-ary (NIF), with fuzzy membership function (MIF), with k-means clustering (KIF), with two fuzzy clusters incorporated (CIF) or two fuzzy clusters and the distance to the cluster center (prototype) utilized (C2DIF). The presented research shows that in comparison to the state-of-the-art methods previously discussed by the authors, the accuracy and false positive rate have improved, while the sensitivity has been improved to reach 100%.
Marcin Badurowicz, Pawel Karczmarek, Jerzy Montusiewicz
FUZZ-IEEE2
2021 Influence of the Fuzzy Robust Gamma Rank Correlation, Fuzzy C-Means, and Fuzzy Cognitive Maps to Predict the Z Generation's Acceptance Attitudes Towards Internet Health Information
abstract
In this study, we propose an approach based on the advanced fuzzy techniques such as Fuzzy C-Means, Fuzzy Robust Gamma Rank Correlation and Fuzzy Cognitive Maps to predict the acceptance attitudes towards Internet health information. To improve the Fuzzy Cognitive Maps efficiency we introduce the setting the values of initial matrix with the use of fuzzy methods and to divide the concepts based on the clustering methods. This allows us to use maps as a tool for prediction the acceptance attitudes of the young people in the area of heath information management. Moreover, this work sheds the light on the novel application of both Fuzzy C-Means and Fuzzy Robust Gamma Rank Correlation as tools for settings the initial values of connections between concepts for Fuzzy Cognitive Maps.
Dariusz Czerwinski, Magdalena Czerwinska, Pawel Karczmarek, Adam Kiersztyn
FUZZ-IEEE3
2021 K-Medoids Clustering and Fuzzy Sets for Isolation Forest
abstract
Capturing anomalies in data is one of the most important problems in modern data analysis. In recent years, scientists have developed many interesting approaches. One of the leading is the Isolation Forest method, which is based on searching a forest of binary trees. This method is extremely effective, especially in the case of relatively small databases. Despite of that, a lot of work has been done for years to improve it. For instance, variants based on rotation or fuzzy sets were developed. In this paper, we propose a very effective method of building search trees based on grouping data using the K-Medoids method. The results of the conducted experiments suggest a significant improvement in the quality of the method in relation to the original Isolation Forest.
Pawel Karczmarek, Adam Kiersztyn, Witold Pedrycz, Marcin Badurowicz, Dariusz Czerwinski, Jerzy Montusiewicz
FUZZ-IEEE1
2021 Classification of Complex Ecological Objects with the Use of Information Granules
abstract
The selection of an appropriate method of data analysis is a key problem for researchers from various fields of applications. They consider different methods of data classification, often based on the thematic scope of the data at their disposal. However, various data characteristics, such as data set size, data type and quality, gaps, outliers and other anomalies, can make proper selection significantly difficult. Therefore, in this study we propose a method based on a very universal classifier designed on the basis of calculations using information granules. The main objective of the work is to present and comprehensively verify the effectiveness of the classifier. As an example of application, we propose complicated yet currently important data coming from widely understood ecological research. Detailed numerical experiments indicate the high efficiency of the proposed method and the possibility of easy application to data appearing in other fields. In addition, various types of aggregation functions of the classification results are considered in order to obtain the most reliable results for the discussed problems,
Adam Kiersztyn, Krystyna Kiersztyn, Pawel Karczmarek, Marek Kaminski, Ignacy Kitowski, Adam Zbyryt, Rafal Lopucki, Grzegorz Pitucha, Witold Pedrycz
FUZZ-IEEE3
2021 The Concept of Granular Representation of the Information Potential of Variables
abstract
With the advent of research into Granular Computing, in particular information granules, the way of thinking about data has changed gradually. Researchers and practitioners do not consider only their specific properties, but also try to look at the data in a more general way, closer to the way people think. This kind of knowledge representation is expressed particularly in approaches based on linguistic modeling or fuzzy techniques such as fuzzy clustering, but also newer approaches related to the explanation of how artificial intelligence works on these data (so-called explainable artificial intelligence). Therefore, especially important from the point of view of the methodology of data research is an attempt to understand their potential as information granules. Such a kind of approach to data presentation and analysis may introduce considerations of a higher, more general level of abstraction, while at the same time reliably describing the network of relationships between the data and the observed information granules. In this study, we tackle this topic with particular emphasis on the problem of choosing a predictive model. In a series of numerical experiments based on both artificially generated data, ecological data on changes in bird arrival dates in the context of climate change, and COVID-19 infections data we demonstrate the effectiveness of the proposed approach built with a novel application of information potential granules.
Adam Kiersztyn, Pawel Karczmarek, Krystyna Kiersztyn, Rafal Lopucki, Stanislaw M. Grzegórski, Witold Pedrycz
FUZZ-IEEE2
2021 A Comprehensive Analysis of the Impact of Selecting the Training Set Elements on the Correctness of Classification for Highly Variable Ecological Data
abstract
Classification of objects in empirical data, especially in biological sciences, is a very complex process and has been a big challenge for researchers who do not specialize in data analysis. Therefore, in this study, we present a comprehensive summary of selected classifiers operating on both exact and fuzzy numbers. The results of performance of specific classifiers are compared on the example of a unique set of empirical data on changes in the behavior of animals in response to environmental factors. This is one of the key challenges in ecological research and it is strictly related to ecosystem changes caused by climate change. Nowadays, changes in behavior are a very popular topic of research because as a result of the COVID-19 pandemic and lower activity of people (lockdown effect). Therefore, various unusual reactions of wild animals were found around the world. A detailed compilation of research results, shortcomings, and strengths of various classification methods may be a compendium of knowledge for biologists and other practitioners as well as researchers working with empirical data.
Adam Kiersztyn, Rafal Lopucki, Krystyna Kiersztyn, Pawel Karczmarek, Pawel Powroznik, Dariusz Czerwinski, Witold Pedrycz
FUZZ-IEEE4
2021 Tennis Multivariate Time Series Clustering
abstract
In tennis there are two basic shots (forehand and backhand), which are two of key elements to win points. Sophisticated equipments, such as motion capture systems, enable one to record both the tennis player's movements and tennis racket. The 3D data may be used to define the perfect shot model or to give the directions to the player how to reach to this model and which aspects of impact should be improved. Clustering analysis can result in understanding the phases of a tennis player move and, as a consequence, the improvement of his/her play. Using the memberships obtained in the fuzzy clustering process one can evaluate the quality of a player's move and potentially estimate the player's progress. The main objective of this study is to apply the fuzzy c-means algorithm utilizing the dynamic time warping-based distance to cluster analysis of tennis shots. Both shots were taken into the consideration. The analysis consists of forty moves. Based on the 3D data of the tennis racket positions, the clustering was performed for subsequent two, three, and four clusters. The obtained results clearly show that clustering with two clusters is the most appropriate for analysing tennis shots. The model of a perfect shot was obtained. It is universal and does not depend on the player's height. Based on the model, it is possible to deduce technical differences in the players' shots. This analysis gives the directions for improvements of the shot technique. The advantage of the clustering of our approach is that we can get information to what degree the athlete should still correct his/her shots. The information is given to what extent the stroke is correct in relation to the ideal model.
Maria Skublewska-Paszkowska, Pawel Karczmarek, Edyta Lukasik
FUZZ-IEEE2
2021 Fuzzy Analytic Hierarchy Process Based on Graphical Components
abstract
Many decision-making problems are solved using well-known Analytic Hierarchy Process (AHP). It is a tool to prioritize, rank, or choose from a set of alternatives. Also, its fuzzy version has been widely propagated and popular as well as still examined in a research community. However, practical application of the fuzzy AHP may bring some difficulties since a user (expert) should provide many values forming a membership function to precisely select a part of a range of a scale used in a pairwise comparison process. Moreover, using numerical or linguistic values can bring some limitation to the users, in particular, those who are not experienced with the scale. Therefore, in this paper, we present a comfortable approach to the fuzzy Analytic Hierarchy Process based on an application of a well-known Graphical User Interface programming component such as multirange slider. We thoroughly examine its application in contrast to other methods based on a numerical scale. In particular, we analyze the influence of an application of the multirange slider on the consistency of experts' opinions.
Michal Walczuk, Pawel Karczmarek
FUZZ-IEEE2
2020 Detection of Road Artefacts Using Fuzzy Adaptive Thresholding
abstract
In this paper the authors are proposing approximate method for road artefacts detection and their location by analyzing acceleration values recorded in the car during driving over the road fragment using the smartphone mounted in the car. The new method called F-THRESH has been introduced, which is adaptively adjusting threshold for road artefacts detection by the fuzzy system means, allowing for outlier detection in chaotic time streams. First, the road quality is being calculated, then the difference between the current data point and mean acceleration is calculated and those two values are used as the input for the fuzzy system, which is calculating threshold to classify data point as an outlier. The proposed method has been compared to the previously implemented method and has an accuracy over 94% with 1.3% of False Positive Rate for the same problem which makes it a great candidate to be implemented in the IoT Edge scenarios, for reducing amount of data being sent to the cloud analyzing system.
Marcin Badurowicz, Jerzy Montusiewicz, Pawel Karczmarek
FUZZ-IEEE3
2020 An Application of Fuzzy C-Means, Fuzzy Cognitive Maps, and Fuzzy Rules to Forecasting First Arrival Date of Avian Spring Migrants
abstract
In this study, we propose an approach based on the advanced fuzzy techniques such as Fuzzy C-Means and Fuzzy Cognitive Maps to cluster the birds species, based on the information of first arrival date, into more coherent and uniform groups. The birds are very suitable subject for modelling the climate changes. Very popular indicator to forecast bird migration dynamic is the first arrival date. In many reported studies, this indicator is shown as very useful. However, there is still a lack of precise methods grouping the birds into the classes in satisfying manner producing detailed information about species and the relations between them. As evidenced in the experimental series section, the proposed approach enables the researchers and practitioners working with that important area of ecology to observe subtle dependencies between various bird species. Moreover, this work sheds the light on the novel application of both Fuzzy C-Means and Fuzzy Cognitive Maps as the efficient tools to analyse the ecological data collected in changing climatic environment.
Dariusz Czerwinski, Adam Kiersztyn, Rafal Lopucki, Pawel Karczmarek, Ignacy Kitowski, Adam Zbyryt
FUZZ-IEEE4
2020 Fuzzy Set-Based Isolation Forest
abstract
One of the main challenges is the analysis of large data sets, in particular those containing various types of data, such as time, place, image, and those assuming categorical values. This type of data may contain numerous outliers. Despite the continuous development of data analysis, many methods can be effectively improved, in particular through the use of efficient solutions based on fuzzy set technologies. In this paper, we analyze the improvement of a well-known method, i.e. Isolation Forest, for which we introduce an innovative modification, referred to as the Fuzzy Set-Based Isolation Forest.
Pawel Karczmarek, Adam Kiersztyn, Witold Pedrycz
FUZZ-IEEE1
2020 The Assessment of Importance of Selected Issues of Software Engineering, IT Project Management, and Programming Paradigms Based on Graphical AHP and Fuzzy C-Means
abstract
In this study, we present the results of surveys conducted in a group of employees and students of IT faculties presenting the answers to the most important, in our opinion, issues related to software engineering (SE), IT project management, and programming paradigms. The above topics are chosen because of their high relevance to the professional community. The participants taking part in the experiments quantified their input through the process of pairwise comparisons (a so-called Analytic Hierarchy Process, AHP) using an innovative highly interactive approach based on a graphic communication means. The generic AHP method was augmented by the optimization mechanisms delivered by the Particle Swarm Optimization (PSO) in order to deliver the highest possible consistency of responses of the participants. Moreover, we demonstrate a method based on Fuzzy C-Means (FCM) filtering highly inconsistent and unreal experts' assessments. In a series of experiments, we demonstrate the accuracy and stability of the AHP method based on graphical environment. We discuss two variants of aggregation of experts' opinions according to their level of experience in the field of interest. Finally, we show the efficiency of the FCM as the method of preselection of experts' evaluations.
Pawel Karczmarek, Witold Pedrycz, Dariusz Czerwinski, Adam Kiersztyn
FUZZ-IEEE1
2020 The Concept of Detecting and Classifying Anomalies in Large Data Sets on a Basis of Information Granules
abstract
Anomaly (outlier) detection is one of the most important problems of modern data analysis. Anomalies can be the results of database users' mistakes, operational errors or just missing values. The problem is important because of fast growth of the large data sets. Therefore, we present the initial results of work on a Granular Computing approach to data imputation and missing data analysis. Our proposal brings intuitive and interpretable solutions. Finally, in a series of experiments, we demonstrate its effectiveness for a large dataset in the area of transport.
Adam Kiersztyn, Pawel Karczmarek, Krystyna Kiersztyn, Witold Pedrycz
FUZZ-IEEE2
2020 Data Imputation in Related Time Series Using Fuzzy Set-Based Techniques
abstract
One of the main challenges faced by people who use data from empirical research in their work is missing data. In many scientific disciplines and industries there are references to time series. The suitability of several methods to imputation of the missing data in the study of mutual links between the analysed time series have been presented and tested in this work. In this paper, known methods of supplementing data in time series were enriched by the use of fuzzy sets and their processing was tested on unique data from experimental research and a transport company database. Fuzzy linguistic descriptors-based methods of missing data imputation in databases containing time series are discussed. The proposed method has a high efficiency, which have been proven in a series of experiments with both artificial and real datasets. The proposed methodologies have been tested on theoretical example and empirical data sets from various fields: (1) ecological data on changes in bird arrival dates in the context of climate change and (2) data describing the transport of containers between ports on the Mediterranean. Moreover, an important novelty of this work is, in particular, an application of fuzzy techniques to the correction of the datasets containing bird migration descriptions.
Adam Kiersztyn, Pawel Karczmarek, Rafal Lopucki, Witold Pedrycz, Ebru Al, Ignacy Kitowski, Adam Zbyryt
FUZZ-IEEE2
2020 K-Means-based isolation forest
abstract
The task of anomaly detection in data is one of the main challenges in data science because of the wide plethora of applications and despite a spectrum of available methods. Unfortunately, many of anomaly detection schemes are still imperfect i.e., they are not effective enough or act in a non-intuitive way or they are focused on a specific type of data. In this study, the classical method of Isolation Forest is thoroughly analyzed and augmented by bringing an innovative approach. This is k-Means-Based Isolation Forest that allows to build a search tree based on many branches in contrast to the only two considered in the original method. k-Means clustering is used to predict the number of divisions on each decision tree node. As supported through experimental studies, the proposed method works effectively for data coming from various application areas including intermodal transport and geographical, spatio-temporal data. In addition, it enables a user to intuitively determine the anomaly score for an individual record of the analyzed dataset. The advantage of the proposed method is that it is able to fit the data at the step of decision tree building. Moreover, it returns more intuitively appealing anomaly score values.
Pawel Karczmarek, Adam Kiersztyn, Witold Pedrycz, Ebru Al
Knowl. Based Syst.1
2017 An application of chain code-based local descriptor and its extension to face recognition
Pawel Karczmarek, Adam Kiersztyn, Witold Pedrycz, Michal Dolecki
Pattern Recognit.1
2017 A study in facial features saliency in face recognition: an analytic hierarchy process approach
abstract
In this study, we develop a process of estimation of importance of features considered in face recognition by making use of the analytic hierarchy process (AHP). The AHP method of pairwise comparisons realized at three levels of hierarchy becomes crucial to realize a comprehensive weighting of cues so that sound estimates of weights associated with the individual features of faces can be formed. We demonstrate how to carry out an efficient process of face description by using a collection of linguistic descriptors of the features and their groups. Numerical dependencies between the features are quantified with the help of experienced criminology and psychology experts. Finally, we present an entropy-based method of evaluation of the relevance of the estimation process completed by the individuals. The intuitively appealing results of experiments are presented and analyzed in detail.
Pawel Karczmarek, Witold Pedrycz, Adam Kiersztyn, Przemyslaw Rutka
Soft Comput.1
2016 Linguistic descriptors and fuzzy sets in face recognition realized by humans
abstract
In this study, we present a new approach to the face retrieval and face classification problem, which exploits available expert's knowledge and introduces a novel way of describing facial features. These features are described by manually assigned weights corresponding to membership grades with respect to the linguistic descriptors such as short, medium, or long. In the series of experiments, we also use weights produced by the Analytic Hierarchy Process aimed at producing saliences of facial cues. We identify a group of the most essential facial features.
Adam Kiersztyn, Pawel Karczmarek, Michal Dolecki, Witold Pedrycz
FUZZ-IEEE2
2016 Face recognition by humans performed on basis of linguistic descriptors and neural networks
abstract
In this study, we present a new approach to the problem of face classification, which relies on the linguistic description of the facial features. In this method, face descriptors are represented through the Analytic Hierarchy Process (AHP) and formalized as information granules. Moreover, neural networks are used to construct efficient classifiers. Furthermore, with usage of AHP we realize a transition from the linguistic description of the facial features to the vectors of numbers that are used by a neural network in the process of matching faces. The results of experiments demonstrate the potential applicability of our proposal to the forensic investigations. Finally, discussed are important aspects of constructing neural networks regarded as a vehicle to perform classification process.
Michal Dolecki, Pawel Karczmarek, Adam Kiersztyn, Witold Pedrycz
IJCNN2
2014 A study in facial regions saliency: a fuzzy measure approach
abstract
People recognize familiar faces in a similar way by using interior facial features (facial regions) such as eyes, nose, mouth, etc. However, the importance of these regions in the realization of face identification and a quantification of the impact of such regions on the recognition process could vary from one region to another. An intuitively appealing observation is that of monotonicity: the more regions are taken into account in the recognition process, the better. From a formal point of view, the relevance of the facial regions and an aggregation of these pieces of experimental evidence can be described in the formal setting of fuzzy measures. Fuzzy measures are of particular interest with this regard given their monotonicity property (which stands in a clear contrast with the more restrictive additivity property inherent to probability–like measures). In this study, we concentrate on the construction of fuzzy measures (more specifically, $$ \lambda $$ λ -fuzzy measure) and characterize their performance in the problem of face recognition using a collection of experimental data.
Pawel Karczmarek, Witold Pedrycz, Marek Z. Reformat, Elaheh Akhoundi
Soft Comput.1
2013 Local descriptors in application to the aging problem in face recognition
Michal Bereta, Pawel Karczmarek, Witold Pedrycz, Marek Z. Reformat
Pattern Recognit.2