Pawel Ksieniewicz

dblp:145/6756 · DBLP profile ↗
← Back
33ranked-venue papers
12as first author
22since 2021 · last 2026
0000-0001-9578-8395ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 9 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Structuring the processing frameworks for data stream evaluation and application
abstract
The following work addresses the problem of frameworks for data stream processing that can be used to evaluate the solutions in an environment that resembles real-world applications. The definition of structured frameworks stems from the need to reliably assess data stream classification methods, considering the constraints of delayed label access, the costs of their acquisition and the costs of model adaptation. The current experimental evaluation often boundlessly exploits the assumption of the immediate label access to monitor the recognition quality and adapt the methods to the changing concepts. The problem is leveraged by reviewing currently described methods and techniques for data stream processing and verifying their outcomes in simulated environment . This work defines a taxonomy of data stream processing frameworks and presents four processing schemes that link the tasks of drift detection and classification while considering a natural phenomenon of label delay . The presented research shows that classification quality is significantly affected not only by the disruptive phenomenon of concept drifts and label delay , but also by the undertaken processing scheme that describes the flow of labels in the recognition system. Considering a specific processing framework depending on real-world constraints proves to be a critical aspect of reliable and realistic experimental evaluation.
Joanna Komorniczak, Pawel Ksieniewicz, Pawel Zyblewski
Pattern Recognit.2
2025 Taking class imbalance into account in open set recognition evaluation
Joanna Komorniczak, Pawel Ksieniewicz
Neural Comput. Appl.2
2024 Involving Society to Protect Society from Fake News and Disinformation: Crowdsourced Datasets and Text Reliability Assessment
Gracjan Katek, Marta Gackowska, Joanna Komorniczak, Pawel Ksieniewicz, Rafal Kozik, Marek Pawlicki, Michal Choras
ACIIDS (2)4
2024 torchosr - A PyTorch extension package for Open Set Recognition models evaluation in Python
abstract
The article presents the torchosr module – a Python package compatible with PyTorch library – offering functionality and models dedicated to Open Set Recognition in Deep Neural Networks. Included software offers two frequently used base recognition methods in the field and a set of functions for handling datasets, enabling the generation of derived datasets, where some classes are considered unknown and used only in the testing process. Code base is enhanced with a set of helper functions, facilitating model validation process. The main goal of the proposal is to simplify and promote the correct experimental evaluation, where experiments are carried out on a large number of derivative sets with various Openness, related to the cardinality of known and unknown classes, and class-to-category assignments. The authors hope that methods available in the package will become a source of a correct and open-source implementation of the relevant baseline and state-of-the-art solutions in the domain.
Joanna Komorniczak, Pawel Ksieniewicz
Neurocomputing2
2024 Distance profile layer for binary classification and density estimation
Joanna Komorniczak, Pawel Ksieniewicz
Neurocomputing2
2024 Towards explainable fake news detection and automated content credibility assessment: Polish internet and digital media use-case
Rafal Kozik, Gracjan Katek, Marta Gackowska, Sebastian Kula, Joanna Komorniczak, Pawel Ksieniewicz, Aleksandra Pawlicka, Marek Pawlicki, Michal Choras
Neurocomputing6
2024 On metafeatures' ability of implicit concept identification
abstract
Abstract Concept drift in data stream processing remains an intriguing challenge and states a popular research topic. Methods that actively process data streams usually employ drift detectors, whose performance is often based on monitoring the variability of different stream properties. This publication provides an overview and analysis of metafeatures variability describing data streams with concept drifts. Five experiments conducted on synthetic, semi-synthetic, and real-world data streams examine the ability of over 160 metafeatures from 9 categories to recognize concepts in non-stationary data streams. The work reveals the distinctions in the considered sources of streams and specifies 17 metafeatures with a high ability of concept identification.
Joanna Komorniczak, Pawel Ksieniewicz
Mach. Learn.2
2023 Neural network architecture with intermediate distribution-driven layer for classification of multidimensional data with low class separability
abstract
Abstract Simple neural network classification tasks are based on performing extraction as transformations of the set simultaneously with optimization of weights on individual layers. In this paper, the Representation 7 architecture is proposed, the primary assumption of which is to divide the inductive procedure into separate blocks – transformation and decision – which may lead to a better generalization ability of the presented model. Architecture is based on the processing context of the typical neural network and unifies datasets into a shared, generically sampled space. It can be applicable in the case of difficult problems – defined not as imbalance or streaming data but by low-class separability and a high dimensionality. This article has tested the hypothesis that – in such conditions – the proposed method could achieve better results than reference algorithms by comparing the R7 architecture with state-of-the-art methods, raw mlp and Tabnet architecture. The contributions of this work are the proposition of the new architecture and complete experiments on synthetic and real datasets with the evaluation of the quality and loss achieved by R7 and by reference methods.
Weronika Borek-Marciniec, Pawel Ksieniewicz
Appl. Intell.2
2023 Processing data stream with chunk-similarity model selection
Pawel Ksieniewicz
Appl. Intell.1
2023 Alphabet Flatting as a variant of n-gram feature extraction method in ensemble classification of fake news
Pawel Ksieniewicz, Pawel Zyblewski, Weronika Borek-Marciniec, Rafal Kozik, Michal Choras, Michal Wozniak 0001
Eng. Appl. Artif. Intell.1
2023 problexity - An open-source Python library for supervised learning problem complexity assessment
Joanna Komorniczak, Pawel Ksieniewicz
Neurocomputing2
2023 Complexity-based drift detection for nonstationary data streams
abstract
This publication presents the Complexity Drift Detector (C2D) – the method for detecting a concept shift in the data stream based on the classification task complexity measures. The method belongs to the group of detectors agnostic to the recognition quality of the base classifier. The possibility of selecting a set of difficulty measures taken into account during the data stream processing allows applying the method to many tasks in which the detection of a classification task complexity change is expected. The publication includes experiments analyzing the hyperparameters’ influence on the operation of the method and a broad comparative experiment comparing the proposed algorithm with state-of-the-art solutions. The experiments were carried out on synthetic data streams of different dimensions and with different concept drift characteristics, also presenting the effects of processing real-world data streams. The results of the conducted research confirm the high efficiency of the method in detecting concept changes, sensitive not only to the fact of drift occurrence but also to its dynamics.
Joanna Komorniczak, Pawel Ksieniewicz
Neurocomputing2
2023 Active Weighted Aging Ensemble for drifted data stream classification
abstract
Purpose One of the significant problems in data stream classification is the concept drift phenomenon, which consists of the change in probabilistic characteristics of the classification task . Such changes in posterior probability destabilize the classification model performance, seriously degrading its quality. It is necessary to design appropriate strategies to counteract this phenomenon, allowing the classifier to adapt to the changing probabilistic characteristics. It is tough to propose such an approach with limited access to data labels. A human bias of high quality is usually costly, so to minimize the expenses related to this process, it is also necessary to propose learning strategies based on semi-supervised learning. Such strategies employ active learning methods indicating which of the incoming objects are valuable to be labeled for improving the classifier's performance. Methods This paper proposes Active Weighted Aging Ensemble algorithm, a novel chunk-based method for non-stationary data stream classification. It employs a classifier ensemble approach and utilizes the changing ensemble lineup to react to concept drift appropriately. It also proposed a new active learning method, considering a limited budget that may be applied to any data stream classifier. Results AWAE has been evaluated through computer experiments using real and synthetic data streams, confirming the proposed algorithm's high quality over state-of-the-art methods. Conclusion The research conducted on benchmark data streams confirmed the effectiveness of the proposed solution and highlighted its strengths in comparison with state-of-the-art methods. The estimated computational complexity is acceptable and comparable to the benchmark algorithms.
Michal Wozniak 0001, Pawel Zyblewski, Pawel Ksieniewicz
Inf. Sci.3
2022 Feature Integration Strategies for Multilingual Fake News Classification
abstract
The abundance of information in digital media, which in today’s world is the main source of knowledge about current events for the masses, makes it possible to spread disinformation on a larger scale than ever before. Consequently, there is a need to develop novel fake news detection approaches capable of adapting to changing factual contexts and generalizing previously or concurrently acquired knowledge. To deal with this problem, we propose a ensemble-based approach, which allows for fake news detection in multiple languages and the mutual transfer of knowledge acquired in each of them. Both classical feature extractors, such as Term frequency-inverse document frequency or Latent Dirichlet Allocation, and integrated deep NLP (Natural Language Processing) BERT (Bidirectional Encoder Representations from Transformers) models paired with MLP (Multilayer Perceptron) classifier, were employed. The results of experiments conducted on two datasets dedicated to the fake news classification t ask ( in English and Spanish, respectively), supported by statistical analysis, confirmed that utilization of additional languages could improve performance for traditional methods. Also, in some cases supplementing the deep learning method with classical ones can positively impact obtained results. The ability of models to generalize the knowledge acquired between the analyzed languages was also observed.
Jedrzej Kozal, Michal Les, Pawel Zyblewski, Pawel Ksieniewicz, Michal Wozniak 0001
IEEE Big Data4
2022 Data stream generation through real concept's interpolation
abstract
Among the recently published works in the field of data stream analysis -both in the context of classification task and concept drift detection -the deficit of real-world data streams is a recurring problem.This article proposes a method for generating data streams with given parameters based on real-world static data.The method uses onedimensional interpolation to generate sudden or incremental concept drifts.The generated streams were subjected to an exemplary analysis in the concept drift detection task with a detector ensemble.The method can potentially contribute to the development of methods focused on data stream processing.
Joanna Komorniczak, Pawel Ksieniewicz
ESANN2
2022 Inductive Parallel Learning for Multiple Classification Problems
abstract
The default approach to the construction of pattern recognition models solving supervised learning problems is fitting the knowledge representation to the samples of a single problem, describing a single, static concept. In recent years, however, the transfer learning approach has been gaining popularity, consisting in training a model initially learned on a different, standard dataset to a specific problem. However, such solutions are mainly used for signal data, in applications typical of deep learning models. This paper proposes a new training procedure for tabular data, in which a serial ensemble of models based on neural networks is trained in parallel to recognize two different problems. In extensive experimental analysis, the approaches based on disjoint models, disjoint learning and - constituting the basic proposal of the work - inductive parallel learning of two problems were compared. The analysis carried out on synthetic and real-world problems shows that the proposed approach is a promising starting point for further research.
Weronika Borek-Marciniec, Pawel Ksieniewicz
IJCNN2
2022 Imbalanced Data Stream Classification Assisted by Prior Probability Estimation
abstract
With the processing of data streams, come inevitable challenges, such as changes in the prior (class drift) and posterior (concept drift) probability distribution over the processing time. Both these phenomena have a negative impact on the quality of the classification. Heavily imbalanced problems, which are often typical for real-world applications, bring additional processing difficulties. Classifiers are often biased towards the majority class and have difficulty identifying instances of categories described with a lower number of objects. The following article proposes a Prior Probability Assisted Classifier (2PAC), a method aiming to improve the classification quality of heavily imbalanced data streams with dynamic changes by using the estimated prior probability value and the correction of the classifier's decision for batch predictions. Presented extensive computer experiments, supported by statistical analysis, show the ability to improve the classification quality using the proposed method.
Joanna Komorniczak, Pawel Zyblewski, Pawel Ksieniewicz
IJCNN3
2022 Stream-learn - open-source Python library for difficult data stream batch analysis
Pawel Ksieniewicz, Pawel Zyblewski
Neurocomputing1
2022 Statistical Drift Detection Ensemble for batch processing of data streams
Joanna Komorniczak, Pawel Zyblewski, Pawel Ksieniewicz
Knowl. Based Syst.3
2021 Prior Probability Estimation in Dynamically Imbalanced Data Streams
abstract
Despite the fact that real-life data streams may often be characterized by the dynamic changes in the prior class probabilities, there is a scarcity of articles trying to clearly describe and classify this problem as well as suggest new methods dedicated to resolving this issue. The following paper aims to fill this gap by proposing a novel data stream taxonomy defined in the context of prior class probability and by introducing the Dynamic Statistical Concept Analysis (DSCA) - prior probability estimation algorithm. The proposed method was evaluated using computer experiments carried out on 100 synthetically generated data streams with various class imbalance characteristics. The obtained results, supported by statistical analysis, confirmed the usefulness of the proposed solution, especially in the case of discrete dynamically imbalanced data streams (DDIS).
Joanna Komorniczak, Pawel Zyblewski, Pawel Ksieniewicz
IJCNN3
2021 The prior probability in the batch classification of imbalanced data streams
Pawel Ksieniewicz
Neurocomputing1
2021 Fusion of linear base classifiers in geometric space
Pawel Ksieniewicz, Pawel Zyblewski, Robert Burduk
Knowl. Based Syst.1
2020 Fake News Detection from Data Streams
abstract
Using fake news as a political or economic tool is not new, but the scale of their use is currently alarming, especially on social media. The authors of misinformation try to influence the users' decisions, both in the economic and political sphere. The facts of using disinformation during elections are well known. Currently, two fake news detection approaches dominate. The first approach, so-called fact or news checker, is based on the knowledge and work of volunteers, the second approach employs artificial intelligence algorithms for news analysis and manipulation detection. In this work, we will focus on using machine learning methods to detect fake news. However, unlike most approaches, we will treat incoming messages as stream data, taking into account the possibility of concept drift occurring, i.e., appearing changes in the probabilistic characteristics of the classification model during the exploitation of the classifier. The developed methods have been evaluated based on computer experiments on benchmark data, and the obtained results prove their usefulness for the problem under consideration. The proposed solutions are part of the distributed platform developed by the H2020 SocialTruth project consortium.
Pawel Ksieniewicz, Pawel Zyblewski, Michal Choras, Rafal Kozik, Agata Gielczyk, Michal Wozniak 0001
IJCNN1
2019 SMOTE Algorithm Variations in Balancing Data Streams
Bogdan Gulowaty, Pawel Ksieniewicz
IDEAL (2)2
2019 A Genetic-Based Ensemble Learning Applied to Imbalanced Data Classification
Jakub Klikowski, Pawel Ksieniewicz, Michal Wozniak 0001
IDEAL (2)2
2019 Imbalance Reduction Techniques Applied to ECG Classification Problem
Jedrzej Kozal, Pawel Ksieniewicz
IDEAL (2)2
2019 Machine Learning Methods for Fake News Classification
Pawel Ksieniewicz, Michal Choras, Rafal Kozik, Michal Wozniak 0001
IDEAL (2)1
2019 Data stream classification using active learned neural networks
Pawel Ksieniewicz, Michal Wozniak 0001, Boguslaw Cyganek, Andrzej Kasprzak, Krzysztof Walkowiak
Neurocomputing1
2018 An Empirical Insight Into Concept Drift Detectors Ensemble Strategies
abstract
Contemporary decision support systems have to take into consideration the fact that most of data gathered is nowadays in motion, i.e., that successive observations, about objects being analyzed, form so-called data streams. Unfortunately, during analytical model utilization, unpredictable changes may appear in data distributions, leading to significant deterioration in the predictive performance and reliability of these learners. This phenomenon is called concept drift and refers to changes in the input data in relation to target variable in supervised learning task. Due to its potentially catastrophic impact on the underlying learner, it must be detected and handled as soon as it occurs. Over the years, many methods have been developed to address this issue. We focus on supervised classification task, aiming at answering the question on how to detect significant changes in data distribution effectively using ensemble of drift detectors. We discuss several models of combined drift detectors, among them the local detector which analyses distribution of each attribute separately. Experimental evaluations confirm the effectiveness of ensemble detectors, making them highly interesting to be used in solving real-world problems.
Andrzej Lapinski, Bartosz Krawczyk, Pawel Ksieniewicz, Michal Wozniak 0001
CEC3
2018 Combined Classifier Based on Quantized Subspace Class Distribution
Pawel Ksieniewicz
IDEAL (1)1
2018 Imbalanced Data Classification Based on Feature Selection Techniques
Pawel Ksieniewicz, Michal Wozniak 0001
IDEAL (2)1
2018 Ensemble of Extreme Learning Machines with trained classifier combination and statistical features for hyperspectral data
Pawel Ksieniewicz, Bartosz Krawczyk, Michal Wozniak 0001
Neurocomputing1
2015 Blurred Labeling Segmentation Algorithm for Hyperspectral Images
Pawel Ksieniewicz, Manuel Graña, Michal Wozniak 0001
ICCCI (2)1