Pawel Zyblewski

dblp:241/5551 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0002-4224-6709ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Structuring the processing frameworks for data stream evaluation and application
abstract
The following work addresses the problem of frameworks for data stream processing that can be used to evaluate the solutions in an environment that resembles real-world applications. The definition of structured frameworks stems from the need to reliably assess data stream classification methods, considering the constraints of delayed label access, the costs of their acquisition and the costs of model adaptation. The current experimental evaluation often boundlessly exploits the assumption of the immediate label access to monitor the recognition quality and adapt the methods to the changing concepts. The problem is leveraged by reviewing currently described methods and techniques for data stream processing and verifying their outcomes in simulated environment . This work defines a taxonomy of data stream processing frameworks and presents four processing schemes that link the tasks of drift detection and classification while considering a natural phenomenon of label delay . The presented research shows that classification quality is significantly affected not only by the disruptive phenomenon of concept drifts and label delay , but also by the undertaken processing scheme that describes the flow of labels in the recognition system. Considering a specific processing framework depending on real-world constraints proves to be a critical aspect of reliable and realistic experimental evaluation.
Joanna Komorniczak, Pawel Ksieniewicz, Pawel Zyblewski
Pattern Recognit.3
2025 How to RETIRE Tabular Data in Favor of Discrete Digital Signal Representation
Pawel Zyblewski, Szymon Wojciechowski
ECML/PKDD (2)1
2023 Alphabet Flatting as a variant of n-gram feature extraction method in ensemble classification of fake news
Pawel Ksieniewicz, Pawel Zyblewski, Weronika Borek-Marciniec, Rafal Kozik, Michal Choras, Michal Wozniak 0001
Eng. Appl. Artif. Intell.2
2023 Active Weighted Aging Ensemble for drifted data stream classification
abstract
Purpose One of the significant problems in data stream classification is the concept drift phenomenon, which consists of the change in probabilistic characteristics of the classification task . Such changes in posterior probability destabilize the classification model performance, seriously degrading its quality. It is necessary to design appropriate strategies to counteract this phenomenon, allowing the classifier to adapt to the changing probabilistic characteristics. It is tough to propose such an approach with limited access to data labels. A human bias of high quality is usually costly, so to minimize the expenses related to this process, it is also necessary to propose learning strategies based on semi-supervised learning. Such strategies employ active learning methods indicating which of the incoming objects are valuable to be labeled for improving the classifier's performance. Methods This paper proposes Active Weighted Aging Ensemble algorithm, a novel chunk-based method for non-stationary data stream classification. It employs a classifier ensemble approach and utilizes the changing ensemble lineup to react to concept drift appropriately. It also proposed a new active learning method, considering a limited budget that may be applied to any data stream classifier. Results AWAE has been evaluated through computer experiments using real and synthetic data streams, confirming the proposed algorithm's high quality over state-of-the-art methods. Conclusion The research conducted on benchmark data streams confirmed the effectiveness of the proposed solution and highlighted its strengths in comparison with state-of-the-art methods. The estimated computational complexity is acceptable and comparable to the benchmark algorithms.
Michal Wozniak 0001, Pawel Zyblewski, Pawel Ksieniewicz
Inf. Sci.2
2022 Feature Integration Strategies for Multilingual Fake News Classification
abstract
The abundance of information in digital media, which in today’s world is the main source of knowledge about current events for the masses, makes it possible to spread disinformation on a larger scale than ever before. Consequently, there is a need to develop novel fake news detection approaches capable of adapting to changing factual contexts and generalizing previously or concurrently acquired knowledge. To deal with this problem, we propose a ensemble-based approach, which allows for fake news detection in multiple languages and the mutual transfer of knowledge acquired in each of them. Both classical feature extractors, such as Term frequency-inverse document frequency or Latent Dirichlet Allocation, and integrated deep NLP (Natural Language Processing) BERT (Bidirectional Encoder Representations from Transformers) models paired with MLP (Multilayer Perceptron) classifier, were employed. The results of experiments conducted on two datasets dedicated to the fake news classification t ask ( in English and Spanish, respectively), supported by statistical analysis, confirmed that utilization of additional languages could improve performance for traditional methods. Also, in some cases supplementing the deep learning method with classical ones can positively impact obtained results. The ability of models to generalize the knowledge acquired between the analyzed languages was also observed.
Jedrzej Kozal, Michal Les, Pawel Zyblewski, Pawel Ksieniewicz, Michal Wozniak 0001
IEEE Big Data3
2022 Imbalanced Data Stream Classification Assisted by Prior Probability Estimation
abstract
With the processing of data streams, come inevitable challenges, such as changes in the prior (class drift) and posterior (concept drift) probability distribution over the processing time. Both these phenomena have a negative impact on the quality of the classification. Heavily imbalanced problems, which are often typical for real-world applications, bring additional processing difficulties. Classifiers are often biased towards the majority class and have difficulty identifying instances of categories described with a lower number of objects. The following article proposes a Prior Probability Assisted Classifier (2PAC), a method aiming to improve the classification quality of heavily imbalanced data streams with dynamic changes by using the estimated prior probability value and the correction of the classifier's decision for batch predictions. Presented extensive computer experiments, supported by statistical analysis, show the ability to improve the classification quality using the proposed method.
Joanna Komorniczak, Pawel Zyblewski, Pawel Ksieniewicz
IJCNN2
2022 Stream-learn - open-source Python library for difficult data stream batch analysis
Pawel Ksieniewicz, Pawel Zyblewski
Neurocomputing2
2022 Statistical Drift Detection Ensemble for batch processing of data streams
Joanna Komorniczak, Pawel Zyblewski, Pawel Ksieniewicz
Knowl. Based Syst.2
2021 Prior Probability Estimation in Dynamically Imbalanced Data Streams
abstract
Despite the fact that real-life data streams may often be characterized by the dynamic changes in the prior class probabilities, there is a scarcity of articles trying to clearly describe and classify this problem as well as suggest new methods dedicated to resolving this issue. The following paper aims to fill this gap by proposing a novel data stream taxonomy defined in the context of prior class probability and by introducing the Dynamic Statistical Concept Analysis (DSCA) - prior probability estimation algorithm. The proposed method was evaluated using computer experiments carried out on 100 synthetically generated data streams with various class imbalance characteristics. The obtained results, supported by statistical analysis, confirmed the usefulness of the proposed solution, especially in the case of discrete dynamically imbalanced data streams (DDIS).
Joanna Komorniczak, Pawel Zyblewski, Pawel Ksieniewicz
IJCNN2
2021 Fusion of linear base classifiers in geometric space
Pawel Ksieniewicz, Pawel Zyblewski, Robert Burduk
Knowl. Based Syst.2
2020 Fake News Detection from Data Streams
abstract
Using fake news as a political or economic tool is not new, but the scale of their use is currently alarming, especially on social media. The authors of misinformation try to influence the users' decisions, both in the economic and political sphere. The facts of using disinformation during elections are well known. Currently, two fake news detection approaches dominate. The first approach, so-called fact or news checker, is based on the knowledge and work of volunteers, the second approach employs artificial intelligence algorithms for news analysis and manipulation detection. In this work, we will focus on using machine learning methods to detect fake news. However, unlike most approaches, we will treat incoming messages as stream data, taking into account the possibility of concept drift occurring, i.e., appearing changes in the probabilistic characteristics of the classification model during the exploitation of the classifier. The developed methods have been evaluated based on computer experiments on benchmark data, and the obtained results prove their usefulness for the problem under consideration. The proposed solutions are part of the distributed platform developed by the H2020 SocialTruth project consortium.
Pawel Ksieniewicz, Pawel Zyblewski, Michal Choras, Rafal Kozik, Agata Gielczyk, Michal Wozniak 0001
IJCNN2
2020 Novel clustering-based pruning algorithms
abstract
Abstract One of the crucial problems of designing a classifier ensemble is the proper choice of the base classifier line-up. Basically, such an ensemble is formed on the basis of individual classifiers, which are trained in such a way to ensure their high diversity or they are chosen on the basis of pruning which reduces the number of predictive models in order to improve efficiency and predictive performance of the ensemble. This work is focusing on clustering-based ensemble pruning, which looks for the group of similar classifiers which are replaced by their representatives. We propose a novel pruning criterion based on well-known diversity measures and describe three algorithms using classifier clustering. The first method selects the model with the best predictive performance from each cluster to form the final ensemble, the second one employs the multistage organization, where instead of removing the classifiers from the ensemble each classifier cluster makes the decision independently, while the third proposition combines multistage organization and sampling with replacement. The proposed approaches were evaluated using 30 datasets with different characteristics. Experimentation results validated through statistical tests confirmed the usefulness of the proposed approaches.
Pawel Zyblewski, Michal Wozniak 0001
Pattern Anal. Appl.1