Jerzy Stefanowski

dblp:98/6412 · DBLP profile ↗
← Back
21ranked-venue papers in the field
5as first author
6since 2021 · last 2025
0000-0002-4949-8271ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 10 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 7 (1 first)Database Systems & Data Management · 2 (1 first)Other / Interdisciplinary · 2 (1 first)
YearPublicationVenuePosition
2025 Counterfactual Explanations with Probabilistic Guarantees on their Robustness to Model Change
Ignacy Stepka, Jerzy Stefanowski, Mateusz Lango
KDD (1)2
2024 Properties of Fairness Measures in the Context of Varying Class Imbalance and Protected Group Ratios
abstract
Society is increasingly relying on predictive models in fields like criminal justice, credit risk management, and hiring. To prevent such automated systems from discriminating against people belonging to certain groups, fairness measures have become a crucial component in socially relevant applications of machine learning. However, existing fairness measures have been designed to assess the bias between predictions for protected groups without considering the imbalance in the classes of the target variable. Current research on the potential effect of class imbalance on fairness focuses on practical applications rather than dataset-independent measure properties. In this article, we study the general properties of fairness measures for changing class and protected group proportions. For this purpose, we analyze the probability mass functions of six of the most popular group fairness measures. We also measure how the probability of achieving perfect fairness changes for varying class imbalance ratios. Moreover, we relate the dataset-independent properties of fairness measures described in this work to classifier fairness in real-life tasks. Our results show that measures such as Equal Opportunity and Positive Predictive Parity are more sensitive to changes in class imbalance than Accuracy Equality. These findings can help guide researchers and practitioners in choosing the most appropriate fairness measures for their classification problems.
Dariusz Brzezinski, Julia Stachowiak, Jerzy Stefanowski, Izabela Szczech, Robert Susmaga, Sofya Aksenyuk, Uladzimir Ivashka, Oleksandr Yasinskyi
ACM Trans. Knowl. Discov. Data3
2023 Multi-criteria Approaches to Explaining Black Box Machine Learning Models
Jerzy Stefanowski
ACIIDS (2)1
2022 Quality Versus Speed in Energy Demand Prediction - Experience Report from an R &D project
Witold Andrzejewski, Jedrzej Potoniec, Maciej Drozdowski, Jerzy Stefanowski, Robert Wrembel, Pawel Stapf
DEXA (1)4
2021 Time Aspect in Making an Actionable Prediction of a Conversation Breakdown
Piotr Janiszewski, Mateusz Lango, Jerzy Stefanowski
ECML/PKDD (5)3
2021 The impact of data difficulty factors on classification of imbalanced and concept drifting data streams
abstract
Abstract Class imbalance introduces additional challenges when learning classifiers from concept drifting data streams. Most existing work focuses on designing new algorithms for dealing with the global imbalance ratio and does not consider other data complexities. Independent research on static imbalanced data has highlighted the influential role of local data difficulty factors such as minority class decomposition and presence of unsafe types of examples. Despite often being present in real-world data, the interactions between concept drifts and local data difficulty factors have not been investigated in concept drifting data streams yet. We thoroughly study the impact of such interactions on drifting imbalanced streams. For this purpose, we put forward a new categorization of concept drifts for class imbalanced problems. Through comprehensive experiments with synthetic and real data streams, we study the influence of concept drifts, global class imbalance, local data difficulty factors, and their combinations, on predictions of representative online classifiers. Experimental results reveal the high influence of new considered factors and their local drifts, as well as differences in existing classifiers’ reactions to such factors. Combinations of multiple factors are the most challenging for classifiers. Although existing classifiers are partially capable of coping with global class imbalance, new approaches are needed to address challenges posed by imbalanced data streams.
Dariusz Brzezinski, Leandro L. Minku, Tomasz Pewinski, Jerzy Stefanowski, Artur Szumaczuk
Knowl. Inf. Syst.4
2018 Visual-based analysis of classification measures and their properties for class imbalanced problems
Dariusz Brzezinski, Jerzy Stefanowski, Robert Susmaga, Izabela Szczech
Inf. Sci.2
2018 Multi-class and feature selection extensions of Roughly Balanced Bagging for imbalanced data
abstract
Roughly Balanced Bagging is one of the most efficient ensembles specialized for class imbalanced data. In this paper, we study its basic properties that may influence its good classification performance. We experimentally analyze them with respect to bootstrap construction, deciding on the number of component classifiers, their diversity, and ability to deal with the most difficult types of the minority examples. Then, we introduce two generalizations of this ensemble for dealing with a higher number of attributes and for adapting it to handle multiple minority classes. Experiments with synthetic and real life data confirm usefulness of both proposals.
Mateusz Lango, Jerzy Stefanowski
J. Intell. Inf. Syst.2
2017 Tetrahedron: Barycentric Measure Visualizer
Dariusz Brzezinski, Jerzy Stefanowski, Robert Susmaga, Izabela Szczech
ECML/PKDD (3)2
2017 Prequential AUC: properties of the area under the ROC curve for data streams with concept drift
abstract
Modern data-driven systems often require classifiers capable of dealing with streaming imbalanced data and concept changes. The assessment of learning algorithms in such scenarios is still a challenge, as existing online evaluation measures focus on efficiency, but are susceptible to class ratio changes over time. In case of static data, the area under the receiver operating characteristics curve, or simply AUC, is a popular measure for evaluating classifiers both on balanced and imbalanced class distributions. However, the characteristics of AUC calculated on time-changing data streams have not been studied. This paper analyzes the properties of our recent proposal, an incremental algorithm that uses a sorted tree structure with a sliding window to compute AUC with forgetting. The resulting evaluation measure, called prequential AUC, is studied in terms of: visualization over time, processing speed, differences compared to AUC calculated on blocks of examples, and consistency with AUC calculated traditionally. Simulation results show that the proposed measure is statistically consistent with AUC computed traditionally on streams without drift and comparably fast to existing evaluation procedures. Finally, experiments on real-world and synthetic data showcase characteristic properties of prequential AUC compared to classification accuracy, G-mean, Kappa, Kappa M, and recall when used to evaluate classifiers on imbalanced streams with various difficulty factors.
Dariusz Brzezinski, Jerzy Stefanowski
Knowl. Inf. Syst.2
2016 Types of minority class examples and their influence on learning classifiers from imbalanced data
abstract
Many real-world applications reveal difficulties in learning classifiers from imbalanced data. Although several methods for improving classifiers have been introduced, the identification of conditions for the efficient use of the particular method is still an open research problem. It is also worth to study the nature of imbalanced data, characteristics of the minority class distribution and their influence on classification performance. However, current studies on imbalanced data difficulty factors have been mainly done with artificial datasets and their conclusions are not easily applicable to the real-world problems, also because the methods for their identification are not sufficiently developed. In our paper, we capture difficulties of class distribution in real datasets by considering four types of minority class examples: safe, borderline, rare and outliers. First, we confirm their occurrence in real data by exploring multidimensional visualizations of selected datasets. Then, we introduce a method for an identification of these types of examples, which is based on analyzing a class distribution in a local neighbourhood of the considered example. Two ways of modeling this neighbourhood are presented: with k-nearest examples and with kernel functions. Experiments with artificial datasets show that these methods are able to re-discover simulated types of examples. Next contributions of this paper include carrying out a comprehensive experimental study with 26 real world imbalanced datasets, where (1) we identify new data characteristics basing on the analysis of types of minority examples; (2) we demonstrate that considering the results of this analysis allow to differentiate classification performance of popular classifiers and pre-processing methods and to evaluate their areas of competence. Finally, we highlight directions of exploiting the results of our analysis for developing new algorithms for learning classifiers and pre-processing methods.
Krystyna Napierala, Jerzy Stefanowski
J. Intell. Inf. Syst.2
2015 SMOTE-IPF: Addressing the noisy and borderline examples problem in imbalanced classification by a re-sampling method with filtering
José A. Sáez, Julián Luengo, Jerzy Stefanowski, Francisco Herrera
Inf. Sci.3
2014 Combining block-based and online methods in learning ensembles from concept drifting data streams
Dariusz Brzezinski, Jerzy Stefanowski
Inf. Sci.2
2014 Processing and mining complex data streams
Jerzy Stefanowski, Alfredo Cuzzocrea, Dominik Slezak
Inf. Sci.1
2012 BRACID: a comprehensive approach to learning rules from imbalanced data
abstract
In this paper we consider induction of rule-based classifiers from imbalanced data, where one class (a minority class) is under-represented in comparison to the remaining majority classes. The minority class is usually of primary interest. However, most rule-based classifiers are biased towards the majority classes and they have difficulties with correct recognition of the minority class. In this paper we discuss sources of these difficulties related to data characteristics or to an algorithm itself. Among the problems related to the data distribution we focus on the role of small disjuncts, overlapping of classes and presence of noisy examples. Then, we show that standard techniques for induction of rule-based classifiers, such as sequential covering, top-down induction of rules or classification strategies, were created with the assumption of balanced data distribution, and we explain why they are biased towards the majority classes. Some modifications of rule-based classifiers have been already introduced, but they usually concentrate on individual problems. Therefore, we propose a novel algorithm, BRACID, which more comprehensively addresses the issues associated with imbalanced data. Its main characteristics includes a hybrid representation of rules and single examples, bottom-up learning of rules and a local classification strategy using nearest rules. The usefulness of BRACID has been evaluated in experiments on several imbalanced datasets. The results show that BRACID significantly outperforms the well known rule-based classifiers C4.5rules, RIPPER, PART, CN2, MODLEM as well as other related classifiers as RISE or K-NN. Moreover, it is comparable or better than the studied approaches specialized for imbalanced data such as generalizations of rule algorithms or combinations of SMOTE + ENN preprocessing with PART. Finally, it improves the support of minority class rules, leading to better recognition of the minority class examples.
Krystyna Napierala, Jerzy Stefanowski
J. Intell. Inf. Syst.2
2011 Local neighbourhood extension of SMOTE for mining imbalanced data
abstract
In this paper we discuss problems of inducing classifiers from imbalanced data and improving recognition of minority class using focused resampling techniques. We are particularly interested in SMOTE over-sampling method that generates new synthetic examples from the minority class between the closest neighbours from this class. However, SMOTE could also overgeneralize the minority class region as it does not consider distribution of other neighbours from the majority classes. Therefore, we introduce a new generalization of SMOTE, called LN-SMOTE, which exploits more precisely information about the local neighbourhood of the considered examples. In the experiments we compare this method with original SMOTE and its two, the most related, other generalizations Borderline and Safe-Level SMOTE. All these pre-processing methods are applied together with either decision tree or Naive Bayes classifiers. The results show that the new LN-SMOTE method improves evaluation measures for the minority class.
Tomasz Maciejewski, Jerzy Stefanowski
CIDM2
2008 Selective Pre-processing of Imbalanced Data for Improving Classification Performance
Jerzy Stefanowski, Szymon Wilk
DaWaK1
2001 Three discretization methods for rule induction
abstract
We discuss problems associated with induction of decision rules from data with numerical attributes. Real-life data frequently contain numerical attributes. Rule induction from numerical data requires an additional step called discretization. In this step numerical values are converted into intervals. Most existing discretization methods are used before rule induction, as a part of data preprocessing. Some methods discretize numerical attributes while learning decision rules. We compare the classification accuracy of a discretization method based on conditional entropy, applied before rule induction, with two newly proposed methods, incorporated directly into the rule induction algorithm LEM2, where discretization and rule induction are performed at the same time. In all three approaches the same system is used for classification of new, unseen data. As a result, we conclude that an error rate for all three methods does not show significant difference, however, rules induced by the two new methods are simpler and stronger. © 2001 John Wiley & Sons, Inc.
Jerzy W. Grzymala-Busse, Jerzy Stefanowski
Int. J. Intell. Syst.2
2001 Induction of decision rules in classification and discovery-oriented perspectives
abstract
This paper discusses induction of decision rules from data tables representing information about a set of objects described by a set of attributes. If the input data contains inconsistencies, rough sets theory can be used to handle them. The most popular perspectives of rule induction are classification and knowledge discovery. The evaluation of decision rules is quite different depending on the perspective. Criteria for evaluating the quality of a set of rules are presented and discussed. The degree of conflict and the possibility of achieving a satisfying compromise between criteria relevant to classification and criteria relevant to discovery are then analyzed. For this purpose, we performed an extensive experimental study on several well-known data sets where we compared two different approaches: (1) the popular rough set based rule induction algorithm LEM2 generating classification rules, (2) our own algorithm Explore—specific for discovery perspective. © 2001 John Wiley & Sons, Inc.
Jerzy Stefanowski, Daniel Vanderpooten
Int. J. Intell. Syst.1
1998 Experiments on Solving Multiclass Learning Problems by n2-classifier
Jacek Jelonek, Jerzy Stefanowski
ECML2
1997 Rough Set Theory and Rule Induction Techniques for Discovery of Attribute Dependencies in Medical Information Systems
Jerzy Stefanowski, Krzysztof Slowinski
PKDD1