Dariusz Brzezinski

dblp:27/9686 · DBLP profile ↗
← Back
10ranked-venue papers in the field
6as first author
5since 2021 · last 2025
0000-0001-9723-525XORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 8 (4 first)Knowledge Engineering, Semantic Web & Information Systems · 2 (2 first)
YearPublicationVenuePosition
2025 Random Similarity Isolation Forests
Sebastian Chwilczynski, Dariusz Brzezinski
PAKDD (1)2
2025 Among Them: A Game-Based Framework for Assessing Persuasion Capabilities of LLMs
Mateusz Idziejczak, Vasyl Korzavatykh, Mateusz Stawicki, Andrii Chmutov, Marcin Korcz, Iwo Bladek, Dariusz Brzezinski
PAKDD (5)7
2024 Properties of Fairness Measures in the Context of Varying Class Imbalance and Protected Group Ratios
abstract
Society is increasingly relying on predictive models in fields like criminal justice, credit risk management, and hiring. To prevent such automated systems from discriminating against people belonging to certain groups, fairness measures have become a crucial component in socially relevant applications of machine learning. However, existing fairness measures have been designed to assess the bias between predictions for protected groups without considering the imbalance in the classes of the target variable. Current research on the potential effect of class imbalance on fairness focuses on practical applications rather than dataset-independent measure properties. In this article, we study the general properties of fairness measures for changing class and protected group proportions. For this purpose, we analyze the probability mass functions of six of the most popular group fairness measures. We also measure how the probability of achieving perfect fairness changes for varying class imbalance ratios. Moreover, we relate the dataset-independent properties of fairness measures described in this work to classifier fairness in real-life tasks. Our results show that measures such as Equal Opportunity and Positive Predictive Parity are more sensitive to changes in class imbalance than Accuracy Equality. These findings can help guide researchers and practitioners in choosing the most appropriate fairness measures for their classification problems.
Dariusz Brzezinski, Julia Stachowiak, Jerzy Stefanowski, Izabela Szczech, Robert Susmaga, Sofya Aksenyuk, Uladzimir Ivashka, Oleksandr Yasinskyi
ACM Trans. Knowl. Discov. Data1
2022 Random Similarity Forests
Maciej Piernik, Dariusz Brzezinski, Pawel Zawadzki
ECML/PKDD (5)2
2021 The impact of data difficulty factors on classification of imbalanced and concept drifting data streams
abstract
Abstract Class imbalance introduces additional challenges when learning classifiers from concept drifting data streams. Most existing work focuses on designing new algorithms for dealing with the global imbalance ratio and does not consider other data complexities. Independent research on static imbalanced data has highlighted the influential role of local data difficulty factors such as minority class decomposition and presence of unsafe types of examples. Despite often being present in real-world data, the interactions between concept drifts and local data difficulty factors have not been investigated in concept drifting data streams yet. We thoroughly study the impact of such interactions on drifting imbalanced streams. For this purpose, we put forward a new categorization of concept drifts for class imbalanced problems. Through comprehensive experiments with synthetic and real data streams, we study the influence of concept drifts, global class imbalance, local data difficulty factors, and their combinations, on predictions of representative online classifiers. Experimental results reveal the high influence of new considered factors and their local drifts, as well as differences in existing classifiers’ reactions to such factors. Combinations of multiple factors are the most challenging for classifiers. Although existing classifiers are partially capable of coping with global class imbalance, new approaches are needed to address challenges posed by imbalanced data streams.
Dariusz Brzezinski, Leandro L. Minku, Tomasz Pewinski, Jerzy Stefanowski, Artur Szumaczuk
Knowl. Inf. Syst.1
2018 Visual-based analysis of classification measures and their properties for class imbalanced problems
Dariusz Brzezinski, Jerzy Stefanowski, Robert Susmaga, Izabela Szczech
Inf. Sci.1
2017 Tetrahedron: Barycentric Measure Visualizer
Dariusz Brzezinski, Jerzy Stefanowski, Robert Susmaga, Izabela Szczech
ECML/PKDD (3)1
2017 Prequential AUC: properties of the area under the ROC curve for data streams with concept drift
abstract
Modern data-driven systems often require classifiers capable of dealing with streaming imbalanced data and concept changes. The assessment of learning algorithms in such scenarios is still a challenge, as existing online evaluation measures focus on efficiency, but are susceptible to class ratio changes over time. In case of static data, the area under the receiver operating characteristics curve, or simply AUC, is a popular measure for evaluating classifiers both on balanced and imbalanced class distributions. However, the characteristics of AUC calculated on time-changing data streams have not been studied. This paper analyzes the properties of our recent proposal, an incremental algorithm that uses a sorted tree structure with a sliding window to compute AUC with forgetting. The resulting evaluation measure, called prequential AUC, is studied in terms of: visualization over time, processing speed, differences compared to AUC calculated on blocks of examples, and consistency with AUC calculated traditionally. Simulation results show that the proposed measure is statistically consistent with AUC computed traditionally on streams without drift and comparably fast to existing evaluation procedures. Finally, experiments on real-world and synthetic data showcase characteristic properties of prequential AUC compared to classification accuracy, G-mean, Kappa, Kappa M, and recall when used to evaluate classifiers on imbalanced streams with various difficulty factors.
Dariusz Brzezinski, Jerzy Stefanowski
Knowl. Inf. Syst.1
2016 Clustering XML documents by patterns
abstract
Now that the use of XML is prevalent, methods for mining semi-structured documents have become even more important. In particular, one of the areas that could greatly benefit from in-depth analysis of XML’s semi-structured nature is cluster analysis. Most of the XML clustering approaches developed so far employ pairwise similarity measures. In this paper, we study clustering algorithms, which use patterns to cluster documents without the need for pairwise comparisons. We investigate the shortcomings of existing approaches and establish a new pattern-based clustering framework called XPattern, which tries to address these shortcomings. The proposed framework consists of four steps: choosing a pattern definition, pattern mining, pattern clustering, and document assignment. The framework’s distinguishing feature is the combination of pattern clustering and document-cluster assignment, which allows to group documents according to their characteristic features rather than their direct similarity. We experimentally evaluate the proposed approach by implementing an algorithm called PathXP, which mines maximal frequent paths and groups them into profiles. PathXP was found to match, in terms of accuracy, other XML clustering approaches, while requiring less parametrization and providing easily interpretable cluster representatives. Additionally, the results of an in-depth experimental study lead to general suggestions concerning pattern-based XML clustering.
Maciej Piernik, Dariusz Brzezinski, Tadeusz Morzy
Knowl. Inf. Syst.2
2014 Combining block-based and online methods in learning ensembles from concept drifting data streams
Dariusz Brzezinski, Jerzy Stefanowski
Inf. Sci.1