Dariusz Brzezinski

dblp:27/9686 · DBLP profile ↗
← Back
20ranked-venue papers
10as first author
8since 2021 · last 2025
0000-0001-9723-525XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 10 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Random Similarity Isolation Forests
Sebastian Chwilczynski, Dariusz Brzezinski
PAKDD (1)2
2025 Among Them: A Game-Based Framework for Assessing Persuasion Capabilities of LLMs
Mateusz Idziejczak, Vasyl Korzavatykh, Mateusz Stawicki, Andrii Chmutov, Marcin Korcz, Iwo Bladek, Dariusz Brzezinski
PAKDD (5)7
2025 Ligand identification in CryoEM and X-ray maps using deep learning
abstract
MOTIVATION: Accurately identifying ligands plays a crucial role in the process of structure-guided drug design. Based on density maps from X-ray diffraction or cryogenic-sample electron microscopy (cryoEM), scientists verify whether small-molecule ligands bind to active sites of interest. However, the interpretation of density maps is challenging, and cognitive bias can sometimes mislead investigators into modeling fictitious compounds. Ligand identification can be aided by automatic methods, but existing approaches are available only for X-ray diffraction and are based on iterative fitting or feature-engineered machine learning rather than end-to-end deep learning. RESULTS: Here, we propose to identify ligands using a deep-learning approach that treats density maps as 3D point clouds. We show that the proposed model is on par with existing machine learning methods for X-ray crystallography while also being applicable to cryoEM density maps. Our study demonstrates that electron density map fragments can aid the training of models that can later be applied to cryoEM structures but also highlights challenges associated with the standardization of electron microscopy maps and the quality assessment of cryoEM ligands. AVAILABILITY AND IMPLEMENTATION: Code and model weights are available on GitHub at https://github.com/jkarolczak/ligands-classification. An accompanying ChimeraX bundle is available at https://github.com/wtaisner/chimerax-ligand-recognizer.
Jacek Karolczak, Anna Przybylowska, Konrad Szewczyk, Witold Taisner, John M. Heumann, Michael H. B. Stowell, Michal R. Nowicki, Dariusz Brzezinski
Bioinform.8
2024 Properties of Fairness Measures in the Context of Varying Class Imbalance and Protected Group Ratios
abstract
Society is increasingly relying on predictive models in fields like criminal justice, credit risk management, and hiring. To prevent such automated systems from discriminating against people belonging to certain groups, fairness measures have become a crucial component in socially relevant applications of machine learning. However, existing fairness measures have been designed to assess the bias between predictions for protected groups without considering the imbalance in the classes of the target variable. Current research on the potential effect of class imbalance on fairness focuses on practical applications rather than dataset-independent measure properties. In this article, we study the general properties of fairness measures for changing class and protected group proportions. For this purpose, we analyze the probability mass functions of six of the most popular group fairness measures. We also measure how the probability of achieving perfect fairness changes for varying class imbalance ratios. Moreover, we relate the dataset-independent properties of fairness measures described in this work to classifier fairness in real-life tasks. Our results show that measures such as Equal Opportunity and Positive Predictive Parity are more sensitive to changes in class imbalance than Accuracy Equality. These findings can help guide researchers and practitioners in choosing the most appropriate fairness measures for their classification problems.
Dariusz Brzezinski, Julia Stachowiak, Jerzy Stefanowski, Izabela Szczech, Robert Susmaga, Sofya Aksenyuk, Uladzimir Ivashka, Oleksandr Yasinskyi
ACM Trans. Knowl. Discov. Data1
2022 Random Similarity Forests
Maciej Piernik, Dariusz Brzezinski, Pawel Zawadzki
ECML/PKDD (5)2
2022 DBFE: distribution-based feature extraction from structural variants in whole-genome data
abstract
MOTIVATION: Whole-genome sequencing has revolutionized biosciences by providing tools for constructing complete DNA sequences of individuals. With entire genomes at hand, scientists can pinpoint DNA fragments responsible for oncogenesis and predict patient responses to cancer treatments. Machine learning plays a paramount role in this process. However, the sheer volume of whole-genome data makes it difficult to encode the characteristics of genomic variants as features for learning algorithms. RESULTS: In this article, we propose three feature extraction methods that facilitate classifier learning from sets of genomic variants. The core contributions of this work include: (i) strategies for determining features using variant length binning, clustering and density estimation; (ii) a programing library for automating distribution-based feature extraction in machine learning pipelines. The proposed methods have been validated on five real-world datasets using four different classification algorithms and a clustering approach. Experiments on genomes of 219 ovarian, 61 lung and 929 breast cancer patients show that the proposed approaches automatically identify genomic biomarkers associated with cancer subtypes and clinical response to oncological treatment. Finally, we show that the extracted features can be used alongside unsupervised learning methods to analyze genomic samples. AVAILABILITY AND IMPLEMENTATION: The source code of the presented algorithms and reproducible experimental scripts are available on Github at https://github.com/MNMdiagnostics/dbfe. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Maciej Piernik, Dariusz Brzezinski, Pawel Sztromwasser, Klaudia Pacewicz, Weronika Majer-Burman, Michal Gniot, Dawid Sielski, Oleksii Bryzghalov, Alicja Wozna, Pawel Zawadzki
Bioinform.2
2021 Detecting anomalies in X-ray diffraction images using convolutional neural networks
Adam Czyzewski, Faustyna Krawiec, Dariusz Brzezinski, Przemyslaw J. Porebski, Wladek Minor
Expert Syst. Appl.3
2021 The impact of data difficulty factors on classification of imbalanced and concept drifting data streams
abstract
Abstract Class imbalance introduces additional challenges when learning classifiers from concept drifting data streams. Most existing work focuses on designing new algorithms for dealing with the global imbalance ratio and does not consider other data complexities. Independent research on static imbalanced data has highlighted the influential role of local data difficulty factors such as minority class decomposition and presence of unsafe types of examples. Despite often being present in real-world data, the interactions between concept drifts and local data difficulty factors have not been investigated in concept drifting data streams yet. We thoroughly study the impact of such interactions on drifting imbalanced streams. For this purpose, we put forward a new categorization of concept drifts for class imbalanced problems. Through comprehensive experiments with synthetic and real data streams, we study the influence of concept drifts, global class imbalance, local data difficulty factors, and their combinations, on predictions of representative online classifiers. Experimental results reveal the high influence of new considered factors and their local drifts, as well as differences in existing classifiers’ reactions to such factors. Combinations of multiple factors are the most challenging for classifiers. Although existing classifiers are partially capable of coping with global class imbalance, new approaches are needed to address challenges posed by imbalanced data streams.
Dariusz Brzezinski, Leandro L. Minku, Tomasz Pewinski, Jerzy Stefanowski, Artur Szumaczuk
Knowl. Inf. Syst.1
2020 On the Dynamics of Classification Measures for Imbalanced and Streaming Data
abstract
As each imbalanced classification problem comes with its own set of challenges, the measure used to evaluate classifiers must be individually selected. To help researchers make this decision in an informed manner, experimental and theoretical investigations compare general properties of measures. However, existing studies do not analyze changes in measure behavior imposed by different imbalance ratios. Moreover, several characteristics of imbalanced data streams, such as the effect of dynamically changing class proportions, have not been thoroughly investigated from the perspective of different metrics. In this paper, we study measure dynamics by analyzing changes of measure values, distributions, and gradients with diverging class proportions. For this purpose, we visualize measure probability mass functions and gradients. In addition, we put forward a histogram-based normalization method that provides a unified, probabilistic interpretation of any measure over data sets with different class distributions. The results of analyzing eight popular classification measures show that the effect class proportions have on each measure is different and should be taken into account when evaluating classifiers. Apart from highlighting imbalance-related properties of each measure, our study shows a direct connection between class ratio changes and certain types of concept drift, which could be influential in designing new types of classifiers and drift detectors for imbalanced data streams.
Dariusz Brzezinski, Jerzy Stefanowski, Robert Susmaga, Izabela Szczech
IEEE Trans. Neural Networks Learn. Syst.1
2019 Automatic recognition of ligands in electron density by machine learning
abstract
Motivation: The correct identification of ligands in crystal structures of protein complexes is the cornerstone of structure-guided drug design. However, cognitive bias can sometimes mislead investigators into modeling fictitious compounds without solid support from the electron density maps. Ligand identification can be aided by automatic methods, but existing approaches are based on time-consuming iterative fitting. Results: Here we report a new machine learning algorithm called CheckMyBlob that identifies ligands from experimental electron density maps. In benchmark tests on portfolios of up to 219 931 ligand binding sites containing the 200 most popular ligands found in the Protein Data Bank, CheckMyBlob markedly outperforms the existing automatic methods for ligand identification, in some cases doubling the recognition rates, while requiring significantly less time. Our work shows that machine learning can improve the automation of structure modeling and significantly accelerate the drug screening process of macromolecule-ligand complexes. Availability and implementation: Code and data are available on GitHub at https://github.com/dabrze/CheckMyBlob. Supplementary information: Supplementary data are available at Bioinformatics online.
Marcin Kowiel, Dariusz Brzezinski, Przemyslaw J. Porebski, Ivan G. Shabalin, Mariusz Jaskolski, Wladek Minor
Bioinform.2
2018 Visual-based analysis of classification measures and their properties for class imbalanced problems
Dariusz Brzezinski, Jerzy Stefanowski, Robert Susmaga, Izabela Szczech
Inf. Sci.1
2017 Discovering Minority Sub-clusters and Local Difficulty Factors from Imbalanced Data
Mateusz Lango, Dariusz Brzezinski, Sebastian Firlik, Jerzy Stefanowski
DS2
2017 Using Network Analysis to Improve Nearest Neighbor Classification of Non-network Data
Maciej Piernik, Dariusz Brzezinski, Tadeusz Morzy, Mikolaj Morzy
ISMIS2
2017 Tetrahedron: Barycentric Measure Visualizer
Dariusz Brzezinski, Jerzy Stefanowski, Robert Susmaga, Izabela Szczech
ECML/PKDD (3)1
2017 Prequential AUC: properties of the area under the ROC curve for data streams with concept drift
abstract
Modern data-driven systems often require classifiers capable of dealing with streaming imbalanced data and concept changes. The assessment of learning algorithms in such scenarios is still a challenge, as existing online evaluation measures focus on efficiency, but are susceptible to class ratio changes over time. In case of static data, the area under the receiver operating characteristics curve, or simply AUC, is a popular measure for evaluating classifiers both on balanced and imbalanced class distributions. However, the characteristics of AUC calculated on time-changing data streams have not been studied. This paper analyzes the properties of our recent proposal, an incremental algorithm that uses a sorted tree structure with a sliding window to compute AUC with forgetting. The resulting evaluation measure, called prequential AUC, is studied in terms of: visualization over time, processing speed, differences compared to AUC calculated on blocks of examples, and consistency with AUC calculated traditionally. Simulation results show that the proposed measure is statistically consistent with AUC computed traditionally on streams without drift and comparably fast to existing evaluation procedures. Finally, experiments on real-world and synthetic data showcase characteristic properties of prequential AUC compared to classification accuracy, G-mean, Kappa, Kappa M, and recall when used to evaluate classifiers on imbalanced streams with various difficulty factors.
Dariusz Brzezinski, Jerzy Stefanowski
Knowl. Inf. Syst.1
2016 Ensemble Diversity in Evolving Data Streams
Dariusz Brzezinski, Jerzy Stefanowski
DS1
2016 Clustering XML documents by patterns
abstract
Now that the use of XML is prevalent, methods for mining semi-structured documents have become even more important. In particular, one of the areas that could greatly benefit from in-depth analysis of XML’s semi-structured nature is cluster analysis. Most of the XML clustering approaches developed so far employ pairwise similarity measures. In this paper, we study clustering algorithms, which use patterns to cluster documents without the need for pairwise comparisons. We investigate the shortcomings of existing approaches and establish a new pattern-based clustering framework called XPattern, which tries to address these shortcomings. The proposed framework consists of four steps: choosing a pattern definition, pattern mining, pattern clustering, and document assignment. The framework’s distinguishing feature is the combination of pattern clustering and document-cluster assignment, which allows to group documents according to their characteristic features rather than their direct similarity. We experimentally evaluate the proposed approach by implementing an algorithm called PathXP, which mines maximal frequent paths and groups them into profiles. PathXP was found to match, in terms of accuracy, other XML clustering approaches, while requiring less parametrization and providing easily interpretable cluster representatives. Additionally, the results of an in-depth experimental study lead to general suggestions concerning pattern-based XML clustering.
Maciej Piernik, Dariusz Brzezinski, Tadeusz Morzy
Knowl. Inf. Syst.2
2014 Adaptive XML Stream Classification Using Partial Tree-Edit Distance
Dariusz Brzezinski, Maciej Piernik
ISMIS1
2014 Combining block-based and online methods in learning ensembles from concept drifting data streams
Dariusz Brzezinski, Jerzy Stefanowski
Inf. Sci.1
2014 Reacting to Different Types of Concept Drift: The Accuracy Updated Ensemble Algorithm
abstract
Data stream mining has been receiving increased attention due to its presence in a wide range of applications, such as sensor networks, banking, and telecommunication. One of the most important challenges in learning from data streams is reacting to concept drift, i.e., unforeseen changes of the stream's underlying data distribution. Several classification algorithms that cope with concept drift have been put forward, however, most of them specialize in one type of change. In this paper, we propose a new data stream classifier, called the Accuracy Updated Ensemble (AUE2), which aims at reacting equally well to different types of drift. AUE2 combines accuracy-based weighting mechanisms known from block-based ensembles with the incremental nature of Hoeffding Trees. The proposed algorithm is experimentally compared with 11 state-of-the-art stream methods, including single classifiers, block-based and online ensembles, and hybrid approaches in different drift scenarios. Out of all the compared algorithms, AUE2 provided best average classification accuracy while proving to be less memory consuming than other ensemble approaches. Experimental results show that AUE2 can be considered suitable for scenarios, involving many types of drift as well as static environments.
Dariusz Brzezinski, Jerzy Stefanowski
IEEE Trans. Neural Networks Learn. Syst.1