Bartosz Krawczyk

dblp:26/11077 · DBLP profile ↗
← Back
99ranked-venue papers
45as first author
23since 2021 · last 2024
0000-0002-9774-0106ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 81 · 39 first-author · 16 since 2021Databases, data management, data science and information retrieval · 24 · 12 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 6 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3 · 2 first-authorComputer networks · 2 · 2 since 2021Theory of computation · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Improved KD-tree based imbalanced big data classification and oversampling for MapReduce platforms
William C. Sleeman IV, Martha I. Roseberry, Preetam Ghosh, Alberto Cano 0001, Bartosz Krawczyk
Appl. Intell.5
2024 ClassyNet: Class-Aware Early-Exit Neural Networks for Edge Devices
abstract
Edge-based and IoT devices have seen phenomenal growth in recent years, driven by the surge in demand for emerging applications that leverage machine learning models, such as Deep Neural Networks (DNNs). However, a primary drawback of DNNs is their substantial storage/memory needs and high computational overhead, making their adoption in edge devices challenging. This limitation prompted the development of early-exit models like BranchyNet, which enable decisions to be made at earlier stages by incorporating dedicated exits within the architecture’s inner layers. Nonetheless, these existing early-exit models lack control over the specific class that should exit and when. The necessity for such class-aware models is evident in numerous edge applications, where particular high-priority classes must be detected earlier due to their time-sensitive nature. In this paper, we introduce ClassyNet, the first early-exit architecture designed to return only selected classes at each exit. This feature facilitates faster inference times for critical classes, allowing the initial layers to operate on edge devices. This strategy conserves considerable computational time and resources on the edge without compromising accuracy. Through extensive experiments, we show the effectiveness of ClassyNet compared to other models under various scenarios.
Mohammed Ayyat, Tamer Nadeem, Bartosz Krawczyk
IEEE Internet Things J.3
2024 A survey on learning from imbalanced data streams: taxonomy, challenges, empirical study, and reproducible experimental framework
Gabriel Aguiar, Bartosz Krawczyk, Alberto Cano 0001
Mach. Learn.2
2024 Understanding imbalanced data: XAI & interpretable ML framework
abstract
Abstract There is a gap between current methods that explain deep learning models that work on imbalanced image data and the needs of the imbalanced learning community. Existing methods that explain imbalanced data are geared toward binary classification, single layer machine learning models and low dimensional data. Current eXplainable Artificial Intelligence (XAI) techniques for vision data mainly focus on mapping predictions of specific instances to inputs, instead of examining global data properties and complexities of entire classes. Therefore, there is a need for a framework that is tailored to modern deep networks, that incorporates large, high dimensional, multi-class datasets, and uncovers data complexities commonly found in imbalanced data. We propose a set of techniques that can be used by both deep learning model users to identify, visualize and understand class prototypes, sub-concepts and outlier instances; and by imbalanced learning algorithm developers to detect features and class exemplars that are key to model performance. The components of our framework can be applied sequentially in their entirety or individually, making it fully flexible to the user’s specific needs ( https://github.com/dd1github/XAI_for_Imbalanced_Learning ).
Damien Dablain, Colin Bellinger, Bartosz Krawczyk, David W. Aha, Nitesh V. Chawla
Mach. Learn.3
2024 The class imbalance problem in deep learning
Kushankur Ghosh, Colin Bellinger, Roberto Corizzo, Paula Branco, Bartosz Krawczyk, Nathalie Japkowicz
Mach. Learn.5
2024 Correction: Adversarial concept drift detection under poisoning attacks for robust data stream mining
Lukasz Korycki, Bartosz Krawczyk
Mach. Learn.2
2023 Efficient Augmentation for Imbalanced Deep Learning
abstract
Deep learning models may not effectively generalize across under-represented or minority classes. We empirically study a convolutional neural network’s (CNN) internal representation of imbalanced image data and measure the generalization gap between a model’s feature embeddings in the training and test sets, showing that the gap is wider for minority classes. This insight enables us to design an efficient three-phase CNN training framework for imbalanced data. The framework involves training the network end-to-end on imbalanced data to learn feature embeddings, performing data augmentation in the learned embedding space to balance the training data distribution, and fine-tuning the classifier head on the embedded balanced training data. We develop Expansive Over-Sampling (EOS) as a data augmentation technique to utilize in the training framework. EOS forms synthetic training instances as convex combinations between the minority class samples and their nearest adversaries in the embedding space to reduce the generalization gap. The proposed framework improves the accuracy over leading cost-sensitive and resampling methods commonly used in imbalanced learning. Moreover, it is more computationally efficient than standard data pre-processing methods, such as SMOTE and GAN-based over-sampling, as it requires fewer parameters and less training time. The source code for the proposed framework is available at: https://github.com/dd1github/EOS.
Damien Dablain, Colin Bellinger, Bartosz Krawczyk, Nitesh V. Chawla
ICDE3
2023 Class-Aware Neural Networks for Efficient Intrusion Detection on Edge Devices
abstract
The exponential growth of IoT and edge devices has led to their widespread use across various applications. However, the security of these devices remains a significant concern due to their vulnerability to a broad spectrum of cyber-attacks. Network Intrusion Detection Systems (NIDS) are crucial for identifying and mitigating such threats. Traditional NIDS approaches, while effective, struggle to detect sophisticated modern attacks and often require substantial computational power and memory, which may not be feasible for edge devices. Machine learning and neural network-based methods have demonstrated promising improvements in NIDS detection accuracy. Yet, their deployment on resource-constrained edge devices presents a challenge. This has led to the development of Dynamic Neural Networks, an approach that allows models to adapt according to the input, making them more efficient and lightweight. However, these networks are class-agnostic, rendering them unsuitable for handling cases with uneven classification priorities. In this paper, we introduce ClassyNet, a platform designed for efficient, classaware NIDS on edge devices. ClassyNet leverages class-specific feature extraction and a class-specific neural network architecture to enhance intrusion detection efficiency. Experimental results indicate that our proposed approach matches the detection accuracy of traditional machine learning and neural network-based methods while significantly improving resource efficiency.
Mohammed Ayyat, Tamer Nadeem, Bartosz Krawczyk
SECON3
2023 Adversarial concept drift detection under poisoning attacks for robust data stream mining
Lukasz Korycki, Bartosz Krawczyk
Mach. Learn.2
2023 DeepSMOTE: Fusing Deep Learning and SMOTE for Imbalanced Data
abstract
Despite over two decades of progress, imbalanced data is still considered a significant challenge for contemporary machine learning models. Modern advances in deep learning have further magnified the importance of the imbalanced data problem, especially when learning from images. Therefore, there is a need for an oversampling method that is specifically tailored to deep learning models, can work on raw images while preserving their properties, and is capable of generating high-quality, artificial images that can enhance minority classes and balance the training set. We propose Deep synthetic minority oversampling technique (SMOTE), a novel oversampling algorithm for deep learning models that leverages the properties of the successful SMOTE algorithm. It is simple, yet effective in its design. It consists of three major components: 1) an encoder/decoder framework; 2) SMOTE-based oversampling; and 3) a dedicated loss function that is enhanced with a penalty term. An important advantage of DeepSMOTE over generative adversarial network (GAN)-based oversampling is that DeepSMOTE does not require a discriminator, and it generates high-quality artificial images that are both information-rich and suitable for visual inspection. DeepSMOTE code is publicly available at https://github.com/dd1github/DeepSMOTE.
Damien Dablain, Bartosz Krawczyk, Nitesh V. Chawla
IEEE Trans. Neural Networks Learn. Syst.2
2022 ROSE: robust online self-adjusting ensemble for continual learning on imbalanced drifting data streams
Alberto Cano 0001, Bartosz Krawczyk
Mach. Learn.2
2022 Instance exploitation for learning temporary concepts from sparsely labeled drifting data streams
Lukasz Korycki, Bartosz Krawczyk
Pattern Recognit.2
2021 On the combined effect of class imbalance and concept complexity in deep learning
abstract
Structural concept complexity, class overlap, and data scarcity are some of the most important factors influencing the performance of classifiers under class imbalance conditions. When these effects were uncovered in the early 2000s, understandably, the classifiers on which they were demonstrated belonged to the classical rather than Deep Learning categories of approaches. As Deep Learning is gaining ground over classical machine learning and is beginning to be used in critical applied settings, it is important to assess systematically how well they respond to the kind of challenges their classical counterparts have struggled with in the past two decades. The purpose of this paper is to study the behavior of deep learning systems in settings that have previously been deemed challenging to classical machine learning systems to find out whether the depth of the systems is an asset in such settings. The results in both artificial and real-world image datasets show that these settings remain mostly challenging for Deep Learning systems. Deeper architectures help with structural concept complexity but not with data scarcity and class overlap.
Kushankur Ghosh, Colin Bellinger, Roberto Corizzo, Bartosz Krawczyk, Nathalie Japkowicz
IEEE BigData4
2021 Tensor Decision Trees for Continual Learning from Drifting Data Streams
abstract
Data stream classification is one of the most vital areas of contemporary machine learning, as many real-life problems generate data continuously and in large volumes. However, most of research in this area focuses on vector-based representations, which are unsuitable for capturing properties of more complex multi-dimensional structures, such as images and video sequences. In this paper, we propose a novel methodology for learning adaptive decision trees from data streams of tensors. We introduce Chordal Kernel Decision Tree for continual learning from tensor data streams. In order to maintain the tensor characteristics, we propose to train and update classifiers in the kernel space designed to work with tensor representation. We use chordal distance to compute similarities between tensors and then apply it as a new feature space in which decision trees are trained. This allows for a direct decision tree induction on tensors. In order to accommodate the streaming and drifting nature of data, we propose a concept drift detection scheme based on tensor representation. It allows us to reconstruct the kernel feature space every time when change is detected. The proposed approach allows for fast and efficient induction of decision trees on streaming data with tensor representation. Experimental study, conducted on 4 real-world and 52 artificial large-scale tensor data streams, shows that using the native tensor feature space leads to more accurate classification than outperforms the vectorized representations.
Bartosz Krawczyk
DSAA1
2021 Concept Drift Detection from Multi-Class Imbalanced Data Streams
abstract
Continual learning from data streams is among the most important topics in contemporary machine learning. One of the biggest challenges in this domain lies in creating algorithms that can continuously adapt to arriving data. However, previously learned knowledge may become outdated, as streams evolve over time. This phenomenon is known as concept drift and must be detected to facilitate efficient adaptation of the learning model. While there exists a plethora of drift detectors, all of them assume that we are dealing with roughly balanced classes. In the case of imbalanced data streams, those detectors will be biased towards the majority classes, ignoring changes happening in the minority ones. Furthermore, class imbalance may evolve over time and classes may change their roles (majority becoming minority and vice versa). This is especially challenging in the multi-class setting, where relationships among classes become complex. In this paper, we propose a detailed taxonomy of challenges posed by concept drift in multi-class imbalanced data streams, as well as a novel trainable concept drift detector based on Restricted Boltzmann Machine. It is capable of monitoring multiple classes at once and using reconstruction error to detect changes in each of them independently. Our detector utilizes a skew-insensitive loss function that allows it to handle multiple imbalanced distributions. Due to its trainable nature, it is capable of following changes in a stream and evolving class roles, as well as it can deal with local concept drift occurring in minority classes. Extensive experimental study on multi-class drifting data streams, enriched with a detailed analysis of the impact of local drifts and changing imbalance ratios, confirms the high efficacy of our approach.
Lukasz Korycki, Bartosz Krawczyk
ICDE2
2021 Undersampling with Support Vectors for Multi-Class Imbalanced Data Classification
abstract
Learning from imbalanced data poses significant challenges for the classifier. This becomes even more difficult, when dealing with multi-class problems. Here relationships among classes are no longer well-defined and it is easy to loose performance on one of the classes while gaining on other. In last years this topic has gained increased interest from the machine learning community - however, still there is a need for developing new and efficient algorithms to handle this challenge. In this paper we propose a new approach for balancing multi-class imbalanced problems. It is based on a two-step undersampling methodology. In the first step, a one-class classifier is being trained on each of the classes, achieving skew-insensitive data description. Support vectors for each class are extracted and used as new class representatives, thus achieving significant reduction in the terms of used instances. In the second step, an evolutionary undersampling approach is being used on these support vectors in order to further balance the training set. By applying this technique on a set of support vectors and not on a full dataset, we achieve a significant reduction of the computational time and increased accuracy. Finally, a standard multi-class classifier is being trained on the balanced data set. A thorough experimental study proves the usefulness of the proposed approach in comparison with state-of-the-art approaches for handling multi-class imbalanced data.
Bartosz Krawczyk, Colin Bellinger, Roberto Corizzo, Nathalie Japkowicz
IJCNN1
2021 Low-Dimensional Representation Learning from Imbalanced Data Streams
Lukasz Korycki, Bartosz Krawczyk
PAKDD (1)2
2021 Locally Linear Support Vector Machines for Imbalanced Data Classification
Bartosz Krawczyk, Alberto Cano 0001
PAKDD (1)1
2021 Streaming Decision Trees for Lifelong Learning
Lukasz Korycki, Bartosz Krawczyk
ECML/PKDD (1)2
2021 XRRpred: accurate predictor of crystal structure quality from protein sequence
abstract
MOTIVATION: X-ray crystallography was used to produce nearly 90% of protein structures. These efforts were supported by numerous sequence-based tools that accurately predict crystallizable proteins. However, protein structures vary widely in their quality, typically measured with resolution and R-free. This impacts the ability to use these structures for some applications including rational drug design and molecular docking and motivates development of methods that accurately predict structure quality from sequence. RESULTS: We introduce XRRpred, the first predictor of the resolution and R-free values from protein sequences. XRRpred relies on original sequence profiles, hand-crafted features, empirically selected and parametrized regressors and modern resampling techniques. Using an independent test dataset, we show that XRRpred provides accurate predictions of resolution and R-free. We demonstrate that XRRpred's predictions correctly model relationship between the resolution and R-free and reproduce structure quality relations between structural classes of proteins. We also show that XRRpred significantly outperforms indirect alternative ways to predict the structure quality that include predictors of crystallization propensity and an alignment-based approach. XRRpred is available as a convenient webserver that allows batch predictions and offers informative visualization of the results. AVAILABILITY AND IMPLEMENTATION: http://biomine.cs.vcu.edu/servers/XRRPred/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sina Ghadermarzi, Bartosz Krawczyk, Jiangning Song, Lukasz A. Kurgan
Bioinform.2
2021 Self-adjusting k nearest neighbors for continual learning from multi-label drifting data streams
Martha I. Roseberry, Bartosz Krawczyk, Youcef Djenouri, Alberto Cano 0001
Neurocomputing2
2021 Multi-class imbalanced big data classification on Spark
William C. Sleeman IV, Bartosz Krawczyk
Knowl. Based Syst.2
2021 Tensor decision trees for continual learning from drifting data streams
Bartosz Krawczyk
Mach. Learn.1
2020 Online Oversampling for Sparsely Labeled Imbalanced and Non-Stationary Data Streams
abstract
Learning from imbalanced data and data stream mining are among most popular areas in contemporary machine learning. There is a strong interplay between these domains, as data streams are frequently characterized by skewed distributions. However, most of existing works focus on binary problems, omitting significantly more challenging multi-class imbalanced data. In this paper, we propose a novel framework for learning from multi-class imbalanced data streams that simultaneously tackles three major problems in this area: (i) changing imbalance ratios among multiple classes; (ii) concept drift; and (iii) limited access to ground truth. We use active learning combined with streaming-based oversampling that uses both information about current class ratios and classifier errors on each class to create new instances in a meaningful way. Conducted experimental study shows that our single-classifier framework is capable of outperforming state-of-the-art ensembles dedicated to multi-class imbalanced data streams in both fully supervised and sparsely labeled learning scenarios.
Lukasz Korycki, Bartosz Krawczyk
IJCNN2
2020 A Machine Learning method for relabeling arbitrary DICOM structure sets to TG-263 defined labels
William C. Sleeman IV, Joseph Nalluri, Khajamoinuddin Syed, Preetam Ghosh, Bartosz Krawczyk, Michael Hagan, Jatinder Palta, Rishabh Kapoor
J. Biomed. Informatics5
2020 Combined Cleaning and Resampling algorithm for multi-class imbalanced data with label noise
Michal Koziarski, Michal Wozniak 0001, Bartosz Krawczyk
Knowl. Based Syst.3
2020 Kappa Updated Ensemble for drifting data stream mining
Alberto Cano 0001, Bartosz Krawczyk
Mach. Learn.2
2020 Radial-Based Oversampling for Multiclass Imbalanced Data Classification
abstract
Learning from imbalanced data is among the most popular topics in the contemporary machine learning. However, the vast majority of attention in this field is given to binary problems, while their much more difficult multiclass counterparts are relatively unexplored. Handling data sets with multiple skewed classes poses various challenges and calls for a better understanding of the relationship among classes. In this paper, we propose multiclass radial-based oversampling (MC-RBO), a novel data-sampling algorithm dedicated to multiclass problems. The main novelty of our method lies in using potential functions for generating artificial instances. We take into account information coming from all of the classes, contrary to existing multiclass oversampling approaches that use only minority class characteristics. The process of artificial instance generation is guided by exploring areas where the value of the mutual class distribution is very small. This way, we ensure a smart oversampling procedure that can cope with difficult data distributions and alleviate the shortcomings of existing methods. The usefulness of the MC-RBO algorithm is evaluated on the basis of extensive experimental study and backed-up with a thorough statistical analysis. Obtained results show that by taking into account information coming from all of the classes and conducting a smart oversampling, we can significantly improve the process of learning from multiclass imbalanced data.
Bartosz Krawczyk, Michal Koziarski, Michal Wozniak 0001
IEEE Trans. Neural Networks Learn. Syst.1
2019 Active Learning with Abstaining Classifiers for Imbalanced Drifting Data Streams
abstract
Learning from data streams is one of the most promising and challenging domains in modern machine learning. Proliferating online data sources provide us access to real-time knowledge we have never had before. At the same time, new obstacles emerge and we have to overcome them in order to fully and effectively utilize the potential of the data. Prohibitive time and memory constraints or non-stationary distributions are only some of the problems. When dealing with classification tasks, one has to remember that effective adaptation has to be achieved on weak foundations of partially labeled and often imbalanced data. In our work, we propose an online framework for binary classification, that aims to handle the complex problem of working with dynamic, sparsely labeled and imbalanced streams. The main part of it is a novel active learning strategy (MD-OAL) that is able to prioritize labeling of minority instances and, as a result, improve the balance of the learning process. We combine the strategy with a dynamic ensemble of base learners that can abstain from making decisions, if they are very uncertain. We adjust the abstaining mechanism in favor of minority instances, providing an effective method for handling remaining imbalance and a concept drift simultaneously. The conducted evaluation shows that in the challenging and realistic scenarios our framework outperforms state-of-the-art algorithms, providing higher resilience to the combined effect of limited labeling and imbalance.
Lukasz Korycki, Alberto Cano 0001, Bartosz Krawczyk
IEEE BigData3
2019 Bagging Using Instance-Level Difficulty for Multi-Class Imbalanced Big Data Classification on Spark
abstract
Most machine learning methods work under the assumption that classes have a roughly balanced number of instances. However, in many real-life problems we may have some types of instances appearing predominantly more frequently than the others which causes a bias towards the majority class during classifier training. This becomes even more challenging when dealing with multiple classes, where relationships between them are not easily defined. Learning from multi-class imbalanced data has not been widely considered in the context of big data mining, despite the fact that this is a learning difficulty frequently appearing in this domain. In this paper, we address this challenge by proposing a comprehensive ensemble-based framework. We propose to analyze each class to extract instance-level characteristics describing their difficulty levels. We embed this information into the existing UnderBagging framework. Our ensemble samples instances with probabilities proportional to their difficulty levels. This allows us to focus the learning process on the most difficult instances, better capturing the properties of multi-class imbalanced problems. We implemented our framework on Apache Spark to allow for high-performance computing over big data sets. This experimental study shows that taking into account the instance-level difficulty leads to training of significantly more accurate ensembles.
William C. Sleeman IV, Bartosz Krawczyk
IEEE BigData2
2019 Unsupervised Drift Detector Ensembles for Data Stream Mining
abstract
Data stream mining is among the most contemporary branches of machine learning. The potentially infinite sources give us many opportunities and at the same time pose new challenges. To properly handle streaming data we need to improve our well-established methods, so they can work with dynamic data and under strict constraints. Supervised streaming machine learning algorithms require a certain number of labeled instances in order to stay up-to-date. Since high budgets dedicated for this purpose are usually infeasible, we have to limit the supervision as much as we can. One possible approach is to trigger labeling, only if a change is explicitly indicated by a detector. While there are several supervised algorithms dedicated for this purpose, the more practical unsupervised ones are still lacking a proper attention. In this paper, we propose a novel unsupervised ensemble drift detector that recognizes local changes in feature subspaces (EDFS) without additional supervision, using specialized committees of incremental Kolmogorov-Smirnov tests. We combine it with an adaptive classifier and update it, only if the drift detector signalizes a change. Conducted experiments show that our framework is able to efficiently adapt to various concept drifts and outperform other unsupervised algorithms.
Lukasz Korycki, Bartosz Krawczyk
DSAA2
2019 Adaptive Ensemble Active Learning for Drifting Data Stream Mining
abstract
Learning from data streams is among the most vital contemporary fields in machine learning and data mining. Streams pose new challenges to learning systems, due to their volume and velocity, as well as ever-changing nature caused by concept drift. Vast majority of works for data streams assume a fully supervised learning scenario, having an unrestricted access to class labels. This assumption does not hold in real-world applications, where obtaining ground truth is costly and time-consuming. Therefore, we need to carefully select which instances should be labeled, as usually we are working under a strict label budget. In this paper, we propose a novel active learning approach based on ensemble algorithms that is capable of using multiple base classifiers during the label query process. It is a plug-in solution, capable of working with most of existing streaming ensemble classifiers. We realize this process as a Multi-Armed Bandit problem, obtaining an efficient and adaptive ensemble active learning procedure by selecting the most competent classifier from the pool for each query. In order to better adapt to concept drifts, we guide our instance selection by measuring the generalization capabilities of our classifiers. This adaptive solution leads not only to better instance selection under sparse access to class labels, but also to improved adaptation to various types of concept drift and increasing the diversity of the underlying ensemble classifier.
Bartosz Krawczyk, Alberto Cano 0001
IJCAI1
2019 Towards highly accurate coral texture images classification using deep convolutional neural networks and data augmentation
Anabel Gómez-Ríos, Siham Tabik, Julián Luengo, A. S. M. Shihavuddin, Bartosz Krawczyk, Francisco Herrera
Expert Syst. Appl.5
2019 Monotonic classification: An overview on algorithms, performance measures and data sets
José Ramón Cano, Pedro Antonio Gutiérrez, Bartosz Krawczyk, Michal Wozniak 0001, Salvador García 0001
Neurocomputing3
2019 Radial-Based oversampling for noisy imbalanced data classification
Michal Koziarski, Bartosz Krawczyk, Michal Wozniak 0001
Neurocomputing2
2019 Speeding up k-Nearest Neighbors classifier for large-scale multi-label learning on GPUs
Przemyslaw Skryjomski, Bartosz Krawczyk, Alberto Cano 0001
Neurocomputing2
2019 Instance reduction for one-class classification
Bartosz Krawczyk, Isaac Triguero, Salvador García 0001, Michal Wozniak 0001, Francisco Herrera
Knowl. Inf. Syst.1
2019 Evolving rule-based classifiers with genetic programming on GPUs for drifting data streams
Alberto Cano 0001, Bartosz Krawczyk
Pattern Recognit.2
2019 Multi-Label Punitive kNN with Self-Adjusting Memory for Drifting Data Streams
abstract
In multi-label learning, data may simultaneously belong to more than one class. When multi-label data arrives as a stream, the challenges associated with multi-label learning are joined by those of data stream mining, including the need for algorithms that are fast and flexible, able to match both the speed and evolving nature of the stream. This article presents a punitive k nearest neighbors algorithm with a self-adjusting memory (MLSAMPkNN) for multi-label, drifting data streams. The memory adjusts in size to contain only the current concept and a novel punitive system identifies and penalizes errant data examples early, removing them from the window. By retaining and using only data that are both current and beneficial, MLSAMPkNN is able to adapt quickly and efficiently to changes within the data stream while still maintaining a low computational complexity. Additionally, the punitive removal mechanism offers increased robustness to various data-level difficulties present in data streams, such as class imbalance and noise. The experimental study compares the proposal to 24 algorithms using 30 real-world and 15 artificial multi-label data streams on six multi-label metrics, evaluation time, and memory consumption. The superior performance of the proposed method is validated through non-parametric statistical analysis, proving both high accuracy and low time complexity. MLSAMPkNN is a versatile classifier, capable of returning excellent performance in diverse stream scenarios.
Martha I. Roseberry, Bartosz Krawczyk, Alberto Cano 0001
ACM Trans. Knowl. Discov. Data2
2018 Clustering-Driven and Dynamically Diversified Ensemble for Drifting Data Streams
abstract
Data stream mining is a rapidly developing branch of contemporary machine learning. Ensemble approaches have proven themselves to be highly effective in this domain, due to their predictive power and capabilities for handling evolving data. One of the key aspects of ensemble learning is diversity among base classifiers - it improves accuracy and allows for anticipating and recovering from concept drifts. It has been shown that while diversity is desirable during changes, it may impede learning when data becomes stationary. In this paper, we present a novel ensemble technique that exploits the idea of dynamic diversification, which increases diversity during changes and reduces it when a stream becomes stable. The algorithm uses online clustering for this task by creating locally specialized base learners trained on spatially related instances. Three control strategies based on the novel range heuristic for managing a trade-off between error (a change indicator) and diversity are utilized. Additionally, two intensification strategies are proposed for exploitation of newly arriving instances, allowing for faster adaptation. Experimental study evaluates the general performance and diversity of the proposed algorithm, proving its capabilities to outperform state-of-the-art ensembles dedicated to drifting data stream mining.
Lukasz Korycki, Bartosz Krawczyk
IEEE BigData2
2018 Combining active learning with concept drift detection for data stream mining
abstract
Most of data stream classifier learning methods assume that a true class of an incoming object is available right after the instance has been processed and new and labeled instance may be used to update a classifier's model, drift detection or capturing novel concepts. However, assumption that we have an unlimited and infinite access to class labels is very naive and usually would require a very high labeling cost. Therefore the applicability of many supervised techniques is limited in real-life stream analytics scenarios. Active learning emerges as a potential solution to this problem, concentrating on selecting only the most valuable instances and learning an accurate predictive model with as few labeling queries as possible. However learning from data streams differ from online learning as distribution of examples may change over time. Therefore, an active learning strategy must be able to handle concept drift and quickly adapt to evolving nature of data. In this paper we present novel active learning strategies that are designed for effective tackling of such changes. We assume that most labeling effort is required when concept drift occurs, as we need a representative sample of new concept to retrain properly the predictive model. Therefore, we propose active learning strategies that are guided by drift detection module to save budget for difficult and evolving instances. Three proposed strategies are based on learner uncertainty, dynamic allocation of budget over time and search space randomization. Experimental evaluation of the proposed methods prove their usefulness for reducing labeling effort in learning from drifting data streams.
Bartosz Krawczyk, Bernhard Pfahringer, Michal Wozniak 0001
IEEE BigData1
2018 Learning Classification Rules with Differential Evolution for High-Speed Data Stream Mining on GPU s
abstract
High-speed data streams are potentially infinite sequences of rapidly arriving instances that may be subject to concept drift phenomenon. Hence, dedicated learning algorithms must be able to update themselves with new data and provide an accurate prediction in a limited amount of time. This requirement was considered as prohibitive for using evolutionary algorithms for high-speed data stream mining. This paper introduces a massively parallel implementation on GPUs of a differential evolution algorithm for learning classification rules in the presence of concept drift. The proposal based on the DE /rand - to - best/1/bin strategy takes advantage of up to four nested levels of parallelism to maximize the performance of the algorithm. Efficient GPU kernels parallelize the evolution of the populations, rules, conditional clauses, and evaluation on instances. The proposed method is evaluated on 25 data stream benchmarks considering different types of concept drifts. Results are compared with other publicly available streaming rule learners. Obtained results and their statistical analysis proves an excellent performance of the proposed classifier that offers improved predictive accuracy, model update time, decision time, and a compact rule set.
Alberto Cano 0001, Bartosz Krawczyk
CEC2
2018 An Empirical Insight Into Concept Drift Detectors Ensemble Strategies
abstract
Contemporary decision support systems have to take into consideration the fact that most of data gathered is nowadays in motion, i.e., that successive observations, about objects being analyzed, form so-called data streams. Unfortunately, during analytical model utilization, unpredictable changes may appear in data distributions, leading to significant deterioration in the predictive performance and reliability of these learners. This phenomenon is called concept drift and refers to changes in the input data in relation to target variable in supervised learning task. Due to its potentially catastrophic impact on the underlying learner, it must be detected and handled as soon as it occurs. Over the years, many methods have been developed to address this issue. We focus on supervised classification task, aiming at answering the question on how to detect significant changes in data distribution effectively using ensemble of drift detectors. We discuss several models of combined drift detectors, among them the local detector which analyses distribution of each attribute separately. Experimental evaluations confirm the effectiveness of ensemble detectors, making them highly interesting to be used in solving real-world problems.
Andrzej Lapinski, Bartosz Krawczyk, Pawel Ksieniewicz, Michal Wozniak 0001
CEC2
2018 Addressing Local Class Imbalance in Balanced Datasets with Dynamic Impurity Decision Trees
Andriy Mulyar, Bartosz Krawczyk
DS2
2018 Synthetic Oversampling with the Majority Class: A New Perspective on Handling Extreme Imbalance
abstract
The class imbalance problem is a pervasive issue in many real-world domains. Oversampling methods that inflate the rare class by generating synthetic data are amongst the most popular techniques for resolving class imbalance. However, they concentrate on the characteristics of the minority class and use them to guide the oversampling process. By completely overlooking the majority class, they lose a global view on the classification problem and, while alleviating the class imbalance, may negatively impact learnability by generating borderline or overlapping instances. This becomes even more critical when facing extreme class imbalance, where the minority class is strongly underrepresented and on its own does not contain enough information to conduct the oversampling process. We propose a novel method for synthetic oversampling that uses the rich information inherent in the majority class to synthesize minority class data. This is done by generating synthetic data that is at the same Mahalanbois distance from the majority class as the known minority instances. We evaluate over 26 benchmark datasets, and show that our method offers a distinct performance improvement over the existing state-of-the-art in oversampling techniques.
Shiven Sharma, Colin Bellinger, Bartosz Krawczyk, Osmar R. Zaïane, Nathalie Japkowicz
ICDM3
2018 Selecting local ensembles for multi-class imbalanced data classification
abstract
Learning from imbalanced data is a challenge that machine learning community is facing over last decades, due to its ever-growing presence in real-life problems. While there is a significant number of works addressing the issue of handling binary and skewed datasets, its multi-class counterpart have not received as much attention. This problem is much more difficult, as presence of multiple imbalanced classes can significantly deteriorate the predictive power of any classifier. The relationship among classes are no longer clearly established and there are many difficulties embedded in the nature of such data that needs to be properly addressed. In this work, we discuss the issue of forming effective ensembles for multi-class imbalanced data based on static classifier selection approach. We propose a fully adaptive learning scheme that splits the original feature space into a number of competence areas and modifies their size and location in order to most effectively exploit the supplied pool of base classifiers. Additionally, for each established cluster we perform a weighted classifier combination, where weights are set individually for each cluster and each considered class. This allows for exploiting local competencies of each base learner in given part of feature space, as well as for each of considered classes. These two tasks are combined together in a single hybrid training scheme guided by an evolutionary algorithm. The optimization criterion is formulated in order to achieve skew-insensitive ensemble of local ensembles able to tackle highly imbalanced and multi-class problems. Experimental study proves the high efficacy of the proposed method and its superiority to other ensemble selection methods.
Bartosz Krawczyk, Alberto Cano 0001, Michal Wozniak 0001
IJCNN1
2018 Leveraging Ensemble Pruning for Imbalanced Data Classification
abstract
The effectiveness of machine learning algorithms depends on the quality of the supplied training data. Any problems embedded in the nature of data will result in obtaining incorrect classification models, especially imbalanced data distribution is among the most significant learning difficulties that can affect classifiers. As one of the classes has much more instances than the other, the learning process becomes biased towards it. Therefore, methods for alleviating the impact of skewed distributions are highly sought after. Ensemble learning has emerged as one of the leading paradigms for imbalanced data. Creation of an efficient pool of classifiers is not a trivial task and one needs to carefully select which classifiers should be combined to obtain the best predictive power. In this paper, we propose a compound ensemble pruning algorithm for imbalanced data. It aims to retain classifiers that offer the best performance on both minority and majority classes, and display a high level of diversity. Remaining learners are discarded from the pool. This is achieved by the means of a multi-criteria evolutionary algorithm. Extensive experimental study show that our proposal is able to create smaller ensembles than the state-of-the-art methods, while offering an improved robustness to imbalanced class distributions.
Bartosz Krawczyk, Michal Wozniak 0001
SMC1
2018 Ensemble of Extreme Learning Machines with trained classifier combination and statistical features for hyperspectral data
Pawel Ksieniewicz, Bartosz Krawczyk, Michal Wozniak 0001
Neurocomputing2
2018 Dynamic ensemble selection for multi-class classification with one-class classifiers
Bartosz Krawczyk, Mikel Galar, Michal Wozniak 0001, Humberto Bustince, Francisco Herrera
Pattern Recognit.1
2018 Local ensemble learning from imbalanced and noisy data for word sense disambiguation
Bartosz Krawczyk, Bridget T. McInnes
Pattern Recognit.1
2017 Online query by committee for active learning from drifting data streams
abstract
Most of data stream learning methods assume that a true class of an incoming instance is available right after it has been processed. However, assumption that we have an unlimited access to class labels is unrealistic and is directly connected with a very high labeling cost. This is a driving force behind growing development of methods that require reduced or no access to class labels. Among several potential directions active learning emerges as a promising solution, by allowing for a selection of most valuable instances from the stream and using as few label queries. Despite numerous proposals of active learning methods for static data, this domain is still developing for data streams. Here, non-stationary nature of data must be taken into consideration and proposed algorithms must accommodate potential occurrences of concept drift. In this paper we propose a Query by Committee active learning strategy that is adapted to online learning from drifting data streams. A decision regarding label query is made by an ensemble of classifiers instead of a single learner, leading to an improved instance selection. We present four different approaches for online Query by Committee and evaluate their usefulness on the basis of obtained accuracy with limited budgets and ability to handle concept drift. We introduce Budget Loss of Accuracy, a novel measure for evaluating active learning algorithms. Finally, we investigate the relationships between the efficacy of Query by Committee models and diversity of underlying ensembles. Based on thorough experimental investigation we are able to show the usefulness of proposed algorithms for reducing labeling effort in learning from drifting data streams.
Bartosz Krawczyk, Michal Wozniak 0001
IJCNN1
2017 Cost-Sensitive Perceptron Decision Trees for Imbalanced Drifting Data Streams
Bartosz Krawczyk, Przemyslaw Skryjomski
ECML/PKDD (2)1
2017 Fault diagnosis of marine 4-stroke diesel engines using a one-vs-one extreme learning ensemble
Jerzy Kowalski, Bartosz Krawczyk, Michal Wozniak 0001
Eng. Appl. Artif. Intell.2
2017 A survey on data preprocessing for data stream mining: Current status and future directions
Sergio Ramírez-Gallego, Bartosz Krawczyk, Salvador García 0001, Michal Wozniak 0001, Francisco Herrera
Neurocomputing2
2017 Active and adaptive ensemble learning for online activity recognition from data streams
Bartosz Krawczyk
Knowl. Based Syst.1
2017 The deterministic subspace method for constructing classifier ensembles
abstract
Ensemble classification remains one of the most popular techniques in contemporary machine learning, being characterized by both high efficiency and stability. An ideal ensemble comprises mutually complementary individual classifiers which are characterized by the high diversity and accuracy. This may be achieved, e.g., by training individual classification models on feature subspaces. Random Subspace is the most well-known method based on this principle. Its main limitation lies in stochastic nature, as it cannot be considered as a stable and a suitable classifier for real-life applications. In this paper, we propose an alternative approach, Deterministic Subspace method, capable of creating subspaces in guided and repetitive manner. Thus, our method will always converge to the same final ensemble for a given dataset. We describe general algorithm and three dedicated measures used in the feature selection process. Finally, we present the results of the experimental study, which prove the usefulness of the proposed method.
Michal Koziarski, Bartosz Krawczyk, Michal Wozniak 0001
Pattern Anal. Appl.2
2017 Selecting locally specialised classifiers for one-class classification ensembles
abstract
One-class classification belongs to the one of the novel and very promising topics in contemporary machine learning. In recent years ensemble approaches have gained significant attention due to increasing robustness to unknown outliers and reducing the complexity of the learning process. In our previous works, we proposed a highly efficient one-class classifier ensemble, based on input data clustering and training weighted one-class classifiers on clustered subsets. However, the main drawback of this approach lied in difficult and time consuming selection of a number of competence areas which indirectly affects a number of members in the ensemble. In this paper, we investigate ten different methodologies for an automatic determination of the optimal number of competence areas for the proposed ensemble. They have roots in model selection for clustering, but can be also effectively applied to the classification task. In order to select the most useful technique, we investigate their performance in a number of one-class and multi-class problems. Numerous experimental results, backed-up with statistical testing, allows us to propose an efficient and fully automatic method for tuning the one-class clustering-based ensembles.
Bartosz Krawczyk, Boguslaw Cyganek
Pattern Anal. Appl.1
2017 Nearest Neighbor Classification for High-Speed Big Data Streams Using Spark
abstract
Mining massive and high-speed data streams among the main contemporary challenges in machine learning. This calls for methods displaying a high computational efficacy, with ability to continuously update their structure and handle ever-arriving big number of instances. In this paper, we present a new incremental and distributed classifier based on the popular nearest neighbor algorithm, adapted to such a demanding scenario. This method, implemented in Apache Spark, includes a distributed metric-space ordering to perform faster searches. Additionally, we propose an efficient incremental instance selection method for massive data streams that continuously update and remove outdated examples from the case-base. This alleviates the high computational requirements of the original classifier, thus making it suitable for the considered problem. Experimental study conducted on a set of real-life massive data streams proves the usefulness of the proposed solution and shows that we are able to provide the first efficient nearest neighbor solution for high-speed big and streaming data.
Sergio Ramírez-Gallego, Bartosz Krawczyk, Salvador García 0001, Michal Wozniak 0001, José Manuel Benítez 0001, Francisco Herrera
IEEE Trans. Syst. Man Cybern. Syst.2
2016 Hybrid One-Class Ensemble for High-Dimensional Data Classification
Bartosz Krawczyk
ACIIDS (2)1
2016 Forming Classifier Ensembles with Deterministic Feature Subspaces
abstract
Ensemble learning is being considered as one of the most well-established and efficient techniques in the contemporary machine learning.The key to the satisfactory performance of such combined models lies in the supplied base learners and selected combination strategy.In this paper we will focus on the former issue.Having classifiers that are of high individual quality and complementary to each other is a desirable property.Among several ways to ensure diversity feature space division deserves attention.The most popular method employed here is Random Subspace approach.However, due to its random nature one cannot consider this approach as stable one or suitable for reallife applications.Therefore, we propose a new approach called Deterministic Subspace that constructs feature subspaces in a guided and repetitive manner.We present a general framework and three dedicated measures that can be used for selecting diverse and uncorrelated features for each base learner.This way we will always obtain identical sets of features, leading to creation of stable ensembles.Experimental study backed-up with statistical analysis prove the usefulness of our method in comparison to popular randomized solution.
Michal Koziarski, Bartosz Krawczyk, Michal Wozniak 0001
FedCSIS2
2016 Tackling label noise with multi-class decomposition using fuzzy one-class support vector machines
abstract
Class label noise is a data-level difficulty associated with training objects with incorrectly assigned labels. This problem may originate from poorly documented historic data, errors during data generation process or mistakes made by human experts. Inclusion of such examples during the training process will mislead the classifier by presenting a falsified class distribution and consequently lead to degradation of models' generalization abilities. This phenomenon becomes even more troublesome in multi-class scenarios that may be affected by highly complex intra-class noise. Decomposition strategies with binary classifiers were proven to alleviate this difficulty by using simplified binary subtasks that are less affected by the noise. In this paper we propose to extend this approach by using the one-class classification decomposition. In this scenario each class has assigned individual one-class classifier that aims at capturing its distinguishing characteristics. This allows to create a robust data description and then apply a dedicated classifier combination in order to reconstruct the original multi-class task. We further extend this concept by using fuzzy one-class classifiers that allow to associate membership values with each training objects. This allows us to reduce the influence of uncertain and potentially noisy samples on the shape of learned decision boundary. Experimental study backed-up with statistical analysis shows that fuzzy one-class classifier decomposition offers an excellent robustness to noise in multi-class classification.
Bartosz Krawczyk, José A. Sáez, Michal Wozniak 0001
FUZZ-IEEE1
2016 Cost-sensitive one-vs-one ensemble for multi-class imbalanced data
abstract
Learning from imbalanced data poses significant challenges for machine learning algorithms, as they need to deal with uneven distribution of examples in the training set. As standard classifiers will be biased towards the majority class there exist a need for specific methods than can overcome this single-class dominance. Most of works concentrated on binary problems, where majority and minority class can be distinguished. But a more challenging problem arises when imbalance is present within multi-class datasets, as relations between classes tend to complicate. One class can be a minority class for some, while a majority for others. In this paper, we propose an efficient method for handling such scenarios, that combines the problem decomposition with cost-sensitive learning. According to divide-and-conquer rule, we decompose our multi-class data into a number of binary subproblems using one-versus-one approach. To each simplified task we delegate a cost-sensitive neural network with moving threshold. It relies on scaling the output of the classifier with a given cost function. This way, we adjust our support functions towards the minority class. We propose a novel method for automatically determining the cost, based on the Receiver Operating Characteristic (ROC) curve analysis. This way we can estimate the cost matrix for each class pair independently. Then using a dedicated classifier fusion approach, we reconstruct the original multi-class problem. Experimental analysis backed-up with statistical testing clearly proves that such an approach is superior to state-of-the art ad-hoc and decomposition methods used in the literature.
Bartosz Krawczyk
IJCNN1
2016 Untrained weighted classifier combination with embedded ensemble pruning
Bartosz Krawczyk, Michal Wozniak 0001
Neurocomputing1
2016 Dynamic classifier selection for one-class classification
Bartosz Krawczyk, Michal Wozniak 0001
Knowl. Based Syst.1
2016 Empowering one-vs-one decomposition with ensemble learning for multi-class imbalanced data
Zhongliang Zhang 0001, Bartosz Krawczyk, Salvador García 0001, Alejandro Rosales-Pérez, Francisco Herrera
Knowl. Based Syst.2
2016 Analyzing the oversampling of different classes and types of examples in multi-class imbalanced datasets
José A. Sáez, Bartosz Krawczyk, Michal Wozniak 0001
Pattern Recognit.2
2015 Data Classification with Ensembles of One-Class Support Vector Machines and Sparse Nonnegative Matrix Factorization
Boguslaw Cyganek, Bartosz Krawczyk
ACIIDS (1)2
2015 Pruning Ensembles of One-Class Classifiers with X-means Clustering
Bartosz Krawczyk, Michal Wozniak 0001
ACIIDS (1)1
2015 Pruning Ensembles with Cost Constraints
Bartosz Krawczyk, Michal Wozniak 0001
ACIIDS (1)1
2015 Cost-Sensitive Neural Network with ROC-Based Moving Threshold for Imbalanced Classification
Bartosz Krawczyk, Michal Wozniak 0001
IDEAL1
2015 Weighted Naïve Bayes Classifier with Forgetting for Drifting Data Streams
abstract
Mining massive data streams in real-time is one of the contemporary challenges for machine learning systems. Such a domain encompass many of difficulties hidden beneath the term of Big Data. We deal with massive, incoming information that must be processed on-the-fly, with lowest possible response delay. We are forced to take into account time, memory and quality constraints. Our models must be able to quickly process large collection of data and swiftly adapt themselves to occurring changes (shifts and drifts) in data streams. In this paper, we propose a novel version of simple, yet effective Naïve Bayes classifier for mining streams. We add a weighting module, that automatically assigns an importance factor to each object extracted from the stream. The higher the weight, the bigger influence given object exerts on the classifier training procedure. We assume, that our model works in the non-stationary environment with the presence of concept drift phenomenon. To allow our classifier to quickly adapt its properties to evolving data, we imbue it with forgetting principle implemented as weight decay. With each passing iteration, the level of importance of previous objects is decreased until they are discarded from the data collection. We propose an efficient sigmoidal function for modeling the forgetting rate. Experimental analysis, carried out on a number of large data streams with concept drift prove that our weighted Naïve Bayes classifier displays highly satisfactory performance in comparison with state-of-the-art stream classifiers.
Bartosz Krawczyk, Michal Wozniak 0001
SMC1
2015 A hybrid cost-sensitive ensemble for imbalanced breast thermogram classification
Bartosz Krawczyk, Gerald Schaefer, Michal Wozniak 0001
Artif. Intell. Medicine1
2015 Multidimensional data classification with chordal distance based kernel and Support Vector Machines
Boguslaw Cyganek, Bartosz Krawczyk, Michal Wozniak 0001
Eng. Appl. Artif. Intell.2
2015 One-class classifier ensemble pruning and weighting with firefly algorithm
Bartosz Krawczyk
Neurocomputing1
2015 Data stream classification and big data analytics
Bartosz Krawczyk, Jerzy Stefanowski, Michal Wozniak 0001
Neurocomputing1
2015 On the usefulness of one-class classifier ensembles for decomposition of multi-class problems
Bartosz Krawczyk, Michal Wozniak 0001, Francisco Herrera
Pattern Recognit.1
2015 One-class classifiers with incremental learning and forgetting for data streams with concept drift
abstract
One of the most important challenges for machine learning community is to develop efficient classifiers which are able to cope with data streams, especially with the presence of the so-called concept drift. This phenomenon is responsible for the change of classification task characteristics, and poses a challenge for the learning model to adapt itself to the current state of the environment. So there is a strong belief that one-class classification is a promising research direction for data stream analysis—it can be used for binary classification without an access to counterexamples, decomposing a multi-class data stream, outlier detection or novel class recognition. This paper reports a novel modification of weighted one-class support vector machine, adapted to the non-stationary streaming data analysis. Our proposition can deal with the gradual concept drift, as the introduced one-class classifier model can adapt its decision boundary to new, incoming data and additionally employs a forgetting mechanism which boosts the ability of the classifier to follow the model changes. In this work, we propose several different strategies for incremental learning and forgetting, and additionally we evaluate them on the basis of several real data streams. Obtained results confirmed the usability of proposed classifier to the problem of data stream classification with the presence of concept drift. Additionally, implemented forgetting mechanism assures the limited memory consumption, because only quite new and valuable examples should be memorized.
Bartosz Krawczyk, Michal Wozniak 0001
Soft Comput.1
2014 Optimization Algorithms for One-Class Classification Ensemble Pruning
Bartosz Krawczyk, Michal Wozniak 0001
ACIIDS (2)1
2014 A first attempt on evolutionary prototype reduction for nearest neighbor one-class classification
abstract
Evolutionary prototype reduction techniques are data preprocessing methods originally developed to enhance the nearest neighbor rule. They reduce the training data by selecting or generating representative examples of a given problem. These algorithms have been designed and widely analyzed in standard classification providing very competitive results. However, its application scope can be extended to many other specific domains, such as one-class classification, in which its way of working is very interesting in order to reduce computational complexity and sensitivity to noisy data. In this contribution, we perform a first study on the usefulness of evolutionary prototype reduction methods for one-class classification. To do so, we will focus on two recent evolutionary approaches that follow very different strategies: selection and generation of examples from the training data. Both alternatives provide a resulting preprocessed data set that will be used later by a nearest neighbor one-class classifier as its training data. The results achieved support that these data reduction techniques are suitable tools to improve the performance of the nearest neighbor one-class classification.
Bartosz Krawczyk, Isaac Triguero, Salvador García 0001, Michal Wozniak 0001, Francisco Herrera
IEEE Congress on Evolutionary Computation1
2014 Cost-sensitive texture classification
abstract
Texture recognition plays an important role in many computer vision tasks including segmentation, scene understanding and interpretation, medical imaging and object recognition. In some situations, the correct identification of particular textures is more important compared to others, for example recognition of enemy uniforms for automatic defense systems, or isolation of textures related to tumors in medical images. Such cost-sensitive texture classification is the focus of this paper, which we address by reformulating the classification problem as a cost minimisation problem. We do this by constructing a cost-sensitive classifier ensemble that is tuned using a genetic algorithm. Based on experimental results obtained on several Outex datasets with cost definitions, we show our approach to work well in comparison with canonical classification methods and the ensemble approach to lead to better performance compared to single predictors.
Gerald Schaefer, Bartosz Krawczyk, Niraj P. Doshi, Tomoharu Nakashima
IEEE Congress on Evolutionary Computation2
2014 Weighted one-class classification for different types of minority class examples in imbalanced data
abstract
Imbalanced classification is one of the most challenging machine learning problem. Recent studies show, that often the uneven ratio of objects in classes is not the biggest factor, determining the drop of classification accuracy. It is also related to some difficulties embedded in the nature of the data. In this paper we study the different types of minority class examples and distinguish four groups of objects - safe, borderline, rare and outliers. To deal with the imbalance problem, we use a one-class classification, that is focused on a proper identification of the minority class samples. We further augment this model by incorporating the knowledge about the minority object types in the training dataset. This is done applying weighted one-class classifier and adjusting weights assigned to minority class objects, depending on their type. A strategy for calculating the new weights for minority examples is proposed. Experimental analysis, carried on a set of benchmark datasets, confirms that the proposed model can achieve a satisfactory recognition rate and often outperform other state-of-the-art methods, dedicated to the imbalanced classification.
Bartosz Krawczyk, Michal Wozniak 0001, Francisco Herrera
CIDM1
2014 Designing a compact Genetic fuzzy rule-based system for one-class classification
abstract
This paper proposes a method for designing Fuzzy Rule-Based Classification Systems to deal with One-Class Classification, where during the training phase we have access only to objects originating from a single class. However, the trained model must be prepared to deal with new, unseen adversarial objects, known as outliers. We use a Genetic Algorithm for learning the granularity, domains and fuzzy partitions of the model and we propose an ad-hoc rule generation method specific for One-Class Classification. Several datasets from UCI repository, previously transformed to one-class problems, are used in the experiments and we compare with two of the classical methods used in the One-Class community, one-class Support Vector Machines and Support Vector Data Description. Our proposal of fuzzy model obtains similar results than the other methods but presents a high interpretability due its reduced number of rules.
Pedro Villar, Bartosz Krawczyk, Rosana Montes-Soldado, Francisco Herrera
FUZZ-IEEE2
2014 Weighted One-Class Classifier Ensemble Based on Fuzzy Feature Space Partitioning
abstract
This paper introduces a novel method for forming efficient one-class classifier ensembles. A common problem in one-class classification is a complex structure of the target class, which often leads to creation of a too expanded decision boundary. We propose to employ a clustering step in order to partition the target class into atomic subsets and using these as input for one-class classifiers. By this, we are able to detect sub-structures in the target concept. Additionally, to increase the diversity and robustness of our method weighted one-class classifiers are used. We introduce a novel scheme for calculating weights for training objects. Membership functions, obtained from the fuzzy clustering, are used to initialize the weighted classifiers. Based on the results of a number of computational experiments we show that the proposed method outperforms both the single one-class methods, as well as popular one-class ensembles. Other advantages are the highly parallel structure of the proposed solution, which facilitates parallel training and execution stages, and the relatively small number of control parameters.
Bartosz Krawczyk, Michal Wozniak 0001, Boguslaw Cyganek
ICPR1
2014 New untrained aggregation methods for classifier combination
abstract
The combined classification is a promising direction in pattern recognition and there are numerous methods that deal with forming classifier ensembles. The most popular approaches employ voting, where the final decision of compound classifier is a combination of individual classifiers' outputs, i.e., class labels or support functions. This paper concentrates on the problem how to design an effective combination rule, which takes into consideration the values of support functions returned by the individual classifiers. Because in many practical tasks we do not have a training set at our disposal, then we express our interest in aggregation methods which do not require learning. A special attention is paid to weighted aggregation, especially when the different weights depend on particular support function of a given individual classifier. We propose a novel approach for untrained combination of support functions using the Gaussian function to assign mentioned above weights. The computer experiments carried out on the set of benchmark data sets confirm the advantages of the proposed approach for particular cases, especially when the number of class labels is high.
Bartosz Krawczyk, Michal Wozniak 0001
IJCNN1
2014 Untrained Method for Ensemble Pruning and Weighted Combination
Bartosz Krawczyk, Michal Wozniak 0001
ISNN1
2014 One-Class Classification Ensemble with Dynamic Classifier Selection
Bartosz Krawczyk, Michal Wozniak 0001
ISNN1
2014 Cytological image analysis with firefly nuclei detection and hybrid one-class classification decomposition
Bartosz Krawczyk, Pawel Filipczuk
Eng. Appl. Artif. Intell.1
2014 Improved Adaptive Splitting and Selection: the Hybrid Training Method of a Classifier Based on a Feature Space Partitioning
abstract
Currently, methods of combined classification are the focus of intense research. A properly designed group of combined classifiers exploiting knowledge gathered in a pool of elementary classifiers can successfully outperform a single classifier. There are two essential issues to consider when creating combined classifiers: how to establish the most comprehensive pool and how to design a fusion model that allows for taking full advantage of the collected knowledge. In this work, we address the issues and propose an AdaSS+, training algorithm dedicated for the compound classifier system that effectively exploits local specialization of the elementary classifiers. An effective training procedure consists of two phases. The first phase detects the classifier competencies and adjusts the respective fusion parameters. The second phase boosts classification accuracy by elevating the degree of local specialization. The quality of the proposed algorithms are evaluated on the basis of a wide range of computer experiments that show that AdaSS+ can outperform the original method and several reference classifiers.
Konrad Jackowski, Bartosz Krawczyk, Michal Wozniak 0001
Int. J. Neural Syst.2
2014 Diversity measures for one-class classifier ensembles
Bartosz Krawczyk, Michal Wozniak 0001
Neurocomputing1
2014 Clustering-based ensembles for one-class classification
Bartosz Krawczyk, Michal Wozniak 0001, Boguslaw Cyganek
Inf. Sci.1
2013 Adaptive Splitting and Selection Method for Noninvasive Recognition of Liver Fibrosis Stage
Bartosz Krawczyk, Michal Wozniak 0001, Tomasz Orczyk, Piotr Porwik
ACIIDS (2)1
2013 Combining One-Class Support Vector Machines for Microarray Classification
Bartosz Krawczyk
FedCSIS1
2013 Improved LBP texture classification using ensemble learning
abstract
Texture analysis and classification play an important role in many multimedia and computer vision applications. Local binary patterns (LBP) form a simple yet powerful texture descriptor characterising local neighbourhood properties, and consequently LBP variants are widely employed. In this paper, we demonstrate that through appropriate construction of a multiple classifier system, improved texture classification based on LBP features is possible. In particular, we employ a classifier ensemble where each classifier (a support vector machine) is trained in conjunction with a different feature selection method. The ensemble is then pruned based on a diversity measure, and the remaining models are combined using a neural fuser. Experimental results, obtained on Outex benchmark datasets and employing four LBP variants, confirm that our proposed approach leads to statistically significantly improved texture classification.
Gerald Schaefer, Bartosz Krawczyk, Niraj P. Doshi
ICME2
2013 Application of Adaptive Splitting and Selection Classifier to the Spam Filtering Problem
abstract
E-Mail spam is one of the major problems plaguing the contemporary Internet, causing an inconvenience to an individual user and financial loss to a company. Spam filtering allows for early detection of unwanted messages and separates them from the incoming e-mail. Nonetheless, designing an effective spam detection system is not a trivial task, due to the problems connected with the analysis of the e-mail content and the occurrence of variation in spam characteristics. This article presents an application of a novel ensemble classifier system for spam detection. The system is an extension of the adaptive splitting and selection (AdaSS) framework. The idea of the ensemble is based on the assumption that high effectiveness of detection can be obtained by exploitation of the local competency of a set of diverse elementary classifiers. Therefore, the ensemble training algorithm divides the feature space into several disjoint subspaces and assigns an area classifier to each of them. The area classifier consists of elementary classifiers that make a collective decision based on the weighted fusion of their support functions. The weight reflects the local competency of the classifier. To maintain the diversity of the pool of elementary classifiers, we exploit different e-mail feature extraction methods while filling the pool. There are two main extensions of the presented algorithm over original AdaSS: the aforementioned weighted fusion model used for decision making and adaptation of the AdaSS training procedure to process data streams featuring the concept drift. The effectiveness of the classifier model in spam recognition was verified in a series of experiments on two sets of spam databases. Comparison of the algorithm with some other state-of-the-art ensemble methods showed that the presented AdaSS extension can effectively recognize local competences of elementary classifiers and result in very high effectiveness of spam recognition outperforming competing methods.
Konrad Jackowski, Bartosz Krawczyk, Michal Wozniak 0001
Cybern. Syst.2
2013 Classifier ensemble for an effective cytological image analysis
Pawel Filipczuk, Bartosz Krawczyk, Michal Wozniak 0001
Pattern Recognit. Lett.2
2012 Experiments on distance measures for combining one-class classifiers
Bartosz Krawczyk, Michal Wozniak 0001
FedCSIS1
2012 Adaptive Splitting and Selection Algorithm for Classification of Breast Cytology Images
Bartosz Krawczyk, Pawel Filipczuk, Michal Wozniak 0001
ICCCI (1)1
2012 Effective multiple classifier systems for breast thermogram analysis
Bartosz Krawczyk, Gerald Schaefer
ICPR1
2012 Cost-Sensitive Splitting and Selection Method for Medical Decision Support System
Konrad Jackowski, Bartosz Krawczyk, Michal Wozniak 0001
IDEAL2