EDBT 2026 Demo / reviewers in the wild / expert
Michal Wozniak 0001
dblp:37/5714-1
· DBLP profile ↗
122ranked-venue papers
13as first author
24since 2021 · last 2026
0000-0003-0146-4205ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 92 · 7 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 26 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 22 · 6 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 2Theory of computation · 2 · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MFmamba: A Multi-function Network for Panchromatic Image Resolution Restoration Based on State-Space ModelabstractRemote sensing images are becoming increasingly widespread in military, earth resource exploration. Because of the limitation of a single sensor, we can obtain high spatial resolution grayscale panchromatic (PAN) images and low spatial resolution color multispectral (MS) images. Therefore, an important issue is to obtain a color image with high spatial resolution when there is only a PAN image at the input. The existing methods improve spatial resolution using super-resolution (SR) technology and spectral recovery using colorization technology. However, the SR technique cannot improve the spectral resolution, and the colorization technique cannot improve the spatial resolution. Moreover, the pansharpening method needs two registered inputs and can not achieve SR. As a result, an integrated approach is expected. We designed a novel multi-function model (MFmamba) to realize the tasks of SR, spectral recovery, joint SR and spectral recovery through three different inputs. Firstly, MFmamba utilizes UNet++ as the backbone, and a Mamba Upsample Block (MUB) is combined with UNet++. Secondly, a Dual Pool Attention (DPA) is designed to replace the skip connection in UNet++. Finally, a Multi-scale Hybrid Cross Block (MHCB) is proposed for initial feature extraction. Many experiments show that MFmamba is competitive in evaluation metrics and visual results and performs well in the three tasks when only the input PAN image is used. Qianqian Wang 0013, Xin Jin 0005, Michal Wozniak 0001, Shaowen Yao 0001, Wei Zhou 0011 |
AAAI | 4 |
| 2026 | Dynamic self-paced ensemble for imbalance-aware learning
Amgad M. Mohammed, Enrique Onieva, Michal Wozniak 0001 |
Pattern Anal. Appl. | 3 |
| 2025 | SR_ColorNet: Multi-path attention aggregated and mask enhanced network for the super resolution and colorization of panchromatic image
Qianqian Wang 0013, Shengfa Miao, Xin Jin 0005, Shin-Jye Lee, Michal Wozniak 0001, Shaowen Yao 0001 |
Expert Syst. Appl. | 6 |
| 2025 | DS-GAN: a dual sub-structure GAN for thermal infrared image colorization using U-Net with ConvNeXt and multi-scale large kernel attention
Guoliang Yao, Xin Jin 0005, Michal Wozniak 0001, Shengfa Miao, Shaowen Yao 0001, Wei Zhou 0011 |
Vis. Comput. | 4 |
| 2024 | Modeling configuration-performance relation in a mobile network: a data-driven approachabstractMobile network performance modeling typically assumes either a fixed cell’s configuration or only considers a limited number of parameters. This prohibits the exploration of multidimensional, diverse configuration space for, e.g., optimization purposes. This paper presents a method for performance predictions based on a network cell’s configuration and network conditions, which utilizes neural network architecture. We evaluate the idea by extensive experiments, with data from more than $\mathrm{5 0, 0 0 0} \mathrm{5 G}$ cells. The assessment included a comparison of the proposed method against models developed for fixed configuration. Results show that combined configuration-performance modeling outperforms single-configuration models and allows for performance prediction of unknown configurations, i.e., it is not used for model training. A substantially lower mean absolute error was achieved ($\mathrm{0 . 2 5}$ vs. $\mathrm{0 . 4 5}$ for fixed-configuration MLP-based models). Michal Panek, Ireneusz Jablonski, Michal Wozniak 0001 |
PIMRC | 3 |
| 2024 | A natural gas consumption forecasting system for continual learning scenarios based on Hoeffding trees with change point detection mechanism
Radek Svoboda, Sebastián Basterrech, Jedrzej Kozal, Jan Platos, Michal Wozniak 0001 |
Knowl. Based Syst. | 5 |
| 2024 | Local neighborhood encodings for imbalanced data classificationabstractAbstract This paper aims to propose Local Neighborhood Encodings (LNE)-a hybrid data preprocessing method dedicated to skewed class distribution balancing. The proposed LNE algorithm uses both over- and undersampling methods. The intensity of the methods is chosen separately for each fraction of minority and majority class objects. It is selected depending on the type of neighborhoods of objects of a given class, understood as the number of neighbors from the same class closest to a given object. The process of selecting the over- and undersampling intensities is treated as an optimization problem for which an evolutionary algorithm is used. The quality of the proposed method was evaluated through computer experiments. Compared with SOTA resampling strategies, LNE shows very good results. In addition, an experimental analysis of the algorithms behavior was performed, i.e., the determination of data preprocessing parameters depending on the selected characteristics of the decision problem, as well as the type of classifier used. An ablation study was also performed to evaluate the influence of components on the quality of the obtained classifiers. The evaluation of how the quality of classification is influenced by the evaluation of the objective function in an evolutionary algorithm is presented. In the considered task, the objective function is not de facto deterministic and its value is subject to estimation. Hence, it was important from the point of view of computational efficiency to investigate the possibility of using for quality assessment the so-called proxy classifier, i.e., a classifier of low computational complexity, although the final model was learned using a different model. The proposed data preprocessing method has high quality compared to SOTA, however, it should be noted that it requires significantly more computational effort. Nevertheless, it can be successfully applied to the case as no very restrictive model building time constraints are imposed. Michal Koziarski, Michal Wozniak 0001 |
Mach. Learn. | 2 |
| 2023 | A Continual Learning System with Self Domain Shift Adaptation for Fake News DetectionabstractDetecting fake news is currently one of the critical challenges facing modern societies. The problem is particularly relevant, as disinformation is readily used for political warfare but can also cause significant harm to the health of citizens, such as by promoting false data on the harmfulness of selected therapies. One way to combat disinformation is to treat fake news detection as a machine learning task. This paper presents such an approach, which additionally addresses an important problem related to the non-stationarity characteristics of the fake news. We elaborated a stream data with the simulation of domain shift based on two popular benchmark datasets dedicated to the fake news classification problem (Kaggle Fake News and Constraint@AAAI2021–COVID19 Fake News Detection). The proposed learning system works in a Continual Learning (CL) framework and integrates a self domain shift adaptation in a machine learning scheme. The method was built following state-of-the-art techniques, that includes Word2Vec as a feature extractor and the LSTM model as a classifier. The performance of the approach has been evaluated over the generated data stream. The convenience of our approach is showed in the results, where the accuracy gain with respect to a CL approach without domain adaptation is observed to be significant. Sebastián Basterrech, Andrzej Kasprzak, Jan Platos, Michal Wozniak 0001 |
DSAA | 4 |
| 2023 | DE-Forest - Optimized Decision Tree Ensemble
Joanna Klikowska, Michal Wozniak 0001 |
ICCCI | 2 |
| 2023 | SVM ensemble training for imbalanced data classification using multi-objective optimization techniquesabstractAbstract One of the main problems with classifier training for imbalanced data is defining the correct learning criterion. On the one hand, we want the minority class to be correctly recognized, and on the other hand, we do not want to make too many mistakes in the majority class. Commonly used metrics focus either on the predictive quality of the distinguished class or propose an aggregation of simple metrics. The aggregate metrics, such asGmeanorAUC, are primarily ambiguous, i.e., they do not indicate the specific values of errors made on the minority or majority class. Additionally, improper use of aggregate metrics results in solutions selected with their help that may favor the majority class. The authors realize that a solution to this problem is using overall risk. However, this requires knowledge of the costs associated with errors made between classes, which is often unavailable. Hence, this paper will propose thesemoosalgorithm - an approach based on multi-objective optimization that optimizes criteria related to the prediction quality of both minority and majority classes.semoosreturns a pool of non-dominated solutions from which the user can choose the model that best suits him. Automatic solution selection formulas with a so-called Pareto front have also been proposed to comparestate-of-the-artmethods. The proposed approach will train asvmclassifier ensemble dedicated to the imbalanced data classification task. The experimental evaluations carried out on a large number of benchmark datasets confirm its usefulness. Joanna Klikowska, Michal Wozniak 0001 |
Appl. Intell. | 2 |
| 2023 | Alphabet Flatting as a variant of n-gram feature extraction method in ensemble classification of fake news
Pawel Ksieniewicz, Pawel Zyblewski, Weronika Borek-Marciniec, Rafal Kozik, Michal Choras, Michal Wozniak 0001 |
Eng. Appl. Artif. Intell. | 6 |
| 2023 | Active Weighted Aging Ensemble for drifted data stream classificationabstractPurpose One of the significant problems in data stream classification is the concept drift phenomenon, which consists of the change in probabilistic characteristics of the classification task . Such changes in posterior probability destabilize the classification model performance, seriously degrading its quality. It is necessary to design appropriate strategies to counteract this phenomenon, allowing the classifier to adapt to the changing probabilistic characteristics. It is tough to propose such an approach with limited access to data labels. A human bias of high quality is usually costly, so to minimize the expenses related to this process, it is also necessary to propose learning strategies based on semi-supervised learning. Such strategies employ active learning methods indicating which of the incoming objects are valuable to be labeled for improving the classifier's performance. Methods This paper proposes Active Weighted Aging Ensemble algorithm, a novel chunk-based method for non-stationary data stream classification. It employs a classifier ensemble approach and utilizes the changing ensemble lineup to react to concept drift appropriately. It also proposed a new active learning method, considering a limited budget that may be applied to any data stream classifier. Results AWAE has been evaluated through computer experiments using real and synthetic data streams, confirming the proposed algorithm's high quality over state-of-the-art methods. Conclusion The research conducted on benchmark data streams confirmed the effectiveness of the proposed solution and highlighted its strengths in comparison with state-of-the-art methods. The estimated computational complexity is acceptable and comparable to the benchmark algorithms. Michal Wozniak 0001, Pawel Zyblewski, Pawel Ksieniewicz |
Inf. Sci. | 1 |
| 2022 | Employing Generative Adversarial Network in COVID-19 Diagnosis
Jakub Deren, Michal Wozniak 0001 |
ACIIDS (1) | 2 |
| 2022 | Feature Integration Strategies for Multilingual Fake News ClassificationabstractThe abundance of information in digital media, which in today’s world is the main source of knowledge about current events for the masses, makes it possible to spread disinformation on a larger scale than ever before. Consequently, there is a need to develop novel fake news detection approaches capable of adapting to changing factual contexts and generalizing previously or concurrently acquired knowledge. To deal with this problem, we propose a ensemble-based approach, which allows for fake news detection in multiple languages and the mutual transfer of knowledge acquired in each of them. Both classical feature extractors, such as Term frequency-inverse document frequency or Latent Dirichlet Allocation, and integrated deep NLP (Natural Language Processing) BERT (Bidirectional Encoder Representations from Transformers) models paired with MLP (Multilayer Perceptron) classifier, were employed. The results of experiments conducted on two datasets dedicated to the fake news classification t ask ( in English and Spanish, respectively), supported by statistical analysis, confirmed that utilization of additional languages could improve performance for traditional methods. Also, in some cases supplementing the deep learning method with classical ones can positively impact obtained results. The ability of models to generalize the knowledge acquired between the analyzed languages was also observed. Jedrzej Kozal, Michal Les, Pawel Zyblewski, Pawel Ksieniewicz, Michal Wozniak 0001 |
IEEE Big Data | 5 |
| 2022 | Search-based framework for transparent non-overlapping ensemble modelsabstractDue to their generalizing ability, classifier ensembles are considered very powerful predictive models. A typical ensemble consists of a static or dynamic pool of classifiers and a combination method, which translates predictions of many models into one. The combination step is often complex and renders the inner behavior of the whole ensemble incomprehensible to a typical user. In this work, in the light of recent interest in Explainable AI (XAI) research, we are proposing a novel approach to building an interpretable ensemble model. It is based on decision space splitting into non-overlapping regions. Every area has an assigned interpretable classifier and its boundaries are selected using the genetic programming approach. We experimentally evaluate the proposed method and compare it to Decision Tree and Random Forest. The results show that the proposed approach is competitive with the state-of-the-art techniques and prone to further expansion. Bogdan Gulowaty, Michal Wozniak 0001 |
IJCNN | 2 |
| 2022 | Automated identification of systematic performance changes in 5G networks by changepoint trackingabstractTelecommunication 5G performance data exhibit complex behaviours and relations valid for wireless networks, e.g., sudden spikes of demand, poorly performing cells, temporal anomalous patterns invoked by radio signal degradation or systematic changes due to the roll-out of new functionalities or changes of configuration parameters, etc. The zero-touch networks paradigm brings to the challenge of automated management of the resources in a wireless network, which require understanding and following the complex patterns proper for 5G performance data. The paper proposes an approach to decode automatically the events in operator networks associated with the hardware and/or software configuration updates and parametrization changes. The assumption is that there are configuration changes in wireless network which brings systematic differences in observed performance output (intentional or not). For 118 time series of 5G Cell Average Throughput in uplink and downlink direction, with unique abrupt changes in behavior labeled by telco network specialists, we apply a supervised approach to penalty setting of the PELT algorithm, which is designed for identification of the time stamp when reconfiguration occurred in operator network. Comparative studies showed that in application to the 5G performance data, the modified PELT procedure outperforms the other algorithms dedicated for change-point detection, e.g., Binary Segmentation, Bayesian Online Changepoint Detection or PELT with information criterion used as penalty term. The experimental studies exhibited that filtering of the spiky events from 5G data can further improve the reliability of concluding, i.e. the F1 measure estimated for this case was 0.76 as oppose to e.g. PELT algorithm with Modified Bayesian Information Criterion used as penalty term for which the F1 measure was 0.49. Reported results are the preliminary demonstration of the potential of the change-point modeling for the depiction of interrelations spanned between the configuration input and the performance output in the cellular networks. Further extensions of the procedure specified in the paper are open for applications in telecommunication and other domains. Michal Panek, Ireneusz Jablonski, Michal Wozniak 0001 |
IJCNN | 3 |
| 2022 | Experimental Analysis on Dissimilarity Metrics and Sudden Concept Drift Detection
Sebastián Basterrech, Jan Platos, Gerardo Rubino, Michal Wozniak 0001 |
ISDA (3) | 4 |
| 2022 | Tracking changes using Kullback-Leibler divergence for the continual learningabstractRecently, continual learning has received a lot of attention. One of the significant problems is the occurrence of concept drift, which consists of changing probabilistic characteristics of the incoming data. In the case of the classification task, this phenomenon destabilizes the model’s performance and negatively affects the achieved prediction quality. Most current methods apply statistical learning and similarity analysis over the raw data. However, similarity analysis in streaming data remains a complex problem due to time limitation, non-precise values, fast decision speed, scalability, etc. This article introduces a novel method for monitoring changes in the probabilistic distribution of multi-dimensional data streams. As a measure of the rapidity of changes, we analyze the popular Kullback-Leibler divergence. During the experimental study, we show how to use this metric to predict the concept drift occurrence and understand its nature. The obtained results encourage further work on the proposed methods and its application in the real tasks where the prediction of the future appearance of concept drift plays a crucial role, such as predictive maintenance. Sebastián Basterrech, Michal Wozniak 0001 |
SMC | 2 |
| 2022 | Selective ensemble of classifiers trained on selective samples
Amgad M. Mohammed, Enrique Onieva, Michal Wozniak 0001 |
Neurocomputing | 3 |
| 2022 | Cybersecurity applications of computational intelligence
Álvaro Herrero 0001, Emilio Corchado, Michal Wozniak 0001, Sung-Bae Cho, Slobodan Petrovic |
Neural Comput. Appl. | 3 |
| 2022 | An analysis of heuristic metrics for classifier ensemble pruning based on ordered aggregationabstractClassifier ensemble pruning is a strategy through which a subensemble can be identified via optimizing a predefined performance criterion. Choosing the optimum or suboptimum subensemble decreases the initial ensemble size and increases its predictive performance. In this article, a set of heuristic metrics will be analyzed to guide the pruning process. The analyzed metrics are based on modifying the order of the classifiers in the bagging algorithm, with selecting the first set in the queue. Some of these criteria include general accuracy, the complementarity of decisions, ensemble diversity, the margin of samples, minimum redundancy, discriminant classifiers, and margin hybrid diversity. The efficacy of those metrics is affected by the original ensemble size, the required subensemble size, the kind of individual classifiers, and the number of classes. While the efficiency is measured in terms of the computational cost and the memory space requirements. The performance of those metrics is assessed over fifteen binary and fifteen multiclass benchmark classification tasks, respectively. In addition, the behavior of those metrics against randomness is measured in terms of the distribution of their accuracy around the median. Results show that ordered aggregation is an efficient strategy to generate subensembles that improve both predictive performance as well as computational and memory complexities of the whole bagging ensemble. Amgad M. Mohammed, Enrique Onieva, Michal Wozniak 0001, Gonzalo Martínez-Muñoz |
Pattern Recognit. | 3 |
| 2021 | RB-CCR: Radial-Based Combined Cleaning and Resampling algorithm for imbalanced data classification
Michal Koziarski, Colin Bellinger, Michal Wozniak 0001 |
DSAA | 3 |
| 2021 | Extracting Interpretable Decision Tree Ensemble from Random ForestabstractMachine learning predictive models are nowadays widely applied in various systems in both commercial and public areas. The need to understand and comprehend their behavior has arisen and was, through an extensive increase in research in Explainable AI (XAI), noticed in the last years. One of the problems in the aforementioned XAI spectrum is understanding the large rule sets, such as those mined from big datasets or induced through Random Forest algorithms. This work explores the possibility of tackling such an issue using multiple nonoverlapping decision trees created with incorporated knowledge provided by the set of rules. Random Forest is being used as a source of such rules. Performance evaluation conducted on commonly used datasets shows that such a model could be competitive with an equally deep Random Forest while providing explanations that can be adjusted and interpreted as a rule list. Bogdan Gulowaty, Michal Wozniak 0001 |
IJCNN | 2 |
| 2021 | RB-CCR: Radial-Based Combined Cleaning and Resampling algorithm for imbalanced data classificationabstractAbstract Real-world classification domains, such as medicine, health and safety, and finance, often exhibit imbalanced class priors and have asynchronous misclassification costs. In such cases, the classification model must achieve a high recall without significantly impacting precision. Resampling the training data is the standard approach to improving classification performance on imbalanced binary data. However, the state-of-the-art methods ignore the local joint distribution of the data or correct it as a post-processing step. This can causes sub-optimal shifts in the training distribution, particularly when the target data distribution is complex. In this paper, we propose Radial-Based Combined Cleaning and Resampling (RB-CCR). RB-CCR utilizes the concept of class potential to refine the energy-based resampling approach of CCR. In particular, RB-CCR exploits the class potential to accurately locate sub-regions of the data-space for synthetic oversampling. The category sub-region for oversampling can be specified as an input parameter to meet domain-specific needs or be automatically selected via cross-validation. Our $$5\times 2$$ 5 × 2 cross-validated results on 57 benchmark binary datasets with 9 classifiers show that RB-CCR achieves a better precision-recall trade-off than CCR and generally out-performs the state-of-the-art resampling methods in terms of AUC and G-mean. Michal Koziarski, Colin Bellinger, Michal Wozniak 0001 |
Mach. Learn. | 3 |
| 2020 | Data Preprocessing for des-knn and Its Application to Imbalanced Medical Data Classification
Maciej Kinal, Michal Wozniak 0001 |
ACIIDS (1) | 2 |
| 2020 | Employing dropout regularization to classify recurring drifted data streamsabstractStreaming data analysis is currently a rapidly growing research direction. One of the serious problems hindering the data stream classification is the fact that during the exploitation of the model, its probabilistic characteristics may change. This phenomenon is called concept drift. Until today, multiple methods have been proposed to overcome their negative influence on model performance during learning in dynamic environments. This work introduces a new streaming data classifier based on a dropout technique that can significantly reduce model restoration time and performance loss and can improve its overall score in the presence of recurring concept drifts. The usefulness of the proposed algorithm is evaluated based on extensive experimental study and backed-up with thorough statistical analysis. Filip Guzy, Michal Wozniak 0001 |
IJCNN | 2 |
| 2020 | Fake News Detection from Data StreamsabstractUsing fake news as a political or economic tool is not new, but the scale of their use is currently alarming, especially on social media. The authors of misinformation try to influence the users' decisions, both in the economic and political sphere. The facts of using disinformation during elections are well known. Currently, two fake news detection approaches dominate. The first approach, so-called fact or news checker, is based on the knowledge and work of volunteers, the second approach employs artificial intelligence algorithms for news analysis and manipulation detection. In this work, we will focus on using machine learning methods to detect fake news. However, unlike most approaches, we will treat incoming messages as stream data, taking into account the possibility of concept drift occurring, i.e., appearing changes in the probabilistic characteristics of the classification model during the exploitation of the classifier. The developed methods have been evaluated based on computer experiments on benchmark data, and the obtained results prove their usefulness for the problem under consideration. The proposed solutions are part of the distributed platform developed by the H2020 SocialTruth project consortium. Pawel Ksieniewicz, Pawel Zyblewski, Michal Choras, Rafal Kozik, Agata Gielczyk, Michal Wozniak 0001 |
IJCNN | 6 |
| 2020 | Special Issue SOCO 2017: New trends in soft computing and its application in industrial and environmental problems
Francisco Herrera, Ajith Abraham, Michal Wozniak 0001, Hilde Pérez 0001, Emilio Corchado |
Neurocomputing | 3 |
| 2020 | Combined Cleaning and Resampling algorithm for multi-class imbalanced data with label noise
Michal Koziarski, Michal Wozniak 0001, Bartosz Krawczyk |
Knowl. Based Syst. | 2 |
| 2020 | Special issue SOCO 2017: AI and ML applied to Health Sciences (MLHS)
Francisco Javier de Cos Juez, Michal Wozniak 0001, Juan A. Méndez, José R. Villar 0001 |
Neural Comput. Appl. | 2 |
| 2020 | Novel clustering-based pruning algorithmsabstractAbstract One of the crucial problems of designing a classifier ensemble is the proper choice of the base classifier line-up. Basically, such an ensemble is formed on the basis of individual classifiers, which are trained in such a way to ensure their high diversity or they are chosen on the basis of pruning which reduces the number of predictive models in order to improve efficiency and predictive performance of the ensemble. This work is focusing on clustering-based ensemble pruning, which looks for the group of similar classifiers which are replaced by their representatives. We propose a novel pruning criterion based on well-known diversity measures and describe three algorithms using classifier clustering. The first method selects the model with the best predictive performance from each cluster to form the final ensemble, the second one employs the multistage organization, where instead of removing the classifiers from the ensemble each classifier cluster makes the decision independently, while the third proposition combines multistage organization and sampling with replacement. The proposed approaches were evaluated using 30 datasets with different characteristics. Experimentation results validated through statistical tests confirmed the usefulness of the proposed approaches. Pawel Zyblewski, Michal Wozniak 0001 |
Pattern Anal. Appl. | 2 |
| 2020 | Radial-Based Oversampling for Multiclass Imbalanced Data ClassificationabstractLearning from imbalanced data is among the most popular topics in the contemporary machine learning. However, the vast majority of attention in this field is given to binary problems, while their much more difficult multiclass counterparts are relatively unexplored. Handling data sets with multiple skewed classes poses various challenges and calls for a better understanding of the relationship among classes. In this paper, we propose multiclass radial-based oversampling (MC-RBO), a novel data-sampling algorithm dedicated to multiclass problems. The main novelty of our method lies in using potential functions for generating artificial instances. We take into account information coming from all of the classes, contrary to existing multiclass oversampling approaches that use only minority class characteristics. The process of artificial instance generation is guided by exploring areas where the value of the mutual class distribution is very small. This way, we ensure a smart oversampling procedure that can cope with difficult data distributions and alleviate the shortcomings of existing methods. The usefulness of the MC-RBO algorithm is evaluated on the basis of extensive experimental study and backed-up with a thorough statistical analysis. Obtained results show that by taking into account information coming from all of the classes and conducting a smart oversampling, we can significantly improve the process of learning from multiclass imbalanced data. Bartosz Krawczyk, Michal Koziarski, Michal Wozniak 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | A Genetic-Based Ensemble Learning Applied to Imbalanced Data Classification
Jakub Klikowski, Pawel Ksieniewicz, Michal Wozniak 0001 |
IDEAL (2) | 3 |
| 2019 | Machine Learning Methods for Fake News Classification
Pawel Ksieniewicz, Michal Choras, Rafal Kozik, Michal Wozniak 0001 |
IDEAL (2) | 4 |
| 2019 | Special issue on hybrid artificial intelligence systems from the HAIS 2017 conference - Editorial
Francisco J. Martínez de Pisón Ascacibar, Francisco Herrera, Ajith Abraham, Michal Wozniak 0001, Emilio Corchado |
Neurocomputing | 4 |
| 2019 | Monotonic classification: An overview on algorithms, performance measures and data sets
José Ramón Cano, Pedro Antonio Gutiérrez, Bartosz Krawczyk, Michal Wozniak 0001, Salvador García 0001 |
Neurocomputing | 4 |
| 2019 | Radial-Based oversampling for noisy imbalanced data classification
Michal Koziarski, Bartosz Krawczyk, Michal Wozniak 0001 |
Neurocomputing | 3 |
| 2019 | Data stream classification using active learned neural networks
Pawel Ksieniewicz, Michal Wozniak 0001, Boguslaw Cyganek, Andrzej Kasprzak, Krzysztof Walkowiak |
Neurocomputing | 2 |
| 2019 | Special issue HAIS 2014: Recent advancements in hybrid artificial intelligence systems and its application to real-world problems
Marios M. Polycarpou, André C. P. L. F. de Carvalho, Jeng-Shyang Pan 0001, Michal Wozniak 0001, Héctor Quintián, Emilio Corchado |
Neurocomputing | 4 |
| 2019 | Instance reduction for one-class classification
Bartosz Krawczyk, Isaac Triguero, Salvador García 0001, Michal Wozniak 0001, Francisco Herrera |
Knowl. Inf. Syst. | 4 |
| 2018 | Combining active learning with concept drift detection for data stream miningabstractMost of data stream classifier learning methods assume that a true class of an incoming object is available right after the instance has been processed and new and labeled instance may be used to update a classifier's model, drift detection or capturing novel concepts. However, assumption that we have an unlimited and infinite access to class labels is very naive and usually would require a very high labeling cost. Therefore the applicability of many supervised techniques is limited in real-life stream analytics scenarios. Active learning emerges as a potential solution to this problem, concentrating on selecting only the most valuable instances and learning an accurate predictive model with as few labeling queries as possible. However learning from data streams differ from online learning as distribution of examples may change over time. Therefore, an active learning strategy must be able to handle concept drift and quickly adapt to evolving nature of data. In this paper we present novel active learning strategies that are designed for effective tackling of such changes. We assume that most labeling effort is required when concept drift occurs, as we need a representative sample of new concept to retrain properly the predictive model. Therefore, we propose active learning strategies that are guided by drift detection module to save budget for difficult and evolving instances. Three proposed strategies are based on learner uncertainty, dynamic allocation of budget over time and search space randomization. Experimental evaluation of the proposed methods prove their usefulness for reducing labeling effort in learning from drifting data streams. Bartosz Krawczyk, Bernhard Pfahringer, Michal Wozniak 0001 |
IEEE BigData | 3 |
| 2018 | An Empirical Insight Into Concept Drift Detectors Ensemble StrategiesabstractContemporary decision support systems have to take into consideration the fact that most of data gathered is nowadays in motion, i.e., that successive observations, about objects being analyzed, form so-called data streams. Unfortunately, during analytical model utilization, unpredictable changes may appear in data distributions, leading to significant deterioration in the predictive performance and reliability of these learners. This phenomenon is called concept drift and refers to changes in the input data in relation to target variable in supervised learning task. Due to its potentially catastrophic impact on the underlying learner, it must be detected and handled as soon as it occurs. Over the years, many methods have been developed to address this issue. We focus on supervised classification task, aiming at answering the question on how to detect significant changes in data distribution effectively using ensemble of drift detectors. We discuss several models of combined drift detectors, among them the local detector which analyses distribution of each attribute separately. Experimental evaluations confirm the effectiveness of ensemble detectors, making them highly interesting to be used in solving real-world problems. Andrzej Lapinski, Bartosz Krawczyk, Pawel Ksieniewicz, Michal Wozniak 0001 |
CEC | 4 |
| 2018 | Imbalanced Data Classification Based on Feature Selection Techniques
Pawel Ksieniewicz, Michal Wozniak 0001 |
IDEAL (2) | 2 |
| 2018 | Selecting local ensembles for multi-class imbalanced data classificationabstractLearning from imbalanced data is a challenge that machine learning community is facing over last decades, due to its ever-growing presence in real-life problems. While there is a significant number of works addressing the issue of handling binary and skewed datasets, its multi-class counterpart have not received as much attention. This problem is much more difficult, as presence of multiple imbalanced classes can significantly deteriorate the predictive power of any classifier. The relationship among classes are no longer clearly established and there are many difficulties embedded in the nature of such data that needs to be properly addressed. In this work, we discuss the issue of forming effective ensembles for multi-class imbalanced data based on static classifier selection approach. We propose a fully adaptive learning scheme that splits the original feature space into a number of competence areas and modifies their size and location in order to most effectively exploit the supplied pool of base classifiers. Additionally, for each established cluster we perform a weighted classifier combination, where weights are set individually for each cluster and each considered class. This allows for exploiting local competencies of each base learner in given part of feature space, as well as for each of considered classes. These two tasks are combined together in a single hybrid training scheme guided by an evolutionary algorithm. The optimization criterion is formulated in order to achieve skew-insensitive ensemble of local ensembles able to tackle highly imbalanced and multi-class problems. Experimental study proves the high efficacy of the proposed method and its superiority to other ensemble selection methods. Bartosz Krawczyk, Alberto Cano 0001, Michal Wozniak 0001 |
IJCNN | 3 |
| 2018 | Leveraging Ensemble Pruning for Imbalanced Data ClassificationabstractThe effectiveness of machine learning algorithms depends on the quality of the supplied training data. Any problems embedded in the nature of data will result in obtaining incorrect classification models, especially imbalanced data distribution is among the most significant learning difficulties that can affect classifiers. As one of the classes has much more instances than the other, the learning process becomes biased towards it. Therefore, methods for alleviating the impact of skewed distributions are highly sought after. Ensemble learning has emerged as one of the leading paradigms for imbalanced data. Creation of an efficient pool of classifiers is not a trivial task and one needs to carefully select which classifiers should be combined to obtain the best predictive power. In this paper, we propose a compound ensemble pruning algorithm for imbalanced data. It aims to retain classifiers that offer the best performance on both minority and majority classes, and display a high level of diversity. Remaining learners are discarded from the pool. This is achieved by the means of a multi-criteria evolutionary algorithm. Extensive experimental study show that our proposal is able to create smaller ensembles than the state-of-the-art methods, while offering an improved robustness to imbalanced class distributions. Bartosz Krawczyk, Michal Wozniak 0001 |
SMC | 2 |
| 2018 | Ensemble of Extreme Learning Machines with trained classifier combination and statistical features for hyperspectral data
Pawel Ksieniewicz, Bartosz Krawczyk, Michal Wozniak 0001 |
Neurocomputing | 3 |
| 2018 | Dynamic ensemble selection for multi-class classification with one-class classifiers
Bartosz Krawczyk, Mikel Galar, Michal Wozniak 0001, Humberto Bustince, Francisco Herrera |
Pattern Recognit. | 3 |
| 2017 | Online query by committee for active learning from drifting data streamsabstractMost of data stream learning methods assume that a true class of an incoming instance is available right after it has been processed. However, assumption that we have an unlimited access to class labels is unrealistic and is directly connected with a very high labeling cost. This is a driving force behind growing development of methods that require reduced or no access to class labels. Among several potential directions active learning emerges as a promising solution, by allowing for a selection of most valuable instances from the stream and using as few label queries. Despite numerous proposals of active learning methods for static data, this domain is still developing for data streams. Here, non-stationary nature of data must be taken into consideration and proposed algorithms must accommodate potential occurrences of concept drift. In this paper we propose a Query by Committee active learning strategy that is adapted to online learning from drifting data streams. A decision regarding label query is made by an ensemble of classifiers instead of a single learner, leading to an improved instance selection. We present four different approaches for online Query by Committee and evaluate their usefulness on the basis of obtained accuracy with limited budgets and ability to handle concept drift. We introduce Budget Loss of Accuracy, a novel measure for evaluating active learning algorithms. Finally, we investigate the relationships between the efficacy of Query by Committee models and diversity of underlying ensembles. Based on thorough experimental investigation we are able to show the usefulness of proposed algorithms for reducing labeling effort in learning from drifting data streams. Bartosz Krawczyk, Michal Wozniak 0001 |
IJCNN | 2 |
| 2017 | Fault diagnosis of marine 4-stroke diesel engines using a one-vs-one extreme learning ensemble
Jerzy Kowalski, Bartosz Krawczyk, Michal Wozniak 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2017 | SCR: simulated concept recurrence - a non-supervised tool for dealing with shifting conceptabstractAbstract Most of the approaches coping with concept drift described in the machine learning literature are focused solely on detecting the concept changes or adapting the classification system, and hardly any works exist, which try to also describe the changes in concept. Nowadays, we desire methods, which are able to detect the concept drift in the absence of information about class labels. In this article, we present a semi‐supervised method and analyze the possibilities of using the simulated concept recurrence against concept drift and also expanding the previously presented functionality of the algorithm from the sole concept characterization to both concept drift detection and concept characterization. The supervision is limited to the system setup phase, and during the evaluation of the algorithm, we assume no support from the experts. Further comparing this work to our previous publication, the scope of experiments has been extended from the single concept drift problems to the multi‐concept scenarios, and also, the new method is evaluated on three different levels of the prior knowledge presented to the system beforehand. Piotr Sobolewski, Michal Wozniak 0001 |
Expert Syst. J. Knowl. Eng. | 2 |
| 2017 | A survey on data preprocessing for data stream mining: Current status and future directions
Sergio Ramírez-Gallego, Bartosz Krawczyk, Salvador García 0001, Michal Wozniak 0001, Francisco Herrera |
Neurocomputing | 4 |
| 2017 | The deterministic subspace method for constructing classifier ensemblesabstractEnsemble classification remains one of the most popular techniques in contemporary machine learning, being characterized by both high efficiency and stability. An ideal ensemble comprises mutually complementary individual classifiers which are characterized by the high diversity and accuracy. This may be achieved, e.g., by training individual classification models on feature subspaces. Random Subspace is the most well-known method based on this principle. Its main limitation lies in stochastic nature, as it cannot be considered as a stable and a suitable classifier for real-life applications. In this paper, we propose an alternative approach, Deterministic Subspace method, capable of creating subspaces in guided and repetitive manner. Thus, our method will always converge to the same final ensemble for a given dataset. We describe general algorithm and three dedicated measures used in the feature selection process. Finally, we present the results of the experimental study, which prove the usefulness of the proposed method. Michal Koziarski, Bartosz Krawczyk, Michal Wozniak 0001 |
Pattern Anal. Appl. | 3 |
| 2017 | Nearest Neighbor Classification for High-Speed Big Data Streams Using SparkabstractMining massive and high-speed data streams among the main contemporary challenges in machine learning. This calls for methods displaying a high computational efficacy, with ability to continuously update their structure and handle ever-arriving big number of instances. In this paper, we present a new incremental and distributed classifier based on the popular nearest neighbor algorithm, adapted to such a demanding scenario. This method, implemented in Apache Spark, includes a distributed metric-space ordering to perform faster searches. Additionally, we propose an efficient incremental instance selection method for massive data streams that continuously update and remove outdated examples from the case-base. This alleviates the high computational requirements of the original classifier, thus making it suitable for the considered problem. Experimental study conducted on a set of real-life massive data streams proves the usefulness of the proposed solution and shows that we are able to provide the first efficient nearest neighbor solution for high-speed big and streaming data. Sergio Ramírez-Gallego, Bartosz Krawczyk, Salvador García 0001, Michal Wozniak 0001, José Manuel Benítez 0001, Francisco Herrera |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2016 | Forming Classifier Ensembles with Deterministic Feature SubspacesabstractEnsemble learning is being considered as one of the most well-established and efficient techniques in the contemporary machine learning.The key to the satisfactory performance of such combined models lies in the supplied base learners and selected combination strategy.In this paper we will focus on the former issue.Having classifiers that are of high individual quality and complementary to each other is a desirable property.Among several ways to ensure diversity feature space division deserves attention.The most popular method employed here is Random Subspace approach.However, due to its random nature one cannot consider this approach as stable one or suitable for reallife applications.Therefore, we propose a new approach called Deterministic Subspace that constructs feature subspaces in a guided and repetitive manner.We present a general framework and three dedicated measures that can be used for selecting diverse and uncorrelated features for each base learner.This way we will always obtain identical sets of features, leading to creation of stable ensembles.Experimental study backed-up with statistical analysis prove the usefulness of our method in comparison to popular randomized solution. Michal Koziarski, Bartosz Krawczyk, Michal Wozniak 0001 |
FedCSIS | 3 |
| 2016 | Tackling label noise with multi-class decomposition using fuzzy one-class support vector machinesabstractClass label noise is a data-level difficulty associated with training objects with incorrectly assigned labels. This problem may originate from poorly documented historic data, errors during data generation process or mistakes made by human experts. Inclusion of such examples during the training process will mislead the classifier by presenting a falsified class distribution and consequently lead to degradation of models' generalization abilities. This phenomenon becomes even more troublesome in multi-class scenarios that may be affected by highly complex intra-class noise. Decomposition strategies with binary classifiers were proven to alleviate this difficulty by using simplified binary subtasks that are less affected by the noise. In this paper we propose to extend this approach by using the one-class classification decomposition. In this scenario each class has assigned individual one-class classifier that aims at capturing its distinguishing characteristics. This allows to create a robust data description and then apply a dedicated classifier combination in order to reconstruct the original multi-class task. We further extend this concept by using fuzzy one-class classifiers that allow to associate membership values with each training objects. This allows us to reduce the influence of uncertain and potentially noisy samples on the shape of learned decision boundary. Experimental study backed-up with statistical analysis shows that fuzzy one-class classifier decomposition offers an excellent robustness to noise in multi-class classification. Bartosz Krawczyk, José A. Sáez, Michal Wozniak 0001 |
FUZZ-IEEE | 3 |
| 2016 | Recent advancements in hybrid artificial intelligence systems and its application to real-world problems
Emilio Corchado, Ajith Abraham, André C. P. L. F. de Carvalho, Michal Wozniak 0001, Sung-Bae Cho, Héctor Quintián |
Neurocomputing | 4 |
| 2016 | Untrained weighted classifier combination with embedded ensemble pruning
Bartosz Krawczyk, Michal Wozniak 0001 |
Neurocomputing | 2 |
| 2016 | Dynamic classifier selection for one-class classification
Bartosz Krawczyk, Michal Wozniak 0001 |
Knowl. Based Syst. | 2 |
| 2016 | Analyzing the oversampling of different classes and types of examples in multi-class imbalanced datasets
José A. Sáez, Bartosz Krawczyk, Michal Wozniak 0001 |
Pattern Recognit. | 3 |
| 2015 | Pruning Ensembles of One-Class Classifiers with X-means Clustering
Bartosz Krawczyk, Michal Wozniak 0001 |
ACIIDS (1) | 2 |
| 2015 | Pruning Ensembles with Cost Constraints
Bartosz Krawczyk, Michal Wozniak 0001 |
ACIIDS (1) | 2 |
| 2015 | Selected aspects of electronic health record analysis from the big data perspectiveabstractThe electronic health record (EHR) groups all digital documents related to a given patient as anamnesis, results of the laboratory tests, prescriptions, recorded medical signals as ECG or images etc. Dealing with such data representation we face with plethora of problems as different form of data, unstructured data (as doctor's notes), huge and fast growing volume, etc. It causes that EHR should be considered as the complex data representation. Accordingly, taking into consideration its complexity, hetorogenousity, fast growing and size we need special tools to analyse such medical big data. Such tools should be able to analyse datasets characterized by so-called 4Vs (volume, velocity, variety, and veracity). Notwithstandingly, we should also add the fifth V-value, because the only analytics tool deployment makes sense if it leads to healthcare improvement (as personalised patient's care, unnecessary hospitalization decreasing or reducing the patient's readmissions). In this paper we focus on the selected aspects EHR analysis from the big data perspective. Boguslaw Cyganek, Manuel Graña, Andrzej Kasprzak, Krzysztof Walkowiak, Michal Wozniak 0001 |
BIBM | 5 |
| 2015 | Tensor based representation and analysis of the electronic healthcare record dataabstractThe paper addresses the problem of multidimensional data representation and analysis in electronic healthcare records. Our methodology is based on the best all-rank tensor decomposition which allows data compression and simultaneous classification in the tensor subspaces. Experiments were run on the MRI brain signals. The obtained results show high compression ratios which do not sacrifice reconstruction accuracies. Also, the method allows fast and highly discriminative matching of the MRI signals to the models built with the proposed method. Boguslaw Cyganek, Michal Wozniak 0001 |
BIBM | 2 |
| 2015 | Joint optimization of multicast and unicast flows in elastic optical networksabstractNowadays, multicasting is increasingly popular, due to the ability to provision in efficient way such desirable services like streaming, IP television, software distribution, etc. At the same time, the research in the field of optical networks concentrates on the very promising elastic optical networks (EONs) approach. In this paper, we focus on the joint optimization of multicast and unicast flows in EONs. We propose to model multicast flows in EONs using pre-generated candidate trees. We present two new ILP models and a heuristic method, dedicated to solve the optimization problem with joint multicast and unicast flows. We report results of the numerical experiments carried out to compare proposed ILP models and the heuristic method as well as to evaluate potential benefits of multicasting in EONs. We show that the proposed modeling approach and heuristic method are highly elastic in terms of their applicability to different network scenarios, whilst multicasting can bring significant spectrum savings compared to unicasting. Krzysztof Walkowiak, Róza Goscien, Michal Wozniak 0001, Miroslaw Klinkowski |
ICC | 3 |
| 2015 | Blurred Labeling Segmentation Algorithm for Hyperspectral Images
Pawel Ksieniewicz, Manuel Graña, Michal Wozniak 0001 |
ICCCI (2) | 3 |
| 2015 | Cost-Sensitive Neural Network with ROC-Based Moving Threshold for Imbalanced Classification
Bartosz Krawczyk, Michal Wozniak 0001 |
IDEAL | 2 |
| 2015 | Weighted Naïve Bayes Classifier with Forgetting for Drifting Data StreamsabstractMining massive data streams in real-time is one of the contemporary challenges for machine learning systems. Such a domain encompass many of difficulties hidden beneath the term of Big Data. We deal with massive, incoming information that must be processed on-the-fly, with lowest possible response delay. We are forced to take into account time, memory and quality constraints. Our models must be able to quickly process large collection of data and swiftly adapt themselves to occurring changes (shifts and drifts) in data streams. In this paper, we propose a novel version of simple, yet effective Naïve Bayes classifier for mining streams. We add a weighting module, that automatically assigns an importance factor to each object extracted from the stream. The higher the weight, the bigger influence given object exerts on the classifier training procedure. We assume, that our model works in the non-stationary environment with the presence of concept drift phenomenon. To allow our classifier to quickly adapt its properties to evolving data, we imbue it with forgetting principle implemented as weight decay. With each passing iteration, the level of importance of previous objects is decreased until they are discarded from the data collection. We propose an efficient sigmoidal function for modeling the forgetting rate. Experimental analysis, carried out on a number of large data streams with concept drift prove that our weighted Naïve Bayes classifier displays highly satisfactory performance in comparison with state-of-the-art stream classifiers. Bartosz Krawczyk, Michal Wozniak 0001 |
SMC | 2 |
| 2015 | A hybrid cost-sensitive ensemble for imbalanced breast thermogram classification
Bartosz Krawczyk, Gerald Schaefer, Michal Wozniak 0001 |
Artif. Intell. Medicine | 3 |
| 2015 | Multidimensional data classification with chordal distance based kernel and Support Vector Machines
Boguslaw Cyganek, Bartosz Krawczyk, Michal Wozniak 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2015 | Data stream classification and big data analytics
Bartosz Krawczyk, Jerzy Stefanowski, Michal Wozniak 0001 |
Neurocomputing | 3 |
| 2015 | Special issue HAIS 2012: Recent advancements in hybrid artificial intelligence systems and its application to real-world problemsabstractDealing with distributed data is one of the challenges for clustering, as most clustering techniques require the data to be centralized. One of them, k-means, has been elected as one of the most influential data mining algorithms for being simple, scalable, and easily modifiable to a variety of contexts and application domains. However, exact distributed versions of k-means are still sensitive to the selection of the initial cluster prototypes and require the number of clusters to be specified in advance. Additionally, preserving data privacy among repositories may be a complicating factor. In order to overcome k-means limitations, two different approaches were adopted in this paper: the first obtains a final model identical to the centralized version of the clustering algorithm and the second generates and selects clusters for each distributed data subset and combines them afterwards. It is also described how to apply the algorithms compared while preserving data privacy. The algorithms are compared experimentally from two perspectives: the theoretical one, through asymptotic complexity analyses, and the experimental one, through a comparative evaluation of results obtained from a collection of experiments and statistical tests. The results obtained indicate which algorithm is more suitable for each application scenario. Héctor Quintián, Emilio Corchado, Ajith Abraham, André C. P. L. F. de Carvalho, Michal Wozniak 0001, Václav Snásel, Sung-Bae Cho |
Neurocomputing | 5 |
| 2015 | On the usefulness of one-class classifier ensembles for decomposition of multi-class problems
Bartosz Krawczyk, Michal Wozniak 0001, Francisco Herrera |
Pattern Recognit. | 2 |
| 2015 | One-class classifiers with incremental learning and forgetting for data streams with concept driftabstractOne of the most important challenges for machine learning community is to develop efficient classifiers which are able to cope with data streams, especially with the presence of the so-called concept drift. This phenomenon is responsible for the change of classification task characteristics, and poses a challenge for the learning model to adapt itself to the current state of the environment. So there is a strong belief that one-class classification is a promising research direction for data stream analysis—it can be used for binary classification without an access to counterexamples, decomposing a multi-class data stream, outlier detection or novel class recognition. This paper reports a novel modification of weighted one-class support vector machine, adapted to the non-stationary streaming data analysis. Our proposition can deal with the gradual concept drift, as the introduced one-class classifier model can adapt its decision boundary to new, incoming data and additionally employs a forgetting mechanism which boosts the ability of the classifier to follow the model changes. In this work, we propose several different strategies for incremental learning and forgetting, and additionally we evaluate them on the basis of several real data streams. Obtained results confirmed the usability of proposed classifier to the problem of data stream classification with the presence of concept drift. Additionally, implemented forgetting mechanism assures the limited memory consumption, because only quite new and valuable examples should be memorized. Bartosz Krawczyk, Michal Wozniak 0001 |
Soft Comput. | 2 |
| 2014 | Vehicle Logo Recognition with an Ensemble of Classifiers
Boguslaw Cyganek, Michal Wozniak 0001 |
ACIIDS (2) | 2 |
| 2014 | Optimization Algorithms for One-Class Classification Ensemble Pruning
Bartosz Krawczyk, Michal Wozniak 0001 |
ACIIDS (2) | 2 |
| 2014 | The Influence of a Classifiers' Diversity on the Quality of Weighted Aging Ensemble
Michal Wozniak 0001, Piotr Cal, Boguslaw Cyganek |
ACIIDS (2) | 1 |
| 2014 | A first attempt on evolutionary prototype reduction for nearest neighbor one-class classificationabstractEvolutionary prototype reduction techniques are data preprocessing methods originally developed to enhance the nearest neighbor rule. They reduce the training data by selecting or generating representative examples of a given problem. These algorithms have been designed and widely analyzed in standard classification providing very competitive results. However, its application scope can be extended to many other specific domains, such as one-class classification, in which its way of working is very interesting in order to reduce computational complexity and sensitivity to noisy data. In this contribution, we perform a first study on the usefulness of evolutionary prototype reduction methods for one-class classification. To do so, we will focus on two recent evolutionary approaches that follow very different strategies: selection and generation of examples from the training data. Both alternatives provide a resulting preprocessed data set that will be used later by a nearest neighbor one-class classifier as its training data. The results achieved support that these data reduction techniques are suitable tools to improve the performance of the nearest neighbor one-class classification. Bartosz Krawczyk, Isaac Triguero, Salvador García 0001, Michal Wozniak 0001, Francisco Herrera |
IEEE Congress on Evolutionary Computation | 4 |
| 2014 | Weighted one-class classification for different types of minority class examples in imbalanced dataabstractImbalanced classification is one of the most challenging machine learning problem. Recent studies show, that often the uneven ratio of objects in classes is not the biggest factor, determining the drop of classification accuracy. It is also related to some difficulties embedded in the nature of the data. In this paper we study the different types of minority class examples and distinguish four groups of objects - safe, borderline, rare and outliers. To deal with the imbalance problem, we use a one-class classification, that is focused on a proper identification of the minority class samples. We further augment this model by incorporating the knowledge about the minority object types in the training dataset. This is done applying weighted one-class classifier and adjusting weights assigned to minority class objects, depending on their type. A strategy for calculating the new weights for minority examples is proposed. Experimental analysis, carried on a set of benchmark datasets, confirms that the proposed model can achieve a satisfactory recognition rate and often outperform other state-of-the-art methods, dedicated to the imbalanced classification. Bartosz Krawczyk, Michal Wozniak 0001, Francisco Herrera |
CIDM | 2 |
| 2014 | Weighted One-Class Classifier Ensemble Based on Fuzzy Feature Space PartitioningabstractThis paper introduces a novel method for forming efficient one-class classifier ensembles. A common problem in one-class classification is a complex structure of the target class, which often leads to creation of a too expanded decision boundary. We propose to employ a clustering step in order to partition the target class into atomic subsets and using these as input for one-class classifiers. By this, we are able to detect sub-structures in the target concept. Additionally, to increase the diversity and robustness of our method weighted one-class classifiers are used. We introduce a novel scheme for calculating weights for training objects. Membership functions, obtained from the fuzzy clustering, are used to initialize the weighted classifiers. Based on the results of a number of computational experiments we show that the proposed method outperforms both the single one-class methods, as well as popular one-class ensembles. Other advantages are the highly parallel structure of the proposed solution, which facilitates parallel training and execution stages, and the relatively small number of control parameters. Bartosz Krawczyk, Michal Wozniak 0001, Boguslaw Cyganek |
ICPR | 2 |
| 2014 | New untrained aggregation methods for classifier combinationabstractThe combined classification is a promising direction in pattern recognition and there are numerous methods that deal with forming classifier ensembles. The most popular approaches employ voting, where the final decision of compound classifier is a combination of individual classifiers' outputs, i.e., class labels or support functions. This paper concentrates on the problem how to design an effective combination rule, which takes into consideration the values of support functions returned by the individual classifiers. Because in many practical tasks we do not have a training set at our disposal, then we express our interest in aggregation methods which do not require learning. A special attention is paid to weighted aggregation, especially when the different weights depend on particular support function of a given individual classifier. We propose a novel approach for untrained combination of support functions using the Gaussian function to assign mentioned above weights. The computer experiments carried out on the set of benchmark data sets confirm the advantages of the proposed approach for particular cases, especially when the number of class labels is high. Bartosz Krawczyk, Michal Wozniak 0001 |
IJCNN | 2 |
| 2014 | Untrained Method for Ensemble Pruning and Weighted Combination
Bartosz Krawczyk, Michal Wozniak 0001 |
ISNN | 2 |
| 2014 | One-Class Classification Ensemble with Dynamic Classifier Selection
Bartosz Krawczyk, Michal Wozniak 0001 |
ISNN | 2 |
| 2014 | Improved Adaptive Splitting and Selection: the Hybrid Training Method of a Classifier Based on a Feature Space PartitioningabstractCurrently, methods of combined classification are the focus of intense research. A properly designed group of combined classifiers exploiting knowledge gathered in a pool of elementary classifiers can successfully outperform a single classifier. There are two essential issues to consider when creating combined classifiers: how to establish the most comprehensive pool and how to design a fusion model that allows for taking full advantage of the collected knowledge. In this work, we address the issues and propose an AdaSS+, training algorithm dedicated for the compound classifier system that effectively exploits local specialization of the elementary classifiers. An effective training procedure consists of two phases. The first phase detects the classifier competencies and adjusts the respective fusion parameters. The second phase boosts classification accuracy by elevating the degree of local specialization. The quality of the proposed algorithms are evaluated on the basis of a wide range of computer experiments that show that AdaSS+ can outperform the original method and several reference classifiers. Konrad Jackowski, Bartosz Krawczyk, Michal Wozniak 0001 |
Int. J. Neural Syst. | 3 |
| 2014 | Recent trends in intelligent data analysis
Emilio Corchado, Michal Wozniak 0001, Ajith Abraham, André C. P. L. F. de Carvalho, Václav Snásel |
Neurocomputing | 2 |
| 2014 | Diversity measures for one-class classifier ensembles
Bartosz Krawczyk, Michal Wozniak 0001 |
Neurocomputing | 2 |
| 2014 | Clustering-based ensembles for one-class classification
Bartosz Krawczyk, Michal Wozniak 0001, Boguslaw Cyganek |
Inf. Sci. | 2 |
| 2013 | Adaptive Splitting and Selection Method for Noninvasive Recognition of Liver Fibrosis Stage
Bartosz Krawczyk, Michal Wozniak 0001, Tomasz Orczyk, Piotr Porwik |
ACIIDS (2) | 2 |
| 2013 | A framework for image analysis and object recognition in industrial applications with the ensemble of classifiersabstractThe paper presents a work-in-progress on a classification system for object detection in vision based industrial applications. The main idea of the presented system is application of the ensemble of one-class classifiers trained with specific features of objects of the whole scene. During recognition stage the ensemble tries to recognize the trained patterns. The system is well suited for parallel processing and therefore it can operate in real-time conditions. Preliminary experiments in automotive applications show promising results. Boguslaw Cyganek, Michal Wozniak 0001 |
ETFA | 2 |
| 2013 | Weighted Aging Classifier Ensemble for the Incremental Drifted Data Streams
Michal Wozniak 0001, Andrzej Kasprzak, Piotr Cal |
FQAS | 1 |
| 2013 | Application of Adaptive Splitting and Selection Classifier to the Spam Filtering ProblemabstractE-Mail spam is one of the major problems plaguing the contemporary Internet, causing an inconvenience to an individual user and financial loss to a company. Spam filtering allows for early detection of unwanted messages and separates them from the incoming e-mail. Nonetheless, designing an effective spam detection system is not a trivial task, due to the problems connected with the analysis of the e-mail content and the occurrence of variation in spam characteristics. This article presents an application of a novel ensemble classifier system for spam detection. The system is an extension of the adaptive splitting and selection (AdaSS) framework. The idea of the ensemble is based on the assumption that high effectiveness of detection can be obtained by exploitation of the local competency of a set of diverse elementary classifiers. Therefore, the ensemble training algorithm divides the feature space into several disjoint subspaces and assigns an area classifier to each of them. The area classifier consists of elementary classifiers that make a collective decision based on the weighted fusion of their support functions. The weight reflects the local competency of the classifier. To maintain the diversity of the pool of elementary classifiers, we exploit different e-mail feature extraction methods while filling the pool. There are two main extensions of the presented algorithm over original AdaSS: the aforementioned weighted fusion model used for decision making and adaptation of the AdaSS training procedure to process data streams featuring the concept drift. The effectiveness of the classifier model in spam recognition was verified in a series of experiments on two sets of spam databases. Comparison of the algorithm with some other state-of-the-art ensemble methods showed that the presented AdaSS extension can effectively recognize local competences of elementary classifiers and result in very high effectiveness of spam recognition outperforming competing methods. Konrad Jackowski, Bartosz Krawczyk, Michal Wozniak 0001 |
Cybern. Syst. | 3 |
| 2013 | Guest Editorial: Intelligent Network Security and SurvivabilityabstractWe are living in a digital world consisting of a large number of computers and other equipment creating ubiquitous computer networks. Because most human activities depend on computer networks, secu... Krzysztof Walkowiak, Michal Wozniak 0001 |
Cybern. Syst. | 2 |
| 2013 | Special section on invited papers from NetCoM-2009
Sanguthevar Rajasekaran, Michal Wozniak 0001 |
Future Gener. Comput. Syst. | 2 |
| 2013 | Classifier ensemble for an effective cytological image analysis
Pawel Filipczuk, Bartosz Krawczyk, Michal Wozniak 0001 |
Pattern Recognit. Lett. | 3 |
| 2013 | Special issue on "Innovative knowledge based techniques in pattern recognition"
Manuel Graña, Michal Wozniak 0001, Nima Hatami |
Pattern Recognit. Lett. | 2 |
| 2012 | Data with Shifting Concept Classification Using Simulated Recurrence
Piotr Sobolewski, Michal Wozniak 0001 |
ACIIDS (1) | 2 |
| 2012 | Experiments on distance measures for combining one-class classifiers
Bartosz Krawczyk, Michal Wozniak 0001 |
FedCSIS | 2 |
| 2012 | Pixel-Based Object Detection and Tracking with Ensemble of Support Vector Machines and Extended Structural Tensor
Boguslaw Cyganek, Michal Wozniak 0001 |
ICCCI (1) | 2 |
| 2012 | Adaptive Splitting and Selection Algorithm for Classification of Breast Cytology Images
Bartosz Krawczyk, Pawel Filipczuk, Michal Wozniak 0001 |
ICCCI (1) | 3 |
| 2012 | Comparison of Fuzzy Combiner Training Methods
Tomasz Wilk, Michal Wozniak 0001 |
ICCCI (1) | 2 |
| 2012 | Performance Evaluation of Hybrid Implementation of Support Vector Machine
Konrad Gajewski, Michal Wozniak 0001 |
IDEAL | 2 |
| 2012 | Cost-Sensitive Splitting and Selection Method for Medical Decision Support System
Konrad Jackowski, Bartosz Krawczyk, Michal Wozniak 0001 |
IDEAL | 3 |
| 2012 | Combining classifiers under probabilistic models: experimental comparative analysis of methodsabstractAbstract This work will present a review of the concept of classifier combination based on the combined discriminant function. We will present a Bayesian approach, in which the discriminant function assumes the role of the posterior probability. We will propose a probabilistic interpretation of expert rules and conditions of knowledge consistency for expert rules and learning sets. We will suggest how to measure the quality of learning materials and we will use the measure mentioned above for an algorithm that eliminates contradictions in the rule set. In this work several recognition algorithms will be described, based on either: (i) pure rules, or; (ii) rules together with learning sets. Furthermore, the original concept of information unification, which enables the formation of rules on the basis of learning set or learning set on the basis of rules will be proposed. The obtained conclusions will serve as a spring‐board for the formulation of new project guidelines for this type of decision‐making system. At the end, experimental results of the proposed algorithms will be presented, both from computer generated data and for a real problem from the medical diagnostics field. Marek Kurzynski, Michal Wozniak 0001 |
Expert Syst. J. Knowl. Eng. | 2 |
| 2012 | Editorial: New trends and applications on hybrid artificial intelligence systems
Emilio Corchado, Manuel Graña, Michal Wozniak 0001 |
Neurocomputing | 3 |
| 2012 | Soft computing methods applied to combination of one-class classifiers
Tomasz Wilk, Michal Wozniak 0001 |
Neurocomputing | 2 |
| 2011 | Multiple Classifier Method for Structured Output Prediction Based on Error Correcting Output Codes
Tomasz Kajdanowicz, Michal Wozniak 0001, Przemyslaw Kazienko |
ACIIDS (2) | 2 |
| 2011 | Knowledge Source Confidence Measure Applied to a Rule-Based Recognition System
Michal Wozniak 0001 |
ACIIDS (1) | 1 |
| 2011 | Combining Classifier with a Fuser Implemented as a One Layer Perceptron
Michal Wozniak 0001, Marcin Zmyslony |
ACIIDS (2) | 1 |
| 2011 | Decentralized Distributed Computing System for Privacy-Preserving Combined Classifiers - Modeling and Optimization
Krzysztof Walkowiak, Szymon Sztajer, Michal Wozniak 0001 |
ICCSA (1) | 3 |
| 2011 | A hybrid decision tree training method using data streamsabstractClassical classification methods usually assume that pattern recognition models do not depend on the timing of the data. However, this assumption is not valid in cases where new data frequently become available. Such situations are common in practice, for example, spam filtering or fraud detection, where dependencies between feature values and class numbers are continually changing. Unfortunately, most classical machine learning methods (such as decision trees) do not take into consideration the possibility of the model changing, as a result of so-called concept drift and they cannot adapt to a new classification model. This paper focuses on the problem of concept drift, which is a very important issue, especially in data mining methods that use complex structures (such as decision trees) for making decisions. We propose an algorithm that is able to co-train decision trees using a modified NGE (Nested Generalized Exemplar) algorithm. The potential for adaptation of the proposed algorithm and the quality thereof are evaluated through computer experiments, carried out on benchmark datasets from the UCI Machine Learning Repository. Michal Wozniak 0001 |
Knowl. Inf. Syst. | 1 |
| 2010 | Simple combining classifiers for a special case of incremental concept drift problemabstractPaper deals with the problem of designing efficient classifiers for a special case of incremental concept drift. We focus on its classification based on the multiple classifier system. For the problem under consideration we propose four simple methods of combining classification and evaluate them via computer experiments. Bartosz Kurlej, Piotr Sobolewski, Michal Wozniak 0001 |
ISDA | 3 |
| 2010 | Designing combining classifier with trained fuser - Analytical and experimental evaluationabstractCombining pattern recognition is the promising direction in designing an effective classifier systems. There are several approaches of collective decision-making, among them voting methods, where the decision is a combination of individual classifiers' outputs are quite popular. This article focuses on the problem of fuser design which uses continuous outputs of individual classifiers to make a decision. We formulate problem of fuser design as an optimization task and use neural approach as its solver. We propose a taxonomy of aforementioned fusers and their main features are presented for some of them. The results of computer experiments carried out on benchmark datasets confirm quality of proposed concept. Michal Wozniak 0001, Marcin Zmyslony |
ISDA | 1 |
| 2010 | Method of classifier selection using the genetic approachabstractAbstract:The paper presents a novel machine learning algorithm used for training a compound classifier system that consists of a set of area classifiers. Area classifiers recognize objects derived from the respective competence area. Splitting feature space into areas and selecting area classifiers are two key processes of the algorithm; both take place simultaneously in the course of an optimization process aimed at maximizing the system performance. An evolutionary algorithm is used to find the optimal solution. A number of experiments have been carried out to evaluate system performance. The results prove that the proposed method outperforms each elementary classifier as well as simple voting. Konrad Jackowski, Michal Wozniak 0001 |
Expert Syst. J. Knowl. Eng. | 2 |
| 2010 | Cost-sensitive methods of constructing hierarchical classifiersabstractAbstract: The cost of a future exploitation of a decision support system plays a key role. The paper deals with the problem of feature value acquisition cost for such systems. We present a modification of a cost‐sensitive learning method for decision‐tree induction with fixed attribute acquisition cost limit. Properties of the concept are established during computer experiments conducted on chosen benchmark databases from the UCI Machine Learning Repository and a real medical decision task. The results of experiments confirm that, for some decision problems, our proposition allows us to obtain a classifier with the same quality as a classifier obtained without cost limit but its exploitation is cheaper. Wojciech Penar, Michal Wozniak 0001 |
Expert Syst. J. Knowl. Eng. | 2 |
| 2010 | Computer recognition systemsabstractThe aim of the recognition task (Duda et al., 2001) is to classify a given object of interest by assigning it to some predefined category, on the basis of observing the features of the object. Depending on the practical application, these objects (so-called patterns) can be images, signal waveforms or any type of measurements that need to be classified (Theodoridis and Koutroumbas, 2003). Pattern recognition has a long history, properly becoming a scientific discipline at the end of the 1950s with the publication of Frank Rosenblatt's work devoted to the perceptron (Rosenblatt, 1958). Since that time, the progress of computer technology has increased the demand for practical applications of pattern recognition and caused the development of new efficient theoretical methods of recognition required by more and more sophisticated decision problems. Nowadays, the worldwide economy is a knowledge economy (Drucker, 1969), which needs the discovery, transfer and better utilization of knowledge, and pattern recognition methods are widely used by today's engineering applications and research. They are an integral part in most machine intelligence systems built for decision making. There is much current research into developing even more efficient and accurate recognition algorithms, based on technologies such as neural networks, statistical and symbolic learning and fuzzy methods to name but a few. Such methods are implemented in the form of computer software and applied in many practical areas, such as character and speech recognition, machine vision, computer-aided medical diagnosis, prediction of customer behaviour, fraud detection and so on. As a result of the call for papers for this issue entitled ‘Computer Recognition Systems’, the articles were submitted and reviewed through a rigorous peer-review process. In the end, four contributions were selected. We hope that the selected papers provide the reader with an excellent discussion of the current issues of computer recognition systems. The selected topics cover the vital parts of modern recognition systems such as information fusion, bioprosthesis decision control, biometrics, and medical decision support. Proença (2010), in his paper ‘An iris recognition approach through structural pattern analysis methods’, proposes a method based on structural (syntactic) pattern recognition. Iris recognition is at present used in several scenarios (airport check-in, refugee control etc.) with good results. In order to achieve acceptable error rates several imaging constraints are enforced, which reduce the fluidity of iris recognition systems. The related papers existing in the literature use statistical pattern recognition and encode iris texture information through phase, zero-crossing or texture analysis based methods. In this report of the experiments performed, three well-known iris image data sets (CASIA, ICE and UBIRIS) are used. The variability of error rates regarding the amount of noise that images contain is also analysed. The experiments show that the proposed method behaves comparably to the statistical approach that constitutes the basis of nearly all deployed systems. Pietka et al. (2010), in their paper ‘Open architecture computer-aided diagnosis system’, extend the traditional goal of a computer-aided diagnosis (CAD) system to assist physicians in performing diagnosis and treatment. The presented platform helps the system designer in developing a new CAD workflow by implementing general-purpose modules as well as problem-dependent procedures. The CAD environment is validated through its use in three systems – a multiple sclerosis CAD, a lung nodule CAD and a pneumothorax CAD – by which it is shown that the various procedures (fuzzy c-means, fuzzy connectedness, labelling, filters) are developed once but employed by many CADs. The results obtained during the CAD evaluation demonstrate the high flexibility of the infrastructure. The trade-offs, well known to CAD designers (e.g. computational cost versus segmentation accuracy), can easily be handled by the operators in a user-friendly manner by choosing various workflow paths. Straszecka (2010), in her paper ‘Combining knowledge from different sources’, deals with the problem of an assessment of symptoms in medical diagnosis. A unified interpretation of symptoms is often necessary to estimate their significance in a diagnosis. This paper shows how to combine evaluations that may originate from an expert or from statistical features of data for diagnostic cases. A new model of diagnostic inference is proposed in the framework, based on Dempster–Shafer theory extended by fuzzy focal elements. An algorithm of the basic probability assignment calculation is suggested and tested for medical data. Wołczowski and Kurzyński (2010), in their paper ‘Human–machine interface in bioprosthesis control using EMG signal classification’, discuss EMG signal characteristics and the problem of processing them, including acquisition, feature extraction and classification. On the basis of a learning set, a fuzzy relation is determined as a solution of an appropriate optimization problem and then the relation in the form of a matrix of membership degrees is used at successive instants of the sequential decision process. The authors describe three algorithms of sequential classification, which differ from one another in the sets of input data and procedure, and infer that the combination of sequential recognition and fuzzy relation brings new possibilities to EMG signal analysis. We would like to thank the Editor-in-Chief of Expert Systems, Jon G. Hall, for his enthusiasm and continuing support for this special issue. Many of the original reviewers helped in the preparation of the special issue, and we thank them greatly for their help. Thanks also to Wiley-Blackwell's Expert Systems' office, who have made this special issue available in good time. Michal Wozniak Michal Wozniak is Professor of Computer Science in the Department of Systems and Computer Networks, Faculty of Electronics, Wroclaw University of Technology, Poland. He received an MS degree in biomedical engineering in 1992 from the Wroclaw University of Technology, and PhD and DSc (habilitation) degrees in computer science in 1996 and 2007, respectively, from the same university. His research focuses on multiple classifier systems, machine learning, data and web mining, Bayes compound theory, distributed algorithms, computer and networks security and teleinformatics. Professor Wozniak has published over 120 papers and two books, and has edited three books. He is Editor-in-Chief of International Journal of Computer Networks and Communications and associate editor of several international journals including Pattern Analysis and Applications, Expert Systems and International Journal of Communication Networks and Distributed Systems. He serves on the program committees of numerous international conferences. His works have been transitioned into commercial applications. Professor Wozniak has been involved in many research projects related to machine learning, computer networks and telemedicine. Moreover, he has been a consultant on several commercial projects for well-known Polish companies and for the Polish public administration. Professor Wozniak is a member of the IEEE (Computational Intelligence Society and Systems, Man and Cybernetics Society) and IBS (International Biometric Society). For a more detailed profile see http://www.kssk.pwr.wroc.pl/pracownicy/michal.wozniak-en. Elif Derya Übeyli Elif Derya Übeyli (http://edubeyli.etu.edu.tr/) is an Associate Professor at the Department of Electrical and Electronics Engineering, TOBB University of Economics and Technology. She obtained her PhD degree in electronics and computer technology from Gazi University in 2004. She has worked on a variety of topics including biomedical signal processing, neural networks, optimization and artificial intelligence. She has worked on several projects related to biomedical signal acquisition, processing and classification. Dr Übeyli has served (or is currently serving) as a program organizing committee member of many national and international conferences. She is editorial board member of several scientific journals (Journal of Engineering and Applied Sciences, International Journal of Soft Computing, Research Journal of Applied Sciences, Research Journal of Medical Sciences, Scientific Journals International/Electrical, Mechanical, Manufacturing, and Aerospace Engineering, The Open Medical Informatics Journal, Bulletin of the International Scientific Surgical Association, International Journal of Real-Time Systems, Journal of Biomedical Science and Engineering, International Journal of Engineering and Applied Sciences). She is Associate Editor of Expert Systems. She has served as a guest editor to Expert Systems on a special issue on ‘Advances in medical decision support systems’. Moreover, she is voluntarily serving as a technical publication reviewer for many respected scientific journals and conferences. She has also published 118 journal and 44 conference papers on her research areas. Michal Wozniak 0001, Elif Derya Übeyli |
Expert Syst. J. Knowl. Eng. | 1 |
| 2009 | Modeling of Network Computing Systems for Decision Tree Induction Tasks
Krzysztof Walkowiak, Michal Wozniak 0001 |
IDEAL | 2 |
| 2009 | Modification of Nested Hyperrectangle Exemplar as a Proposition of Information Fusion Method
Michal Wozniak 0001 |
IDEAL | 1 |
| 2009 | Algorithm of designing compound recognition system on the basis of combining classifiers with simultaneous splitting feature space into competence areas
Konrad Jackowski, Michal Wozniak 0001 |
Pattern Anal. Appl. | 2 |
| 2006 | Proposition of common classifier construction for pattern recognition with context task
Michal Wozniak 0001 |
Knowl. Based Syst. | 1 |
| 2004 | Information Fusion for Probabilistic Reasoning and Its Application to the Medical Decision Support Systems
Michal Wozniak 0001 |
ICCSA (3) | 1 |
| 2003 | Case- and Rule-Based Algorithms for the Contextual Pattern Recognition Problem
Michal Wozniak 0001 |
ICCSA (1) | 1 |
| 2003 | Proposition of the Quality Measure for the Probabilistic Decision Support System
Michal Wozniak 0001 |
IEA/AIE | 1 |
| 1995 | Diagnosis of Human Acid-Base Balance States via Combined Pattern Recognition of Markov Chains
Marek Kurzynski, Michal Wozniak 0001, Alexandra Blinowska |
AIME | 2 |