EDBT 2026 Demo / reviewers in the wild / expert
Shuo Wang 0005
dblp:63/1591-5
· DBLP profile ↗
35ranked-venue papers
15as first author
13since 2021 · last 2025
0000-0003-1380-6428ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 11 first-author · 6 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-author · 3 since 2021Computer networks · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BALM-TSF: Balanced Multimodal Alignment for LLM-Based Time Series ForecastingabstractTime series forecasting is a long-standing and highly challenging research topic. Recently, driven by the rise of large language models (LLMs), research has increasingly shifted from purely time series methods toward harnessing textual modalities to enhance forecasting performance. However, the vast discrepancy between text and temporal data often leads current multimodal architectures to over-emphasise one modality while neglecting the other, resulting in information loss that harms forecasting performance. To address this modality imbalance, we introduce BALM-TSF (Balanced Multimodal Alignment for LLM-Based Time Series Forecasting), a lightweight time series forecasting framework that maintains balance between the two modalities. Specifically, raw time series are processed by the time series encoder, while descriptive statistics of raw time series are fed to an LLM with learnable prompt, producing compact textual embeddings. To ensure balanced cross-modal context alignment of time series and textual embeddings, a simple yet effective scaling strategy combined with a contrastive objective then maps these textual embeddings into the latent space of the time series embeddings. Finally, the aligned textual semantic embeddings and time series embeddings are together integrated for forecasting. Extensive experiments on standard benchmarks show that, with minimal trainable parameters, BALM-TSF achieves state-of-the-art performance in both long-term and few-shot forecasting, confirming its ability to harness complementary information from text and time series. Code is available at https://github.com/ShiqiaoZhou/BALM-TSF. Shiqiao Zhou, Holger Schöner, Huanbo Lyu, Edouard Fouché, Shuo Wang 0005 |
CIKM | 5 |
| 2025 | Automated Class Imbalance Learning via Few-shot Bayesian Optimization with Meta-learned Deep Kernel SurrogatesabstractThe class imbalance problem is a critical challenge in real-world applications, such as fault diagnosis, intrusion detection, and fraud detection, where the data exhibit highly skewed class distributions. Traditional methods to address class imbalance, such as resampling approaches, require careful model selection and hyperparameter tuning, which are complex and time-consuming. Automated Class Imbalance Learning (AutoCIL) has recently emerged as a promising paradigm, leveraging Combined Algorithm Selection and Hyperparameter Optimization (CASH) to automate this process. However, existing methods often suffer from inefficiencies and ineffectiveness, especially under resource constraints. In this paper, we propose a novel method called AutoCILFBO – Automated Class Imbalance Learning via Few-shot Bayesian Optimization with Meta-learned Deep Kernel Surrogates. Our approach introduces few-shot Bayesian optimization with deep kernel Gaussian processes tailored for class imbalance domains. Specifically, we meta-learn a shared probabilistic deep kernel surrogate model from a collection of pre-evaluated class imbalance optimization tasks, enabling rapid adaptation to target tasks. Experimental results demonstrate that our method outperforms existing approaches across 16 tasks with statistically significant improvements in terms of efficiency and effectiveness. Shuo Wang 0005, Damien Ernst |
IJCNN | 2 |
| 2025 | A Two-Stage Multi-Source Transfer Learning Approach for Software Defect PredictionabstractMulti-source transfer learning (MSTL) acquires knowledge from multiple source domains to improve learning in a target domain, offering higher generalization and robustness than single-source transfer. In software defect prediction, MSTL utilizes cross-project data for more accurate prediction. However, existing methods often neglect the scarcity of defect samples in the target domain. To address this, we propose Cross-Domain Consistency Multisource Transfer Learning (CDC-MSTL), a novel ensemble method based on MsTrAdaBoost. CDC-MSTL considers both domain similarity through two-stage instance transfer and the target dataset’s imbalance degree during training. Experiments on four software projects (23 datasets) show CDC-MSTL significantly outperforms state-of-the-art MSTL methods in Balanced Accuracy (BA) and is competitive in AUC, F-measure, and accuracy. Shuo Wang 0005 |
INDIN | 2 |
| 2025 | FedPAC: A Federated Semi-Supervised Learning Approach for Non-IID Data with Feature ShiftabstractFederated Semi-Supervised Learning (FSSL) enables collaborative model training across distributed clients with limited labeled data while preserving data privacy. However, a critical challenge in FSSL is feature shift, where clients exhibit diverse feature distributions despite sharing the same task. To address this issue, we propose FedPAC, a novel FSSL framework that integrates Contrastive Mean-Teacher Regularization and Perturbation-Aware Gradient Descent. Our framework enhances feature representation learning by aligning feature distributions between teacher and student models and mitigates optimization challenges caused by feature heterogeneity through controlled gradient perturbations. Extensive experiments on benchmark datasets demonstrate that FedPAC outperforms existing FSSL methods in feature shift scenarios, making it a practical solution for real-world applications such as medical imaging and industrial fault diagnosis. Shuo Wang 0005 |
SMC | 2 |
| 2024 | A Study of Virtual Concept Drift in Federated Data Stream LearningabstractWith the widespread application of FL across various domains, learning from data streams and addressing concept drift (i.e. distributional changes in data) have emerged as a crucial research focus in this field. As a major type of concept drift, virtual drift, however, has received very little attention so far. This study aims to provide a deep understanding of how virtual drift in streaming data can affect FL models and how it can be detected and overcome effectively. We propose a FL framework FL-HVD, to tackle virtual drifting data. Based on this framework, firstly, we characterise virtual drift with four spatial features and design 11 virtual drifting scenarios to investigate its impact. Secondly, we study and compare six distribution-based and unsupervised drift detection techniques to identify virtual drift. We find that the distribution-based methods outperform the unsupervised ones in terms of accuracy and timeliness in general, among which Ks_2samp is the best. Thirdly, we explore the effectiveness of three adaptation methods to minimize the negative impact of virtual drift once it is detected. The experimental results demonstrate significant improvements and performance stability achieved by applying the FL-HVD framework combined with the Ks_2samp detector and having more local training rounds in training FL models with virtual drifting data. This paper provides valuable insights and guidance for addressing virtual drift in FL. Tengsen Zhang, Guanhui Yang, Shuo Wang 0005 |
MSN | 4 |
| 2024 | A Multi-Model Approach for Handling Concept Drifting Data in Federated LearningabstractFederated learning (FL) is a distributed machine learning framework that trains a global model with local model updates from multiple client devices. When learning from data streams, the performance of FL models can be significantly degraded, because of concept drift and its subsequent issue of data heterogeneity. Concept drift refers to changes in data distribution. Concept drift in FL can be temporal and spatial, leading to outdated models under new data distributions and poorer model performance from exacerbated heterogeneity. We propose a multi-model approach (FedMCD) that keeps two local models at each client to balance the global and local performance discrepancy caused by concept drift. FedMCdalso integrates drift detection and an adaptation mechanism. We perform an extensive range of experiments to test the effectiveness through 17 artificial scenarios and 3 real-world datasets. We demonstrate that FedMCD achieves the highest accuracy over time in 16 out of 17 scenarios and presents the fastest performance recovery after drifts. It also significantly improved average accuracy on real-world data compared to six other drift-handling approaches. Guanhui Yang, Tengsen Zhang, Shuo Wang 0005 |
MSN | 4 |
| 2024 | Incremental Sampling for Class Imbalanced Data Streams in Federated LearningabstractFederated learning is gaining attention for its ability to extract knowledge from distributed data streams without sharing client data. In this process, each client continuously collects samples from its local data stream for local training. The samples collected by clients at different points in time are only the tip of the iceberg of the overall dataset. The imbalance between these data chunks can be dynamically changing. At the same time, different imbalanced distributions may exist among clients, leading to spatial heterogeneity in data. To fill this research gap in federated learning, this paper presents the following innovation: 1) We identify nine characteristics of class imbalance in federated data streams and study their impact on global performance in a series of artificial scenarios; 2) We develop two new techniques, clustering-based incremental sampling and distance-based incremental sampling, to address the dynamic class imbalance in the federated data stream by retaining key samples locally at the client. Our incremental sampling method performs well and has set new records in real datasets regardless of the aggregation method. Tengsen Zhang, Guanhui Yang, Shuo Wang 0005 |
MSN | 4 |
| 2024 | A Novel Neural Network-Based Multiobjective Evolution Lower Upper Bound Estimation Method for Electricity Load Interval ForecastabstractCurrently, an interval prediction model, lower and upper bounds estimation (LUBE) which constructs the prediction intervals (PIs) by using the double outputs of the neural network (NN) is growing popular. However, existing LUBE researches have two problems. One is that the applied NNs are flawed: feedforward NN (FNN) cannot map the dynamic relationship of data and recurrent NN (RNN) is computationally expensive. The other is that most LUBE models are built under a single-objective frame in which the uncertainty cannot be fully quantified. In this article, a novel wavelet NN (WNN) with direct input–output links (DLWNN) is proposed to obtain PIs in a multiobjective LUBE frame. Different from WNN, the proposed DLWNN adds the direct links from the input layer to output layer which can make full use of the information of time series data. Besides, a niched differential evolution nondominated fast sort genetic algorithm (NDENSGA) is proposed to optimize the prediction model, so as to achieve a balance between estimation accuracy and the average width of the PIs. NDENSGA modifies the traditional population renewal mechanism to increase population diversity and adopts a new elite selection strategy for obtaining more extensive and uniform solutions. The effectiveness of DLWNN and NDENSGA is evaluated through a series of experiments with real electricity load data sets. The results show that the proposed model has better performance than others in terms of convergence and diversity of obtained nondominated solutions. Yaoyao He, Shuo Wang 0005 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2023 | An Impact Study of Concept Drift in Federated LearningabstractFederated learning (FL) is a rising distributed machine learning area, which aims to train a high-performing global model with data collected from a number of local clients. Many FL applications receive data over time in the form of data streams. Streaming data are likely to suffer concept drift. It can significantly harm a model’s predictive ability. However, no study has characterized concept drift in FL or investigated how it can affect the global and local models’ performance. This paper aims to provide such understanding by 1) categorizing concept drift in temporal and spatial dimensions with ten features and 2) investigating the impact of the features in depth. We find that: the temporal features degrade FL models to a different extend and do not affect model convergence after the new data concept becomes stable; the spatial features cause data heterogeneity and affect both accuracy and convergence speed. Guanhui Yang, Tengsen Zhang, Shuo Wang 0005, Yun Yang 0003 |
ICDM | 4 |
| 2023 | Online Automated Machine Learning for Class Imbalanced Data StreamsabstractAutomated machine learning (AutoML) has achieved great success in offline class imbalance learning where data are static. However, many real world applications data nowadays tend to evolve over time in the form of data streams and involve class imbalance distributions, e.g., intrusion detection, fault diagnosis systems, and fraud detection. These learning tasks require AutoML processing the instances instantly and adapting to the dynamic data changes. Nevertheless, existing AutoML research either only focuses on class imbalance in static data sets, or discusses data streams with concept drift. No existing work studied the joint learning challenges of class imbalance and online data stream learning in AutoML. To close the gap, this paper focuses on learning dynamic data streams with a skewed class distribution in AutoML. In this paper, we propose two new AutoML approaches, UEvoAutoML and OEvoAutoML, which integrate adaptive resampling techniques into an existing online AutoML framework. Their performance is investigated through a set of synthetic imbalanced data streams under various stationary and non-stationary scenarios and 5 real-world data streams. As the pioneering work of exploring how class imbalance techniques benefit online AutoML, this paper demonstrated that the effectiveness of adaptive resampling in AutoML frameworks. Shuo Wang 0005 |
IJCNN | 2 |
| 2023 | Triplets Oversampling for Class Imbalanced Federated Datasets
Chenguang Xiao, Shuo Wang 0005 |
ECML/PKDD (2) | 2 |
| 2022 | Adaptive Fuzzy Learning Superpixel Representation for PolSAR Image ClassificationabstractThe increasing applications of polarimetric synthetic aperture radar (PolSAR) image classification demand for effective superpixels’ algorithms. Fuzzy superpixels’ algorithms reduce the misclassification rate by dividing pixels into superpixels, which are groups of pixels of homogenous appearance and undetermined pixels. However, two key issues remain to be addressed in designing a fuzzy superpixel algorithm for PolSAR image classification. First, the polarimetric scattering information, which is unique in PolSAR images, is not effectively used. Such information can be utilized to generate superpixels more suitable for PolSAR images. Second, the ratio of undetermined pixels is fixed for each image in the existing techniques, ignoring the fact that the difficulty of classifying different objects varies in an image. To address these two issues, we propose a polarimetric scattering information-based adaptive fuzzy superpixel (AFS) algorithm for PolSAR images classification. In AFS, the correlation between pixels’ polarimetric scattering information, for the first time, is considered through fuzzy rough set theory to generate superpixels. This correlation is further used to dynamically and adaptively update the ratio of undetermined pixels. AFS is evaluated extensively against different evaluation metrics and compared with the state-of-the-art superpixels’ algorithms on three PolSAR images. The experimental results demonstrate the superiority of AFS on PolSAR image classification problems. Yuwei Guo 0001, Licheng Jiao, Rong Qu, Zhuangzhuang Sun, Shuang Wang 0001, Shuo Wang 0005, Fang Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2021 | Uncertainty analysis of wind power probability density forecasting based on cubic spline interpolation and support vector quantile regression
Yaoyao He, Shuo Wang 0005, Xin Yao 0001 |
Neurocomputing | 3 |
| 2020 | Understanding the automated parameter optimization on transfer learning for cross-project defect prediction: an empirical studyabstractData-driven defect prediction has become increasingly important in software engineering process. Since it is not uncommon that data from a software project is insufficient for training a reliable defect prediction model, transfer learning that borrows data/konwledge from other projects to facilitate the model building at the current project, namely cross-project defect prediction (CPDP), is naturally plausible. Most CPDP techniques involve two major steps, i.e., transfer learning and classification, each of which has at least one parameter to be tuned to achieve their optimal performance. This practice fits well with the purpose of automated parameter optimization. However, there is a lack of thorough understanding about what are the impacts of automated parameter optimization on various CPDP techniques. In this paper, we present the first empirical study that looks into such impacts on 62 CPDP techniques, 13 of which are chosen from the existing CPDP literature while the other 49 ones have not been explored before. We build defect prediction models over 20 real-world software projects that are of different scales and characteristics. Our findings demonstrate that: (1) Automated parameter optimization substantially improves the defect prediction performance of 77% CPDP techniques with a manageable computational cost. Thus more efforts on this aspect are required in future CPDP studies. (2) Transfer learning is of ultimate importance in CPDP. Given a tight computational budget, it is more cost-effective to focus on optimizing the parameter configuration of transfer learning algorithms (3) The research on CPDP is far from mature where it is 'not difficult' to find a better alternative by making a combination of existing transfer learning and classification techniques. This finding provides important insights about the future design of CPDP techniques. Ke Li 0001, Zilin Xiang, Tao Chen 0001, Shuo Wang 0005, Kay Chen Tan |
ICSE | 4 |
| 2020 | AUC Estimation and Concept Drift Detection for Imbalanced Data Streams with Multiple ClassesabstractOnline class imbalance learning deals with data streams having very skewed class distributions. When learning from data streams, concept drift is one of the major challenges that deteriorate the classification performance. Although several approaches have been recently proposed to overcome concept drift in imbalanced data, they are all limited to two-class cases. Multi-class imbalance imposes additional challenges in concept drift detection and performance evaluation, such as a more severe imbalanced distribution and the limited choice of performance measures. This paper extends AUC for evaluating classifiers on multi-class imbalanced data in online learning scenarios. The proposed metrics, PMAUC, WAUC and EWAUC, are studied through comprehensive experiments, focusing on their characteristics on time-changing data streams and whether and how they can be used to detect concept drift. The AUC-based metrics show effectiveness in detecting concept drift in a variety of artificial data streams and a real-world data application with multiple classes. In particular, EWAUC is shown to be both effective and efficient. Shuo Wang 0005, Leandro L. Minku |
IJCNN | 1 |
| 2019 | Learning from data streams and class imbalanceabstractWith the wide application of machine learning algorithms to the real world, class imbalance and concept drift have become crucial learning issues. Applications in various domains such as risk manag... Shuo Wang 0005, Leandro L. Minku, Nitesh V. Chawla, Xin Yao 0001 |
Connect. Sci. | 1 |
| 2019 | Learning in the presence of class imbalance and concept drift
Shuo Wang 0005, Leandro L. Minku, Nitesh V. Chawla, Xin Yao 0001 |
Neurocomputing | 1 |
| 2018 | To Adapt or Not to Adapt?: Technical Debt and Learning Driven Self-Adaptation for Managing Runtime PerformanceabstractSelf-adaptive system (SAS) can adapt itself to optimize various key performance indicators in response to the dynamics and uncertainty in environment. In this paper, we present Debt Learning Driven Adaptation (DLDA), an framework that dynamically determines when and whether to adapt the SAS at runtime. DLDA leverages the temporal adaptation debt, a notion derived from the technical debt metaphor, to quantify the time-varying money that the SAS carries in relation to its performance and Service Level Agreements. We designed a temporal net debt driven labeling to label whether it is economically healthier to adapt the SAS (or not) in a circumstance, based on which an online machine learning classifier learns the correlation, and then predicts whether to adapt under the future circumstances. We conducted comprehensive experiments to evaluate DLDA with two different planners, using 5 online machine learning classifiers, and in comparison to 4 state-of-the-art debt-oblivious triggering approaches. The results reveal the effectiveness and superiority of DLDA according to different metrics. Tao Chen 0001, Rami Bahsoon, Shuo Wang 0005, Xin Yao 0001 |
ICPE | 3 |
| 2018 | Fuzzy Sparse Autoencoder Framework for Single Image Per Person Face RecognitionabstractThe issue of single sample per person (SSPP) face recognition has attracted more and more attention in recent years. Patch/local-based algorithm is one of the most popular categories to address the issue, as patch/local features are robust to face image variations. However, the global discriminative information is ignored in patch/local-based algorithm, which is crucial to recognize the nondiscriminative region of face images. To make the best of the advantage of both local information and global information, a novel two-layer local-to-global feature learning framework is proposed to address SSPP face recognition. In the first layer, the objective-oriented local features are learned by a patch-based fuzzy rough set feature selection strategy. The obtained local features are not only robust to the image variations, but also usable to preserve the discrimination ability of original patches. Global structural information is extracted from local features by a sparse autoencoder in the second layer, which reduces the negative effect of nondiscriminative regions. Besides, the proposed framework is a shallow network, which avoids the over-fitting caused by using multilayer network to address SSPP problem. The experimental results have shown that the proposed local-to-global feature learning framework can achieve superior performance than other state-of-the-art feature learning algorithms for SSPP face recognition. Yuwei Guo 0001, Licheng Jiao, Shuang Wang 0001, Shuo Wang 0005, Fang Liu 0001 |
IEEE Trans. Cybern. | 4 |
| 2018 | Fuzzy Superpixels for Polarimetric SAR Images ClassificationabstractSuperpixels technique has drawn much attention in computer vision applications. Each superpixels algorithm has its own advantages. Selecting a more appropriate superpixels algorithm for a specific application can improve the performance of the application. In the last few years, superpixels are widely used in polarimetric synthetic aperture radar (PolSAR) image classification. However, no superpixel algorithm is especially designed for image classification. It is believed that both mixed superpixels and pure superpixels exist in an image. Nevertheless, mixed superpixels have negative effects on classification accuracy. Thus, it is necessary to generate superpixels containing as few mixed superpixels as possible for image classification. In this paper, first, a novel superpixels concept, named fuzzy superpixels, is proposed for reducing the generation of mixed superpixels. In fuzzy superpixels, not all pixels are assigned to a corresponding superpixel. We would rather ignore the pixels than assigning them to improper superpixels. Second, a new algorithm, named FuzzyS (FS), is proposed to generate fuzzy superpixels for PolSAR image classification. Three PolSAR images are used to verify the effect of the proposed FS algorithm. Experimental results demonstrate the superiority of the proposed FS algorithm over several state-of-the-art superpixels algorithms. Yuwei Guo 0001, Licheng Jiao, Shuang Wang 0001, Shuo Wang 0005, Fang Liu 0001, Wenqiang Hua |
IEEE Trans. Fuzzy Syst. | 4 |
| 2018 | A Systematic Study of Online Class Imbalance Learning With Concept DriftabstractAs an emerging research topic, online class imbalance learning often combines the challenges of both class imbalance and concept drift. It deals with data streams having very skewed class distributions, where concept drift may occur. It has recently received increased research attention; however, very little work addresses the combined problem where both class imbalance and concept drift coexist. As the first systematic study of handling concept drift in class-imbalanced data streams, this paper first provides a comprehensive review of current research progress in this field, including current research focuses and open challenges. Then, an in-depth experimental study is performed, with the goal of understanding how to best overcome concept drift in online learning with class imbalance. Shuo Wang 0005, Leandro L. Minku, Xin Yao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2016 | Dealing with Multiple Classes in Online Class Imbalance Learning
Shuo Wang 0005, Leandro L. Minku, Xin Yao 0001 |
IJCAI | 1 |
| 2016 | Online Ensemble Learning of Data Streams with Gradually Evolved ClassesabstractClass evolution, the phenomenon of class emergence and disappearance, is an important research topic for data stream mining. All previous studies implicitly regard class evolution as a transient change, which is not true for many real-world problems. This paper concerns the scenario where classes emerge or disappear gradually. A class-based ensemble approach, namely Class-Based ensemble for Class Evolution (CBCE), is proposed. By maintaining a base learner for each class and dynamically updating the base learners with new data, CBCE can rapidly adjust to class evolution. A novel under-sampling method for the base learners is also proposed to handle the dynamic class-imbalance problem caused by the gradual evolution of classes. Empirical studies demonstrate the effectiveness of CBCE in various class evolution scenarios in comparison to existing class evolution adaptation methods. Yu Sun 0019, Ke Tang 0001, Leandro L. Minku, Shuo Wang 0005, Xin Yao 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2015 | A novel dynamic rough subspace based selective ensemble
Yuwei Guo 0001, Licheng Jiao, Shuang Wang 0001, Shuo Wang 0005, Fang Liu 0001, Kaixuan Rong |
Pattern Recognit. | 4 |
| 2015 | Resampling-Based Ensemble Methods for Online Class Imbalance LearningabstractOnline class imbalance learning is a new learning problem that combines the challenges of both online learning and class imbalance learning. It deals with data streams having very skewed class distributions. This type of problems commonly exists in real-world applications, such as fault diagnosis of real-time control monitoring systems and intrusion detection in computer networks. In our earlier work, we defined class imbalance online, and proposed two learning algorithms OOB and UOB that build an ensemble model overcoming class imbalance in real time through resampling and time-decayed metrics. In this paper, we further improve the resampling strategy inside OOB and UOB, and look into their performance in both static and dynamic data streams. We give the first comprehensive analysis of class imbalance in data streams, in terms of data distributions, imbalance rates and changes in class imbalance status. We find that UOB is better at recognizing minority-class examples in static data streams, and OOB is more robust against dynamic changes in class imbalance status. The data distribution is a major factor affecting their performance. Based on the insight gained, we then propose two new ensemble methods that maintain both OOB and UOB with adaptive weights for final predictions, called WEOB1 and WEOB2. They are shown to possess the strength of OOB and UOB with good accuracy and robustness. Shuo Wang 0005, Leandro L. Minku, Xin Yao 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2014 | A multi-objective ensemble method for online class imbalance learningabstractOnline class imbalance learning is an emerging learning area that combines the challenges of both online learning and class imbalance learning. In addition to the learning difficulty from the imbalanced distribution, another major challenge is that the imbalanced rate in a data stream can be dynamically changing. OOB and UOB are two state-of-the-art methods for online class imbalance problems [1]. UOB is better at recognizing minority-class examples when the imbalance rate does not change much over time, while OOB is more prepared for the case with a dynamic rate. Aiming for an effective method for both static and dynamic cases, this paper proposes a multi-objective ensemble method MOSOB that combines OOB and UOB. MOSOB finds the Pareto-optimal weights for OOB and UOB at each time step, to maximize minority-class recall and majority-class recall simultaneously. Experiments on five real-world data applications show that MOSOB performs well in both static and dynamic data streams. Furthermore, we look into its performance on a group of highly imbalanced data streams. To respond to the minority class within 10000 time steps, the imbalance rate can be as low as 0.1% for easy data streams; at least 3% of imbalance rate is required to classify difficult data streams. Shuo Wang 0005, Leandro L. Minku, Xin Yao 0001 |
IJCNN | 1 |
| 2014 | A multi-population cooperative coevolutionary algorithm for multi-objective capacitated arc routing problem
Ronghua Shang, Licheng Jiao, Shuo Wang 0005, Liping Qi |
Inf. Sci. | 5 |
| 2013 | Concept drift detection for online class imbalance learningabstractConcept drift detection methods are crucial components of many online learning approaches. Accurate drift detections allow prompt reaction to drifts and help to maintain high performance of online models over time. Although many methods have been proposed, no attention has been given to data streams with imbalanced class distributions, which commonly exist in real-world applications, such as fault diagnosis of control systems and intrusion detection in computer networks. This paper studies the concept drift problem for online class imbalance learning. We look into the impact of concept drift on single-class performance of online models based on three types of classifiers, under seven different scenarios with the presence of class imbalance. The analysis reveals that detecting drift in imbalanced data streams is a more difficult task than in balanced ones. Minority-class recall suffers from a significant drop after the drift involving the minority class. Overall accuracy is not suitable for drift detection. Based on the findings, we propose a new detection method DDM-OCI derived from the existing method DDM. DDM-OCI monitors minority-class recall online to capture the drift. The results show a quick response of the online model working with DDM-OCI to the new concept. Shuo Wang 0005, Leandro L. Minku, Davide Ghezzi, Daniele Caltabiano, Peter Tiño, Xin Yao 0001 |
IJCNN | 1 |
| 2013 | Online Class Imbalance Learning and its Applications in Fault DetectionabstractAlthough class imbalance learning and online learning have been extensively studied in the literature separately, online class imbalance learning that considers the challenges of both fields has not drawn much attention. It deals with data streams having very skewed class distributions, such as fault diagnosis of real-time control monitoring systems and intrusion detection in computer networks. To fill in this research gap and contribute to a wide range of real-world applications, this paper first formulates online class imbalance learning problems. Based on the problem formulation, a new online learning algorithm, sampling-based online bagging (SOB), is proposed to tackle class imbalance adaptively. Then, we study how SOB and other state-of-the-art methods can benefit a class of fault detection data under various scenarios and analyze their performance in depth. Through extensive experiments, we find that SOB can balance the performance between classes very well across different data domains and produce stable G-mean when learning constantly imbalanced data streams, but it is sensitive to sudden changes in class imbalance, in which case SOB's predecessor undersampling-based online bagging (UOB) is more robust. Shuo Wang 0005, Leandro L. Minku, Xin Yao 0001 |
Int. J. Comput. Intell. Appl. | 1 |
| 2013 | Relationships between Diversity of Classification Ensembles and Single-Class Performance MeasuresabstractIn class imbalance learning problems, how to better recognize examples from the minority class is the key focus, since it is usually more important and expensive than the majority class. Quite a few ensemble solutions have been proposed in the literature with varying degrees of success. It is generally believed that diversity in an ensemble could help to improve the performance of class imbalance learning. However, no study has actually investigated diversity in depth in terms of its definitions and effects in the context of class imbalance learning. It is unclear whether diversity will have a similar or different impact on the performance of minority and majority classes. In this paper, we aim to gain a deeper understanding of if and when ensemble diversity has a positive impact on the classification of imbalanced data sets. First, we explain when and why diversity measured by Q-statistic can bring improved overall accuracy based on two classification patterns proposed by Kuncheva et al. We define and give insights into good and bad patterns in imbalanced scenarios. Then, the pattern analysis is extended to single-class performance measures, including recall, precision, and F-measure, which are widely used in class imbalance learning. Six different situations of diversity's impact on these measures are obtained through theoretical analysis. Finally, to further understand how diversity affects the single class performance and overall performance in class imbalance problems, we carry out extensive experimental studies on both artificial data sets and real-world benchmarks with highly skewed class distributions. We find strong correlations between diversity and discussed performance measures. Diversity shows a positive impact on the minority class in general. It is also beneficial to the overall performance in terms of AUC and G-mean. Shuo Wang 0005, Xin Yao 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2013 | Using Class Imbalance Learning for Software Defect PredictionabstractTo facilitate software testing, and save testing costs, a wide range of machine learning methods have been studied to predict defects in software modules. Unfortunately, the imbalanced nature of this type of data increases the learning difficulty of such a task. Class imbalance learning specializes in tackling classification problems with imbalanced distributions, which could be helpful for defect prediction, but has not been investigated in depth so far. In this paper, we study the issue of if and how class imbalance learning methods can benefit software defect prediction with the aim of finding better solutions. We investigate different types of class imbalance learning methods, including resampling techniques, threshold moving, and ensemble algorithms. Among those methods we studied, AdaBoost.NC shows the best overall performance in terms of the measures including balance, G-mean, and Area Under the Curve (AUC). To further improve the performance of the algorithm, and facilitate its use in software defect prediction, we propose a dynamic version of AdaBoost.NC, which adjusts its parameter automatically during training. Without the need to pre-define any parameters, it is shown to be more effective and efficient than the original AdaBoost.NC. Shuo Wang 0005, Xin Yao 0001 |
IEEE Trans. Reliab. | 1 |
| 2012 | Multiclass Imbalance Problems: Analysis and Potential SolutionsabstractClass imbalance problems have drawn growing interest recently because of their classification difficulty caused by the imbalanced class distributions. In particular, many ensemble methods have been proposed to deal with such imbalance. However, most efforts so far are only focused on two-class imbalance problems. There are unsolved issues in multiclass imbalance problems, which exist in real-world applications. This paper studies the challenges posed by the multiclass imbalance problems and investigates the generalization ability of some ensemble solutions, including our recently proposed algorithm AdaBoost.NC, with the aim of handling multiclass and imbalance effectively and directly. We first study the impact of multiminority and multimajority on the performance of two basic resampling techniques. They both present strong negative effects. "Multimajority" tends to be more harmful to the generalization performance. Motivated by the results, we then apply AdaBoost.NC to several real-world multiclass imbalance tasks and compare it to other popular ensemble methods. AdaBoost.NC is shown to be better at recognizing minority class examples and balancing the performance among classes in terms of G-mean without using any class decomposition. Shuo Wang 0005, Xin Yao 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2010 | Negative correlation learning for classification ensemblesabstractThis paper proposes a new negative correlation learning (NCL) algorithm, called AdaBoost.NC, which uses an ambiguity term derived theoretically for classification ensembles to introduce diversity explicitly. All existing NCL algorithms, such as CELS and NCCD, and their theoretical backgrounds were studied in the regression context. We focus on classification problems in this paper. First, we study the ambiguity decomposition with the 0-1 error function, which is different from the one proposed by Krogh et al.. It is applicable to both binary-class and multi-class problems. Then, to overcome the identified drawbacks of the existing algorithms, AdaBoost.NC is proposed by exploiting the ambiguity term in the decomposition to improve diversity. Comprehensive experiments are performed on a collection of benchmark data sets. The results show AdaBoost.NC is a promising algorithm to solve classification problems, which gives better performance than the standard AdaBoost and NCCD, and consumes much less computation time than CELS. Shuo Wang 0005, Huanhuan Chen 0001, Xin Yao 0001 |
IJCNN | 1 |
| 2009 | Diversity analysis on imbalanced data sets by using ensemble modelsabstractMany real-world applications have problems when learning from imbalanced data sets, such as medical diagnosis, fraud detection, and text classification. Very few minority class instances cannot provide sufficient information and result in performance degrading greatly. As a good way to improve the classification performance of weak learner, some ensemble-based algorithms have been proposed to solve class imbalance problem. However, it is still not clear that how diversity affects classification performance especially on minority classes, since diversity is one influential factor of ensemble. This paper explores the impact of diversity on each class and overall performance. As the other influential factor, accuracy is also discussed because of the trade-off between diversity and accuracy. Firstly, three popular re-sampling methods are combined into our ensemble model and evaluated for diversity analysis, which includes under-sampling, over-sampling, and SMOTE - a data generation algorithm. Secondly, we experiment not only on two-class tasks, but also those with multiple classes. Thirdly, we improve SMOTE in a novel way for solving multi-class data sets in ensemble model - SMOTEBagging. Shuo Wang 0005, Xin Yao 0001 |
CIDM | 1 |
| 2009 | Diversity exploration and negative correlation learning on imbalanced data setsabstractClass imbalance learning is an important research area in machine learning, where instances in some classes heavily outnumber the instances in other classes. This unbalanced class distribution causes performance degradation. Some ensemble solutions have been proposed for the class imbalance problem. Diversity has been proved to be an influential aspect in ensemble learning, which describes the degree of different decisions made by classifiers. However, none of those proposed solutions explore the impact of diversity on imbalanced data sets. In addition, most of them are based on re-sampling techniques to rebalance class distribution, and over-sampling usually causes overfitting (high generalisation error). This paper investigates if diversity can relieve this problem by using negative correlation learning (NCL) model, which encourages diversity explicitly by adding a penalty term in the error function of neural networks. A variation model of NCL is also proposed - NCLCost. Our study shows that diversity has a direct impact on the measure of recall. It is also a factor that causes the reduction of F-measure. In addition, although NCL-based models with extreme settings do not produce better recall values of minority class than SMOTEBoost [1], they have slightly better performance of F-measure and G-mean than both independent ANNs and SMOTEBoost and better recall than independent ANNs. Shuo Wang 0005, Ke Tang 0001, Xin Yao 0001 |
IJCNN | 1 |