VLDB 2026 Research / reviewers in the wild / expert
Leandro L. Minku
dblp:55/7992 · also Leandro Lei Minku
· DBLP profile ↗
86ranked-venue papers
12as first author
32since 2021 · last 2025
0000-0002-2639-0671ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 2 first-author · 17 since 2021Software engineering, systems software and programming languages · 35 · 8 first-author · 10 since 2021Databases, data management, data science and information retrieval · 14 · 2 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Theory of computation · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Value of Diversity for Dealing with Concept Drift in Class-Imbalanced Data StreamsabstractConcept drift and class imbalance are critical chal-lenges in real-time data stream learning. Existing ensemble methods use homogeneous diversity (models for the same con-cept) to tackle these challenges but often overlook heterogeneous diversity (models from different concepts), which could improve adaptation, especially with scarce minority data. This paper provides the first analysis of when and why each type of diversity is beneficial for class-imbalanced data streams. To enable this analysis, we introduce CDCMS.CIL, a novel class imbalance learning framework for leveraging heterogeneous diversity. Experiments based on 80 artificial and 9 real-world data streams show that heterogeneous diversity can significantly aid concept drift handling in highly imbalanced scenarios, while homogeneous diversity is better during stable periods. These findings provide crucial guidance for designing robust ensembles for drifting class imbalanced data streams. Chun Wai Chiu, Leandro L. Minku |
DSAA | 2 |
| 2025 | Hyperon: An Online Hyperparameter Tuning Approach for Data Stream LearningabstractPredictive models built using machine learning algorithms usually involve a number of hyperparameters that can significantly affect their performance. While many approaches for hyperparameter tuning have been investigated for offline learning, there is little work in the context of online data stream learning. Hyperparameter tuning for online data stream learning can be particularly challenging, due to possible changes in the underlying distribution of the problem. Such changes can result in the best hyperparameter choice varying over time, requiring efficient, real-time online adaptation. However, existing online hyperparameter tuning approaches are limited to specific models, are susceptible to local optima, rely on fixed hyperparameter grids, or on concept drift detection methods. We propose a novel online hyperparameter tuning approch for data stream learning called Hyperon to overcome these issues. Hyperon intertwines online data stream learning with a steady-state evolutionary algorithm, enabling efficient and effective hyperparameter optimisation over time. Experiments on 10 real world data streams show that Hyperon is able to significantly improve the predictive performance of the underlying online data stream learning approach in a computationally efficient manner. Sadia Tabassum, Leandro L. Minku |
DSAA | 2 |
| 2025 | Multi-Label Transfer Learning in Non-Stationary Data StreamsabstractLabel concepts in multi-label data streams often experience drift in non-stationary environments, either independently or in relation to other labels. Transferring knowledge between related labels can accelerate adaptation, yet research on multi-label transfer learning for data streams remains limited. To address this, we propose two novel transfer learning methods: BR-MARLENE leverages knowledge from different labels in both source and target streams for multi-label classification; BRPW-MARLENE builds on this by explicitly modelling and transferring pairwise label dependencies to enhance learning performance. Comprehensive experiments show that both methods outperform state-of-the-art multi-label stream approaches in non-stationary environments, demonstrating the effectiveness of inter-label knowledge transfer for improved predictive performance. The implementation is available at https://github.com/nino2222/MARLENE. Honghui Du, Leandro L. Minku, Aonghus Lawlor, Huiyu Zhou 0001 |
ICDM | 2 |
| 2025 | Online ensemble model compression for nonstationary data stream learning
Rodrigo G. F. Soares, Leandro L. Minku |
Neural Networks | 2 |
| 2025 | Learning to Expand/Contract Pareto Sets in Dynamic Multiobjective Optimization With a Changing Number of ObjectivesabstractDynamic multi-objective optimization problems (DMOPs) with a changing number of objectives may have Pareto-optimal set (PS) manifold expanding or contracting over time. Knowledge transfer has been used for solving DMOPs, since it can transfer useful information from solving one problem instance to solve another related problem instance. However, we show that the state-of-the-art transfer approach based on heuristic lacks diversity on problem with extremely strong bias and loses convergence on problems with multi-modality and variable correlation, after the number of objectives increases and decreases, respectively. Therefore, we propose a novel transfer strategy based on learning, called learning to expand and contract PS (denoted as LEC) for enhancing diversity and convergence after number of objective increases and decreases, respectively. It firstly learns potentially good directions for expansion and contraction separately via principal component analysis. Then, the most promising expansion and contraction directions are selected from their candidates according to whether they help diversity and convergence, respectively. Lastly, PS is learnt to be expanded and contracted based on these most promising directions. Comprehensive studies using 13 DMOP benchmarks with a changing number of objectives demonstrate that our proposed LEC is effective on improving solution quality, not only right after changes but also after optimization of different generations, compared to state-of-the-art algorithms. Gan Ruan, Leandro L. Minku, Stefan Menzel, Bernhard Sendhoff, Xin Yao 0001 |
IEEE Trans. Evol. Comput. | 2 |
| 2024 | Correction to: An investigation of online and offline learning models for online just-in-time software defect predictionabstractWhile the University of Birmingham exercises care and attention in making items available there are rare occasions when an item has been uploaded in error or has been deemed to be commercially or otherwise sensitive.If you believe that this is the case for this document, please contact [email protected] providing details and we will remove George G. Cabral, Leandro L. Minku, Adriano Lorena Inácio de Oliveira, Dinaldo A. Pessoa, Sadia Tabassum |
Empir. Softw. Eng. | 2 |
| 2024 | Smoclust: synthetic minority oversampling based on stream clustering for evolving data streamsabstractAbstract Many real-world data stream applications not only suffer from concept drift but also class imbalance. Yet, very few existing studies investigated this joint challenge. Data difficulty factors, which have been shown to be key challenges in class imbalanced data streams, are not taken into account by existing approaches when learning class imbalanced data streams. In this work, we propose a drift adaptable oversampling strategy to synthesise minority class examples based on stream clustering. The motivation is that stream clustering methods continuously update themselves to reflect the characteristics of the current underlying concept, including data difficulty factors. This nature can potentially be used to compress past information without caching data in the memory explicitly. Based on the compressed information, synthetic examples can be created within the region that recently generated new minority class examples. Experiments with artificial and real-world data streams show that the proposed approach can handle concept drift involving different minority class decomposition better than existing approaches, especially when the data stream is severely class imbalanced and presenting high proportions of safe and borderline minority class examples. Chun Wai Chiu, Leandro L. Minku |
Mach. Learn. | 2 |
| 2024 | Evolving Memristive ReservoirabstractIn light of the dynamic plasticity, nanosize, and energy efficiency of memristors, memristive reservoirs have attracted increasing attention in diverse fields of research recently. However, limited by deterministic hardware implementation, hardware reservoir adaptation is hard to realize. Existing evolutionary algorithms for evolving reservoirs are not designed for hardware implementation. They often ignore the circuit scalability and feasibility of the memristive reservoirs. In this work, based on the reconfigurable memristive units (RMUs), we first propose an evolvable memristive reservoir circuit that is capable of adaptive evolution for varying tasks, where the configuration signals of memristor are evolved directly avoiding the device variance of the memristors. Second, considering the feasibility and scalability of memristive circuits, we propose a scalable algorithm for evolving the proposed reconfigurable memristive reservoir circuit, where the reservoir circuit will not only be valid according to the circuit laws but also has the sparse topology, alleviating the scalability issue and ensuring the circuit feasibility during the evolution. Finally, we apply our proposed scalable algorithm to evolve the reconfigurable memristive reservoir circuits for a wave generation task, six prediction tasks, and one classification task. Through experiments, the feasibility and superiority of our proposed evolvable memristive reservoir circuit are demonstrated. Xinming Shi, Leandro L. Minku, Xin Yao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | A Practical Human Labeling Method for Online Just-in-Time Software Defect PredictionabstractJust-in-Time Software Defect Prediction (JIT-SDP) can be seen as an online learning problem where additional software changes produced over time may be labeled and used to create training examples. These training examples form a data stream that can be used to update JIT-SDP models in an attempt to avoid models becoming obsolete and poorly performing. However, labeling procedures adopted in existing online JIT-SDP studies implicitly assume that practitioners would not inspect software changes upon a defect-inducing prediction, delaying the production of training examples. This is inconsistent with a real-world scenario where practitioners would adopt JIT-SDP models and inspect certain software changes predicted as defect-inducing to check whether they really induce defects. Such inspection means that some software changes would be labeled much earlier than assumed in existing work, potentially leading to different JIT-SDP models and performance results. This paper aims at formulating a more practical human labeling procedure that takes into account the adoption of JIT-SDP models during the software development process. It then analyses whether and to what extent it would impact the predictive performance of JIT-SDP models. We also propose a new method to target the labeling of software changes with the aim of saving human inspection effort. Experiments based on 14 GitHub projects revealed that adopting a more realistic labeling procedure led to significantly higher predictive performance than when delaying the labeling process, meaning that existing work may have been underestimating the performance of JIT-SDP. In addition, our proposed method to target the labeling process was able to reduce human effort while maintaining predictive performance by recommending practitioners to inspect software changes that are more likely to induce defects. We encourage the adoption of more realistic human labeling methods in research studies to obtain an evaluation of JIT-SDP predictive performance that is closer to reality. Liyan Song, Leandro L. Minku, Cong Teng, Xin Yao 0001 |
ESEC/SIGSOFT FSE | 2 |
| 2023 | An investigation of online and offline learning models for online Just-in-Time Software Defect PredictionabstractAbstract Just-in-Time Software Defect Prediction (JIT-SDP) operates in an online scenario where additional training data is received over time. Existing online JIT-SDP studies used online Oza ensemble learning methods with Hoeffding Trees as base learners to learn and update JIT-SDP models over time in this scenario. However, it is unknown how these approaches compare against offline learning approaches adapted to operate in online scenarios, and how the use of any other online or offline base learners would affect online JIT-SDP in terms of predictive performance and computational cost. We therefore propose a new approach called Batch Oversampling Rate Boosting (BORB) that is able to use offline base learners in an online JIT-SDP scenario. Based on 10 open source projects, we provide a comprehensive evaluation of BORB with 5 different base learners and the existing online approach Oversampling Rate Boosting with 4 different base learners, both in within-project and cross-project online JIT-SDP scenarios. The results show that offline learning can lead to better predictive performance than the top performing online learning approaches considered in our study, at a higher computational cost. Cross-project data was helpful to improve predictive performance both for offline and online learning, but especially for online learning. George G. Cabral, Leandro L. Minku, Adriano Lorena Inácio de Oliveira, Dinaldo A. Pessoa, Sadia Tabassum |
Empir. Softw. Eng. | 2 |
| 2023 | On the validity of retrospective predictive performance evaluation procedures in just-in-time software defect predictionabstractAbstract Just-In-Time Software Defect Prediction (JIT-SDP) is concerned with predicting whether software changes are defect-inducing or clean. It operates in scenarios where labels of software changes arrive over time with delay, which in part corresponds to the time we wait to label software changes as clean (waiting time). However, clean labels decided based on waiting time may be different from the true labels of software changes, i.e., there may be label noise. This typically overlooked issue has recently been shown to affect the validity of continuous performance evaluation procedures used to monitor the predictive performance of JIT-SDP models during the software development process. It is still unknown whether this issue could potentially also affect evaluation procedures that rely on retrospective collection of software changes such as those adopted in JIT-SDP research studies, affecting the validity of the conclusions of a large body of existing work. We conduct the first investigation of the extent with which the choice of waiting time and its corresponding label noise would affect the validity of retrospective performance evaluation procedures. Based on 13 GitHub projects, we found that the choice of waiting time did not have a significant impact on the validity and that even small waiting times resulted in high validity. Therefore, (1) the estimated predictive performances in JIT-SDP studies are likely reliable in view of different waiting times, and (2) future studies can make use of not only larger (5k+ software changes), but also smaller (1k software changes) projects for evaluating performance of JIT-SDP models. Liyan Song, Leandro L. Minku, Xin Yao 0001 |
Empir. Softw. Eng. | 2 |
| 2023 | Tackling Virtual and Real Concept Drifts: An Adaptive Gaussian Mixture Model ApproachabstractReal-world applications have been dealing with large amounts of data that arrive over time and generally present changes in their underlying joint probability distribution, i.e., concept drift. Concept drift can be subdivided into two types: virtual drift, which affects the unconditional probability distribution p(x), and real drift, which affects the conditional probability distribution p(y|x). Existing works focuses on real drift. However, strategies to cope with real drift may not be the best suited for dealing with virtual drift, since the real class boundaries remain unchanged. We provide the first in depth analysis of the differences between the impact of virtual and real drifts on classifiers' suitability. We propose an approach to handle both drifts called On-line Gaussian Mixture Model With Noise Filter For Handling Virtual and Real Concept Drifts (OGMMF-VRD). Experiments with seven synthetics and seven real-world datasets show that OGMMF-VRD outperforms other approaches with separate mechanisms to deal with virtual and real drifts. It also has more stable rankings and smaller drops in performance during drifting periods than existing ensemble approaches, thus being more reliable for adoption in practice. Gustavo H. F. M. Oliveira, Leandro L. Minku, Adriano Lorena Inácio de Oliveira |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Guest Editorial: Special Issue on Stream LearningabstractIn recent years, learning from streaming data, commonly known as stream learning, has enjoyed tremendous growth and shown a wealth of development at both the conceptual and application levels. Stream learning is highly visible in both the machine learning and data science fields and has become a hot new direction in research. Advancements in stream learning include learning with concept drift detection, that includes whether a drift has occurred; understanding where, when, and how a drift occurs; adaptation by actively or passively updating models; and online learning, active learning, incremental learning, and reinforcement learning in data streaming situations. Jie Lu 0001, João Gama 0001, Xin Yao 0001, Leandro L. Minku |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | OSNN: An Online Semisupervised Neural Network for Nonstationary Data StreamsabstractLearning from data streams that emerge from nonstationary environments has many real-world applications and poses various challenges. A key characteristic of such a task is the varying nature of target functions and data distributions over time (concept drifts). Most existing work relies solely on labeled data to adapt to concept drifts in classification problems. However, labeling all instances in a potentially life-long data stream is frequently prohibitively expensive, hindering such approaches. Therefore, we propose a novel algorithm to exploit unlabeled instances, which are typically plentiful and easily obtained. The algorithm is an online semisupervised radial basis function neural network (OSNN) with manifold-based training to exploit unlabeled data while tackling concept drifts in classification problems. OSNN employs a novel semisupervised learning vector quantization (SLVQ) to train network centers and learn meaningful data representations that change over time. It uses manifold learning on dynamic graphs to adjust the network weights. Our experiments confirm that OSNN can effectively use unlabeled data to elucidate underlying structures of data streams while its dynamic topology learning provides robustness to concept drifts. Rodrigo G. F. Soares, Leandro L. Minku |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Towards Reliable Online Just-in-Time Software Defect PredictionabstractThroughout its development period, a software project experiences different phases, comprises modules with different complexities and is touched by many different developers. Hence, it is natural that problems such as Just-in-Time Software Defect Prediction (JIT-SDP) are affected by changes in the defect generating process (concept drifts), potentially hindering predictive performance. JIT-SDP also suffers from delays in receiving the labels of training examples (verification latency), potentially exacerbating the challenges posed by concept drift and further hindering predictive performance. However, little is known about what types of concept drift affect JIT-SDP and how they affect JIT-SDP classifiers in view of verification latency. This work performs the first detailed analysis of that. Among others, it reveals that different types of concept drift together with verification latency significantly impair the stability of the predictive performance of existing JIT-SDP approaches, drastically affecting their reliability over time. Based on the findings, a new JIT-SDP approach is proposed, aimed at providing higher and more stable predictive performance (i.e., reliable) over time. Experiments based on ten GitHub open source projects show that our approach was capable of produce significantly more stable predictive performances in all investigated datasets while maintaining or improving the predictive performance obtained by state-of-art methods. George G. Cabral, Leandro L. Minku |
IEEE Trans. Software Eng. | 2 |
| 2023 | A Procedure to Continuously Evaluate Predictive Performance of Just-In-Time Software Defect Prediction Models During Software DevelopmentabstractJust-In-Time Software Defect Prediction (JIT-SDP) uses machine learning to predict whether software changes are defect-inducing or clean. When adopting JIT-SDP, changes in the underlying defect generating process may significantly affect the predictive performance of JIT-SDP models over time. Therefore, being able to continuously track the predictive performance of JIT-SDP models during the software development process is of utmost importance for software companies to decide whether or not to trust the predictions provided by such models over time. However, there has been little discussion on how to continuously evaluate predictive performance in practice, and such evaluation is not straightforward. In particular, labeled software changes that can be used for evaluation arrive over time with a delay, which in part corresponds to the time we have to wait to label software changes as ‘clean’ (waiting time). A clean label assigned based on a given waiting time may not correspond to the true label of the software changes. This can potentially hinder the validity of any continuous predictive performance evaluation procedure for JIT-SDP models. This paper provides the first discussion of how to continuously evaluate predictive performance of JIT-SDP models over time during the software development process, and the first investigation of whether and to what extent waiting time affects the validity of such continuous performance evaluation procedure in JIT-SDP. Based on 13 GitHub projects, we found that waiting time had a significant impact on the validity. Though typically small, the differences in estimated predicted performance were sometimes large, and thus inappropriate choices of waiting time can lead to misleading estimations of predictive performance over time. Such impact did not normally change the ranking between JIT-SDP models, and thus conclusions in terms of which JIT-SDP model performs better are likely reliable independent of the choice of waiting time, especially when considered across projects. Liyan Song, Leandro L. Minku |
IEEE Trans. Software Eng. | 2 |
| 2023 | Cross-Project Online Just-In-Time Software Defect PredictionabstractCross-Project (CP) Just-In-Time Software Defect Prediction (JIT-SDP) makes use of CP data to overcome the lack of data necessary to train well performing JIT-SDP classifiers at the beginning of software projects. However, such approaches have never been investigated in realistic online learning scenarios, where Within-Project (WP) software changes naturally arrive over time and can be used to automatically update the classifiers. We provide the first investigation of when and to what extent CP data are useful for JIT-SDP in such realistic scenarios. For that, we propose three different online CP JIT-SDP approaches that can be updated with incoming CP and WP training examples over time. We also collect data on 9 proprietary software projects and use 10 open source software projects to analyse these approaches. We find that training classifiers with incoming CP+WP data can lead to absolute improvements in G-mean of up to 53.89% and up to 35.02% at the initial stage of the projects compared to classifiers using WP-only and CP-only data, respectively. Using CP+WP data was also shown to be beneficial after a large number of WP data were received. Using CP data to supplement WP data helped the classifiers to reduce or prevent large drops in predictive performance that may occur over time, leading to absolute G-Mean improvements of up to 37.35% and 48.16% compared to WP-only and CP-only data during such periods, respectively. During periods of stable predictive performance, absolute improvements were of up to 29.03% and up to 41.25% compared to WP-only and CP-only classifiers, respectively. Our results highlight the importance of using both CP and WP data together in realistic online JIT-SDP scenarios. Sadia Tabassum, Leandro L. Minku, Danyi Feng |
IEEE Trans. Software Eng. | 2 |
| 2022 | Benchmarking Dynamic Capacitated Arc Routing Algorithms Using Real-World Traffic SimulationabstractThe dynamic capacitated arc routing problem (DCARP) aims at re-scheduling the service plans of agents, such as vehicles in a city scenario, when dynamic events deteriorate the quality of the current schedule. Various algorithms have been proposed to solve DCARP instances in different dynamic scenarios. However, most existing work evaluated their algorithms' performance based on artificially constructed dynamic environments instead of using more realistic traffic simulations which are built on actual traffic data. In this paper, we constructed a novel DCARP benchmarking framework based on the Simulation of Urban MObility (SUMO) transportation simulation software, which allows to include real-world traffic environments for generating a set of DCARP instances from dynamic events, such as road congestion or task changes. The flexibility of the framework allows to develop DCARP optimization algorithms and evaluate their effectiveness more comprehensively. We use the benchmarking framework to generate 12 different dynamic instances using real-world traffic data of Dublin City. We then demonstrate the value of our framework by using these instances to compare our previously proposed hybrid local search algorithm (HyLS) with a state-of-the-art meta-heuristic optimization algorithm. The generated benchmark scenarios indicate that HyLS is a very effective optimizer on DCARP scenarios with real traffic data for reducing the total service cost. They also demonstrate the importance of our DCARP benchmarking framework for the development and benchmarking of optimization algorithms in more realistic scenarios. Leandro L. Minku, Stefan Menzel, Bernhard Sendhoff, Xin Yao 0001 |
CEC | 2 |
| 2022 | What makes the dynamic capacitated Arc routing problem hard to solve: insights from fitness landscape analysisabstractThe Capacitated Arc Routing Problem (CARP) aims at assigning vehicles to serve tasks which are located at different arcs in a graph. However, the originally planned routes are easily affected by different dynamic events like newly added tasks. This gives rise to Dynamic CARP (DCARP) instances, which need to be efficiently optimized for new high-quality service plans in a short time. However, it is unknown which dynamic events make DCARP instances especially hard to solve. Therefore, in this paper, we provide an investigation of the influence of different dynamic events on DCARP instances from the perspective of fitness landscape analysis based on a recently proposed hybrid local search (HyLS) algorithm. We generate a large set of DCARP instances based on a variety of dynamic events and analyze the fitness landscape of these instances using several different measures such as fitness correlation length. From the empirical results we conclude that cost-related events have no significant impact on the difficulty of DCARP instances, but instances which require more new vehicles to serve the remaining tasks are harder to solve. These insights improve our understanding of the DCARP instances and pave the way for future work on improving the performance of DCARP algorithms. Leandro L. Minku, Stefan Menzel, Bernhard Sendhoff, Xin Yao 0001 |
GECCO | 2 |
| 2022 | Split-AE: An Autoencoder-based Disentanglement Framework for 3D Shape-to-shape Feature TransferabstractRecent advancements in machine learning comprise generative models such as autoencoders (AE) for learning and compressing 3D data to generate low-dimensional latent representations of 3D shapes. Learning latent representations that disentangle the underlying factors of variations in 3D shapes is an intuitive way to achieve generalization in generative models. However, it remains an open problem to learn a generative model of 3D shapes such that the latent variables are disentangled and represent different interpretable aspects of 3D shapes. In this paper, we propose Split-AE, which is an autoencoder-based architecture for partitioning the latent space into two sets, named as content and style codes. The content code represents global features of 3D shapes to differentiate between semantic categories of shapes, while style code represents distinct visual features to differentiate between shape categories having similar semantic meaning. We present qualitative and quantitative experiments to verify feature disentanglement using our Split-AE. Further, we demonstrate that, given a source shape as an initial shape and a target shape as a style reference, the trained Split-AE combines the content of a source and style of a target shape to generate a novel augmented shape, that possesses the distinct features of the target shape category yet maintains the similarity of the global features with the source shape. We conduct a qualitative study showing that the augmented shapes exhibit a realistic interpretable mixture of content and style features across different shape classes with similar semantic meaning. Sneha Saha, Leandro L. Minku, Xin Yao 0001, Bernhard Sendhoff, Stefan Menzel |
IJCNN | 2 |
| 2022 | A Novel Data Stream Learning Approach to Tackle One-Sided Label Noise From Verification LatencyabstractMany real-world data stream applications suffer from verification latency, where the labels of the training examples arrive with a delay. In binary classification problems, the labeling process frequently involves waiting for a pre-determined period of time to observe an event that assigns the example to a given class. Once this time passes, if such labeling event does not occur, the example is labeled as belonging to the other class. For example, in software defect prediction, one may wait to see if a defect is associated to a software change implemented by a developer, producing a defect-inducing training example. If no defect is found during the waiting time, the training example is labeled as clean. Such verification latency inherently causes label noise associated to insufficient waiting time. For example, a defect may be observed only after the pre-defined waiting time has passed, resulting in a noisy example of the clean class. Due to the nature of the waiting time, such noise is frequently one-sided, meaning that it only occurs to examples of one of the classes. However, no existing work tackles label noise associated to verification latency. This paper proposes a novel data stream learning approach that estimates the confidence in the labels assigned to the training examples and uses this to improve predictive performance in problems with one-sided label noise. Our experiments with 14 real-world datasets from the domain of software defect prediction demonstrate the effectiveness of the proposed approach compared to existing ones. Liyan Song, Leandro L. Minku, Xin Yao 0001 |
IJCNN | 3 |
| 2022 | Adaptive Memory-Enhanced Time Delay Reservoir and its Memristive ImplementationabstractTime Delay Reservoir (TDR) is a hardware-friendly machine learning approach from two perspectives. First, it can prevent the connection overhead of neural networks with increasing neurons. Second, through its dynamic system representation, TDR can also be implemented in hardware by different systems. However, it performs poorly on tasks that involve long-term dependency. In this work, we first introduce a higher-order delay unit, which is capable of accumulating and transferring the long history states in an adaptive manner to further enhance the reservoir memory. Particle Swarm Optimisation is applied to optimize the enhanced degree of memory adaptivity. Our experiments demonstrate its superiority both for short- and long-term memory datasets over seven existing approaches. In light of the hardware-friendly feature of TDR, we further propose a memristive implementation of our adaptive memory-enhanced TDR, where a dynamic memristor and the memristor-based delay element are applied to construct the reservoir. Through circuit simulation, the feasibility of our proposed memristive implementation is verified. The comparisons with different hardware reservoirs show that our proposed memristive implementation is effective both for short- and long-term memory datasets, while exhibiting benefits in terms of smaller circuit area and lower power consumption compared with traditional hardware reservoirs. Xinming Shi, Leandro L. Minku, Xin Yao 0001 |
IEEE Trans. Computers | 2 |
| 2022 | A Novel Generalized Metaheuristic Framework for Dynamic Capacitated Arc Routing ProblemsabstractThe capacitated arc routing problem (CARP) is a challenging combinatorial optimization problem abstracted from many real-world applications, such as waste collection, road gritting, and mail delivery. However, few studies considered dynamic changes during the vehicles’ service, which can cause the original schedule infeasible or obsolete. The few existing studies are limited by the dynamic scenarios considered, and by overly complicated algorithms that are unable to benefit from the wealth of contributions provided by the existing CARP literature. In this article, we first provide a mathematical formulation of dynamic CARP (DCARP) and design a simulation system that is able to consider dynamic events while a routing solution is already partially executed. We then propose a novel framework which can benefit from the existing static CARP optimization algorithms so that they could be used to handle DCARP instances. The framework is very flexible. In response to a dynamic event, it can use either a simple restart strategy or a sequence transfer strategy that benefits from the past optimization experience. Empirical studies have been conducted on a wide range of DCARP instances to evaluate our proposed framework. The results show that the proposed framework significantly improves over state-of-the-art dynamic optimization algorithms. Leandro L. Minku, Stefan Menzel, Bernhard Sendhoff, Xin Yao 0001 |
IEEE Trans. Evol. Comput. | 2 |
| 2022 | BIAS: A Toolbox for Benchmarking Structural Bias in the Continuous DomainabstractBenchmarking heuristic algorithms is vital to understand under which conditions and on what kind of problems certain algorithms perform well. Most benchmarks are performance based, to test algorithm performance under a wide set of conditions. There is also resource- and behavior-based benchmarks to test the resource consumption and the behavior of algorithms. In this article, we propose a novel behavior-based benchmark toolbox: BIAS (Bias in algorithms, structural). This toolbox can detect structural bias (SB) per dimension and across dimension-based on 39 statistical tests. Moreover, it predicts the type of SB using a random forest model. BIAS can be used to better understand and improve existing algorithms (removing bias) as well as to test novel algorithms for SB in an early phase of development. Experiments with a large set of generated SB scenarios show that BIAS was successful in identifying bias. In addition, we also provide the results of BIAS on 432 existing state-of-the-art optimization algorithms showing that different kinds of SB are present in these algorithms, mostly toward the center of the objective space or showing discretization behavior. The proposed toolbox is made available open-source and recommendations are provided for the sample size and hyper-parameters to be used when applying the toolbox on other algorithms. Diederick Vermetten, Niki van Stein, Fabio Caraffini, Leandro L. Minku, Anna V. Kononova |
IEEE Trans. Evol. Comput. | 4 |
| 2022 | A Diversity Framework for Dealing With Multiple Types of Concept Drift Based on Clustering in the Model SpaceabstractData stream applications usually suffer from multiple types of concept drift. However, most existing approaches are only able to handle a subset of types of drift well, hindering predictive performance. We propose to use diversity as a framework to handle multiple types of drift. The motivation is that a diverse ensemble can not only contain models representing different concepts, which may be useful to handle recurring concepts, but also accelerate the adaptation to different types of concept drift. Our framework innovatively uses clustering in the model space to build a diverse ensemble and identify recurring concepts. The resulting diversity also accelerates adaptation to different types of drift where the new concept shares similarities with past concepts. Experiments with 20 synthetic and three real-world data streams containing different types of drift show that our diversity framework usually achieves similar or better prequential accuracy than existing approaches, especially when there are recurring concepts or when new concepts share similarities with past concepts. Chun Wai Chiu, Leandro L. Minku |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Continuous and Proactive Software Architecture Evaluation: An IoT CaseabstractDesign-time evaluation is essential to build the initial software architecture to be deployed. However, experts’ assumptions made at design-time are unlikely to remain true indefinitely in systems that are characterized by scale, hyperconnectivity, dynamism, and uncertainty in operations (e.g. IoT). Therefore, experts’ design-time decisions can be challenged at run-time. A continuous architecture evaluation that systematically assesses and intertwines design-time and run-time decisions is thus necessary. This paper proposes the first proactive approach to continuous architecture evaluation of the system leveraging the support of simulation. The approach evaluates software architectures by not only tracking their performance over time, but also forecasting their likely future performance through machine learning of simulated instances of the architecture. This enables architects to make cost-effective informed decisions on potential changes to the architecture. We perform an IoT case study to show how machine learning on simulated instances of architecture can fundamentally guide the continuous evaluation process and influence the outcome of architecture decisions. A series of experiments is conducted to demonstrate the applicability and effectiveness of the approach. We also provide the architect with recommendations on how to best benefit from the approach through choice of learners and input parameters, grounded on experimentation and evidence. Dalia Sobhy, Leandro L. Minku, Rami Bahsoon, Rick Kazman |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2021 | Exploiting Linear Interpolation of Variational Autoencoders for Satisfying Preferences in Evolutionary Design OptimizationabstractIn the early design phase of automotive digital development, one of the key challenges for the designer is to consider multiple-criteria like aerodynamics and structural efficiency besides aesthetic aspects for designing a car shape. In our research, we imagine a cooperative design system in the automotive domain which provides guidance to the designer for finding sets of design options or well-performing designs for preferred search areas. In the present paper, we focus on two perspectives for this multi-criteria decision-making problem: First, a scenario without prior information about design preferences, where the designer aims to explore the search space for a diverse set of design alternatives. Second, a scenario where the designer has a prior intuition on preferred solutions of interest. For both scenarios, we assume that historic 3D car shape data exists, which we can utilize to learn a compact low-dimensional design representation based on a variational autoencoder (VAE). In contrast to evolutionary multi-objective optimization approaches where starting populations are randomly initialized, we propose to seed the population more efficiently by exploiting the advantage of linear interpolation in the latent space of the VAE. In our experiments, we demonstrate that the multi-objective optimization converges faster and achieves a diverse set of solutions. For the second scenario, when specifying design preferences by weights, we improve on the weighted-sum method, which simplifies the multi-objective problem and propose a strategy for efficiently adapting the weights towards the preferred design solution. Sneha Saha, Leandro L. Minku, Xin Yao 0001, Bernhard Sendhoff, Stefan Menzel |
CEC | 2 |
| 2021 | Multi-objective software performance optimisation at the architecture level using randomised search rules
Youcong Ni, Xin Du 0003, Peng Ye 0002, Leandro L. Minku, Xin Yao 0001, Mark Harman, Ruliang Xiao |
Inf. Softw. Technol. | 4 |
| 2021 | Surrogate models in evolutionary single-objective optimization: A new taxonomy and experimental studyabstractSurrogate-assisted evolutionary algorithms (SAEAs), which use efficient surrogate models or meta-models to approximate the fitness function in evolutionary algorithms (EAs), are effective and popular methods for solving computationally expensive optimization problems. During the past decades, a number of SAEAs have been proposed by combining different surrogate models and EAs. This paper dedicates to providing a more systematical review and comprehensive empirical study of surrogate models used in single-objective SAEAs. A new taxonomy of surrogate models in SAEAs for single-objective optimization is introduced in this paper. Surrogate models are classified into two major categories: absolute fitness models, which directly approximate the fitness function values of candidate solutions, and relative fitness models, which estimates the relative rank or preference of candidates rather than their fitness values. Then, the characteristics of different models are analyzed and compared by conducting a series of experiments in terms of time complexity (execution time), model accuracy, parameter influence, and the overall performance when used in EAs. The empirical results are helpful for researchers to select suitable surrogate models when designing SAEAs. Open research questions and future work are discussed at the end of the paper. Changwu Huang, Leandro L. Minku, Xin Yao 0001 |
Inf. Sci. | 3 |
| 2021 | The impact of data difficulty factors on classification of imbalanced and concept drifting data streamsabstractAbstract Class imbalance introduces additional challenges when learning classifiers from concept drifting data streams. Most existing work focuses on designing new algorithms for dealing with the global imbalance ratio and does not consider other data complexities. Independent research on static imbalanced data has highlighted the influential role of local data difficulty factors such as minority class decomposition and presence of unsafe types of examples. Despite often being present in real-world data, the interactions between concept drifts and local data difficulty factors have not been investigated in concept drifting data streams yet. We thoroughly study the impact of such interactions on drifting imbalanced streams. For this purpose, we put forward a new categorization of concept drifts for class imbalanced problems. Through comprehensive experiments with synthetic and real data streams, we study the influence of concept drifts, global class imbalance, local data difficulty factors, and their combinations, on predictions of representative online classifiers. Experimental results reveal the high influence of new considered factors and their local drifts, as well as differences in existing classifiers’ reactions to such factors. Combinations of multiple factors are the most challenging for classifiers. Although existing classifiers are partially capable of coping with global class imbalance, new approaches are needed to address challenges posed by imbalanced data streams. Dariusz Brzezinski, Leandro L. Minku, Tomasz Pewinski, Jerzy Stefanowski, Artur Szumaczuk |
Knowl. Inf. Syst. | 2 |
| 2021 | Dynamic Evaluation of Microservice Granularity AdaptationabstractMicroservices have gained acceptance in software industries as an emerging architectural style for autonomic, scalable, and more reliable computing. Among the critical microservice architecture design decisions is when to adapt the granularity of a microservice architecture by merging/decomposing microservices. No existing work investigates the following question: How can we reason about the trade-off between predicted benefits and cost of pursuing microservice granularity adaptation under uncertainty? To address this question, we provide a novel formulation of the decision problem to pursue granularity adaptation as a real options problem. We propose a novel evaluation process for dynamically evaluating granularity adaptation design decisions under uncertainty. Our process is based on a novel combination of real options and the concept of Bayesian surprises. We show the benefits of our evaluation process by comparing it to four representative industrial microservice runtime monitoring tools, which can be used for retrospective evaluation for granularity adaptation decisions. Our comparison shows that our process can supersede and/or complement these tools. We implement a microservice application—Filmflix—using Amazon Web Service Lambda and use this implementation as a case study to show the unique benefit of our process compared to traditional application of real options analysis. Sara Hassan, Rami Bahsoon, Leandro L. Minku, Nour Ali |
ACM Trans. Auton. Adapt. Syst. | 3 |
| 2021 | Evaluation of Software Architectures under Uncertainty: A Systematic Literature ReviewabstractContext: Evaluating software architectures in uncertain environments raises new challenges, which require continuous approaches. We define continuous evaluation as multiple evaluations of the software architecture that begins at the early stages of the development and is periodically and repeatedly performed throughout the lifetime of the software system. Numerous approaches have been developed for continuous evaluation; to handle dynamics and uncertainties at run-time, over the past years, these approaches are still very few, limited, and lack maturity. Objective: This review surveys efforts on architecture evaluation and provides a unified terminology and perspective on the subject. Method: We conducted a systematic literature review to identify and analyse architecture evaluation approaches for uncertainty including continuous and non-continuous, covering work published between 1990–2020. We examined each approach and provided a classification framework for this field. We present an analysis of the results and provide insights regarding open challenges. Major results and conclusions: The survey reveals that most of the existing architecture evaluation approaches typically lack an explicit linkage between design-time and run-time. Additionally, there is a general lack of systematic approaches on how continuous architecture evaluation can be realised or conducted. To remedy this lack, we present a set of necessary requirements for continuous evaluation and describe some examples. Dalia Sobhy, Rami Bahsoon, Leandro L. Minku, Rick Kazman |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2020 | Computational Study on Effectiveness of Knowledge Transfer in Dynamic Multi-objective OptimizationabstractTransfer learning has been used for solving multiple optimization and dynamic multi-objective optimization problems, since transfer learning is believed to be able to transfer useful information from one problem instance to help solving another related problem instance. This paper aims to study how effective transfer learning is in dynamic multi-objective optimization (DMO). Through computation time analysis of transfer learning, we show that the `inner' optimization problem introduced by transfer learning is very time-consuming. In order to enhance the efficiency, two alternatives are computationally investigated on a number of dynamic bi- and tri-objective test problems. Experimental results have shown that the greatly enhanced efficiency does not result in much degeneration on the performance of transfer learning. Considering the high computational cost of transfer learning, it is likely that the original purpose of using transfer learning in DMO might be negated. In other words, the computation time saved in optimization is eaten up by computationally expensive transfer learning. As a result, there is less gain than expected in the overall computational efficiency. To verify this, experiments have been conducted, regarding using computational cost of transfer learning to optimize randomly generated solutions. The results have demonstrated that the convergence and diversity of final solutions generated from the random solutions are significantly better than those generated from transferred solutions under the same total computational budget. Gan Ruan, Leandro L. Minku, Stefan Menzel, Bernhard Sendhoff, Xin Yao 0001 |
CEC | 2 |
| 2020 | MARLINE: Multi-Source Mapping Transfer Learning for Non-Stationary EnvironmentsabstractConcept drift is a major problem in online learning due to its impact on the predictive performance of data stream mining systems. Recent studies have started exploring data streams from different sources as a strategy to tackle concept drift in a given target domain. These approaches make the assumption that at least one of the source models represents a concept similar to the target concept, which may not hold in many real-world scenarios. In this paper, we propose a novel approach called Multi-source mApping with tRansfer LearnIng for Non-stationary Environments (MARLINE). MARLINE can benefit from knowledge from multiple data sources in non-stationary environments even when source and target concepts do not match. This is achieved by projecting the target concept to the space of each source concept, enabling multiple source sub-classifiers to contribute towards the prediction of the target concept as part of an ensemble. Experiments on several synthetic and real-world datasets show that MARLINE was more accurate than several state-of-the-art data stream learning approaches. Honghui Du, Leandro L. Minku, Huiyu Zhou 0001 |
ICDM | 2 |
| 2020 | An investigation of cross-project learning in online just-in-time software defect predictionabstractJust-In-Time Software Defect Prediction (JIT-SDP) is concerned with predicting whether software changes are defect-inducing or clean based on machine learning classifiers. Building such classifiers requires a sufficient amount of training data that is not available at the beginning of a software project. Cross-Project (CP) JIT-SDP can overcome this issue by using data from other projects to build the classifier, achieving similar (not better) predictive performance to classifiers trained on Within-Project (WP) data. However, such approaches have never been investigated in realistic online learning scenarios, where WP software changes arrive continuously over time and can be used to update the classifiers. It is unknown to what extent CP data can be helpful in such situation. In particular, it is unknown whether CP data are only useful during the very initial phase of the project when there is little WP data, or whether they could be helpful for extended periods of time. This work thus provides the first investigation of when and to what extent CP data are useful for JIT-SDP in a realistic online learning scenario. For that, we develop three different CP JIT-SDP approaches that can operate in online mode and be updated with both incoming CP and WP training examples over time. We also collect 2048 commits from three software repositories being developed by a software company over the course of 9 to 10 months, and use 19,8468 commits from 10 active open source GitHub projects being developed over the course of 6 to 14 years. The study shows that training classifiers with incoming CP+WP data can lead to improvements in G-mean of up to 53.90% compared to classifiers using only WP data at the initial stage of the projects. For the open source projects, which have been running for longer periods of time, using CP data to supplement WP data also helped the classifiers to reduce or prevent large drops in predictive performance that may occur over time, leading to up to around 40% better G-Mean during such periods. Such use of CP data was shown to be beneficial even after a large number of WP data were received, leading to overall G-means up to 18.5% better than those of WP classifiers. Sadia Tabassum, Leandro L. Minku, Danyi Feng, George G. Cabral, Liyan Song |
ICSE | 2 |
| 2020 | AUC Estimation and Concept Drift Detection for Imbalanced Data Streams with Multiple ClassesabstractOnline class imbalance learning deals with data streams having very skewed class distributions. When learning from data streams, concept drift is one of the major challenges that deteriorate the classification performance. Although several approaches have been recently proposed to overcome concept drift in imbalanced data, they are all limited to two-class cases. Multi-class imbalance imposes additional challenges in concept drift detection and performance evaluation, such as a more severe imbalanced distribution and the limited choice of performance measures. This paper extends AUC for evaluating classifiers on multi-class imbalanced data in online learning scenarios. The proposed metrics, PMAUC, WAUC and EWAUC, are studied through comprehensive experiments, focusing on their characteristics on time-changing data streams and whether and how they can be used to detect concept drift. The AUC-based metrics show effectiveness in detecting concept drift in a variety of artificial data streams and a real-world data application with multiple classes. In particular, EWAUC is shown to be both effective and efficient. Shuo Wang 0005, Leandro L. Minku |
IJCNN | 2 |
| 2020 | Towards Novel Meta-heuristic Algorithms for Dynamic Capacitated Arc Routing Problems
Leandro L. Minku, Stefan Menzel, Bernhard Sendhoff, Xin Yao 0001 |
PPSN (2) | 2 |
| 2020 | Better software analytics via "DUO": Data mining algorithms using/used-by optimizers
Amritanshu Agrawal, Tim Menzies, Leandro L. Minku, Markus Wagner 0007, Zhe Yu 0002 |
Empir. Softw. Eng. | 3 |
| 2020 | Guest Editorial: Special Issue on Predictive Models and Data Analytics in Software Engineering
Ayse Tosun Misirli, Shane McIntosh, Leandro L. Minku, Burak Turhan |
Empir. Softw. Eng. | 3 |
| 2020 | Run-time evaluation of architectures: A case study of diversification in IoT
Dalia Sobhy, Leandro L. Minku, Rami Bahsoon, Tao Chen 0001, Rick Kazman |
J. Syst. Softw. | 2 |
| 2020 | A heterogeneous online learning ensemble for non-stationary environments
Mobin M. Idrees, Leandro L. Minku, Frederic T. Stahl, Atta Badii |
Knowl. Based Syst. | 2 |
| 2019 | Class imbalance evolution and verification latency in just-in-time software defect predictionabstractJust-in-Time Software Defect Prediction (JIT-SDP) is an SDP approach that makes defect predictions at the software change level. Most existing JIT-SDP work assumes that the characteristics of the problem remain the same over time. However, JIT-SDP may suffer from class imbalance evolution. Specifically, the imbalance status of the problem (i.e., how much underrepresented the defect-inducing changes are) may be intensified or reduced over time. If occurring, this could render existing JIT-SDP approaches unsuitable, including those that re-build classifiers over time using only recent data. This work thus provides the first investigation of whether class imbalance evolution poses a threat to JIT-SDP. This investigation is performed in a realistic scenario by taking into account verification latency -- the often overlooked fact that labeled training examples arrive with a delay. Based on 10 GitHub projects, we show that JIT-SDP suffers from class imbalance evolution, significantly hindering the predictive performance of existing JIT-SDP approaches. Compared to state-of-the-art class imbalance evolution learning approaches, the predictive performance of JIT-SDP approaches was up to 97.2% lower in terms of g-mean. Hence, it is essential to tackle class imbalance evolution in JIT-SDP. We then propose a novel class imbalance evolution approach for the specific context of JIT-SDP. While maintaining top ranked g-means, this approach managed to produce up to 63.59% more balanced recalls on the defect-inducing and clean classes than state-of-the-art class imbalance evolution approaches. We thus recommend it to avoid overemphasizing one class over the other in JIT-SDP. George G. Cabral, Leandro L. Minku, Emad Shihab, Suhaib Mujahid |
ICSE | 2 |
| 2019 | Multi-Source Transfer Learning for Non-Stationary EnvironmentsabstractIn data stream mining, predictive models typically suffer drops in predictive performance due to concept drift. As enough data representing the new concept must be collected for the new concept to be well learnt, the predictive performance of existing models usually takes some time to recover from concept drift. To speed up recovery from concept drift and improve predictive performance in data stream mining, this work proposes a novel approach called Multi-sourcE onLine TrAnsfer learning for Non-statIonary Environments (Melanie). Melanie is the first approach able to transfer knowledge between multiple data streaming sources in non-stationary environments. It creates several sub-classifiers to learn different aspects from different source and target concepts over time. The sub-classifiers that match the current target concept well are identified, and used to compose an ensemble for predicting examples from the target concept. We evaluate Melanie on several synthetic data streams containing different types of concept drift and on real world data streams. The results indicate that Melanie can deal with a variety drifts and improve predictive performance over existing data stream learning algorithms by making use of multiple sources. Honghui Du, Leandro L. Minku, Huiyu Zhou 0001 |
IJCNN | 2 |
| 2019 | GMM-VRD: A Gaussian Mixture Model for Dealing With Virtual and Real Concept DriftsabstractConcept drift is a change in the joint probability distribution of the problem. This term can be subdivided into two types: real drifts that affect the conditional probabilities p(y|x) or virtual drifts that affect the unconditional probability distribution p(x). Most existing work focuses on dealing with real concept drifts. However, virtual drifts can also cause degradation in predictive performance, requiring mechanisms to be tackled. Moreover, as virtual drifts frequently mean that part of the old knowledge remains useful, they require different strategies from real drifts to be effectively tackled. Motivated on this, we propose an approach called Gaussian Mixture Model for Dealing With Virtual and Real Concept Drifts (GMM-VRD), which updates and creates Gaussians to tackle virtual drifts and resets the system to deal with real drifts. The main results show that the proposed approach obtained the best results, in terms of average accuracy, in relation to the literature methods, which propose to solve that same problem. In terms of accuracy over time, the proposed approach showed lower degradation on concept drifts, which indicates that the proposed approach was efficient. Gustavo H. F. M. Oliveira, Leandro L. Minku, Adriano Lorena Inácio de Oliveira |
IJCNN | 2 |
| 2019 | Learning from data streams and class imbalanceabstractWith the wide application of machine learning algorithms to the real world, class imbalance and concept drift have become crucial learning issues. Applications in various domains such as risk manag... Shuo Wang 0005, Leandro L. Minku, Nitesh V. Chawla, Xin Yao 0001 |
Connect. Sci. | 2 |
| 2019 | A novel online supervised hyperparameter tuning procedure applied to cross-company software effort estimationabstractSoftware effort estimation is an online supervised learning problem, where new training projects may become available over time. In this scenario, the Cross-Company (CC) approach Dycom can drastically reduce the number of Within-Company (WC) projects needed for training, saving their collection cost. However, Dycom requires CC projects to be split into subsets. Both the number and composition of such subsets can affect Dycom’s predictive performance. Even though clustering methods could be used to automatically create CC subsets, there are no procedures for automatically tuning the number of clusters over time in online supervised scenarios. This paper proposes the first procedure for that. An investigation of Dycom using six clustering methods and three automated tuning procedures is performed, to check whether clustering with automated tuning can create well performing CC splits. A case study with the ISBSG Repository shows that the proposed tuning procedure in combination with a simple threshold-based clustering method is the most successful in enabling Dycom to drastically reduce (by a factor of 10) the number of required WC training projects, while maintaining (or even improving) predictive performance in comparison with a corresponding WC model. A detailed analysis is provided to understand the conditions under which this approach does or does not work well. Overall, the proposed online supervised tuning procedure was generally successful in enabling a very simple threshold-based clustering approach to obtain the most competitive Dycom results. This demonstrates the value of automatically tuning hyperparameters over time in a supervised way. Leandro L. Minku |
Empir. Softw. Eng. | 1 |
| 2019 | Learning in the presence of class imbalance and concept drift
Shuo Wang 0005, Leandro L. Minku, Nitesh V. Chawla, Xin Yao 0001 |
Neurocomputing | 2 |
| 2019 | Software Effort Interval Prediction via Bayesian Inference and Synthetic Bootstrap ResamplingabstractSoftware effort estimation (SEE) usually suffers from inherent uncertainty arising from predictive model limitations and data noise. Relying on point estimation only may ignore the uncertain factors and lead project managers (PMs) to wrong decision making. Prediction intervals (PIs) with confidence levels (CLs) present a more reasonable representation of reality, potentially helping PMs to make better-informed decisions and enable more flexibility in these decisions. However, existing methods for PIs either have strong limitations or are unable to provide informative PIs. To develop a “better” effort predictor, we propose a novel PI estimator called Synthetic Bootstrap ensemble of Relevance Vector Machines (SynB-RVM) that adopts Bootstrap resampling to produce multiple RVM models based on modified training bags whose replicated data projects are replaced by their synthetic counterparts. We then provide three ways to assemble those RVM models into a final probabilistic effort predictor, from which PIs with CLs can be generated. When used as a point estimator, SynB-RVM can either significantly outperform or have similar performance compared with other investigated methods. When used as an uncertain predictor, SynB-RVM can achieve significantly narrower PIs compared to its base learner RVM. Its hit rates and relative widths are no worse than the other compared methods that can provide uncertain estimation. Liyan Song, Leandro L. Minku, Xin Yao 0001 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2018 | Are 20% of files responsible for 80% of defects?abstractBackground: Over the past two decades a mixture of anecdote from the industry and empirical studies from academia have suggested that the 80:20 rule (otherwise known as the Pareto Principle) applies to the relationship between source code files and the number of defects in the system: a small minority of files (roughly 20%) are responsible for a majority of defects (roughly 80%). Neil Walkinshaw, Leandro L. Minku |
ESEM | 2 |
| 2018 | Diversity-Based Pool of Models for Dealing with Recurring ConceptsabstractSeveral data stream applications involve recurring concepts, i.e., concept drifts that change the underlying distribution of the data to a distribution previously seen in the data stream. Examples include electricity price prediction and tweet topic classification. In such scenario, it is useful to maintain a pool of old models that could be recovered if their knowledge matches the recurring concept well. A few existing online learning approaches maintain such pools. However, there has been a little investigation on what is the best strategy to maintain an online learning pool with a limited size. We propose to make use of diversity to decide which models to keep in the pool once the pool reaches the maximum size. The motivation behind is that a diverse pool is more likely to maintain a set of representative models with considerably different concepts, helping to handle recurring concepts. We perform experiments to investigate if, when and why maintaining a diverse pool is helpful. The results show that the use of diversity to maintain pools can indeed be helpful to handle recurring concepts. However, the relationship between diversity and accuracy in the presence of concept drift is not straightforward. In particular, an initially good accuracy obtained when using diversity can lead to a stronger subsequent drop in accuracy than other strategies. Chun Wai Chiu, Leandro L. Minku |
IJCNN | 2 |
| 2018 | Data-driven search-based software engineeringabstractThis paper introduces Data-Driven Search-based Software Engineering (DSE), which combines insights from Mining Software Repositories (MSR) and Search-based Software Engineering (SBSE). While MSR formulates software engineering problems as data mining problems, SBSE reformulate Software Engineering (SE) problems as optimization problems and use meta-heuristic algorithms to solve them. Both MSR and SBSE share the common goal of providing insights to improve software engineering. The algorithms used in these two areas also have intrinsic relationships. We, therefore, argue that combining these two fields is useful for situations (a) which require learning from a large data source or (b) when optimizers need to know the lay of the land to find better solutions, faster. Vivek Nair, Amritanshu Agrawal, Wei Fu 0002, George Mathew, Tim Menzies, Leandro L. Minku, Markus Wagner 0007, Zhe Yu 0002 |
MSR | 7 |
| 2018 | A novel automated approach for software effort estimation based on data augmentationabstractSoftware effort estimation (SEE) usually suffers from data scarcity problem due to the expensive or long process of data collection. As a result, companies usually have limited projects for effort estimation, causing unsatisfactory prediction performance. Few studies have investigated strategies to generate additional SEE data to aid such learning. We aim to propose a synthetic data generator to address the data scarcity problem of SEE. Our synthetic generator enlarges the SEE data set size by slightly displacing some randomly chosen training examples. It can be used with any SEE method as a data preprocessor. Its effectiveness is justified with 6 state-of-the-art SEE models across 14 SEE data sets. We also compare our data generator against the only existing approach in the SEE literature. Experimental results show that our synthetic projects can significantly improve the performance of some SEE methods especially when the training data is insufficient. When they cannot significantly improve the prediction performance, they are not detrimental either. Besides, our synthetic data generator is significantly superior or perform similarly to its competitor in the SEE literature. Therefore, our data generator plays a non-harmful if not significantly beneficial effect on the SEE methods investigated in this paper. Therefore, it is helpful in addressing the data scarcity problem of SEE. Liyan Song, Leandro L. Minku, Xin Yao 0001 |
ESEC/SIGSOFT FSE | 2 |
| 2018 | A Q-learning-based memetic algorithm for multi-objective dynamic software project scheduling
Xiao-Ning Shen, Leandro L. Minku, Naresh Marturi, Yinan Guo 0001 |
Inf. Sci. | 2 |
| 2018 | Guest editorial: special issue on predictive models for software quality
Leandro L. Minku, Ayse Basar Bener, Burak Turhan |
Softw. Qual. J. | 1 |
| 2018 | A Systematic Study of Online Class Imbalance Learning With Concept DriftabstractAs an emerging research topic, online class imbalance learning often combines the challenges of both class imbalance and concept drift. It deals with data streams having very skewed class distributions, where concept drift may occur. It has recently received increased research attention; however, very little work addresses the combined problem where both class imbalance and concept drift coexist. As the first systematic study of handling concept drift in class-imbalanced data streams, this paper first provides a comprehensive review of current research progress in this field, including current research focuses and open challenges. Then, an in-depth experimental study is performed, with the goal of understanding how to best overcome concept drift in online learning with class imbalance. Shuo Wang 0005, Leandro L. Minku, Xin Yao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Effort and Cost in Software Engineering: A Comparison of Two Industrial Data SetsabstractContext The research literature on software development projects usually assumes that effort is a good proxy for cost. Practice, however, suggests that there are circumstances in which costs and effort should be distinguished. Objectives: We determine similarities and differences between size, effort, cost, duration, and number of defects of software projects. Method: We compare two established repositories (ISBSG and EBSPM) comprising almost 700 projects from industry. Results: We demonstrate a (log)-linear relation between cost on the one hand, and size, duration and number of defects on the other. This justifies conducting linear regression for cost. We establish that ISBSG is substantially different from EBSPM, in terms of cost (cheaper) and duration (faster), and the relation between cost and effort. We show that while in ISBSG effort is the most important cost factor, this is not the case in other repositories, such as EBSPM in which size is the dominant factor. Conclusion: Practitioners and researchers alike should be cautious when drawing conclusions from a single repository. Hennie Huijgens, Arie van Deursen, Leandro L. Minku, Christopher J. Lokan |
EASE | 3 |
| 2017 | Time Series Forecasting in the Presence of Concept Drift: A PSO-based ApproachabstractTime series forecasting is a problem with many applications. However, in many domains, such as stock market, the underlying generating process of the time series observations may change, making forecasting models obsolete. This problem is known as Concept Drift. Approaches for time series forecasting should be able to detect and react to concept drift in a timely manner, so that the forecasting model can be updated as soon as possible. Despite the fact that the concept drift problem is well investigated in the literature, little effort has been made to solve this problem for time series forecasting so far. This work proposes two novel methods for dealing with the time series forecasting problem in the presence of concept drift. The proposed methods benefit from the Particle Swarm Optimization (PSO) technique to detect and react to concept drifts in the time series data stream. It is expected that the use of collective intelligence of PSO makes the proposed method more robust to false positive drift detections while maintaining a low error rate on the forecasting task. Experiments show that the methods achieved competitive results in comparison to state-of-the-art methods. Gustavo H. F. M. Oliveira, Rodolfo Carneiro Cavalcante, George G. Cabral, Leandro L. Minku, Adriano Lorena Inácio de Oliveira |
ICTAI | 4 |
| 2017 | Which models of the past are relevant to the present? A software effort estimation approach to exploiting useful past modelsabstractSoftware Effort Estimation (SEE) models can be used for decision-support by software managers to determine the effort required to develop a software project. They are created based on data describing projects completed in the past. Such data could include past projects from within the company that we are interested in (WC projects) and/or from other companies (cross-company, i.e., CC projects). In particular, the use of CC data has been investigated in an attempt to overcome limitations caused by the typically small size of WC datasets. However, software companies operate in non-stationary environments, where changes may affect the typical effort required to develop software projects. Our previous work showed that both WC and CC models of the past can become more or less useful over time, i.e., they can sometimes be helpful and sometimes misleading. So, how can we know if and when a model created based on past data represents well the current projects being estimated? We propose an approach called Dynamic Cross-company Learning (DCL) to dynamically identify which WC or CC past models are most useful for making predictions to a given company at the present. DCL automatically emphasizes the predictions given by these models in order to improve predictive performance. Our experiments comparing DCL against existing WC and CC approaches show that DCL is successful in improving SEE by emphasizing the most useful past models. A thorough analysis of DCL’s behaviour is provided, strengthening its external validity. Leandro L. Minku, Xin Yao 0001 |
Autom. Softw. Eng. | 1 |
| 2016 | Diversifying Software Architecture for Sustainability: A Value-Based Perspective
Dalia Sobhy, Rami Bahsoon, Leandro L. Minku, Rick Kazman |
ECSA | 3 |
| 2016 | Dealing with Multiple Classes in Online Class Imbalance Learning
Shuo Wang 0005, Leandro L. Minku, Xin Yao 0001 |
IJCAI | 2 |
| 2016 | FEDD: Feature Extraction for Explicit Concept Drift Detection in time seriesabstractA time series is a sequence of observations collected over fixed sampling intervals. Several real-world dynamic processes can be modeled as a time series, such as stock price movements, exchange rates, temperatures, among others. As a special kind of data stream, a time series may present concept drift, which affects negatively time series analysis and forecasting. Explicit drift detection methods based on monitoring the time series features may provide a better understanding of how concepts evolve over time than methods based on monitoring the forecasting error of a base predictor. In this paper, we propose an online explicit drift detection method that identifies concept drifts in time series by monitoring time series features, called Feature Extraction for Explicit Concept Drift Detection (FEDD). Computational experiments showed that FEDD performed better than error-based approaches in several linear and nonlinear artificial time series with abrupt and gradual concept drifts. Rodolfo Carneiro Cavalcante, Leandro L. Minku, Adriano Lorena Inácio de Oliveira |
IJCNN | 2 |
| 2016 | An Evolutionary Hyper-heuristic for the Software Project Scheduling Problem
Xiuli Wu, Pietro A. Consoli, Leandro L. Minku, Gabriela Ochoa, Xin Yao 0001 |
PPSN | 3 |
| 2016 | Dynamic selection of evolutionary operators based on online learning and fitness landscape analysisabstractSelf-adaptive mechanisms for the identification of the most suitable variation operator in evolutionary algorithms rely almost exclusively on the measurement of the fitness of the offspring, which may not be sufficient to assess the optimality of an operator (e.g., in a landscape with an high degree of neutrality). This paper proposes a novel adaptive operator selection mechanism which uses a set of four fitness landscape analysis techniques and an online learning algorithm, dynamic weighted majority, to provide more detailed information about the search space to better determine the most suitable crossover operator. Experimental analysis on the capacitated arc routing problem has demonstrated that different crossover operators behave differently during the search process, and selecting the proper one adaptively can lead to more promising results. Pietro A. Consoli, Yi Mei 0001, Leandro L. Minku, Xin Yao 0001 |
Soft Comput. | 3 |
| 2016 | Online Ensemble Learning of Data Streams with Gradually Evolved ClassesabstractClass evolution, the phenomenon of class emergence and disappearance, is an important research topic for data stream mining. All previous studies implicitly regard class evolution as a transient change, which is not true for many real-world problems. This paper concerns the scenario where classes emerge or disappear gradually. A class-based ensemble approach, namely Class-Based ensemble for Class Evolution (CBCE), is proposed. By maintaining a base learner for each class and dynamically updating the base learners with new data, CBCE can rapidly adjust to class evolution. A novel under-sampling method for the base learners is also proposed to handle the dynamic class-imbalance problem caused by the gradual evolution of classes. Empirical studies demonstrate the effectiveness of CBCE in various class evolution scenarios in comparison to existing class evolution adaptation methods. Yu Sun 0019, Ke Tang 0001, Leandro L. Minku, Shuo Wang 0005, Xin Yao 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2016 | Dynamic Software Project Scheduling through a Proactive-Rescheduling MethodabstractSoftware project scheduling in dynamic and uncertain environments is of significant importance to real-world software development. Yet most studies schedule software projects by considering static and deterministic scenarios only, which may cause performance deterioration or even infeasibility when facing disruptions. In order to capture more dynamic features of software project scheduling than the previous work, this paper formulates the project scheduling problem by considering uncertainties and dynamic events that often occur during software project development, and constructs a mathematical model for the resulting multi-objective dynamic project scheduling problem (MODPSP), where the four objectives of project cost, duration, robustness and stability are considered simultaneously under a variety of practical constraints. In order to solve MODPSP appropriately, a multi-objective evolutionary algorithm based proactive-rescheduling method is proposed, which generates a robust schedule predictively and adapts the previous schedule in response to critical dynamic events during the project execution. Extensive experimental results on 21 problem instances, including three instances derived from real-world software projects, show that our novel method is very effective. By introducing the robustness and stability objectives, and incorporating the dynamic optimization strategies specifically designed for MODPSP, our proactive-rescheduling method achieves a very good overall performance in a dynamic environment. Xiao-Ning Shen, Leandro L. Minku, Rami Bahsoon, Xin Yao 0001 |
IEEE Trans. Software Eng. | 2 |
| 2015 | An evolutionary algorithm for performance optimization at software architecture levelabstractArchitecture-based software performance optimization can not only significantly save time but also reduce cost. A few rule-based performance optimization approaches at software architecture (SA) level have been proposed in recent years. However, in these approaches, the number of rules being used and the order of application of each rule are uncertain in the optimization process and these uncertainties have not been fully considered so far. As a result, the search space for performance improvement is limited, possibly excluding optimal solutions. Aiming to solve this problem, we propose an evolutionary algorithm for rule-based performance optimization at SA level named EA4PO. First, the rule-based software performance optimization at SA level is abstracted into a mathematical model called RPOM. RPOM can precisely characterize the mathematical relation between the usage of rules and the optimal solution in the performance improvement space. Then, a framework named RSEF is designed to support the execution of rule sequences. Based on RPOM and RSEF, EA4PO is proposed to find the optimal performance improvement solution. In EA4PO, an adaptive mutation operator is designed to guide the search direction by fully considering heuristic information of rule usage during the evolution. Finally, the effectiveness of EA4PO is validated by comparing EA4PO with a typical rule-based approach. The results show that EA4PO can explore a relatively larger space and get better solutions. Xin Du 0003, Youcong Ni, Peng Ye 0002, Xin Yao 0001, Leandro L. Minku, Ruliang Xiao |
CEC | 5 |
| 2015 | How to Make Best Use of Cross-Company Data for Web Effort Estimation?abstract[Context]: The numerous challenges that can hinder software companies from gathering their own data have motivated over the past 15 years research on the use of cross-company (CC) datasets for software effort prediction. Part of this research focused on Web effort prediction, given the large increase worldwide in the development of Web applications. Some of these studies indicate that it may be possible to achieve better performance using CC models if some strategy to make the CC data more similar to the within-company (WC) data is adopted. [Goal]: This study investigates the use of a recently proposed approach called Dycom to assess to what extent Web effort predictions obtained using CC datasets are effective in relation to the predictions obtained using WC data when explicitly mapping the CC models to the WC context. [Method]: Data on 125 Web projects from eight different companies part of the Tukutuku database were used to build prediction models. We benchmarked these models against baseline models (mean and median effort) and a WC base learner that does not benefit of the mapping. We also compared Dycom against a competitive CC approach from the literature (NN-filtering). We report a company-by- company analysis. [Results]: Dycom usually managed to achieve similar or better performance than a WC model while using only half of the WC training data. These results are also an improvement over previous studies that investigated the use of different strategies to adapt CC models to the WC data for Web effort estimation. [Conclusions]: We conclude that the use of Dycom for Web effort prediction is quite promising and in general supports previous results when applying Dycom to conventional software datasets. Leandro L. Minku, Federica Sarro, Emilia Mendes, Filomena Ferrucci |
ESEM | 1 |
| 2015 | The Art and Science of Analyzing Software Data; Quantitative MethodsabstractUsing the tools of quantitative data science, software engineers that can predict useful information on new projects based on past projects. This tutorial reflects on the state-of-the-art in quantitative reasoning in this important field. This tutorial discusses the following: (a) when local data is scarce, we show how to adapt data from other organizations to local problems; (b) when working with data of dubious quality, we show how to prune spurious information; (c) when data or models seem too complex, we show how to simplify data mining results; (d) when the world changes, and old models need to be updated, we show how to handle those updates; (e) when the effect is too complex for one model, we show to how reason over ensembles. Tim Menzies, Leandro L. Minku, Fayola Peters |
ICSE (2) | 2 |
| 2015 | 4th International Workshop on Realizing AI Synergies in Software Engineering (RAISE 2015)abstractThis workshop is the fourth in the series and continued to build upon the work carried out at the previous iterations of the International Workshop on Realizing Artificial Intelligence Synergies in Software Engineering, which were held at ICSE in 2012, 2013 and 2014. RAISE 2015 brought together researchers and practitioners from the artificial intelligence (AI) and software engineering (SE) disciplines to build on the interdis- ciplinary synergies that exist and to stimulate further interaction across these disciplines. Mutually beneficial characteristics have appeared in the past few decades and are still evolving due to new challenges and technological advances. Hence, the question that motivates and drives the RAISE Workshop series is: "Are SE and AI researchers ignoring important insights from AI and SE?". To pursue this question, RAISE'15 explored not only the application of AI techniques to SE problems but also the application of SE techniques to AI problems. RAISE not only strengthens the AI- and-SE community but also continues to develop a roadmap of strategic research directions for AI and SE. Burak Turhan, Ayse Basar Bener, Rachel Harrison, Andriy V. Miranskyy, Çetin Meriçli, Leandro L. Minku |
ICSE (2) | 6 |
| 2015 | An empirical evaluation of ensemble adjustment methods for analogy-based effort estimation
Mohammad Azzeh, Ali Bou Nassif, Leandro L. Minku |
J. Syst. Softw. | 3 |
| 2015 | Resampling-Based Ensemble Methods for Online Class Imbalance LearningabstractOnline class imbalance learning is a new learning problem that combines the challenges of both online learning and class imbalance learning. It deals with data streams having very skewed class distributions. This type of problems commonly exists in real-world applications, such as fault diagnosis of real-time control monitoring systems and intrusion detection in computer networks. In our earlier work, we defined class imbalance online, and proposed two learning algorithms OOB and UOB that build an ensemble model overcoming class imbalance in real time through resampling and time-decayed metrics. In this paper, we further improve the resampling strategy inside OOB and UOB, and look into their performance in both static and dynamic data streams. We give the first comprehensive analysis of class imbalance in data streams, in terms of data distributions, imbalance rates and changes in class imbalance status. We find that UOB is better at recognizing minority-class examples in static data streams, and OOB is more robust against dynamic changes in class imbalance status. The data distribution is a major factor affecting their performance. Based on the insight gained, we then propose two new ensemble methods that maintain both OOB and UOB with adaptive weights for final predictions, called WEOB1 and WEOB2. They are shown to possess the strength of OOB and UOB with good accuracy and robustness. Shuo Wang 0005, Leandro L. Minku, Xin Yao 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2014 | How to make best use of cross-company data in software effort estimation?abstractPrevious works using Cross-Company (CC) data for making Within-Company (WC) Software Effort Estimation (SEE) try to use CC data or models directly to provide predictions in the WC context. So, these data or models are only helpful when they match the WC context well. When they do not, a fair amount of WC training data, which are usually expensive to acquire, are still necessary to achieve good performance. We investigate how to make best use of CC data, so that we can reduce the amount of WC data while maintaining or improving performance in comparison to WC SEE models. This is done by proposing a new framework to learn the relationship between CC and WC projects explicitly, allowing CC models to be mapped to the WC context. Such mapped models can be useful even when the CC models themselves do not match the WC context directly. Our study shows that a new approach instantiating this framework is able not only to use substantially less WC data than a corresponding WC model, but also to achieve similar/better performance. This approach can also be used to provide insight into the behaviour of a company in comparison to others. Leandro L. Minku, Xin Yao 0001 |
ICSE | 1 |
| 2014 | Measuring Energy Consumption for Web Service Product ConfigurationabstractBecause of the economies of scale that Cloud provides, there is great interest in hosting web services on the Cloud. Web services are created from components such as Database Management Systems and HTTP servers. There is a wide variety of components that can be used to configure a web service. The choice of components influences the performance and energy consumption. Most current research in the web service technologies focuses on system performance, and only small number of researchers give attention to energy consumption. In this paper, we propose a method to select the web service configurations which reduce energy consumption. Our method has capabilities to manage feature configuration and predict energy consumption of web service systems. To validate, we developed a technique to measure energy consumption of several web service configurations running in a Virtualized environment. Our approach allows Cloud companies to provide choices of web service technology that consumes less energy. I. Made Murwantara, Behzad Bordbar, Leandro L. Minku |
iiWAS | 3 |
| 2014 | A multi-objective ensemble method for online class imbalance learningabstractOnline class imbalance learning is an emerging learning area that combines the challenges of both online learning and class imbalance learning. In addition to the learning difficulty from the imbalanced distribution, another major challenge is that the imbalanced rate in a data stream can be dynamically changing. OOB and UOB are two state-of-the-art methods for online class imbalance problems [1]. UOB is better at recognizing minority-class examples when the imbalance rate does not change much over time, while OOB is more prepared for the case with a dynamic rate. Aiming for an effective method for both static and dynamic cases, this paper proposes a multi-objective ensemble method MOSOB that combines OOB and UOB. MOSOB finds the Pareto-optimal weights for OOB and UOB at each time step, to maximize minority-class recall and majority-class recall simultaneously. Experiments on five real-world data applications show that MOSOB performs well in both static and dynamic data streams. Furthermore, we look into its performance on a group of highly imbalanced data streams. To respond to the minority class within 10000 time steps, the imbalance rate can be as low as 0.1% for easy data streams; at least 3% of imbalance rate is required to classify difficult data streams. Shuo Wang 0005, Leandro L. Minku, Xin Yao 0001 |
IJCNN | 2 |
| 2014 | Less is More: Temporal Fault Predictive Performance over Multiple Hadoop Releases
Mark Harman, Syed S. Islam, Yue Jia 0001, Leandro L. Minku, Federica Sarro, Komsan Srivisut |
SSBSE | 4 |
| 2014 | Improved Evolutionary Algorithm Design for the Project Scheduling Problem Based on Runtime AnalysisabstractSeveral variants of evolutionary algorithms (EAs) have been applied to solve the project scheduling problem (PSP), yet their performance highly depends on design choices for the EA. It is still unclear how and why different EAs perform differently. We present the first runtime analysis for the PSP, gaining insights into the performance of EAs on the PSP in general, and on specific instance classes that are easy or hard. Our theoretical analysis has practical implications-based on it, we derive an improved EA design. This includes normalizing employees' dedication for different tasks to ensure they are not working overtime; a fitness function that requires fewer pre-defined parameters and provides a clear gradient towards feasible solutions; and an improved representation and mutation operator. Both our theoretical and empirical results show that our design is very effective. Combining the use of normalization to a population gave the best results in our experiments, and normalization was a key component for the practical effectiveness of the new design. Not only does our paper offer a new and effective algorithm for the PSP, it also provides a rigorous theoretical analysis to explain the efficiency of the algorithm, especially for increasingly large projects. Leandro L. Minku, Dirk Sudholt, Xin Yao 0001 |
IEEE Trans. Software Eng. | 1 |
| 2013 | Data science for software engineeringabstractTarget audience: Software practitioners and researchers wanting to understand the state of the art in using data science for software engineering (SE). Content: In the age of big data, data science (the knowledge of deriving meaningful outcomes from data) is an essential skill that should be equipped by software engineers. It can be used to predict useful information on new projects based on completed projects. This tutorial offers core insights about the state-of-the-art in this important field. What participants will learn: Before data science: this tutorial discusses the tasks needed to deploy machine-learning algorithms to organizations (Part 1: Organization Issues). During data science: from discretization to clustering to dichotomization and statistical analysis. And the rest: When local data is scarce, we show how to adapt data from other organizations to local problems. When privacy concerns block access, we show how to privatize data while still being able to mine it. When working with data of dubious quality, we show how to prune spurious information. When data or models seem too complex, we show how to simplify data mining results. When data is too scarce to support intricate models, we show methods for generating predictions. When the world changes, and old models need to be updated, we show how to handle those updates. When the effect is too complex for one model, we show how to reason across ensembles of models. Pre-requisites: This tutorial makes minimal use of maths of advanced algorithms and would be understandable by developers and technical managers. Tim Menzies, Ekrem Kocaguneli, Fayola Peters, Burak Turhan, Leandro L. Minku |
ICSE | 5 |
| 2013 | Concept drift detection for online class imbalance learningabstractConcept drift detection methods are crucial components of many online learning approaches. Accurate drift detections allow prompt reaction to drifts and help to maintain high performance of online models over time. Although many methods have been proposed, no attention has been given to data streams with imbalanced class distributions, which commonly exist in real-world applications, such as fault diagnosis of control systems and intrusion detection in computer networks. This paper studies the concept drift problem for online class imbalance learning. We look into the impact of concept drift on single-class performance of online models based on three types of classifiers, under seven different scenarios with the presence of class imbalance. The analysis reveals that detecting drift in imbalanced data streams is a more difficult task than in balanced ones. Minority-class recall suffers from a significant drop after the drift involving the minority class. Overall accuracy is not suitable for drift detection. Based on the findings, we propose a new detection method DDM-OCI derived from the existing method DDM. DDM-OCI monitors minority-class recall online to capture the drift. The results show a quick response of the online model working with DDM-OCI to the new concept. Shuo Wang 0005, Leandro L. Minku, Davide Ghezzi, Daniele Caltabiano, Peter Tiño, Xin Yao 0001 |
IJCNN | 2 |
| 2013 | Online Class Imbalance Learning and its Applications in Fault DetectionabstractAlthough class imbalance learning and online learning have been extensively studied in the literature separately, online class imbalance learning that considers the challenges of both fields has not drawn much attention. It deals with data streams having very skewed class distributions, such as fault diagnosis of real-time control monitoring systems and intrusion detection in computer networks. To fill in this research gap and contribute to a wide range of real-world applications, this paper first formulates online class imbalance learning problems. Based on the problem formulation, a new online learning algorithm, sampling-based online bagging (SOB), is proposed to tackle class imbalance adaptively. Then, we study how SOB and other state-of-the-art methods can benefit a class of fault detection data under various scenarios and analyze their performance in depth. Through extensive experiments, we find that SOB can balance the performance between classes very well across different data domains and produce stable G-mean when learning constantly imbalanced data streams, but it is sensitive to sudden changes in class imbalance, in which case SOB's predecessor undersampling-based online bagging (UOB) is more robust. Shuo Wang 0005, Leandro L. Minku, Xin Yao 0001 |
Int. J. Comput. Intell. Appl. | 2 |
| 2013 | Ensembles and locality: Insight on improving software effort estimationabstractEnsembles of learning machines and locality are considered two important topics for the next research frontier on Software Effort Estimation (SEE). We aim at (1) evaluating whether existing automated ensembles of learning machines generally improve SEEs given by single learning machines and which of them would be more useful; (2) analysing the adequacy of different locality approaches; and getting insight on (3) how to improve SEE and (4) how to evaluate/choose machine learning (ML) models for SEE. A principled experimental framework is used for the analysis and to provide insights that are not based simply on intuition or speculation. A comprehensive experimental study of several automated ensembles, single learning machines and locality approaches, which present features potentially beneficial for SEE, is performed. Additionally, an analysis of feature selection and regression trees (RTs), and an investigation of two tailored forms of combining ensembles and locality are performed to provide further insight on improving SEE. Bagging ensembles of RTs show to perform well, being highly ranked in terms of performance across different data sets, being frequently among the best approaches for each data set and rarely performing considerably worse than the best approach for any data set. They are recommended over other learning machines should an organisation have no resources to perform experiments to chose a model. Even though RTs have been shown to be more reliable locality approaches, other approaches such as k-Means and k-Nearest Neighbours can also perform well, in particular for more heterogeneous data sets. Combining the power of automated ensembles and locality can lead to competitive results in SEE. By analysing such approaches, we provide several insights that can be used by future research in the area. Leandro L. Minku, Xin Yao 0001 |
Inf. Softw. Technol. | 1 |
| 2013 | Software effort estimation as a multiobjective learning problemabstractEnsembles of learning machines are promising for software effort estimation (SEE), but need to be tailored for this task to have their potential exploited. A key issue when creating ensembles is to produce diverse and accurate base models. Depending on how differently different performance measures behave for SEE, they could be used as a natural way of creating SEE ensembles. We propose to view SEE model creation as a multiobjective learning problem. A multiobjective evolutionary algorithm (MOEA) is used to better understand the tradeoff among different performance measures by creating SEE models through the simultaneous optimisation of these measures. We show that the performance measures behave very differently, presenting sometimes even opposite trends. They are then used as a source of diversity for creating SEE ensembles. A good tradeoff among different measures can be obtained by using an ensemble of MOEA solutions. This ensemble performs similarly or better than a model that does not consider these measures explicitly. Besides, MOEA is also flexible, allowing emphasis of a particular measure if desired. In conclusion, MOEA can be used to better understand the relationship among performance measures and has shown to be very effective in creating SEE models. Leandro L. Minku, Xin Yao 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2012 | Evolutionary algorithms for the project scheduling problem: runtime analysis and improved designabstractEven though genetic algorithms (GAs) have been used for solving the project scheduling problem (PSP), it is not well understood which problem characteristics make it difficult/easy for GAs. We present the first runtime analysis for the PSP, revealing what problem features can make PSP easy or hard. This allows to assess the performance of GAs and to make informed design choices. Our theory has inspired a new evolutionary design, including normalisation of employees' dedication for different tasks to eliminate the problem of exceeding their maximum dedication. Theoretical and empirical results show that our design is very effective in terms of hit rate and solution quality. Leandro L. Minku, Dirk Sudholt, Xin Yao 0001 |
GECCO | 1 |
| 2012 | Using unreliable data for creating more reliable online learnersabstractSome machine learning applications involve the question of whether or not to use unreliable data for the learning. Previous work shows that learners trained using unreliable data in addition to reliable data present either similar or worse performance than learners trained solely on reliable data. Such learners frequently use unreliable data as if they were reliable and consider only the offline learning scenario. The present paper shows that it is possible to use unreliable data to improve the performance in online learning scenarios with a pre-existing set of unreliable data. We propose an approach called Dynamic Un+Reliable data learners (DUR) able to determine when unreliable data could be useful by maintaining a fixed size weighted memory of unreliable data learners. The weights represent how well learners perform for the current concept and are updated throughout DUR's lifetime. This approach manages not only to outperform an approach which uses only reliable data, but also an approach which uses unreliable data as if they were reliable. Moreover, the variance in performance is reduced in comparison to the approach which uses only reliable data. In other words, DUR is a more reliable learner. Leandro L. Minku, Xin Yao 0001 |
IJCNN | 1 |
| 2012 | DDD: A New Ensemble Approach for Dealing with Concept DriftabstractOnline learning algorithms often have to operate in the presence of concept drifts. A recent study revealed that different diversity levels in an ensemble of learning machines are required in order to maintain high generalization on both old and new concepts. Inspired by this study and based on a further study of diversity with different strategies to deal with drifts, we propose a new online ensemble learning approach called Diversity for Dealing with Drifts (DDD). DDD maintains ensembles with different diversity levels and is able to attain better accuracy than other approaches. Furthermore, it is very robust, outperforming other drift handling approaches in terms of accuracy when there are false positive drift detections. In all the experimental comparisons we have carried out, DDD always performed at least as well as other drift handling approaches under various conditions, with very few exceptions. Leandro L. Minku, Xin Yao 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2010 | Design of Experiments in Neuro-Fuzzy SystemsabstractInterest in hybrid methods that combine artificial neural networks and fuzzy inference systems has grown in recent years. These systems are robust solutions that search for representations of domain knowledge, reasoning on uncertainty, automatic learning and adaptation. However, the design and definition of the parameter effectiveness of such systems is still a hard task. In the present work, we perform a statistical analysis to verify interactions and interrelations between parameters in the design of neuro-fuzzy systems. The analysis is carried out using a powerful statistical tool, namely, Design of Experiments (DOE), in two neuro-fuzzy models — Adaptive Neuro Fuzzy Inference System (ANFIS) and Evolving Fuzzy Neural Networks (EFuNN). The results show that, for ANFIS, input MFs number and output MFs shape are usually the factors with the largest influence on the system's RMSE. For EFFuNN, the MF shape and the interaction between MF shape and number usually have the largest effect size. Cleber Zanchettin, Leandro L. Minku, Teresa Bernarda Ludermir |
Int. J. Comput. Intell. Appl. | 2 |
| 2010 | The Impact of Diversity on Online Ensemble Learning in the Presence of Concept DriftabstractOnline learning algorithms often have to operate in the presence of concept drift (i.e., the concepts to be learned can change with time). This paper presents a new categorization for concept drift, separating drifts according to different criteria into mutually exclusive and nonheterogeneous categories. Moreover, although ensembles of learning machines have been used to learn in the presence of concept drift, there has been no deep study of why they can be helpful for that and which of their features can contribute or not for that. As diversity is one of these features, we present a diversity analysis in the presence of different types of drifts. We show that, before the drift, ensembles with less diversity obtain lower test errors. On the other hand, it is a good strategy to maintain highly diverse ensembles to obtain lower test errors shortly after the drift independent on the type of drift, even though high diversity is more important for more severe drifts. Longer after the drift, high diversity becomes less important. Diversity by itself can help to reduce the initial increase in error caused by a drift, but does not provide the faster recovery from drifts in long-term. Leandro L. Minku, Allan P. White, Xin Yao 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |