EDBT 2026 Demo / reviewers in the wild / expert
Alberto Cano 0001
dblp:62/8229-1 · also Alberto Cano Rojas
· DBLP profile ↗
63ranked-venue papers
22as first author
22since 2021 · last 2026
0000-0001-9027-298XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 15 first-author · 16 since 2021Databases, data management, data science and information retrieval · 14 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 4 since 2021Systems, architecture and hardware · 4 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing classification on multi-class drifting data streams with one-vs-rest strategies
Gabriel Aguiar, Alberto Cano 0001, Juan Valentín Guerrero Cano |
Knowl. Inf. Syst. | 2 |
| 2026 | Anticipating to Change: A Proactive Approach for Concept Drift Adaptation in Data StreamsabstractAbstract Adapting to drifting data streams remains a key challenge in online learning, where effective model adaptation depends on timely concept drift detection. Most existing approaches respond to drift only after distributional changes occur, reacting to concept drift, limiting their ability to prevent the classifier’s performance degradation. This work introduces a novel methodology to anticipate concept drift and enable proactive adaptation before the data distribution shift negatively impacts the classifier. We propose four proactive adaptation strategies based on the Very Fast Decision Tree (VFDT) algorithm to leverage data trends to estimate proactive changes in the classifier, mitigating or even preventing the performance degradation consequence of the concept drift. We evaluate the proposed methods across four scenarios with diverse data stream configurations. Results demonstrate that proactive adaptation reduces the adverse effects of concept drift and improves classification performance. In particular, the proposed strategies consistently outperformed in settings with incremental drift, underscoring the potential of anticipatory approaches and addressing a notable gap in the current literature. Juan Valentín Guerrero Cano, Gabriel Aguiar, Alberto Cano 0001 |
Mach. Learn. | 3 |
| 2025 | Online Hierarchical Partitioning of the Output Space in Extreme Multi-Label Data StreamsabstractMining data streams with multi-label outputs poses significant challenges due to evolving distributions, high-dimensional label spaces, sparse label occurrences, and complex label dependencies. Moreover, concept drift affects not only input distributions but also label correlations and imbalance ratios over time, complicating model adaptation. To address these challenges, structured learners are categorized into local and global methods. Local methods break down the task into simpler components, while global methods adapt the algorithm to the full output space, potentially yielding better predictions by exploiting label correlations. This work introduces iHOMER (Incremental Hierarchy Of Multi-label Classifiers), an online multi-label learning framework that incrementally partitions the label space into disjoint, correlated clusters without relying on predefined hierarchies. iHOMER leverages online divisive-agglomerative clustering based on Jaccard similarity and a global tree-based learner driven by a multivariate Bernoulli process to guide instance partitioning. To address non-stationarity, it integrates drift detection mechanisms at both global and local levels, enabling dynamic restructuring of label partitions and subtrees. Experiments across 23 real-world datasets show iHOMER outperforms 5 state-of-the-art global baselines, such as MLHAT, MLHT of Pruned Sets and iSOUPT, by 23%, and 12 local baselines, such as binary relevance transformations of kNN, EFDT, ARF, and ADWIN bagging/boosting ensembles, by 32%, establishing its robustness for online multi-label classification. Lara Neves, Afonso Lourenço, Alberto Cano 0001, Goreti Marreiros |
ECAI | 3 |
| 2025 | Shared Knowledge Base for Multi Deep Learning in Defect DetectionabstractIn recent years, there has been growing interest in applying deep learning techniques for visual anomaly detection, particularly in the manufacturing sector. Various models have been developed to identify defects in manufacturing data, yet selecting and optimizing these models for anomaly detection in intelligent manufacturing environments remains a significant challenge. This research focuses on general-purpose visual anomaly detection, aiming to reduce dependence on domain-specific knowledge and create flexible, generic models. We propose a novel deep learning framework in which multiple models are trained for each image. The visual features and loss values from these models are computed and stored during training. During the testing phase, this stored information is used to select the most appropriate model for each new image using a k-Nearest Neighbors (kNN) approach. The proposed method, KGDL-VAD (Knowledge-Guided Deep Learning for Visual Anomaly Detection), was evaluated on the MVTec AD, and standard aerospace defect detection datasets, achieving an area under the curve (AUC) score of 0.96, outperforming baseline methods. In addition, KGDL-VAD surpasses ensemble learning approaches across multiple domain-independent datasets with varying numbers of trained classes. Youcef Djenouri, Asma Belhadi, Gautam Srivastava 0001, Ahmed Nabil Belbachir, Alberto Cano 0001 |
IJCNN | 5 |
| 2025 | Simultaneous fault prediction in evolving industrial environments with ensembles of Hoeffding adaptive treesabstractAbstract Predictive Maintenance (PdM) emerges as a critical task of Industry 4.0, driving operational efficiency, minimizing downtime, and reducing maintenance costs. However, real-world industrial environments present unsolved challenges, especially in predicting simultaneous and correlated faults under evolving conditions. Traditional batch-based and deep learning approaches for simultaneous fault prediction often fall short due to their assumptions of static data distributions and high computational demands, making them unsuitable for dynamic, resource-constrained systems. In response, we propose OEMLHAT (Online Ensemble of Multi-Label Hoeffding Adaptive Trees), a novel model tailored for real-time, multi-label fault prediction in non-stationary industrial settings. OEMLHAT introduces a scalable online ensemble architecture that integrates online bagging, dynamic feature subspacing, and adaptive output weighting. This design allows it to efficiently handle concept drift, high-dimensional input spaces, and label sparsity, key bottlenecks in existing PdM solutions. Experimental results on three public multi-label PdM case studies demonstrate substantial improvements in predictive performance of OEMLHAT over previous batch-based and online proposals for multi-label classification, particularly with an average improvement in micro-averaged F1-score of 18.49% over the second most-accurate batch-based proposal and of 8.56% in the case of the second best online model. By addressing a critical gap in online multi-label learning for PdM, this work provides a robust and interpretable solution for next-generation industrial monitoring systems for fault detection, particularly for rare and concurrent failures. Aurora Esteban, Alberto Cano 0001, Sebastián Ventura, Amelia Zafra |
Appl. Intell. | 2 |
| 2024 | Easing the Prediction of Student Dropout for everyone integration AutoML and Explainable Artificial Intelligence
Pamela Alexandra Buñay-Guisñan, Juan Alfonso Lara, Alberto Cano 0001, Rebeca Cerezo, Cristóbal Romero 0001 |
EDM | 3 |
| 2024 | Vision-based Spatiotemporal Learning for Human Activity RecognitionabstractThis paper introduces a novel concept for Human Activity Recognition (HAR) that allows robust analysis, classification, and understanding of human movements in various environments. It can be applied in various applications such as Health monitoring and analysis, fitness/dance training and performance analysis, interactive gaming, smart homes, and wearable devices. The novel method, coined as STL-HAR (SpatioTemporal Learning for HAR), learns from sensor data jointly represented in space and time to robustify the HAR process. In the new concept, we propose a hybrid model based on GNN (Graph Neural Network), and LSTM (Long Short-Term Memory). GNN first learns the spatial features from different sensor data locations. The learned features will then be injected to LSTM where the temporal information is captured by observing sensor status at different timestamps. We evaluate and analyze the performance of STL-HAR in real use case scenarios of HAR data compared with baseline HAR-based solutions. STL-HAR has achieved a recognition rate of 92% under different scenarios. Youcef Djenouri, Ahmed Nabil Belbachir, Gautam Srivastava 0001, Alberto Cano 0001 |
IJCNN | 4 |
| 2024 | Improved KD-tree based imbalanced big data classification and oversampling for MapReduce platforms
William C. Sleeman IV, Martha I. Roseberry, Preetam Ghosh, Alberto Cano 0001, Bartosz Krawczyk |
Appl. Intell. | 4 |
| 2024 | Dynamic budget allocation for sparsely labeled drifting data streams
Gabriel Aguiar, Alberto Cano 0001 |
Inf. Sci. | 2 |
| 2024 | A comprehensive analysis of concept drift locality in data streams
Gabriel Aguiar, Alberto Cano 0001 |
Knowl. Based Syst. | 2 |
| 2024 | Hoeffding adaptive trees for multi-label classification on data streams
Aurora Esteban, Alberto Cano 0001, Amelia Zafra, Sebastián Ventura |
Knowl. Based Syst. | 2 |
| 2024 | A survey on learning from imbalanced data streams: taxonomy, challenges, empirical study, and reproducible experimental framework
Gabriel Aguiar, Bartosz Krawczyk, Alberto Cano 0001 |
Mach. Learn. | 3 |
| 2023 | Enhancing Concept Drift Detection in Drifting and Imbalanced Data Streams through Meta-LearningabstractOne of the biggest challenges in learning from data streams is adapting the classification model to new data. Due to the evolving nature of data streams, they are subject to a phenomenon known as concept drift that makes previously learned knowledge and model outdated. Therefore, concept drift must be efficiently detected in order to adapt the classification model. While there exists a plethora of drift detectors, with different mechanisms, selecting the most suitable for a new stream is a difficult task, since apriori knowledge may not be available and changes over time can affect the performance of the detector. This paper proposes a framework that exploits statistical and temporal meta-features from sliding windows to dynamically recommend a suitable drift detector in real-time for unseen chunks of streams according to its properties using Meta-Learning. We performed experiments on 10 real-world data streams and 18 synthetic generated data streams that were subject to concept drift and class imbalance in order to evaluate the performance of the proposed framework. Experiments exposed that the proposed approach was able to enhance the concept drift detection in a variety of scenarios demonstrating robustness to class imbalance and the advantages of dynamically selecting the drift detector. Gabriel Aguiar, Alberto Cano 0001 |
IEEE Big Data | 2 |
| 2023 | Meta-learning for dynamic tuning of active learning on stream classificationabstractSupervised data stream learning depends on the incoming sample’s true label to update a classifier’s model. In real life, obtaining the ground truth for each instance is a challenging process; it is highly costly and time consuming. Active Learning has already bridged this gap by finding a reduced set of instances to support the creation of a reliable stream classifier. However, identifying a reduced number of informative instances to support a suitable classifier update and drift adaptation is very tricky. To better adapt to concept drifts using a reduced number of samples, we propose an online tuning of the Uncertainty Sampling threshold using a meta-learning approach. Our approach exploits statistical meta-features from adaptive windows to meta-recommend a suitable threshold to address the trade-off between the number of labelling queries and high accuracy. Experiments exposed that the proposed approach provides the best trade-off between accuracy and query reduction by dynamic tuning the uncertainty threshold using lightweight meta-features. Vinicius Eiji Martins, Alberto Cano 0001, Sylvio Barbon Junior |
Pattern Recognit. | 2 |
| 2022 | An Explainable Classifier based on Genetically Evolved Graph StructuresabstractTrusting an algorithmic decision is much easier if it is understood how it was achieved. Therefore, data mining algorithms with explainable abilities are preferred over complex models for critical applications. Rule-based algorithms are among the easiest data classification models to understand. However, as most interpretable models, rule-based algorithms do not provide the highest accuracy. The Attribute-based Decision Graph (AbDG) is a model to represent a labeled data set as a weighted graph over the attributes. When associated with a proper algorithm, AbDGs can be used for supervised data mining tasks. An important aspect of AbDGs is that the graph encompasses the original attribute values and their relationships, which makes it easily interpretable by extracting rules. The AbDG drawback is defining a suitable graph structure for a given data set. In previous works, the authors proposed GA-AbDG, a genetic algorithm to search for an AbDG by evolving its edge set, keeping the vertex set fixed. In this paper, we propose an evolutionary algorithm and genetic operators to evolve both, vertex and edge sets, enhancing the search space of possible AbDG structures. Moreover, we associate a rule-based classifier with the AbDG to achieve explainable results. Experimental results show the proposed method outperforms GA-AbDG and five other classical interpretable classification algorithms. João Roberto Bertini Jr., Alberto Cano 0001 |
CEC | 2 |
| 2022 | An ontology matching approach for semantic modeling: A case study in smart citiesabstractAbstract This paper investigates the semantic modeling of smart cities and proposes two ontology matching frameworks, called Clustering for Ontology Matching‐based Instances (COMI) and Pattern mining for Ontology Matching‐based Instances (POMI). The goal is to discover the relevant knowledge by investigating the correlations among smart city data based on clustering and pattern mining approaches. The COMI method first groups the highly correlated ontologies of smart‐city data into similar clusters using the generic k‐means algorithm. The key idea of this method is that it clusters the instances of each ontology and then matches two ontologies by matching their clusters and the corresponding instances within the clusters. The POMI method studies the correlations among the data properties and selects the most relevant properties for the ontology matching process. To demonstrate the usefulness and accuracy of the COMI and POMI frameworks, several experiments on the DBpedia, Ontology Alignment Evaluation Initiative, and NOAA ontology databases were conducted. The results show that COMI and POMI outperform the state‐of‐the‐art ontology matching models regarding computational cost without losing the quality during the matching process. Furthermore, these results confirm the ability of COMI and POMI to deal with heterogeneous large‐scale data in smart‐city environments. Youcef Djenouri, Hiba Belhadi, Karima Akli-Astouati, Alberto Cano 0001, Jerry Chun-Wei Lin |
Comput. Intell. | 4 |
| 2022 | Adaptive ensemble of self-adjusting nearest neighbor subspaces for multi-label drifting data streams
Gavin Alberghini, Sylvio Barbon Junior, Alberto Cano 0001 |
Neurocomputing | 3 |
| 2022 | ROSE: robust online self-adjusting ensemble for continual learning on imbalanced drifting data streams
Alberto Cano 0001, Bartosz Krawczyk |
Mach. Learn. | 1 |
| 2022 | Hybrid Group Anomaly Detection for Sequence Data: Application to Trajectory Data AnalyticsabstractMany research areas depend on group anomaly detection. The use of group anomaly detection can maintain and provide security and privacy to the data involved. This research attempts to solve the deficiency of the existing literature in outlier detection thus a novel hybrid framework to identify group anomaly detection from sequence data is proposed in this paper. It proposes two approaches for efficiently solving this problem: i)Hybrid Data Mining-based algorithm, consists of three main phases: first, the clustering algorithm is applied to derive the micro-clusters. Second, the$kNN$algorithm is applied to each micro-cluster to calculate the candidates of the group’s outliers. Third, a pattern mining framework gets applied to the candidates of the group’s outliers as a pruning strategy, to generate the groups of outliers, and ii) aGPU-basedapproach is presented, which benefits from the massively GPU computing to boost the runtime of the hybrid data mining-based algorithm. Extensive experiments were conducted to show the advantages of different sequence databases of our proposed model. Results clearly show the efficiency of a GPU direction when directly compared to a sequential approach by reaching a speedup of451. In addition, both approaches outperform the baseline methods for group detection. Asma Belhadi, Youcef Djenouri, Gautam Srivastava 0001, Alberto Cano 0001, Jerry Chun-Wei Lin |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Locally Linear Support Vector Machines for Imbalanced Data Classification
Bartosz Krawczyk, Alberto Cano 0001 |
PAKDD (1) | 2 |
| 2021 | Self-adjusting k nearest neighbors for continual learning from multi-label drifting data streams
Martha I. Roseberry, Bartosz Krawczyk, Youcef Djenouri, Alberto Cano 0001 |
Neurocomputing | 4 |
| 2021 | A Two-Phase Anomaly Detection Model for Secure Intelligent Transportation Ride-Hailing TrajectoriesabstractThis paper addresses the taxi fraud problem and introduces a new solution to identify trajectory outliers. The approach as presented allows to identify both individual and group outliers and is based on a two phase-based algorithm. The first phase determines the individual trajectory outliers by computing the distance of each point in each trajectory, whereas the second identifies the group trajectory outliers by exploring the individual trajectory outliers using both feature selection and sliding windows strategies. A parallel version of the algorithm is also proposed using a sliding window-based GPU approach to boost the runtime performance. Extensive experiments have been carried out to thoroughly demonstrate the usefulness of our methodology on both synthetic and real trajectory databases. The results show that the GPU approach enables reaching a speed-up of 341 over the sequential algorithm on large synthetic databases. The efficiency of the proposed method to detect both individual and group trajectory outliers on a real-world taxi trajectory database is also demonstrated in comparison with baseline trajectory outlier and group detection algorithms. The results are very promising and show superiority of the proposed method both in reducing computational time and enhancing the quality of returned outliers. Finally, we prime our methodology and results for future refinement using deep learning methodologies. Asma Belhadi, Youcef Djenouri, Gautam Srivastava 0001, Djamel Djenouri, Alberto Cano 0001, Jerry Chun-Wei Lin |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2020 | A general-purpose distributed pattern mining systemabstractAbstract This paper explores five pattern mining problems and proposes a new distributed framework called DT-DPM: Decomposition Transaction for Distributed Pattern Mining. DT-DPM addresses the limitations of the existing pattern mining problems by reducing the enumeration search space. Thus, it derives the relevant patterns by studying the different correlation among the transactions. It first decomposes the set of transactions into several clusters of different sizes, and then explores heterogeneous architectures, including MapReduce, single CPU, and multi CPU, based on the densities of each subset of transactions. To evaluate the DT-DPM framework, extensive experiments were carried out by solving five pattern mining problems (FIM: Frequent Itemset Mining, WIM: Weighted Itemset Mining, UIM: Uncertain Itemset Mining, HUIM: High Utility Itemset Mining, and SPM: Sequential Pattern Mining). Experimental results reveal that by using DT-DPM, the scalability of the pattern mining algorithms was improved on large databases. Results also reveal that DT-DPM outperforms the baseline parallel pattern mining algorithms on big databases. Asma Belhadi, Youcef Djenouri, Jerry Chun-Wei Lin, Alberto Cano 0001 |
Appl. Intell. | 4 |
| 2020 | Distributed multi-label feature selection using individual mutual information measures
Jorge Gonzalez-Lopez, Sebastián Ventura, Alberto Cano 0001 |
Knowl. Based Syst. | 3 |
| 2020 | Kappa Updated Ensemble for drifting data stream mining
Alberto Cano 0001, Bartosz Krawczyk |
Mach. Learn. | 1 |
| 2020 | Blocking Self-Avoiding Walks Stops Cyber-Epidemics: A Scalable GPU-Based ApproachabstractCyber-epidemics, the widespread of fake news or propaganda through social media, can cause devastating economic and political consequences. A common countermeasure against cyber-epidemics is to disable a small subset of suspected social connections or accounts to effectively contain the epidemics. An example is the recent shutdown of 125,000 ISIS-related Twitter accounts. Despite many proposed methods to identify such a subset, none are scalable enough to provide high-quality solutions in nowadays' billion-size networks. To this end, we investigate the Spread Interdiction problems that seek the most effective links (or nodes) for removal under the well-known Linear Threshold model. We propose novel CPU-GPU methods that scale to networks with billions of edges, yet possess rigorous theoretical guarantee on the solution quality. At the core of our methods is an O(1)-space out-of-core algorithm to generate a new type of random walks, called Hitting Self-avoiding Walks (HSAWs). Such a low memory requirement enables handling of big networks and, more importantly, hiding latency via scheduling of millions of threads on GPUs. Comprehensive experiments on real-world networks show that our algorithms provide much higher quality solutions and are several orders of magnitude faster than the state-of-the art. Comparing to the (single-core) CPU counterpart, our GPU implementations achieve significant speedup factors up to 177× on a single GPU and 338× on a GPU pair. Hung T. Nguyen 0003, Alberto Cano 0001, Tam Vu 0001, Thang N. Dinh |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | Distributed Selection of Continuous Features in Multilabel Classification Using Mutual InformationabstractMultilabel learning is a challenging task demanding scalable methods for large-scale data. Feature selection has shown to improve multilabel accuracy while defying the curse of dimensionality of high-dimensional scattered data. However, the increasing complexity of multilabel feature selection, especially on continuous features, requires new approaches to manage data effectively and efficiently in distributed computing environments. This article proposes a distributed model for mutual information (MI) adaptation on continuous features and multiple labels on Apache Spark. Two approaches are presented based on MI maximization, and minimum redundancy and maximum relevance. The former selects the subset of features that maximize the MI between the features and the labels, whereas the latter additionally minimizes the redundancy between the features. Experiments compare the distributed multilabel feature selection methods on 10 data sets and 12 metrics. Results validated through statistical analysis indicate that our methods outperform reference methods for distributed feature selection for multilabel data, while MIM also reduces the runtime in orders of magnitude. Jorge Gonzalez-Lopez, Sebastián Ventura, Alberto Cano 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Active Learning with Abstaining Classifiers for Imbalanced Drifting Data StreamsabstractLearning from data streams is one of the most promising and challenging domains in modern machine learning. Proliferating online data sources provide us access to real-time knowledge we have never had before. At the same time, new obstacles emerge and we have to overcome them in order to fully and effectively utilize the potential of the data. Prohibitive time and memory constraints or non-stationary distributions are only some of the problems. When dealing with classification tasks, one has to remember that effective adaptation has to be achieved on weak foundations of partially labeled and often imbalanced data. In our work, we propose an online framework for binary classification, that aims to handle the complex problem of working with dynamic, sparsely labeled and imbalanced streams. The main part of it is a novel active learning strategy (MD-OAL) that is able to prioritize labeling of minority instances and, as a result, improve the balance of the learning process. We combine the strategy with a dynamic ensemble of base learners that can abstain from making decisions, if they are very uncertain. We adjust the abstaining mechanism in favor of minority instances, providing an effective method for handling remaining imbalance and a concept drift simultaneously. The conducted evaluation shows that in the challenging and realistic scenarios our framework outperforms state-of-the-art algorithms, providing higher resilience to the combined effect of limited labeling and imbalance. Lukasz Korycki, Alberto Cano 0001, Bartosz Krawczyk |
IEEE BigData | 2 |
| 2019 | Adaptive Ensemble Active Learning for Drifting Data Stream MiningabstractLearning from data streams is among the most vital contemporary fields in machine learning and data mining. Streams pose new challenges to learning systems, due to their volume and velocity, as well as ever-changing nature caused by concept drift. Vast majority of works for data streams assume a fully supervised learning scenario, having an unrestricted access to class labels. This assumption does not hold in real-world applications, where obtaining ground truth is costly and time-consuming. Therefore, we need to carefully select which instances should be labeled, as usually we are working under a strict label budget. In this paper, we propose a novel active learning approach based on ensemble algorithms that is capable of using multiple base classifiers during the label query process. It is a plug-in solution, capable of working with most of existing streaming ensemble classifiers. We realize this process as a Multi-Armed Bandit problem, obtaining an efficient and adaptive ensemble active learning procedure by selecting the most competent classifier from the pool for each query. In order to better adapt to concept drifts, we guide our instance selection by measuring the generalization capabilities of our classifiers. This adaptive solution leads not only to better instance selection under sparse access to class labels, but also to improved adaptation to various types of concept drift and increasing the diversity of the underlying ensemble classifier. Bartosz Krawczyk, Alberto Cano 0001 |
IJCAI | 2 |
| 2019 | Speeding Up Classifier Chains in Multi-label ClassificationabstractMulti-label classification has attracted increasing attention of the scientific community in recent years, given its ability to solve problems where each of the examples simultaneously belongs to multiple labels. From all the techniques developed to solve multi-label classification problems, Classifier Chains has been demonstrated to be one of the best performing techniques. However, one of its main drawbacks is its inherently sequential definition. Although many research works aimed to reduce the runtime of multi-label classification algorithms, to the best of our knowledge, there are no proposals to specifically reduce the runtime of Classifier Chains. Therefore, in this paper we propose a method called Parallel Classifier Chains which enables the parallelization of Classifier Chain. In this way, Parallel Classifier Chains builds k binary classifiers in parallel, where each of them includes as extra input features the predictions of those labels that have been previously built. We performed an experimental evaluation over 20 datasets using 5 metrics to analyze both the runtime and the predictive performance of our proposal. The results of the experiments affirmed that our proposal was able to significantly reduce the runtime of Classifier Chains while the predictive performance was not statistically significantly harmed. Jose M. Moyano, Eva Lucrecia Gibaja Galindo, Sebastián Ventura, Alberto Cano 0001 |
IoTBDS | 4 |
| 2019 | Speeding up k-Nearest Neighbors classifier for large-scale multi-label learning on GPUs
Przemyslaw Skryjomski, Bartosz Krawczyk, Alberto Cano 0001 |
Neurocomputing | 3 |
| 2019 | Exploiting GPU and cluster parallelism in single scan frequent itemset mining
Youcef Djenouri, Djamel Djenouri, Asma Belhadi, Alberto Cano 0001 |
Inf. Sci. | 4 |
| 2019 | Evolving rule-based classifiers with genetic programming on GPUs for drifting data streams
Alberto Cano 0001, Bartosz Krawczyk |
Pattern Recognit. | 1 |
| 2019 | Multi-Label Punitive kNN with Self-Adjusting Memory for Drifting Data StreamsabstractIn multi-label learning, data may simultaneously belong to more than one class. When multi-label data arrives as a stream, the challenges associated with multi-label learning are joined by those of data stream mining, including the need for algorithms that are fast and flexible, able to match both the speed and evolving nature of the stream. This article presents a punitive k nearest neighbors algorithm with a self-adjusting memory (MLSAMPkNN) for multi-label, drifting data streams. The memory adjusts in size to contain only the current concept and a novel punitive system identifies and penalizes errant data examples early, removing them from the window. By retaining and using only data that are both current and beneficial, MLSAMPkNN is able to adapt quickly and efficiently to changes within the data stream while still maintaining a low computational complexity. Additionally, the punitive removal mechanism offers increased robustness to various data-level difficulties present in data streams, such as class imbalance and noise. The experimental study compares the proposal to 24 algorithms using 30 real-world and 15 artificial multi-label data streams on six multi-label metrics, evaluation time, and memory consumption. The superior performance of the proposed method is validated through non-parametric statistical analysis, proving both high accuracy and low time complexity. MLSAMPkNN is a versatile classifier, capable of returning excellent performance in diverse stream scenarios. Martha I. Roseberry, Bartosz Krawczyk, Alberto Cano 0001 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2018 | Learning Classification Rules with Differential Evolution for High-Speed Data Stream Mining on GPU sabstractHigh-speed data streams are potentially infinite sequences of rapidly arriving instances that may be subject to concept drift phenomenon. Hence, dedicated learning algorithms must be able to update themselves with new data and provide an accurate prediction in a limited amount of time. This requirement was considered as prohibitive for using evolutionary algorithms for high-speed data stream mining. This paper introduces a massively parallel implementation on GPUs of a differential evolution algorithm for learning classification rules in the presence of concept drift. The proposal based on the DE /rand - to - best/1/bin strategy takes advantage of up to four nested levels of parallelism to maximize the performance of the algorithm. Efficient GPU kernels parallelize the evolution of the populations, rules, conditional clauses, and evaluation on instances. The proposed method is evaluated on 25 data stream benchmarks considering different types of concept drifts. Results are compared with other publicly available streaming rule learners. Obtained results and their statistical analysis proves an excellent performance of the proposed classifier that offers improved predictive accuracy, model update time, decision time, and a compact rule set. Alberto Cano 0001, Bartosz Krawczyk |
CEC | 1 |
| 2018 | Selecting local ensembles for multi-class imbalanced data classificationabstractLearning from imbalanced data is a challenge that machine learning community is facing over last decades, due to its ever-growing presence in real-life problems. While there is a significant number of works addressing the issue of handling binary and skewed datasets, its multi-class counterpart have not received as much attention. This problem is much more difficult, as presence of multiple imbalanced classes can significantly deteriorate the predictive power of any classifier. The relationship among classes are no longer clearly established and there are many difficulties embedded in the nature of such data that needs to be properly addressed. In this work, we discuss the issue of forming effective ensembles for multi-class imbalanced data based on static classifier selection approach. We propose a fully adaptive learning scheme that splits the original feature space into a number of competence areas and modifies their size and location in order to most effectively exploit the supplied pool of base classifiers. Additionally, for each established cluster we perform a weighted classifier combination, where weights are set individually for each cluster and each considered class. This allows for exploiting local competencies of each base learner in given part of feature space, as well as for each of considered classes. These two tasks are combined together in a single hybrid training scheme guided by an evolutionary algorithm. The optimization criterion is formulated in order to achieve skew-insensitive ensemble of local ensembles able to tackle highly imbalanced and multi-class problems. Experimental study proves the high efficacy of the proposed method and its superiority to other ensemble selection methods. Bartosz Krawczyk, Alberto Cano 0001, Michal Wozniak 0001 |
IJCNN | 2 |
| 2018 | Distributed nearest neighbor classification for large-scale multi-label data on spark
Jorge Gonzalez-Lopez, Sebastián Ventura, Alberto Cano 0001 |
Future Gener. Comput. Syst. | 3 |
| 2018 | MIRSVM: Multi-instance support vector machine with bag representatives
Gabriella Melki, Alberto Cano 0001, Sebastián Ventura |
Pattern Recognit. | 2 |
| 2017 | Parsing MetaMap Files in Hadoop
Amy L. Olex, Alberto Cano 0001, Bridget T. McInnes |
AMIA | 2 |
| 2017 | Extremely high-dimensional optimization with MapReduce: Scaling functions and algorithm
Alberto Cano 0001, Carlos García-Martínez, Sebastián Ventura |
Inf. Sci. | 1 |
| 2017 | Multi-target support vector regression via correlation regressor chains
Gabriella Melki, Alberto Cano 0001, Vojislav Kecman, Sebastián Ventura |
Inf. Sci. | 2 |
| 2017 | An ensemble approach to multi-view multi-instance learning
Alberto Cano 0001 |
Knowl. Based Syst. | 1 |
| 2017 | Multi-objective genetic programming for feature extraction and data visualization
Alberto Cano 0001, Sebastián Ventura, Krzysztof J. Cios |
Soft Comput. | 1 |
| 2016 | 100 Million dimensions large-scale global optimization using distributed GPU computingabstractAt this time, many industrial and science problems deal with a large number of decision variables. Classic metaheuristics have shown excellent search abilities on bounded problems, but they often lose their efficacy when applied to large ones. This is known as the curse of dimensionality. To this issue, we have to add the simple fact that the solution evaluation becomes excessively demanding in time. To push the research state forward on this type of problems, the IEEE Congress on Evolutionary Computation regularly organises a competition on large-scale global optimization since 2008. On the other hand, general purpose computing with graphics processing units has become very attractive in the last years, because they may attain very high speed-up ratios on problems with high data parallelism levels. In this work, we study the benefits of exploiting a scalable and distributed computational architecture with multiple GPUs for large scale function optimisation. The study is carried out in terms of 1) evaluation speed-up, 2) quality of the results, and 3) extremely large scale optimisation with real-parameter functions with up to 108variables. Alberto Cano 0001, Carlos García-Martínez |
CEC | 1 |
| 2016 | Early dropout prediction using data mining: a case study with high school studentsabstractAbstract Early prediction of school dropout is a serious problem in education, but it is not an easy issue to resolve. On the one hand, there are many factors that can influence student retention. On the other hand, the traditional classification approach used to solve this problem normally has to be implemented at the end of the course to gather maximum information in order to achieve the highest accuracy. In this paper, we propose a methodology and a specific classification algorithm to discover comprehensible prediction models of student dropout as soon as possible. We used data gathered from 419 high schools students in Mexico. We carried out several experiments to predict dropout at different steps of the course, to select the best indicators of dropout and to compare our proposed algorithm versus some classical and imbalanced well‐known classification algorithms. Results show that our algorithm was capable of predicting student dropout within the first 4–6 weeks of the course and trustworthy enough to be used in an early warning system. Carlos Márquez-Vera, Alberto Cano 0001, Cristóbal Romero 0001, Amin Y. Noaman, Habib Fardoun, Sebastián Ventura |
Expert Syst. J. Knowl. Eng. | 2 |
| 2016 | LAIM discretization for multi-label data
Alberto Cano 0001, José María Luna, Eva Lucrecia Gibaja Galindo, Sebastián Ventura |
Inf. Sci. | 1 |
| 2016 | Discovering useful patterns from multiple instance data
José María Luna, Alberto Cano 0001, Virgilijus Sakalauskas, Sebastián Ventura |
Inf. Sci. | 2 |
| 2016 | ur-CAIM: improved CAIM discretization for unbalanced and balanced data
Alberto Cano 0001, Dat T. Nguyen, Sebastián Ventura, Krzysztof J. Cios |
Soft Comput. | 1 |
| 2016 | Speeding-Up Association Rule Mining With Inverted Index CompressionabstractThe growing interest in data storage has made the data size to be exponentially increased, hampering the process of knowledge discovery from these large volumes of high-dimensional and heterogeneous data. In recent years, many efficient algorithms for mining data associations have been proposed, facing up time and main memory requirements. Nevertheless, this mining process could still become hard when the number of items and records is extremely high. In this paper, the goal is not to propose new efficient algorithms but a new data structure that could be used by a variety of existing algorithms without modifying its original schema. Thus, our aim is to speed up the association rule mining process regardless the algorithm used to this end, enabling the performance of efficient implementations to be enhanced. The structure simplifies, reorganizes, and speeds up the data access by sorting data by means of a shuffling strategy based on the hamming distance, which achieve similar values to be closer, and considering both an inverted index mapping and a run length encoding compression. In the experimental study, we explore the bounds of the algorithms' performance by using a wide number of data sets that comprise either thousands or millions of both items and records. The results demonstrate the utility of the proposed data structure in enhancing the algorithms' runtime orders of magnitude, and substantially reducing both the auxiliary and the main memory requirements. José María Luna, Alberto Cano 0001, Mykola Pechenizkiy, Sebastián Ventura |
IEEE Trans. Cybern. | 2 |
| 2015 | A classification module for genetic programming algorithms in JCLEC
Alberto Cano 0001, José María Luna, Amelia Zafra, Sebastián Ventura |
J. Mach. Learn. Res. | 1 |
| 2015 | Speeding up multiple instance learning classification rules on GPUs
Alberto Cano 0001, Amelia Zafra, Sebastián Ventura |
Knowl. Inf. Syst. | 1 |
| 2014 | GPU-parallel subtree interpreter for genetic programmingabstractGenetic Programming (GP) is a computationally intensive technique but its nature is embarrassingly parallel. Graphic Processing Units (GPUs) are many-core architectures which have been widely employed to speed up the evaluation of GP. In recent years, many works have shown the high performance and efficiency of GPUs on evaluating both the individuals and the fitness cases in parallel. These approaches are known as population parallel and data parallel. This paper presents a parallel GP interpreter which extends these approaches and adds a new parallelization level based on the concurrent evaluation of the individual's subtrees. A GP individual defined by a tree structure with nodes and branches comprises different depth levels in which there are independent subtrees which can be evaluated concurrently. Threads can cooperate to evaluate different subtrees and share the results via GPU's shared memory. The experimental results show the better performance of the proposal in terms of the GP operations per second (GPops/s) that the GP interpreter is capable of processing, achieving up to 21 billion GPops/s using a NVIDIA 480 GPU. However, some issues raised due to limitations of currently available hardware are to be overcomed by the dynamic parallelization capabilities of the next generation of GPUs. Alberto Cano 0001, Sebastián Ventura |
GECCO | 1 |
| 2014 | Parallel evaluation of Pittsburgh rule-based classifiers on GPUs
Alberto Cano 0001, Amelia Zafra, Sebastián Ventura |
Neurocomputing | 1 |
| 2014 | Scalable CAIM discretization on multiple GPUs using concurrent kernels
Alberto Cano 0001, Sebastián Ventura, Krzysztof J. Cios |
J. Supercomput. | 1 |
| 2013 | A Grammar-Guided Genetic Programming Algorithm for Multi-Label Classification
Alberto Cano 0001, Amelia Zafra, Eva Lucrecia Gibaja Galindo, Sebastián Ventura |
EuroGP | 1 |
| 2013 | Predicting student failure at school using genetic programming and different data mining approaches with high dimensional and imbalanced data
Carlos Márquez-Vera, Alberto Cano 0001, Cristóbal Romero 0001, Sebastián Ventura |
Appl. Intell. | 2 |
| 2013 | An interpretable classification rule mining algorithm
Alberto Cano 0001, Amelia Zafra, Sebastián Ventura |
Inf. Sci. | 1 |
| 2013 | Parallel multi-objective Ant Programming for classification using GPUs
Alberto Cano 0001, Juan Luis Olmo, Sebastián Ventura |
J. Parallel Distributed Comput. | 1 |
| 2013 | Weighted Data Gravitation Classification for Standard and Imbalanced DataabstractGravitation is a fundamental interaction whose concept and effects applied to data classification become a novel data classification technique. The simple principle of data gravitation classification (DGC) is to classify data samples by comparing the gravitation between different classes. However, the calculation of gravitation is not a trivial problem due to the different relevance of data attributes for distance computation, the presence of noisy or irrelevant attributes, and the class imbalance problem. This paper presents a gravitation-based classification algorithm which improves previous gravitation models and overcomes some of their issues. The proposed algorithm, called DGC+, employs a matrix of weights to describe the importance of each attribute in the classification of each class, which is used to weight the distance between data samples. It improves the classification performance by considering both global and local data information, especially in decision boundaries. The proposal is evaluated and compared to other well-known instance-based classification techniques, on 35 standard and 44 imbalanced data sets. The results obtained from these experiments show the great performance of the proposed gravitation model, and they are validated using several nonparametric statistical tests. Alberto Cano 0001, Amelia Zafra, Sebastián Ventura |
IEEE Trans. Cybern. | 1 |
| 2013 | High performance evaluation of evolutionary-mined association rules on GPUs
Alberto Cano 0001, José María Luna, Sebastián Ventura |
J. Supercomput. | 1 |
| 2012 | Binary and multiclass imbalanced classification using multi-objective ant programmingabstractClassification in imbalanced domains is a challenging task, since most of its real domain applications present skewed distributions of data. However, there are still some open issues in this kind of problem. This paper presents a multi-objective grammar-based ant programming algorithm for imbalanced classification, capable of addressing this task from both the binary and multiclass sides, unlike most of the solutions presented so far. We carry out two experimental studies comparing our algorithm against binary and multiclass solutions, demonstrating that it achieves an excellent performance for both binary and multiclass imbalanced data sets. Juan Luis Olmo, Alberto Cano 0001, José Raúl Romero, Sebastián Ventura |
ISDA | 2 |
| 2012 | Speeding up the evaluation phase of GP classification algorithms on GPUs
Alberto Cano 0001, Amelia Zafra, Sebastián Ventura |
Soft Comput. | 1 |
| 2011 | An EP algorithm for learning highly interpretable classifiersabstractThis paper introduces an Evolutionary Programming algorithm for solving classification problems using highly interpretable IF-THEN classification rules. It is an algorithm aimed to maximize the comprehensibility of the classifier by minimizing the number of rules and employing only relevant attributes. The proposal is evaluated and compared to other 5 well-known classification techniques over 18 datasets. The results obtained from the experiments show its competitive accuracy and the significantly better interpretability of the classifiers provided in terms of number of rules, number of conditions and a complexity metric. Alberto Cano 0001, Amelia Zafra, Sebastián Ventura |
ISDA | 1 |