VLDB 2026 Research / reviewers in the wild / expert
Sebastián Ventura
dblp:00/4717 · also Sebastián Ventura Soto
· DBLP profile ↗
162ranked-venue papers
4as first author
29since 2021 · last 2026
0000-0003-4216-6378ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 98 · 2 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 40 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 31 · 9 since 2021Human-computer interaction and ubiquitous computing · 15 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Systems, architecture and hardware · 4Security and privacy · 2Software engineering, systems software and programming languages · 2Theory of computation · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DANTIS library: Detection of ANomalies in TIme seriesabstractAnomaly detection in time series is essential in domains such as predictive maintenance, cybersecurity, health monitoring, and quality control. Although several software libraries provide anomaly detection algorithms, building complete and reproducible workflows still requires considerable expertise in data preprocessing, model configuration, training, evaluation, visualization, and statistical comparison. This complexity often limits the accessibility and reproducibility of anomaly detection workflows. This paper introduces DANTIS (Detection of ANomalies in TIme Series), an open-source Python library and desktop application designed to simplify the end-to-end development and comparison of anomaly detection models for time series. DANTIS provides a unified framework for data import, preprocessing, model training, testing, result visualization, and experiment export through both a programmatic interface and a graphical user interface. The library includes representative statistical, machine learning, and deep learning detectors, and its modular architecture allows new models to be incorporated easily. To validate its practical usefulness, we report a comparative experiment on labelled datasets from the UCR Anomaly Archive, involving nine representative detectors evaluated under a common protocol. The results show that DANTIS can generate predictive metrics, runtime information, and structured result matrices suitable for statistical analysis with Friedman and Nemenyi tests. Overall, DANTIS provides an accessible, extensible, and reproducible environment for researchers and practitioners who need to compare and interpret anomaly detection methods in time-series applications. Christian Luna, Elena Álvarez, Rafael Egea, José María Luna, Sebastián Ventura |
Neurocomputing | 5 |
| 2026 | A Multi-Fidelity Genetic Algorithm for Hyperparameter Optimization of Deep Neural NetworksabstractHyperparameter optimization on Machine Learning models is crucial for their correct refinement. For complex big models (such as Deep Learning models), in which a single training model is supposed to have a very high computational cost, this optimization sometimes becomes unfeasible. Multi-fidelity optimization algorithms are a solution to alleviate this computational cost of optimizing hyperparameters of Deep Learning models. In this scope, we propose GAMF2O, a new multi-fidelity algorithm that relies on a genetic algorithm. This paper clearly defines how to adapt the evolutionary scenario to follow the multi-fidelity approach, and we propose a new scheme to evaluate each individual based on the use of two objectives: the result of the low-fidelity evaluation and the learning capacity, with the use of the latter being novel during the evaluation process. Our experimental section allows us to show how our proposal improves the state-of-the-art in different classification and regression problems. Antonio R. Moya, Sebastián Ventura |
IEEE Trans. Evol. Comput. | 2 |
| 2025 | Deep Learning Approaches to Assessing University Students' Health-Related Quality of Life: A Comparative Study of MLP and GNNabstractHealth-related quality of life (HRQoL) is a multidimensional construct reflecting individuals' overall well-being, including physical, mental, emotional, and social aspects. This study focuses on assessing HRQoL among university students. A comprehensive survey, incorporating several validated instruments was administered to a stratified sample of undergraduate and graduate students. To analyze the collected data, we developed two deep learning models: a Fully Connected Neural Network (FCNN) and a Graph Neural Network (GNN). The GNN model represents each student's responses as a tree-like graph and processes this structured data through multiple convolutional layers enhanced by a sort-pooling layer, which standardizes graphs of varying dimensions. Experimental results indicate that both models effectively classify students based on their HRQoL profiles. Notably, the GNN model achieves higher test accuracy and F-Score compared to the FCNN, demonstrating robust performance even when analyzing incomplete surveys. These findings underscore the potential of graph-based deep learning methods by integrating heterogeneous data and capturing complex relationships, these approaches offer a promising solution for monitoring and improving students' well-being. Index Terms-Deep Learning, Graph Neural Networks, HRQoL José Luis Ávila-Jiménez, Manuel Rich-Ruiz, Francisco J. Rodríguez-Lozano, Vanesa Cantón-Habas, Sebastián Ventura |
CBMS | 5 |
| 2025 | Enhancing Medical Diagnosis with Instance Hardness-Guided Multi-Level Cross-Validation for Imbalanced LearningabstractCross-validation is a critical component for robust machine learning evaluation. In imbalanced learning, stratified cross-validation is commonly recommended to preserve class distribution. However, it neglects the underlying distribution of instance hardness, which can introduce distribution shifts between training and testing folds, ultimately compromising the validity of performance evaluation. This paper proposes a stratified cross-validation informed by the hardness distribution for robust imbalanced medical diagnosis. The proposed multi-level cross-validation (MLCV) retains jointly the class distribution and instance hardness, maintaining equivalent distribution of hardness levels across folds. This strategy enables the model to encounter a more realistic version of the medical data for a reliable performance evaluation. Experimental work demonstrates that the hardness distribution shift exists; the (MLCV) not only enhances classification performance in imbalanced medical data but also improves the results of balancing methods, as measured by classification performance indicators such as recall, precision, and F1-measure. Mabrouka Salmi, Dalia Atif, Sebastián Ventura |
DSAA | 3 |
| 2025 | MIHT: A Hoeffding Tree for Time Series Classification Using Multiple Instance Learning
Aurora Esteban, Amelia Zafra, Sebastián Ventura |
IDEAL (1) | 3 |
| 2025 | Simultaneous fault prediction in evolving industrial environments with ensembles of Hoeffding adaptive treesabstractAbstract Predictive Maintenance (PdM) emerges as a critical task of Industry 4.0, driving operational efficiency, minimizing downtime, and reducing maintenance costs. However, real-world industrial environments present unsolved challenges, especially in predicting simultaneous and correlated faults under evolving conditions. Traditional batch-based and deep learning approaches for simultaneous fault prediction often fall short due to their assumptions of static data distributions and high computational demands, making them unsuitable for dynamic, resource-constrained systems. In response, we propose OEMLHAT (Online Ensemble of Multi-Label Hoeffding Adaptive Trees), a novel model tailored for real-time, multi-label fault prediction in non-stationary industrial settings. OEMLHAT introduces a scalable online ensemble architecture that integrates online bagging, dynamic feature subspacing, and adaptive output weighting. This design allows it to efficiently handle concept drift, high-dimensional input spaces, and label sparsity, key bottlenecks in existing PdM solutions. Experimental results on three public multi-label PdM case studies demonstrate substantial improvements in predictive performance of OEMLHAT over previous batch-based and online proposals for multi-label classification, particularly with an average improvement in micro-averaged F1-score of 18.49% over the second most-accurate batch-based proposal and of 8.56% in the case of the second best online model. By addressing a critical gap in online multi-label learning for PdM, this work provides a robust and interpretable solution for next-generation industrial monitoring systems for fault detection, particularly for rare and concurrent failures. Aurora Esteban, Alberto Cano 0001, Sebastián Ventura, Amelia Zafra |
Appl. Intell. | 3 |
| 2024 | Improving hyper-parameter self-tuning for data streams by adapting an evolutionary approach
Antonio R. Moya, Bruno M. Veloso, João Gama 0001, Sebastián Ventura |
Data Min. Knowl. Discov. | 4 |
| 2024 | Introduction to the special issue on recent advances on digital economy-oriented artificial intelligence
Yu-Lin He, Philippe Fournier-Viger, Sebastián Ventura |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | StaTDS library: Statistical tests for Data ScienceabstractIn Data Science, there is a continual demand for statistical comparison to identify the most advantageous algorithms. Finding a software tool that facilitates the execution of multiple tests on different Data Science experiments without relying on additional libraries poses a challenge. This paper introduces StaTDS, an open-source library and web application implemented entirely in pure Python, designed to analyze, test, and compare Data Science algorithms. StaTDS implements all statistical tests without external dependencies. It ensures its durability and avoids future uncontrolled deprecated dependencies. With support for a wide variety of statistical tests (24 in total), StaTDS surpasses existing libraries dedicated to statistical testing. Moreover, the library incorporates tests to guide users in determining whether to employ parametric or non-parametric tests, such as the assessment of normality and homoscedasticity. This platform-independent library is available on GitHub under the GNU General Public License. Christian Luna, Antonio R. Moya, José María Luna, Sebastián Ventura |
Neurocomputing | 4 |
| 2024 | Data heterogeneity's impact on the performance of frequent itemset mining algorithmsabstractFrequent itemset mining (FIM) is a widely used task that extracts frequently occurring itemsets from data. Plenty of deterministic algorithms are available for this daunting task. However, experimental studies have not considered that data heterogeneity significantly impacts the algorithms' performance, giving rise to unfair comparisons and biased conclusions. This paper seeks to advance by comparing cutting-edge algorithms using various frequency thresholds, considering the resulting data heterogeneity. An extensive experimental study is carried out, including the number of itemsets mined per second as the performance quality measure to compare algorithms. The experiments include defining eight metrics to quantify data heterogeneity, and their values vary the algorithms' performance. The results revealed that some techniques (hypercube decomposition and k-items machine) are essential to achieve excellent performance on any dataset, and most algorithms behave similarly well when they include those techniques. As a final important point, different threshold values produce dissimilar data subsets (data heterogeneity is not an immutable data characteristic), so a previous study on the database characteristics with a few minimum support thresholds could be beneficial to select the best-suited FIM algorithm beforehand. Antonio Manuel Trasierras, José María Luna, Philippe Fournier-Viger, Sebastián Ventura |
Inf. Sci. | 4 |
| 2024 | Hoeffding adaptive trees for multi-label classification on data streams
Aurora Esteban, Alberto Cano 0001, Amelia Zafra, Sebastián Ventura |
Knowl. Based Syst. | 4 |
| 2023 | A Comparison of Neural Network-Based Super-Resolution Models on 3D Rendered Images
Rafael Berral-Soler, Francisco José Madrid-Cuevas, Sebastián Ventura, Rafael Muñoz-Salinas, Manuel J. Marín-Jiménez |
CAIP (1) | 3 |
| 2023 | Radiomics Software Tools: A comparative Analysis on Breast CancerabstractRadiomics is an emerging and promising field used to describe visual information from medical images by means of numerical features. Several Radiomics software tools are available in the literature, but they return different features and make dissimilar calculations. Choosing one tool or another is not easy so a comparison for classification tasks is required. This paper compares three of these frameworks (3D Slicer, LIFEx and MaZda) on breast cancer data. In this analysis, we tested the features extracted from each tool using different pre-processing techniques and machine learning algorithms to classify the lesion as benign or malignant on more than 350 registers. Two different projections were considered, that is, craniocaudal (183 registers) and mediolateral oblique (172 registers). The results demonstrated that 3D Slicer obtained the best performance in the craniocaudal projection, while MaZda and LIFEx are more appropriate for the mediolateral oblique projection. The results are really promising for classification tasks, exceeding 85% in F1-score. Eduardo Almeda Luna, José María Luna, Sebastián Ventura |
CBMS | 3 |
| 2023 | Progressive growing of Generative Adversarial Networks for improving data augmentation and skin cancer diagnosisabstractEarly melanoma diagnosis is the most important factor in the treatment of skin cancer and can effectively reduce mortality rates. Recently, Generative Adversarial Networks have been used to augment data, prevent overfitting and improve the diagnostic capacity of models. However, its application remains a challenging task due to the high levels of inter and intra-class variance seen in skin images, limited amounts of data, and model instability. We present a more robust Progressive Growing of Adversarial Networks based on residual learning, which is highly recommended to ease the training of deep networks. The stability of the training process was increased by receiving additional inputs from preceding blocks. The architecture is able to produce plausible photorealistic synthetic 512 × 512 skin images, even with small dermoscopic and non-dermoscopic skin image datasets as problem domains. In this manner, we tackle the lack of data and the imbalance problems. Additionally, the proposed approach leverages a skin lesion boundary segmentation algorithm and transfer learning to enhance the diagnosis of melanoma. Inception score and Matthews Correlation Coefficient were used to measure the performance of the models. The architecture was evaluated qualitatively and quantitatively through the use of an extensive experimental study on sixteen datasets, illustrating its effectiveness in the diagnosis of melanoma. Finally, four state-of-the-art data augmentation techniques applied in five convolutional neural network models were significantly outperformed. The results indicated that a bigger number of trainable parameters will not necessarily obtain a better performance in melanoma diagnosis. Sebastián Ventura |
Artif. Intell. Medicine | 2 |
| 2023 | A contrast set mining based approach for cancer subtype analysisabstractThe task of detecting common and unique characteristics among different cancer subtypes is an important focus of research that aims to improve personalized therapies. Unlike current approaches mainly based on predictive techniques, our study aims to improve the knowledge about the molecular mechanisms that descriptively led to cancer, thus not requiring previous knowledge to be validated. Here, we propose an approach based on contrast set mining to capture high-order relationships in cancer transcriptomic data. In this way, we were able to extract valuable insights from several cancer subtypes in the form of highly specific genetic relationships related to functional pathways affected by the disease. To this end, we have divided several cancer gene expression databases by the subtype associated with each sample to detect which gene groups are related to each cancer subtype. To demonstrate the potential and usefulness of the proposed approach we have extensively analysed RNA-Seq gene expression data from breast, kidney, and colon cancer subtypes. The possible role of the obtained genetic relationships was further evaluated through extensive literature research, while its prognosis was assessed via survival analysis, finding gene expression patterns related to survival in various cancer subtypes. Some gene associations were described in the literature as potential cancer biomarkers while other results have been not described yet and could be a starting point for future research. Antonio Manuel Trasierras, José María Luna, Sebastián Ventura |
Artif. Intell. Medicine | 3 |
| 2023 | Efficient mining of top-k high utility itemsets through genetic algorithms
José María Luna, R. Uday Kiran, Philippe Fournier-Viger, Sebastián Ventura |
Inf. Sci. | 4 |
| 2023 | Eight years of AutoML: categorisation, review and trendsabstractAbstract Knowledge extraction through machine learning techniques has been successfully applied in a large number of application domains. However, apart from the required technical knowledge and background in the application domain, it usually involves a number of time-consuming and repetitive steps. Automated machine learning (AutoML) emerged in 2014 as an attempt to mitigate these issues, making machine learning methods more practicable to both data scientists and domain experts. AutoML is a broad area encompassing a wide range of approaches aimed at addressing a diversity of tasks over the different phases of the knowledge discovery process being automated with specific techniques. To provide a big picture of the whole area, we have conducted a systematic literature review based on a proposed taxonomy that permits categorising 447 primary studies selected from a search of 31,048 papers. This review performs an extensive and rigorous analysis of the AutoML field, scrutinising how the primary studies have addressed the dimensions of the taxonomy, and identifying any gaps that remain unexplored as well as potential future trends. The analysis of these studies has yielded some intriguing findings. For instance, we have observed a significant growth in the number of publications since 2018. Additionally, it is noteworthy that the algorithm selection problem has gradually been superseded by the challenge of workflow composition, which automates more than one phase of the knowledge discovery process simultaneously. Of all the tasks in AutoML, the growth of neural architecture search is particularly noticeable. Rafael Barbudo, Sebastián Ventura, José Raúl Romero |
Knowl. Inf. Syst. | 2 |
| 2023 | A framework to build accurate Convolutional Neural Network models for melanoma diagnosisabstractIn the past few years, Convolutional Neural Networks have achieved performance levels similar to those achieved by dermatologists. However, the diagnosis of melanoma remains a challenging task, mainly due to the high levels of inter and intra-class variability present in images of moles. With the aim of new methods for an effective melanoma diagnosis, a new framework is proposed. The training process is guided by an expert within an active learning approach where the architectures implicitly learn about the complexity of individual images through query strategies, which allows us to adjust the training process and achieve better performance. In addition, we propose a batch-based query strategy that enables a more stable and faster training process. Besides, the framework leverages segmentation, data augmentation and transfer learning to enhance melanoma diagnosis. The framework is composed by several specialized blocks, which allow us to measure how the diagnosis is improved after each step. In this sense, blocks could be customized and do not depend on specific models. An extensive experimental study was conducted on 16 skin image datasets, where five state-of-the-art models were significantly outperformed. This study corroborated that new active learning query strategies can be employed to effectively train neural networks architectures for the diagnosis of melanoma, achieving 182% better predictive performance in Xception, and an overall 11% and 20% better predictive performance in dermoscopic and non-dermoscopic images, respectively. It is worth mentioning that the informativeness value of each image is shown, which leads to identify the hardest images for the predictive models. Finally, the proposal required 2% of the total training time, and needed 61% less training epochs. Sebastián Ventura |
Knowl. Based Syst. | 2 |
| 2022 | Smart Operators for Inducing Colorectal Cancer Classification Trees with PonyGE2 Grammatical Evolution Python PackageabstractColorectal cancer is a disease that affects many people and requires a multidisciplinary approach, involving significant human and economic resources. We have been provided with a tabular dataset with 1.5 thousand cases of this disease. We are interested in producing interpretable classifiers for predicting the occurrence of complications. Grammatical Evolution has extensively been used for machine learning problems. In particular, it can be used to induce interpretable decision trees, with the advantage of allowing the practitioner to easily control the language by means of the grammar. PonyGE2 [1], [2] is a Python package that provides data scientists with Grammatical Evolution algorithms, which can be configured to their needs quite easily. In addition, and thanks to the benefits of the Python programming language, PonyGE2 is currently becoming more and more popular. However, the capabilities of PonyGE2 for inducing classification trees are still subject of improvement. In particular, it only uses simple equality conditions and requires to encode feature names and values with numbers. We have developed some smart operators for PonyGE2, which, not only enhance the framework in interpretability and performance when dealing with our colorectal cancer dataset, but also allows to produce results comparable to those of the widely known heuristic methods C4.5 and CART. We show how they could be applied to other datasets, and how they affect performance in our case. José A. Delgado-Osuna, Carlos García-Martínez, Sebastián Ventura |
CEC | 3 |
| 2022 | Improving the understanding of cancer in a descriptive way: An emerging pattern mining-based approachabstractThis paper presents an approach based on emerging pattern mining to analyse cancer through genomic data. Unlike existing approaches, mainly focused on predictive purposes, the proposal aims to improve the understanding of cancer descriptively, not requiring either any prior knowledge or hypothesis to be validated. Additionally, it enables to consider high-order relationships, so not only essential genes related to the disease are considered, but also the combined effect of various secondary genes that can influence different pathways directly or indirectly related to the disease. The prime hypothesis is that splitting genomic cancer data into two subsets, that is, cases and controls, will allow us to determine which genes, and their expressions, are associated with different cancer types. The possibilities of the proposal are demonstrated by analyzing RNA-Seq data for six different types of cancer: breast, colon, lung, thyroid, prostate, and kidney. Some of the extracted insights were already described in the related literature as good cancer bio-markers, while others have not been described yet mainly due to existing techniques are biased by prior knowledge provided by biological databases. Antonio Manuel Trasierras, José María Luna, Sebastián Ventura |
Int. J. Intell. Syst. | 3 |
| 2022 | Modeling and predicting students' engagement behaviors using mixture Markov models
Rabia Maqsood, Paolo Ceravolo, Cristóbal Romero 0001, Sebastián Ventura |
Knowl. Inf. Syst. | 4 |
| 2022 | An ensemble-based convolutional neural network model powered by a genetic algorithm for melanoma diagnosisabstractAbstract Melanoma is one of the main causes of cancer-related deaths. The development of new computational methods as an important tool for assisting doctors can lead to early diagnosis and effectively reduce mortality. In this work, we propose a convolutional neural network architecture for melanoma diagnosis inspired by ensemble learning and genetic algorithms. The architecture is designed by a genetic algorithm that finds optimal members of the ensemble. Additionally, the abstract features of all models are merged and, as a result, additional prediction capabilities are obtained. The diagnosis is achieved by combining all individual predictions. In this manner, the training process is implicitly regularized, showing better convergence, mitigating the overfitting of the model, and improving the generalization performance. The aim is to find the models that best contribute to the ensemble. The proposed approach also leverages data augmentation, transfer learning, and a segmentation algorithm. The segmentation can be performed without training and with a central processing unit, thus avoiding a significant amount of computational power, while maintaining its competitive performance. To evaluate the proposal, an extensive experimental study was conducted on sixteen skin image datasets, where state-of-the-art models were significantly outperformed. This study corroborated that genetic algorithms can be employed to effectively find suitable architectures for the diagnosis of melanoma, achieving in overall 11% and 13% better prediction performances compared to the closest model in dermoscopic and non-dermoscopic images, respectively. Finally, the proposal was implemented in a web application in order to assist dermatologists and it can be consulted at http://skinensemble.com . Sebastián Ventura |
Neural Comput. Appl. | 2 |
| 2021 | A semantically enriched text mining system for clinical decision supportabstractAbstract Existing systems to support decision‐taking process based on textual information of clinical reports are insufficient. Currently, there are few systems that unify different subtasks in a single and user‐friendly framework, easing therefore the clinical work by automating complex and arduous tasks such as the detection of clinical alerts as well as clinical information coding. To address this issue, MiNerDoc is proposed as a new text mining (TM) system whose main objective is to support clinical decision‐taking processes by analyzing textual clinical reports in a unified framework. MiNerDoc is a really alluring TM system that includes two relevant tasks in the medical field, that is, detection of risk factors according to five medical entities (disease, pharmacologic, region/part body, procedure/test, and finding/sign) and automatic prediction of standardized diagnostic codes (MeSH descriptors associated with diseases). MiNerDoc integrates a combination of techniques from the TM discipline along with the terminological and semantic enrichment provided by the MetaMap tool and UMLS metathesaurus. Some study cases as well as a wide experimental analysis on real clinical reports have been carried out to demonstrate the effectiveness and promising performance of MiNerDoc on two different tasks, that is, medical entities recognition (FMeasure 81.54%) and diagnostic classification (FMeasuremic 81.04%). Carmen Luque, José Miguel Madueño Luna, Sebastián Ventura |
Comput. Intell. | 3 |
| 2021 | EDITORIAL
Sebastián Ventura, Paolo Soda, Alejandro Rodríguez González |
Comput. Intell. | 1 |
| 2021 | A propositionalization method of multi-relational data based on Grammar-Guided Genetic Programming
Luis A. Quintero-Domínguez, Carlos Morell 0001, Sebastián Ventura |
Expert Syst. Appl. | 3 |
| 2021 | Performing multi-target regression via gene expression programming-based ensemble models
Jose M. Moyano, Oscar Gabriel Reyes Pupo, Habib Fardoun, Sebastián Ventura |
Neurocomputing | 4 |
| 2021 | Introduction to the special issue on Methods and applications in the analysis of social data in healthcare
Alejandro Rodríguez González, Sebastián Ventura, Paolo Soda, Jesualdo Tomás Fernández-Breis |
Inf. Process. Manag. | 2 |
| 2021 | Mining local periodic patterns in a discrete sequence
Philippe Fournier-Viger, R. Uday Kiran, Sebastián Ventura, José María Luna |
Inf. Sci. | 4 |
| 2021 | Convolutional neural networks for the automatic diagnosis of melanoma: An extensive experimental study
Oscar Gabriel Reyes Pupo, Sebastián Ventura |
Medical Image Anal. | 3 |
| 2020 | A Preliminary Study on Evolutionary Clustering for Multiple Instance LearningabstractSince its beginnings, multiple instance learning studies have shown an excellent performance in the areas where it has been applied. This efficiency is due to multiple instance learning allows to represent a complex object by a set of feature vectors, being a more flexible representation to preserve more information than one based on single feature vector. This paper attempts to progress in this area carrying out a first study that introduces evolutionary algorithms for solving multiple instance cluster analysis. Specifically, we present four proposals of genetic algorithms for multi-instance partitional clustering: three of them are adaptations of existing algorithms for single-instance clustering, while the last one is a novel approach based on CHC evolutionary algorithm. Moreover, two classic non-genetic partitional algorithms are included in the final comparison. Experimental results considering ten representative datasets show promising results for our proposal. Aurora Esteban, Amelia Zafra, Sebastián Ventura |
CEC | 3 |
| 2020 | Tree-Shaped Ensemble of Multi-Label Classifiers using Grammar-Guided Genetic ProgrammingabstractMulti-label classification paradigm has had a growing interest because of the emergence of a large number of classification problems where each of the instances of the data can be associated with several output labels simultaneously. Several ensemble methods were proposed to solve the multilabel classification problem. However, most of them simply create diversity in the ensemble by following a random procedure and give the same importance to all members. In this paper, we propose a Grammar-Guided Genetic Programming algorithm to build ensembles of multi-label classifiers. Given a pool of multilabel classifiers, each of them modeling dependencies among a subset of k labels, they are combined into a tree-shaped ensemble. At each node of the tree, predictions of its children nodes are combined, while each leaf represents a classifier from the pool. We propose two configurations for the method: using a fixed value of k for all classifiers in the pool, or using a variable value of k for each classifier, thus being able to capture relationships among groups of labels of different size in the ensemble. The experiments performed over sixteen multi-label dataset and using five evaluation metrics demonstrated that our method performs significantly better than the state-of-the-art ensembles of multilabel classifiers. Jose M. Moyano, Eva Lucrecia Gibaja Galindo, Krzysztof J. Cios, Sebastián Ventura |
CEC | 4 |
| 2020 | Generating Ensembles of Multi-Label Classifiers Using Cooperative Coevolutionary Algorithms
Jose M. Moyano, Eva Lucrecia Gibaja Galindo, Krzysztof J. Cios, Sebastián Ventura |
ECAI | 4 |
| 2020 | Mining Cross-Level High Utility Itemsets
Philippe Fournier-Viger, Jerry Chun-Wei Lin, José María Luna, Sebastián Ventura |
IEA/AIE | 5 |
| 2020 | Fast Convergence of Competitive Spiking Neural Networks with Sample-Based Weight Initialization
Paolo Gabriel Cachi, Sebastián Ventura, Krzysztof J. Cios |
IPMU (3) | 2 |
| 2020 | A supervised machine learning-based methodology for analyzing dysregulation in splicing machinery: An application in cancer diagnosis
Oscar Gabriel Reyes Pupo, Raúl M. Luque, Justo Castaño, Sebastián Ventura |
Artif. Intell. Medicine | 5 |
| 2020 | Heuristics for interesting class association rule mining a colorectal cancer database
José A. Delgado-Osuna, Carlos García-Martínez, Jose Gómez Barbadillo, Sebastián Ventura |
Inf. Process. Manag. | 4 |
| 2020 | Distributed multi-label feature selection using individual mutual information measures
Jorge Gonzalez-Lopez, Sebastián Ventura, Alberto Cano 0001 |
Knowl. Based Syst. | 2 |
| 2020 | Predicting literature's early impact with sentiment analysis in Twitter
Saeed-Ul Hassan, Naif R. Aljohani, Nimra Idrees, Raheem Sarwar, Raheel Nawaz, Eugenio Martínez-Cámara, Sebastián Ventura, Francisco Herrera |
Knowl. Based Syst. | 7 |
| 2020 | Combining multi-label classifiers based on projections of the output space using Evolutionary algorithms
Jose M. Moyano, Eva Lucrecia Gibaja Galindo, Krzysztof J. Cios, Sebastián Ventura |
Knowl. Based Syst. | 4 |
| 2020 | LAC: Library for associative classification
Francisco Padillo, José María Luna, Sebastián Ventura |
Knowl. Based Syst. | 3 |
| 2020 | Distributed Selection of Continuous Features in Multilabel Classification Using Mutual InformationabstractMultilabel learning is a challenging task demanding scalable methods for large-scale data. Feature selection has shown to improve multilabel accuracy while defying the curse of dimensionality of high-dimensional scattered data. However, the increasing complexity of multilabel feature selection, especially on continuous features, requires new approaches to manage data effectively and efficiently in distributed computing environments. This article proposes a distributed model for mutual information (MI) adaptation on continuous features and multiple labels on Apache Spark. Two approaches are presented based on MI maximization, and minimum redundancy and maximum relevance. The former selects the subset of features that maximize the MI between the features and the labels, whereas the latter additionally minimizes the redundancy between the features. Experiments compare the distributed multilabel feature selection methods on 10 data sets and 12 metrics. Results validated through statistical analysis indicate that our methods outperform reference methods for distributed feature selection for multilabel data, while MIM also reduces the runtime in orders of magnitude. Jorge Gonzalez-Lopez, Sebastián Ventura, Alberto Cano 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Obtaining Tractable and Interpretable Descriptions for Cases with Complications from a Colorectal Cancer DatabaseabstractColorectal cancer affects to a significant portion of the population and is one of the leading causes of cancer-related deaths in many countries. Professionals of the Reina Sofia University Hospital have fed a database about this pathology for more than 10 years. In this work, we apply classification and association rule learning tools, including a new methodology, to obtain tractable and interpretable descriptions of those cases where complications appeared, which is one of the attributes. José A. Delgado-Osuna, Carlos García-Martínez, Sebastián Ventura, Jose Gómez Barbadillo |
CBMS | 3 |
| 2019 | MiNerDoc: a Semantically Enriched Text Mining System to Transform Clinical Text into KnowledgeabstractExisting systems to support the daily decision taking process carried out by health professionals need to be used independently to perform different text mining subtasks. In practice, there are few systems that unify all the subtasks into an unique framework, easing therefore the clinical work by automating complex clinical tasks such as the detection of clinical alerts as well as clinical information coding. In this sense, the MiNerDoc system is proposed, whose main objective is to support clinical decision-taking process by analysing tons of textual clinical reports in an unified framework. MiNerDoc performs two basic functions that are of great importance in the medical field: detection of risk factors based on the recognition of five medical entities (Disease, Pharmacologic, Region/Part Body, Procedure/Test, Finding/Sign), and automatic prediction of standardized diagnostic codes (MeSH descriptors). A major feature of MiNerDoc is it includes external knowledge sources such as MetaMap and UMLS to terminologically and semantically enrich the interpretation of clinical texts. Some study cases are considered in this work to demonstrate the power of MiNerDoc. Carmen Luque, José María Luna, Sebastián Ventura |
CBMS | 3 |
| 2019 | A Supervised Methodology for Analyzing Dysregulation in Splicing Machinery: An Application in Cancer DiagnosisabstractDeregulated splicing factors have shown to be associated with the development of several types of cancer and, therefore, the determination of such alterations can help the development of tumor-specific molecular targets for early prognosis and therapy. Determining the relevant splicing factors, however, is not a straightforward task mainly due to the heterogeneity of tumors and the variability across samples. In this work, a methodology based on supervised machine learning methods is proposed, allowing the determination of subsets of relevant factors that best discriminate samples. The methodology comprises three main phases: first, a ranking of splicing factors is determined by means of applying feature weighting algorithms; second, the best subset of factors that allows the induction of an accurate classifier is detected; then the confidence over the induced classifier is assessed by means of explaining the individual predictions. Finally, the utility and benefit of the proposed methodology are illustrated by means of analyzing a small dataset of neuroendocrine lung carcinoids, and the results showed that there exist small subsets of deregulated factors which can effectively distinguish between tumor samples and their respective adjacent non-tumor tissues. Oscar Gabriel Reyes Pupo, Raúl M. Luque, Justo Castaño, Sebastián Ventura |
CBMS | 4 |
| 2019 | Discovering Students' Engagement Behaviors in Confidence-based AssessmentabstractConsidering the usefulness of monitoring students' response to available task-level feedback in confidence-based assessment, in this paper, we introduce a novel approach to classify students problem-solving activities into various engagement and disengagement behaviors and study their occurrences during complete learning sessions. Then by clustering these sessions, we obtained four distinct groups which varied both in terms of students' (dis)engagement behaviors and their quantitative performance scores in confidence-based assessment. Moreover, a qualitative analysis shows that high and low performance students (determined based on their final scores in the course) relate differently to the obtained clusters. Based on these findings we highlight that our approach of investigating students' engagement by observing traces of performed problem-solving activities is promising and opens new avenues of research. Also, our approach is more generic as it does not contain human-expert defined time limits which are usually determined by analyzing students' data who participated in the experimental study. Rabia Maqsood, Paolo Ceravolo, Sebastián Ventura |
EDUCON | 3 |
| 2019 | Associative Classification in Big Data through a G3P ApproachabstractThe associative classification field includes really interesting approaches for building reliable classifiers and any of these approaches generally work on four different phases (data discretization, pattern mining, rule mining, and classifier building). This number of phases is a handicap when big datasets are analysed. The aim of this work is to propose a novel evolutionary algorithm for efficiently building associative classifiers in Big Data. The proposed model works in only two phases (a grammar-guided genetic programming framework is performed in each phase): 1) mining reliable association rules; 2) building an accurate classifier by ranking and combining the previously mined rules. The proposal has been implemented on Apache Spark to take advantage of the distributed computing. The experimental analysis was performend on 40 well-known datasets and considering 13 algorithms taken from literature. A series of non-parametric tests has also been carried out to determine statistical differences. Results are quite promising in terms of reliability and efficiency on high-dimensional data. José María Luna, Francisco Padillo, Sebastián Ventura |
IoTBDS | 3 |
| 2019 | Speeding Up Classifier Chains in Multi-label ClassificationabstractMulti-label classification has attracted increasing attention of the scientific community in recent years, given its ability to solve problems where each of the examples simultaneously belongs to multiple labels. From all the techniques developed to solve multi-label classification problems, Classifier Chains has been demonstrated to be one of the best performing techniques. However, one of its main drawbacks is its inherently sequential definition. Although many research works aimed to reduce the runtime of multi-label classification algorithms, to the best of our knowledge, there are no proposals to specifically reduce the runtime of Classifier Chains. Therefore, in this paper we propose a method called Parallel Classifier Chains which enables the parallelization of Classifier Chain. In this way, Parallel Classifier Chains builds k binary classifiers in parallel, where each of them includes as extra input features the predictions of those labels that have been previously built. We performed an experimental evaluation over 20 datasets using 5 metrics to analyze both the runtime and the predictive performance of our proposal. The results of the experiments affirmed that our proposal was able to significantly reduce the runtime of Classifier Chains while the predictive performance was not statistically significantly harmed. Jose M. Moyano, Eva Lucrecia Gibaja Galindo, Sebastián Ventura, Alberto Cano 0001 |
IoTBDS | 3 |
| 2019 | JCLEC-MO: A Java suite for solving many-objective optimization engineering problems
Aurora Ramírez 0001, José Raúl Romero, Carlos García-Martínez, Sebastián Ventura |
Eng. Appl. Artif. Intell. | 4 |
| 2019 | Virtual learning environment to predict withdrawal by leveraging deep learningabstractThe current evolution in multidisciplinary learning analytics research poses significant challenges for the exploitation of behavior analysis by fusing data streams toward advanced decision-making. The identification of students that are at risk of withdrawals in higher education is connected to numerous educational policies, to enhance their competencies and skills through timely interventions by academia. Predicting student performance is a vital decision-making problem including data from various environment modules that can be fused into a homogenous vector to ascertain decision-making. This research study exploits a temporal sequential classification problem to predict early withdrawal of students, by tapping the power of actionable smart data in the form of students' interactional activities with the online educational system, using the freely available Open University Learning Analytics data set by employing deep long short-term memory (LSTM) model. The deployed LSTM model outperforms baseline logistic regression and artificial neural networks by 10.31% and 6.48% respectively with 97.25% learning accuracy, 92.79% precision, and 85.92% recall. Saeed-Ul Hassan, Hajra Waheed, Naif R. Aljohani, Mohsen Ali, Sebastián Ventura, Francisco Herrera |
Int. J. Intell. Syst. | 5 |
| 2019 | Performing Multi-Target Regression via a Parameter Sharing-Based Deep NetworkabstractMulti-target regression (MTR) comprises the prediction of multiple continuous target variables from a common set of input variables. There are two major challenges when addressing the MTR problem: the exploration of the inter-target dependencies and the modeling of complex input-output relationships. This paper proposes a neural network model that is able to simultaneously address these two challenges in a flexible way. A deep architecture well suited for learning multiple continuous outputs is designed, providing some flexibility to model the inter-target relationships by sharing network parameters as well as the possibility to exploit target-specific patterns by learning a set of nonshared parameters for each target. The effectiveness of the proposal is analyzed through an extensive experimental study on 18 datasets, demonstrating the benefits of using a shared representation that exploits the commonalities between target variables. According to the experimental results, the proposed model is competitive with respect to the state-of-the-art in MTR. Oscar Gabriel Reyes Pupo, Sebastián Ventura |
Int. J. Neural Syst. | 2 |
| 2019 | A survey of many-objective optimisation in search-based software engineering
Aurora Ramírez 0001, José Raúl Romero, Sebastián Ventura |
J. Syst. Softw. | 3 |
| 2019 | LEAC: An efficient library for clustering with evolutionary algorithms
Hermes Robles-Berumen, Amelia Zafra, Habib Fardoun, Sebastián Ventura |
Knowl. Based Syst. | 4 |
| 2018 | Distributed nearest neighbor classification for large-scale multi-label data on spark
Jorge Gonzalez-Lopez, Sebastián Ventura, Alberto Cano 0001 |
Future Gener. Comput. Syst. | 2 |
| 2018 | Effective active learning strategy for multi-label learning
Oscar Gabriel Reyes Pupo, Carlos Morell 0001, Sebastián Ventura |
Neurocomputing | 3 |
| 2018 | Interactive multi-objective evolutionary optimization of software architectures
Aurora Ramírez 0001, José Raúl Romero, Sebastián Ventura |
Inf. Sci. | 3 |
| 2018 | Statistical comparisons of active learning strategies over multiple datasets
Oscar Gabriel Reyes Pupo, Abdulrahman H. Altalhi, Sebastián Ventura |
Knowl. Based Syst. | 3 |
| 2018 | MIRSVM: Multi-instance support vector machine with bag representatives
Gabriella Melki, Alberto Cano 0001, Sebastián Ventura |
Pattern Recognit. | 3 |
| 2018 | Mining Context-Aware Association Rules Using Grammar-Based Genetic ProgrammingabstractReal-world data usually comprise features whose interpretation depends on some contextual information. Such contextual-sensitive features and patterns are of high interest to be discovered and analyzed in order to obtain the right meaning. This paper formulates the problem of mining context-aware association rules, which refers to the search for associations between itemsets such that the strength of their implication depends on a contextual feature. For the discovery of this type of associations, a model that restricts the search space and includes syntax constraints by means of a grammar-based genetic programming methodology is proposed. Grammars can be considered as a useful way of introducing subjective knowledge to the pattern mining process as they are highly related to the background knowledge of the user. The performance and usefulness of the proposed approach is examined by considering synthetically generated datasets. A posteriori analysis on different domains is also carried out to demonstrate the utility of this kind of associations. For example, in educational domains, it is essential to identify and understand contextual and context-sensitive factors that affect overall and individual student behavior and performance. The results of the experiments suggest that the approach is feasible and it automatically identifies interesting context-aware associations from real-world datasets. José María Luna, Mykola Pechenizkiy, María José del Jesus, Sebastián Ventura |
IEEE Trans. Cybern. | 4 |
| 2018 | Apriori Versions Based on MapReduce for Mining Frequent Patterns on Big DataabstractPattern mining is one of the most important tasks to extract meaningful and useful information from raw data. This task aims to extract item-sets that represent any type of homogeneity and regularity in data. Although many efficient algorithms have been developed in this regard, the growing interest in data has caused the performance of existing pattern mining techniques to be dropped. The goal of this paper is to propose new efficient pattern mining algorithms to work in big data. To this aim, a series of algorithms based on the MapReduce framework and the Hadoop open-source implementation have been proposed. The proposed algorithms can be divided into three main groups. First, two algorithms [Apriori MapReduce (AprioriMR) and iterative AprioriMR] with no pruning strategy are proposed, which extract any existing itemset in data. Second, two algorithms (space pruning AprioriMR and top AprioriMR) that prune the search space by means of the well-known anti-monotone property are proposed. Finally, a last algorithm (maximal AprioriMR) is also proposed for mining condensed representations of frequent patterns. To test the performance of the proposed algorithms, a varied collection of big data datasets have been considered, comprising up to 3·1018 transactions and more than 5 million of distinct single-items. The experimental stage includes comparisons against highly efficient and well-known pattern mining algorithms. Results reveal the interest of applying MapReduce versions when complex problems are considered, and also the unsuitability of this paradigm when dealing with small data. José María Luna, Francisco Padillo, Mykola Pechenizkiy, Sebastián Ventura |
IEEE Trans. Cybern. | 4 |
| 2018 | Evolutionary Strategy to Perform Batch-Mode Active Learning on Multi-Label DataabstractMulti-label learning has become an important area of research owing to the increasing number of real-world problems that contain multi-label data. Data labeling is an expensive process that requires expert handling. The annotation of multi-label data is laborious since a human expert needs to consider the presence/absence of each possible label. Consequently, numerous modern multi-label problems may involve a small number of labeled examples and plentiful unlabeled examples simultaneously. Active learning methods allow us to induce better classifiers by selecting the most useful unlabeled data, thus considerably reducing the labeling effort and the cost of training an accurate model. Batch-mode active learning methods focus on selecting a set of unlabeled examples in each iteration in such a way that the selected examples are informative and as diverse as possible. This article presents a strategy to perform batch-mode active learning on multi-label data. The batch-mode active learning is formulated as a multi-objective problem, and it is solved by means of an evolutionary algorithm. Extensive experiments were conducted in a large collection of datasets, and the experimental results confirmed the effectiveness of our proposal for better batch-mode multi-label active learning. Oscar Gabriel Reyes Pupo, Sebastián Ventura |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2017 | Multi-view semi-supervised learning using genetic programming interpretable classification rulesabstractMulti-view learning is a novel paradigm that aims at obtaining better results by examining the information from several perspectives instead of by analysing the same information from a single viewpoint. The multi-view methodology has widely been used for semi-supervised learning, where just some patterns were previously classified by an expert and there is a large amount of unlabelled ones. However to our knowledge, the multi-view learning paradigm has not been applied to produce interpretable rule-based classifiers before. In this work, we present a multi-view extension of a grammar-based genetic programming model for inducing rules for semi-supervised contexts. Its idea is to evolve several populations, and their corresponding views, favouring both the accuracy of the predictions for the labelled patterns and the prediction agreement with the other views for unlabelled ones. We have carried out experiments with two to five views, on six common datasets for fully-supervised learning that have been partially anonymised for our semi-supervised study. Our results show that the multi-view paradigm allows to obtain slightly better rule-based classifiers, and that two views becomes preferred. Carlos García-Martínez, Sebastián Ventura |
CEC | 2 |
| 2017 | An evolutionary algorithm for optimizing the target ordering in Ensemble of Regressor ChainsabstractIn this article we present an evolutionary algorithm for the optimization of sequences of targets for the multi-target regression algorithm Ensemble of Regressor Chains. This algorithm selects several random sequences or chains of targets where to predict each target, the values of previous targets in the chain are included as features, considering in this way the relationship among them. Under the assumption that a target may be better predicted if it is highly correlated with the targets which were included as feature, our proposal, called CCO-ERC, looks for chains where each target is highly correlated with previous targets in the chain. Several methods for the combination of predictions in the ensemble and for the selection of the chains which forms the ensemble are also proposed. CCO-ERC is compared to other state-of-the-art algorithms in multi-target regression, presenting statistically better performance than them. Jose M. Moyano, Eva Lucrecia Gibaja Galindo, Sebastián Ventura |
CEC | 3 |
| 2017 | An evolutionary algorithm for mining rare association rules: A Big Data approachabstractAssociation rule mining is one of the most wellknown techniques to discover interesting relations between items in data. To date, this task has been mainly focused on the discovery of frequent relationships. However, it is often interesting to focus on those that do not occur frequently. Rare association rule mining is an alluring field aiming at describing rare cases or unexpected behavior. This field is really useful over Big Data where abnormal endeavor are more curious than common behavior. In this sense, our aim is to propose a new evolutionary algorithm based on grammars to obtain rare association rules on Big Data. The novelty of our work is that it is eminently designed to be parallel, enabling its use over emerging technologies as Spark and Flink. Furthermore, while other algorithms focus on maximizing a couple of quality measure ignoring the rest, our fitness function has been precisely designed to obtain a trade-off while maximizing a set of well-known quality measures. The experimental study includes more than 70 datasets revealing alluring results in efficiency when more than 300 million of instances and file sizes up to 250 GBytes are considered, and proving that it is able to run efficiently in huge volumes of data. Francisco Padillo, José María Luna, Sebastián Ventura |
CEC | 3 |
| 2017 | On the effect of local search in the multi-objective evolutionary discovery of software architecturesabstractSoftware architects devote substantial efforts to find the most fitting architectural description for their system, which should not only specify its structure, but is also required to meet multiple, simultaneous quality criteria. Evolutionary computation has recently demonstrated to provide insightful support during the design phase by automatically deciding how to organise internal software components and how they should interact each other. Observed from a multi-objective perspective, particular care has to be taken in order to reach an appropriate trade-off among design metrics, while providing the software engineer with diverse alternatives to choose among. However, multi-objective evolutionary algorithms may find difficulties to control both aspects and, at the same time, to explore the entire search space in depth. Under these circumstances, local search can be applied to complement the evolution by scrutinising the most promising search directions. This paper proposes two different approaches that take advantage of the benefits of local search within the multi-objective evolutionary discovery of component-based software architectures. A detailed analysis and comparative study provides interesting findings like the importance of assigning a sufficient number of evaluations to the local improvement. The way in which local search explores and compares solutions for acceptance is a relevant aspect to promote diversity during the discovery process as well. Aurora Ramírez 0001, José Raúl Romero, Sebastián Ventura |
CEC | 3 |
| 2017 | Extremely high-dimensional optimization with MapReduce: Scaling functions and algorithm
Alberto Cano 0001, Carlos García-Martínez, Sebastián Ventura |
Inf. Sci. | 3 |
| 2017 | Multi-target support vector regression via correlation regressor chains
Gabriella Melki, Alberto Cano 0001, Vojislav Kecman, Sebastián Ventura |
Inf. Sci. | 4 |
| 2017 | MLDA: A tool for analyzing multi-label datasets
Jose M. Moyano, Eva Lucrecia Gibaja Galindo, Sebastián Ventura |
Knowl. Based Syst. | 3 |
| 2017 | Multi-objective genetic programming for feature extraction and data visualization
Alberto Cano 0001, Sebastián Ventura, Krzysztof J. Cios |
Soft Comput. | 2 |
| 2016 | Subgroup discovery on big data: Pruning the search space on exhaustive search algorithmsabstractSubgroup Discovery is a broadly applicable supervised local pattern mining method to search relations between different properties with respect to a target variable. With the exponential growth in data storage, the massive data gathered has hampered the performance of current techniques. In this regard, our aim is to propose two new algorithms to discover subgroups on Big Data by using MapReduce. Apache Spark was used to tackle the Big Data requirements. The experimental study includes more than 50 large datasets and a set of efficient algorithms. Search spaces bigger than 1.276 · 1015subgroups are used. The experimental study reveals the alluring results in efficiency when optimistic estimates are considered, as well as demonstrating the usefulness of using Apache Spark to tackle Big Data. Francisco Padillo, José María Luna, Sebastián Ventura |
IEEE BigData | 3 |
| 2016 | Mining Perfectly Rare Itemsets on Big Data: An Approach Based on Apriori-Inverse and MapReduce
Francisco Padillo, José María Luna, Sebastián Ventura |
ISDA | 3 |
| 2016 | Memetic Algorithms for the Automatic Discovery of Software Architectures
Aurora Ramírez 0001, Rafael Barbudo, José Raúl Romero, Sebastián Ventura |
ISDA | 4 |
| 2016 | Early dropout prediction using data mining: a case study with high school studentsabstractAbstract Early prediction of school dropout is a serious problem in education, but it is not an easy issue to resolve. On the one hand, there are many factors that can influence student retention. On the other hand, the traditional classification approach used to solve this problem normally has to be implemented at the end of the course to gather maximum information in order to achieve the highest accuracy. In this paper, we propose a methodology and a specific classification algorithm to discover comprehensible prediction models of student dropout as soon as possible. We used data gathered from 419 high schools students in Mexico. We carried out several experiments to predict dropout at different steps of the course, to select the best indicators of dropout and to compare our proposed algorithm versus some classical and imbalanced well‐known classification algorithms. Results show that our algorithm was capable of predicting student dropout within the first 4–6 weeks of the course and trustworthy enough to be used in an early warning system. Carlos Márquez-Vera, Alberto Cano 0001, Cristóbal Romero 0001, Amin Y. Noaman, Habib Fardoun, Sebastián Ventura |
Expert Syst. J. Knowl. Eng. | 6 |
| 2016 | A comparative study of many-objective evolutionary algorithms for the discovery of software architectures
Aurora Ramírez 0001, José Raúl Romero, Sebastián Ventura |
Empir. Softw. Eng. | 3 |
| 2016 | LAIM discretization for multi-label data
Alberto Cano 0001, José María Luna, Eva Lucrecia Gibaja Galindo, Sebastián Ventura |
Inf. Sci. | 4 |
| 2016 | Discovering useful patterns from multiple instance data
José María Luna, Alberto Cano 0001, Virgilijus Sakalauskas, Sebastián Ventura |
Inf. Sci. | 4 |
| 2016 | Effective lazy learning algorithm based on a data gravitation model for multi-label learning
Oscar Gabriel Reyes Pupo, Carlos Morell 0001, Sebastián Ventura |
Inf. Sci. | 3 |
| 2016 | JCLAL: A Java Framework for Active LearningabstractActive Learning has become an important area of research owing to the increasing number of real-world problems which contain labelled and unlabelled examples at the same time. JCLAL is a Java Class Library for Active Learning which has an architecture that follows strong principles of object-oriented design. It is easy to use, and it allows the developers to adapt, modify and extend the framework according to their needs. The library offers a variety of active learning methods that have been proposed in the literature. The software is available under the GPL license. Oscar Gabriel Reyes Pupo, María del Carmen Rodríguez-Hernández, Habib Fardoun, Sebastián Ventura |
J. Mach. Learn. Res. | 5 |
| 2016 | Mining exceptional relationships with grammar-guided genetic programming
José María Luna, Mykola Pechenizkiy, Sebastián Ventura |
Knowl. Inf. Syst. | 3 |
| 2016 | ur-CAIM: improved CAIM discretization for unbalanced and balanced data
Alberto Cano 0001, Dat T. Nguyen, Sebastián Ventura, Krzysztof J. Cios |
Soft Comput. | 3 |
| 2016 | Speeding-Up Association Rule Mining With Inverted Index CompressionabstractThe growing interest in data storage has made the data size to be exponentially increased, hampering the process of knowledge discovery from these large volumes of high-dimensional and heterogeneous data. In recent years, many efficient algorithms for mining data associations have been proposed, facing up time and main memory requirements. Nevertheless, this mining process could still become hard when the number of items and records is extremely high. In this paper, the goal is not to propose new efficient algorithms but a new data structure that could be used by a variety of existing algorithms without modifying its original schema. Thus, our aim is to speed up the association rule mining process regardless the algorithm used to this end, enabling the performance of efficient implementations to be enhanced. The structure simplifies, reorganizes, and speeds up the data access by sorting data by means of a shuffling strategy based on the hamming distance, which achieve similar values to be closer, and considering both an inverted index mapping and a run length encoding compression. In the experimental study, we explore the bounds of the algorithms' performance by using a wide number of data sets that comprise either thousands or millions of both items and records. The results demonstrate the utility of the proposed data structure in enhancing the algorithms' runtime orders of magnitude, and substantially reducing both the auxiliary and the main memory requirements. José María Luna, Alberto Cano 0001, Mykola Pechenizkiy, Sebastián Ventura |
IEEE Trans. Cybern. | 4 |
| 2015 | Discovering clues to avoid middle school failure at early stagesabstractThe use of data mining techniques in educational domains helps to find new knowledge about how students learn and how to improve the resources management. Using these techniques for predicting school failure is very useful in order to carry out actions to avoid drop out. With this purpose, we try to determine the earliest stage when the quality of the results allows for clarifying the possibility of school failure. We process real information from a Spanish high school by structuring the whole data in incremental datasets, which represent how students' academic records grow. Our study reveals an early and robust detection of the risky cases of school failure at the end of the first out of four courses. Manuel Ángel Jiménez-Gómez, José María Luna, Cristóbal Romero 0001, Sebastián Ventura |
LAK | 4 |
| 2015 | An evolutionary algorithm for the discovery of rare class association rules in learning management systems
José María Luna, Cristóbal Romero 0001, José Raúl Romero, Sebastián Ventura |
Appl. Intell. | 4 |
| 2015 | Scalable extensions of the ReliefF algorithm for weighting and selecting features on the multi-label learning context
Oscar Gabriel Reyes Pupo, Carlos Morell 0001, Sebastián Ventura |
Neurocomputing | 3 |
| 2015 | An approach for the evolutionary discovery of software architectures
Aurora Ramírez 0001, José Raúl Romero, Sebastián Ventura |
Inf. Sci. | 3 |
| 2015 | A classification module for genetic programming algorithms in JCLEC
Alberto Cano 0001, José María Luna, Amelia Zafra, Sebastián Ventura |
J. Mach. Learn. Res. | 4 |
| 2015 | Speeding up multiple instance learning classification rules on GPUs
Alberto Cano 0001, Amelia Zafra, Sebastián Ventura |
Knowl. Inf. Syst. | 3 |
| 2014 | Accepting or Rejecting Students_ Self-grading in their Final Marks by using Data Mining
Javier Fuentes, Cristóbal Romero 0001, Carlos García-Martínez, Sebastián Ventura |
EDM | 4 |
| 2014 | GPU-parallel subtree interpreter for genetic programmingabstractGenetic Programming (GP) is a computationally intensive technique but its nature is embarrassingly parallel. Graphic Processing Units (GPUs) are many-core architectures which have been widely employed to speed up the evaluation of GP. In recent years, many works have shown the high performance and efficiency of GPUs on evaluating both the individuals and the fitness cases in parallel. These approaches are known as population parallel and data parallel. This paper presents a parallel GP interpreter which extends these approaches and adds a new parallelization level based on the concurrent evaluation of the individual's subtrees. A GP individual defined by a tree structure with nodes and branches comprises different depth levels in which there are independent subtrees which can be evaluated concurrently. Threads can cooperate to evaluate different subtrees and share the results via GPU's shared memory. The experimental results show the better performance of the proposal in terms of the GP operations per second (GPops/s) that the GP interpreter is capable of processing, achieving up to 21 billion GPops/s using a NVIDIA 480 GPU. However, some issues raised due to limitations of currently available hardware are to be overcomed by the dynamic parallelization capabilities of the next generation of GPUs. Alberto Cano 0001, Sebastián Ventura |
GECCO | 2 |
| 2014 | On the performance of multiple objective evolutionary algorithms for software architecture discoveryabstractDuring the design of complex systems, software architects have to deal with a tangle of abstract artefacts, measures and ideas to discover the most fitting underlying architecture. A common way to structure these systems is in terms of their interacting software components, whose composition and connections need to be properly adjusted. Its abstract and highly combinatorial nature increases the complexity of the problem. In this scenario, Search-based Software Engineering (SBSE) may serve to support this decision making process from initial analysis models, since the discovery of component-based architectures can be formulated as a challenging multiple optimisation problem, where different metrics and configurations can be applied depending on the design requirements and its specific domain. Many-objective optimisation evolutionary algorithms can provide an interesting alternative to classical multi-objective approaches. This paper presents a comparative study of five different algorithms, including an empirical analysis of their behaviour in terms of quality and variety of the returned solutions. Results are also discussed considering those aspects of concern to the expert in the decision making process, like the number and type of architectures found. The analysis of many-objectives algorithms constitutes an important challenge, since some of them have never been explored before in SBSE. Aurora Ramírez 0001, José Raúl Romero, Sebastián Ventura |
GECCO | 3 |
| 2014 | Parallel evaluation of Pittsburgh rule-based classifiers on GPUs
Alberto Cano 0001, Amelia Zafra, Sebastián Ventura |
Neurocomputing | 3 |
| 2014 | Special issue: Advances in learning schemes for function approximation
Emilio Corchado, Ajith Abraham, Pedro Antonio Gutiérrez, José Manuel Benítez 0001, Sebastián Ventura |
Neurocomputing | 5 |
| 2014 | Foreword: Intelligent data analysis
Sebastián Ventura, Cristóbal Romero 0001, Ajith Abraham |
J. Comput. Syst. Sci. | 1 |
| 2014 | On the adaptability of G3PARM to the extraction of rare association rules
José María Luna, José Raúl Romero, Sebastián Ventura |
Knowl. Inf. Syst. | 3 |
| 2014 | On the Use of Genetic Programming for Mining Comprehensible Rules in Subgroup DiscoveryabstractThis paper proposes a novel grammar-guided genetic programming algorithm for subgroup discovery. This algorithm, called comprehensible grammar-based algorithm for subgroup discovery (CGBA-SD), combines the requirements of discovering comprehensible rules with the ability to mine expressive and flexible solutions owing to the use of a context-free grammar. Each rule is represented as a derivation tree that shows a solution described using the language denoted by the grammar. The algorithm includes mechanisms to adapt the diversity of the population by self-adapting the probabilities of recombination and mutation. We compare the approach with existing evolutionary and classic subgroup discovery algorithms. CGBA-SD appears to be a very promising algorithm that discovers comprehensible subgroups and behaves better than other algorithms as measures by complexity, interest, and precision indicate. The results obtained were validated by means of a series of nonparametric tests. José María Luna, José Raúl Romero, Cristóbal Romero 0001, Sebastián Ventura |
IEEE Trans. Cybern. | 4 |
| 2014 | Scalable CAIM discretization on multiple GPUs using concurrent kernels
Alberto Cano 0001, Sebastián Ventura, Krzysztof J. Cios |
J. Supercomput. | 2 |
| 2013 | ReliefF-ML: An Extension of ReliefF Algorithm to Multi-label Learning
Oscar Gabriel Reyes Pupo, Carlos Morell 0001, Sebastián Ventura |
CIARP (2) | 3 |
| 2013 | A Moodle Block for Selecting, Visualizing and Mining Students' Usage Data
Cristóbal Romero 0001, Cristobal Castro, Sebastián Ventura |
EDM | 3 |
| 2013 | A meta-learning approach for recommending a subset of white-box classification algorithms for Moodle datasets
Cristóbal Romero 0001, Juan Luis Olmo, Sebastián Ventura |
EDM | 3 |
| 2013 | A Grammar-Guided Genetic Programming Algorithm for Multi-Label Classification
Alberto Cano 0001, Amelia Zafra, Eva Lucrecia Gibaja Galindo, Sebastián Ventura |
EuroGP | 4 |
| 2013 | Discovering Subgroups by Means of Genetic Programming
José María Luna, José Raúl Romero, Cristóbal Romero 0001, Sebastián Ventura |
EuroGP | 4 |
| 2013 | Predicting student failure at school using genetic programming and different data mining approaches with high dimensional and imbalanced data
Carlos Márquez-Vera, Alberto Cano 0001, Cristóbal Romero 0001, Sebastián Ventura |
Appl. Intell. | 4 |
| 2013 | Grammar-based multi-objective algorithms for mining association rules
José María Luna, José Raúl Romero, Sebastián Ventura |
Data Knowl. Eng. | 3 |
| 2013 | Association rule mining using genetic programming to provide feedback to instructors from multiple-choice quiz dataabstractAbstract This paper proposes the application of association rule mining to improve quizzes and courses. First, the paper shows how to preprocess quiz data and how to create several data matrices for use in the process of knowledge discovery. Next, the proposed algorithm that uses grammar‐guided genetic programming is described and compared with both classical and recent soft‐computing association rule mining algorithms. Then, different objective and subjective rule evaluation measures are used to select the most interesting and useful rules. Experiments have been carried out by using real data of university students enrolled on an artificial intelligence practice Moodle's course on the CLIPS programming language. Some examples of these rules are shown, together with the feedback that they provide to instructors making decisions about how to improve quizzes and courses. Finally, starting with the information provided by the rules, the CLIPS quiz and course have been updated. These innovations have been evaluated by comparing the performance achieved by students before and after applying the changes using one control group and two different experimental groups. Cristóbal Romero 0001, Amelia Zafra, José María Luna, Sebastián Ventura |
Expert Syst. J. Knowl. Eng. | 4 |
| 2013 | An interpretable classification rule mining algorithm
Alberto Cano 0001, Amelia Zafra, Sebastián Ventura |
Inf. Sci. | 3 |
| 2013 | HyDR-MI: A hybrid algorithm to reduce dimensionality in multiple instance learning
Amelia Zafra, Mykola Pechenizkiy, Sebastián Ventura |
Inf. Sci. | 3 |
| 2013 | Parallel multi-objective Ant Programming for classification using GPUs
Alberto Cano 0001, Juan Luis Olmo, Sebastián Ventura |
J. Parallel Distributed Comput. | 3 |
| 2013 | DRAL: a tool for discovering relevant e-activities for learners
Amelia Zafra, Cristóbal Romero 0001, Sebastián Ventura |
Knowl. Inf. Syst. | 3 |
| 2013 | Weighted Data Gravitation Classification for Standard and Imbalanced DataabstractGravitation is a fundamental interaction whose concept and effects applied to data classification become a novel data classification technique. The simple principle of data gravitation classification (DGC) is to classify data samples by comparing the gravitation between different classes. However, the calculation of gravitation is not a trivial problem due to the different relevance of data attributes for distance computation, the presence of noisy or irrelevant attributes, and the class imbalance problem. This paper presents a gravitation-based classification algorithm which improves previous gravitation models and overcomes some of their issues. The proposed algorithm, called DGC+, employs a matrix of weights to describe the importance of each attribute in the classification of each class, which is used to weight the distance between data samples. It improves the classification performance by considering both global and local data information, especially in decision boundaries. The proposal is evaluated and compared to other well-known instance-based classification techniques, on 35 standard and 44 imbalanced data sets. The results obtained from these experiments show the great performance of the proposed gravitation model, and they are validated using several nonparametric statistical tests. Alberto Cano 0001, Amelia Zafra, Sebastián Ventura |
IEEE Trans. Cybern. | 3 |
| 2013 | High performance evaluation of evolutionary-mined association rules on GPUs
Alberto Cano 0001, José María Luna, Sebastián Ventura |
J. Supercomput. | 3 |
| 2012 | Classification via clustering for predicting final marks starting from the student participation in Forums
Manuel Ignacio López, Cristóbal Romero 0001, Sebastián Ventura, José María Luna |
EDM | 3 |
| 2012 | Meta-learning Approach for Automatic Parameter Tuning: A case of study with educational datasets
María De Mar Molina, Cristóbal Romero 0001, Sebastián Ventura, José María Luna |
EDM | 3 |
| 2012 | Multi-Objective Ant Programming for Mining Classification Rules
Juan Luis Olmo, José Raúl Romero, Sebastián Ventura |
EuroGP | 3 |
| 2012 | VisualJCLEC: A visual framework for evolutionary computationabstractThis paper presents VisualJCLEC, a visual framework based on JCLEC for Evolutionary Computing. In order to have a high degree of adaptability, the architecture and pattern design followed are focused on enhancing the f exibility and scalability. For illustrative purposes, a case study of an optimization classical problem (the knapsack problem) using this framework is presented, as well as some guidelines on how to add new elements to the environment by means of CDL descriptors. Juan Ignacio Jaen, José Raúl Romero, Sebastián Ventura |
ISDA | 3 |
| 2012 | A genetic programming free-parameter algorithm for mining association rulesabstractThis paper presents a free-parameter grammar-guided genetic programming algorithm for mining association rules. This algorithm uses a contex-free grammar to represent individuals, encoding the solutions in a tree-shape conformant to the grammar, so they are more expressive and flexible. The algorithm here presented has the advantages of using evolutionary algorithms for mining association rules, and it also solves the problem of tuning the huge number of parameters required by these algorithms. The main feature of this algorithm is the small number of parameters required, providing the possibility of discovering association rules in an easy way for non-expert users. We compare our approach to existing evolutionary and exhaustive search algorithms, obtaining important results and overcoming the drawbacks of both exhaustive search and evolutionary algorithms. The experimental stage reveals that this approach discovers frequent and reliable rules without a parameter tuning. José María Luna, José Raúl Romero, Cristóbal Romero 0001, Sebastián Ventura |
ISDA | 4 |
| 2012 | Binary and multiclass imbalanced classification using multi-objective ant programmingabstractClassification in imbalanced domains is a challenging task, since most of its real domain applications present skewed distributions of data. However, there are still some open issues in this kind of problem. This paper presents a multi-objective grammar-based ant programming algorithm for imbalanced classification, capable of addressing this task from both the binary and multiclass sides, unlike most of the solutions presented so far. We carry out two experimental studies comparing our algorithm against binary and multiclass solutions, demonstrating that it achieves an excellent performance for both binary and multiclass imbalanced data sets. Juan Luis Olmo, Alberto Cano 0001, José Raúl Romero, Sebastián Ventura |
ISDA | 4 |
| 2012 | Learning similarity metric to improve the performance of lazy multi-label ranking algorithmsabstractThe definition of similarity metrics is one of the most important tasks in the development of nearest neighbours and instance based learning methods. Furthermore, the performance of lazy algorithms can be significantly improved with the use of an appropriate weight vector. In the last years, the learning from multi-label data has attracted significant attention from a lot of researchers, motivated from an increasing number of modern applications that contain this type of data. This paper presents a new method for feature weighting, defining a similarity metric as heuristic to estimate the feature weights, and improving the performance of lazy multi-label ranking algorithms. The experimental stage shows the effectiveness of our proposal. Oscar Gabriel Reyes Pupo, Carlos Morell 0001, Sebastián Ventura |
ISDA | 3 |
| 2012 | ReliefF-MI: An extension of ReliefF to multiple instance learning
Amelia Zafra, Mykola Pechenizkiy, Sebastián Ventura |
Neurocomputing | 3 |
| 2012 | Design and behavior study of a grammar-guided genetic programming algorithm for mining association rules
José María Luna, José Raúl Romero, Sebastián Ventura |
Knowl. Inf. Syst. | 3 |
| 2012 | Speeding up the evaluation phase of GP classification algorithms on GPUs
Alberto Cano 0001, Amelia Zafra, Sebastián Ventura |
Soft Comput. | 3 |
| 2012 | Classification rule mining using ant programming guided by grammar with multiple Pareto fronts
Juan Luis Olmo, José Raúl Romero, Sebastián Ventura |
Soft Comput. | 3 |
| 2012 | Multi-objective approach based on grammar-guided genetic programming for solving multiple instance problems
Amelia Zafra, Sebastián Ventura |
Soft Comput. | 2 |
| 2011 | Predicting School Failure Using Data Mining
Carlos Márquez-Vera, Cristóbal Romero 0001, Sebastián Ventura |
EDM | 3 |
| 2011 | A Java Desktop Tool for Mining Moodle Data
Rafael Pedraza Perez, Cristóbal Romero 0001, Sebastián Ventura |
EDM | 3 |
| 2011 | An EP algorithm for learning highly interpretable classifiersabstractThis paper introduces an Evolutionary Programming algorithm for solving classification problems using highly interpretable IF-THEN classification rules. It is an algorithm aimed to maximize the comprehensibility of the classifier by minimizing the number of rules and employing only relevant attributes. The proposal is evaluated and compared to other 5 well-known classification techniques over 18 datasets. The results obtained from the experiments show its competitive accuracy and the significantly better interpretability of the classifiers provided in terms of number of rules, number of conditions and a complexity metric. Alberto Cano 0001, Amelia Zafra, Sebastián Ventura |
ISDA | 3 |
| 2011 | Association rule mining using a multi-objective grammar-based ant programming algorithmabstractThis paper presents a method for extracting association rules by means of a multi-objective grammar guided ant programming algorithm. Solution construction is guided by a context-free grammar specifically suited for association rule mining, which defines the search space of all possible expressions or programs. Evaluation of individuals is considered from a Pareto-based point of view, measuring support and confidence of rules mined, and assigning them a ranking fitness. The proposed algorithm is verified over 10 varied data sets and compared to other association rule mining algorithms from several paradigms such as exhaustive search, genetic algorithms and genetic programming, showing that ant programming is a good technique at addressing the association task of data mining as well. Juan Luis Olmo, José María Luna, José Raúl Romero, Sebastián Ventura |
ISDA | 4 |
| 2011 | Multiple instance learning for classifying students in learning management systems
Amelia Zafra, Cristóbal Romero 0001, Sebastián Ventura |
Expert Syst. Appl. | 3 |
| 2011 | Using Ant Programming Guided by Grammar for Building Rule-Based ClassifiersabstractThe extraction of comprehensible knowledge is one of the major challenges in many domains. In this paper, an ant programming (AP) framework, which is capable of mining classification rules easily comprehensible by humans, and, therefore, capable of supporting expert-domain decisions, is presented. The algorithm proposed, called grammar based ant programming (GBAP), is the first AP algorithm developed for the extraction of classification rules, and it is guided by a context-free grammar that ensures the creation of new valid individuals. To compute the transition probability of each available movement, this new model introduces the use of two complementary heuristic functions, instead of just one, as typical ant-based algorithms do. The selection of a consequent for each rule mined and the selection of the rules that make up the classifier are based on the use of a niching approach. The performance of GBAP is compared against other classification techniques on 18 varied data sets. Experimental results show that our approach produces comprehensible rules and competitive or better accuracy values than those achieved by the other classification algorithms compared with it. Juan Luis Olmo, José Raúl Romero, Sebastián Ventura |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2011 | Preface to the special issue on data mining for personalised educational systems
Cristóbal Romero 0001, Sebastián Ventura |
User Model. User Adapt. Interact. | 2 |
| 2010 | G3PARM: A Grammar Guided Genetic Programming algorithm for mining association rulesabstractThis paper presents the G3PARM algorithm for mining representative association rules. G3PARM is an evolutionary algorithm that uses G3P (Grammar Guided Genetic Programming) and an auxiliary population made up of its best individuals who will then act as parents for the next generation. Due to the nature of G3P, the G3PARM algorithm allows us to obtain valid individuals by defining them through a context-free grammar and, furthermore, this algorithm is generic with respect to data type. We compare our algorithm to two multiobjective algorithms frequently used in literature and known as NSGA2 (Non dominated Sort Genetic Algorithm) and SPEA2 (Strength Pareto Evolutionary Algorithm) and demonstrate the efficiency of our algorithm in terms of running-time, coverage and average support, providing the user with high representative rules. José María Luna, José Raúl Romero, Sebastián Ventura |
IEEE Congress on Evolutionary Computation | 3 |
| 2010 | A grammar based Ant Programming algorithm for mining classification rulesabstractThis paper focuses on the application of a new ACO-based automatic programming algorithm to the classification task of data mining. This new model, called GBAP algorithm, is based on a context-free grammar that properly guides the creation of new valid individuals. Moreover, its most differentiating factors, such as the use of two complementary heuristic measures for every transition rule, as well as the way it assigns a consequent and evaluates the extracted rules, are also discussed. These features enhance the final rule compilation from the output classifier. The performance of the proposed algorithm is evaluated and compared against other top algorithms, and the results obtained over 17 diverse data sets show that our approach reaches pretty competitive and even better accuracy values than those resulting from the other algorithms considered in the experimentation. Juan Luis Olmo, José Raúl Romero, Sebastián Ventura |
IEEE Congress on Evolutionary Computation | 3 |
| 2010 | Mining Rare Association Rules from e-Learning Data
Cristóbal Romero 0001, José Raúl Romero, José María Luna, Sebastián Ventura |
EDM | 4 |
| 2010 | Class Association Rules Mining from Students' Test Data
Cristóbal Romero 0001, Sebastián Ventura, Ekaterina Vasilyeva, Mykola Pechenizkiy |
EDM | 2 |
| 2010 | Grammar guided genetic programming for multiple instance learning: an experimental studyabstractThis paper introduces a new Grammar-Guided Genetic Programming algorithm for solving multi-instance Learning problems. This algorithm, called G3P-MI, is evaluated and compared with other Multi-Instance classification techniques on different application domains. Computational experiments show that the G3P-MI often obtains consistently better results than other algorithms in terms of accuracy, sensitivity and specificity. Moreover, it adds comprehensibility and clarity into the knowledge discovery process, expressing the information in the form of IF-THEN rules. Our results confirm that evolutionary algorithms are appropriate for dealing with multi-instance learning problems. Amelia Zafra, Sebastián Ventura |
GECCO | 2 |
| 2010 | Web Usage Mining for Improving Students Performance in Learning Management Systems
Amelia Zafra, Sebastián Ventura |
IEA/AIE (3) | 2 |
| 2010 | A TDIDT technique for multi-label classificationabstractThere are numerous problems of increasing significance where a pattern can have several classes simultaneously associated. This kind of problems, usually called multi-label problems, should be tackled with specific techniques in order to generate models more accurate than those obtained with classical classification algorithms. This work presents the adaptation of the J48 algorithm to multi-label classification. The developed algorithm allows the generation of interpretable models and has been tested over several datasets and experiments show that it has a performance which is similar to other multi-label tree-based approaches being specially suitable to be used as base-classifier in an ensemble. Eva Lucrecia Gibaja Galindo, Manuel Victoriano, José Luis Ávila-Jiménez, Sebastián Ventura |
ISDA | 4 |
| 2010 | An intruder detection approach based on infrequent rating pattern miningabstractThis work presents a novel proposal for incremental intruder detection in collaborative recommender systems. We explore the use of rare association rule mining to reveal the existence of a suspected raid of attackers that would alter the normal behaviour of a rating-based system. In this position paper we have extended our previous G3PARM algorithm, which has already proven to serve as a solid method for extracting frequent association rules. G3PARM is an evolutionary algorithm that uses G3P (Grammar Guided Genetic Programming), which provides expressiveness and flexibility enough to adapt and apply the base context-free grammar to each specific problem or domain. We fully outline, moreover, the complete exploration and detection model, which includes some further post-analysis steps. Finally, as a proof of concept, we validate the scalability, efficiency and accuracy of our proposal showing the results obtained when different malicious intruders want to attack an on line recommender system. José María Luna, Aurora Ramírez 0001, José Raúl Romero, Sebastián Ventura |
ISDA | 4 |
| 2010 | Feature selection is the ReliefF for multiple instance learningabstractDimensionality reduction and feature selection in particular are known to be of a great help for making supervised learning more effective and efficient. Many different feature selection techniques have been proposed for the traditional settings, where each instance is expected to have a label. In multiple instance learning (MIL) each example or bag consists of a variable set of instances, and the label is known for the bag as a whole, but not for the individual instances it consists of. Therefore, utilizing class labels for feature selection in MIL is not that straightforward and traditional approaches for feature selection are not directly applicable. This paper proposes a filter feature selection approach based on the ReliefF technique. It allows any previously designed MIL method to benefit from our feature selection approach, which helps to cope with the curse of dimensionality. Experimental results show the effectiveness of the proposed approach in MIL - different MIL algorithms tend to perform better when applied after the dimensionality reduction. Amelia Zafra, Mykola Pechenizkiy, Sebastián Ventura |
ISDA | 3 |
| 2010 | G3P-MI: A genetic programming algorithm for multiple instance learning
Amelia Zafra, Sebastián Ventura |
Inf. Sci. | 2 |
| 2010 | A Survey on the Application of Genetic Programming to ClassificationabstractClassification is one of the most researched questions in machine learning and data mining. A wide range of real problems have been stated as classification problems, for example credit scoring, bankruptcy prediction, medical diagnosis, pattern recognition, text categorization, software quality assessment, and many more. The use of evolutionary algorithms for training classifiers has been studied in the past few decades. Genetic programming (GP) is a flexible and powerful evolutionary technique with some features that can be very valuable and suitable for the evolution of classifiers. This paper surveys existing literature about the application of genetic programming to classification, to show the different ways in which this evolutionary algorithm can help in the construction of accurate and reliable classifiers. Pedro G. Espejo, Sebastián Ventura, Francisco Herrera |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2010 | Educational Data Mining: A Review of the State of the ArtabstractEducational data mining (EDM) is an emerging interdisciplinary research area that deals with the development of methods to explore data originating in an educational context. EDM uses computational approaches to analyze educational data in order to study educational questions. This paper surveys the most relevant studies carried out in this field to date. First, it introduces EDM and describes the different groups of user, types of educational environments, and the data they provide. It then goes on to list the most typical/common tasks in the educational environment that have been resolved through data-mining techniques, and finally, some of the most promising future lines of research are discussed. Cristóbal Romero 0001, Sebastián Ventura |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2009 | Collaborative Data Mining Tool for Education
Cristóbal Romero 0001, Sebastián Ventura, Enrique García 0001, Carlos de Castro, Miguel Gea Megías |
EDM | 2 |
| 2009 | Predicting Student Grades in Learning Management Systems with Multiple Instance Learning Genetic Programming
Amelia Zafra, Sebastián Ventura |
EDM | 2 |
| 2009 | A Niching Algorithm to Learn Discriminant Functions with Multi-Label Patterns
José Luis Ávila-Jiménez, Eva Lucrecia Gibaja Galindo, Amelia Zafra, Sebastián Ventura |
IDEAL | 4 |
| 2009 | Predicting Academic Achievement Using Multiple Instance Genetic ProgrammingabstractThe ability to predict a student's performance could be useful in a great number of different ways associated with university-level learning. In this paper, a grammar guided genetic programming algorithm, G3P-MI, has been applied to predict if the student will fail or pass a certain course and identifies activities to promote learning in a positive or negative way from the perspective of MIL. Computational experiments compare our proposal with the most popular techniques of multiple instance learning (MIL). Results show that G3P-MI achieves better performance with more accurate models and a better trade-off between such contradictory metrics as sensitivity and specificity. Moreover, it adds comprehensibility to the knowledge discovered and finds interesting relationships that correlate certain tasks and the time devoted to solving exercises with the final marks obtained in the course. Amelia Zafra, Cristóbal Romero 0001, Sebastián Ventura |
ISDA | 3 |
| 2009 | Evaluating Web Based Instructional Models Using Association Rule Mining
Enrique García 0001, Cristóbal Romero 0001, Sebastián Ventura, Carlos de Castro |
UMAP | 3 |
| 2009 | Evolutionary algorithms for subgroup discovery in e-learning: A practical application using Moodle data
Cristóbal Romero 0001, Pedro González 0001, Sebastián Ventura, María José del Jesus, Francisco Herrera |
Expert Syst. Appl. | 3 |
| 2009 | Multi-instance genetic programming for web index recommendation
Amelia Zafra, Cristóbal Romero 0001, Sebastián Ventura, Enrique Herrera-Viedma |
Expert Syst. Appl. | 3 |
| 2009 | KEEL: a software tool to assess evolutionary algorithms for data mining problems
Jesús Alcalá-Fdez, Luciano Sánchez, Salvador García 0001, María José del Jesus, Sebastián Ventura, Josep Maria Garrell i Guiu, José Otero, Cristóbal Romero 0001, Jaume Bacardit, Víctor Manuel Rivas Santos, Juan Carlos Fernández 0001, Francisco Herrera |
Soft Comput. | 5 |
| 2009 | An architecture for making recommendations to courseware authors using association rule mining and collaborative filtering
Enrique García 0001, Cristóbal Romero 0001, Sebastián Ventura, Carlos de Castro |
User Model. User Adapt. Interact. | 3 |
| 2008 | Mining and Visualizing Visited Trails in Web-Based Educational Systems
Cristóbal Romero 0001, Sergio Gutiérrez Santos, Manuel Freire-Morán, Sebastián Ventura |
EDM | 4 |
| 2008 | Data Mining Algorithms to Classify Students
Cristóbal Romero 0001, Sebastián Ventura, Pedro G. Espejo, César Hervás-Martínez |
EDM | 2 |
| 2008 | Analyzing Rule Evaluation Measures with Educational Datasets: A Framework to Help the Teacher
Sebastián Ventura, Cristóbal Romero 0001, César Hervás-Martínez |
EDM | 1 |
| 2008 | Multiple Instance Learning with MultiObjective Genetic Programming for Web MiningabstractThis paper introduces a multiobjective grammar based genetic programming algorithm to solve a Web Mining problem from multiple instance perspective. This algorithm, called MOG3P-MI, is evaluated and compared with other available algorithms which extend a well-known neighborhood-based algorithm (k-nearest neighbour algorithm) and with a mono objective version of grammar guided genetic programming G3P-MI. Computational experiments show that, the MOG3PMI algorithm obtains the best results, solves problems of k-nearest neighbour algorithms, such as sparsity and scalability, adds comprehensibility and clarity in the knowledge discovery process and overcomes the results of monoobjective version. Amelia Zafra, Eva Lucrecia Gibaja Galindo, Sebastián Ventura |
HIS | 3 |
| 2008 | JCLEC: a Java framework for evolutionary computation
Sebastián Ventura, Cristóbal Romero 0001, Amelia Zafra, José A. Delgado-Osuna, César Hervás-Martínez |
Soft Comput. | 1 |
| 2007 | Multi-objective Genetic Programming for Multiple Instance Learning
Amelia Zafra, Sebastián Ventura |
ECML | 2 |
| 2007 | Personalized Links Recommendation Based on Data Mining in Adaptive Educational Hypermedia Systems
Cristóbal Romero 0001, Sebastián Ventura, José A. Delgado-Osuna, Paul De Bra |
EC-TEL | 2 |
| 2007 | Educational data mining: A survey from 1995 to 2005
Cristóbal Romero 0001, Sebastián Ventura |
Expert Syst. Appl. | 2 |
| 2006 | Using Rules Discovery for the Continuous Improvement of e-Learning Courses
Enrique García 0001, Cristóbal Romero 0001, Sebastián Ventura, Carlos de Castro |
IDEAL | 3 |
| 2006 | Evolutionary Product-Unit Neural Networks for Classification
Francisco J. Martínez-Estudillo, César Hervás-Martínez, Pedro Antonio Gutiérrez, Alfonso C. Martínez-Estudillo, Sebastián Ventura |
IDEAL | 5 |
| 2006 | Web-based adaptive training simulator system for cardiac life support
Cristóbal Romero 0001, Sebastián Ventura, Eva Lucrecia Gibaja Galindo, César Hervás-Martínez, Francisco Romero Morales |
Artif. Intell. Medicine | 2 |
| 2004 | Knowledge Discovery with Genetic Programming for Providing Feedback to Courseware Authors
Cristóbal Romero 0001, Sebastián Ventura, Paul De Bra |
User Model. User Adapt. Interact. | 2 |
| 2001 | A two steps method: non linear regression and pruning neural network for analyzing multicomponent mixtures
César Hervás-Martínez, José Antonio Martinez Heras, Sebastián Ventura, Manuel Silva 0002 |
ESANN | 3 |