Sebastián Ventura

dblp:00/4717 · also Sebastián Ventura Soto · DBLP profile ↗
← Back
31ranked-venue papers in the field
0as first author
9since 2021 · last 2025
0000-0003-4216-6378ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 13Data Mining & Knowledge Discovery · 11Other / Interdisciplinary · 3Information Retrieval & Web Search · 2Database Systems & Data Management · 1Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2025 Enhancing Medical Diagnosis with Instance Hardness-Guided Multi-Level Cross-Validation for Imbalanced Learning
abstract
Cross-validation is a critical component for robust machine learning evaluation. In imbalanced learning, stratified cross-validation is commonly recommended to preserve class distribution. However, it neglects the underlying distribution of instance hardness, which can introduce distribution shifts between training and testing folds, ultimately compromising the validity of performance evaluation. This paper proposes a stratified cross-validation informed by the hardness distribution for robust imbalanced medical diagnosis. The proposed multi-level cross-validation (MLCV) retains jointly the class distribution and instance hardness, maintaining equivalent distribution of hardness levels across folds. This strategy enables the model to encounter a more realistic version of the medical data for a reliable performance evaluation. Experimental work demonstrates that the hardness distribution shift exists; the (MLCV) not only enhances classification performance in imbalanced medical data but also improves the results of balancing methods, as measured by classification performance indicators such as recall, precision, and F1-measure.
Mabrouka Salmi, Dalia Atif, Sebastián Ventura
DSAA3
2024 Improving hyper-parameter self-tuning for data streams by adapting an evolutionary approach
Antonio R. Moya, Bruno M. Veloso, João Gama 0001, Sebastián Ventura
Data Min. Knowl. Discov.4
2024 Data heterogeneity's impact on the performance of frequent itemset mining algorithms
abstract
Frequent itemset mining (FIM) is a widely used task that extracts frequently occurring itemsets from data. Plenty of deterministic algorithms are available for this daunting task. However, experimental studies have not considered that data heterogeneity significantly impacts the algorithms' performance, giving rise to unfair comparisons and biased conclusions. This paper seeks to advance by comparing cutting-edge algorithms using various frequency thresholds, considering the resulting data heterogeneity. An extensive experimental study is carried out, including the number of itemsets mined per second as the performance quality measure to compare algorithms. The experiments include defining eight metrics to quantify data heterogeneity, and their values vary the algorithms' performance. The results revealed that some techniques (hypercube decomposition and k-items machine) are essential to achieve excellent performance on any dataset, and most algorithms behave similarly well when they include those techniques. As a final important point, different threshold values produce dissimilar data subsets (data heterogeneity is not an immutable data characteristic), so a previous study on the database characteristics with a few minimum support thresholds could be beneficial to select the best-suited FIM algorithm beforehand.
Antonio Manuel Trasierras, José María Luna, Philippe Fournier-Viger, Sebastián Ventura
Inf. Sci.4
2023 Efficient mining of top-k high utility itemsets through genetic algorithms
José María Luna, R. Uday Kiran, Philippe Fournier-Viger, Sebastián Ventura
Inf. Sci.4
2023 Eight years of AutoML: categorisation, review and trends
abstract
Abstract Knowledge extraction through machine learning techniques has been successfully applied in a large number of application domains. However, apart from the required technical knowledge and background in the application domain, it usually involves a number of time-consuming and repetitive steps. Automated machine learning (AutoML) emerged in 2014 as an attempt to mitigate these issues, making machine learning methods more practicable to both data scientists and domain experts. AutoML is a broad area encompassing a wide range of approaches aimed at addressing a diversity of tasks over the different phases of the knowledge discovery process being automated with specific techniques. To provide a big picture of the whole area, we have conducted a systematic literature review based on a proposed taxonomy that permits categorising 447 primary studies selected from a search of 31,048 papers. This review performs an extensive and rigorous analysis of the AutoML field, scrutinising how the primary studies have addressed the dimensions of the taxonomy, and identifying any gaps that remain unexplored as well as potential future trends. The analysis of these studies has yielded some intriguing findings. For instance, we have observed a significant growth in the number of publications since 2018. Additionally, it is noteworthy that the algorithm selection problem has gradually been superseded by the challenge of workflow composition, which automates more than one phase of the knowledge discovery process simultaneously. Of all the tasks in AutoML, the growth of neural architecture search is particularly noticeable.
Rafael Barbudo, Sebastián Ventura, José Raúl Romero
Knowl. Inf. Syst.2
2022 Improving the understanding of cancer in a descriptive way: An emerging pattern mining-based approach
abstract
This paper presents an approach based on emerging pattern mining to analyse cancer through genomic data. Unlike existing approaches, mainly focused on predictive purposes, the proposal aims to improve the understanding of cancer descriptively, not requiring either any prior knowledge or hypothesis to be validated. Additionally, it enables to consider high-order relationships, so not only essential genes related to the disease are considered, but also the combined effect of various secondary genes that can influence different pathways directly or indirectly related to the disease. The prime hypothesis is that splitting genomic cancer data into two subsets, that is, cases and controls, will allow us to determine which genes, and their expressions, are associated with different cancer types. The possibilities of the proposal are demonstrated by analyzing RNA-Seq data for six different types of cancer: breast, colon, lung, thyroid, prostate, and kidney. Some of the extracted insights were already described in the related literature as good cancer bio-markers, while others have not been described yet mainly due to existing techniques are biased by prior knowledge provided by biological databases.
Antonio Manuel Trasierras, José María Luna, Sebastián Ventura
Int. J. Intell. Syst.3
2022 Modeling and predicting students' engagement behaviors using mixture Markov models
Rabia Maqsood, Paolo Ceravolo, Cristóbal Romero 0001, Sebastián Ventura
Knowl. Inf. Syst.4
2021 Introduction to the special issue on Methods and applications in the analysis of social data in healthcare
Alejandro Rodríguez González, Sebastián Ventura, Paolo Soda, Jesualdo Tomás Fernández-Breis
Inf. Process. Manag.2
2021 Mining local periodic patterns in a discrete sequence
Philippe Fournier-Viger, R. Uday Kiran, Sebastián Ventura, José María Luna
Inf. Sci.4
2020 Fast Convergence of Competitive Spiking Neural Networks with Sample-Based Weight Initialization
Paolo Gabriel Cachi, Sebastián Ventura, Krzysztof J. Cios
IPMU (3)2
2020 Heuristics for interesting class association rule mining a colorectal cancer database
José A. Delgado-Osuna, Carlos García-Martínez, Jose Gómez Barbadillo, Sebastián Ventura
Inf. Process. Manag.4
2019 Virtual learning environment to predict withdrawal by leveraging deep learning
abstract
The current evolution in multidisciplinary learning analytics research poses significant challenges for the exploitation of behavior analysis by fusing data streams toward advanced decision-making. The identification of students that are at risk of withdrawals in higher education is connected to numerous educational policies, to enhance their competencies and skills through timely interventions by academia. Predicting student performance is a vital decision-making problem including data from various environment modules that can be fused into a homogenous vector to ascertain decision-making. This research study exploits a temporal sequential classification problem to predict early withdrawal of students, by tapping the power of actionable smart data in the form of students' interactional activities with the online educational system, using the freely available Open University Learning Analytics data set by employing deep long short-term memory (LSTM) model. The deployed LSTM model outperforms baseline logistic regression and artificial neural networks by 10.31% and 6.48% respectively with 97.25% learning accuracy, 92.79% precision, and 85.92% recall.
Saeed-Ul Hassan, Hajra Waheed, Naif R. Aljohani, Mohsen Ali, Sebastián Ventura, Francisco Herrera
Int. J. Intell. Syst.5
2018 Interactive multi-objective evolutionary optimization of software architectures
Aurora Ramírez 0001, José Raúl Romero, Sebastián Ventura
Inf. Sci.3
2018 Evolutionary Strategy to Perform Batch-Mode Active Learning on Multi-Label Data
abstract
Multi-label learning has become an important area of research owing to the increasing number of real-world problems that contain multi-label data. Data labeling is an expensive process that requires expert handling. The annotation of multi-label data is laborious since a human expert needs to consider the presence/absence of each possible label. Consequently, numerous modern multi-label problems may involve a small number of labeled examples and plentiful unlabeled examples simultaneously. Active learning methods allow us to induce better classifiers by selecting the most useful unlabeled data, thus considerably reducing the labeling effort and the cost of training an accurate model. Batch-mode active learning methods focus on selecting a set of unlabeled examples in each iteration in such a way that the selected examples are informative and as diverse as possible. This article presents a strategy to perform batch-mode active learning on multi-label data. The batch-mode active learning is formulated as a multi-objective problem, and it is solved by means of an evolutionary algorithm. Extensive experiments were conducted in a large collection of datasets, and the experimental results confirmed the effectiveness of our proposal for better batch-mode multi-label active learning.
Oscar Gabriel Reyes Pupo, Sebastián Ventura
ACM Trans. Intell. Syst. Technol.2
2017 Extremely high-dimensional optimization with MapReduce: Scaling functions and algorithm
Alberto Cano 0001, Carlos García-Martínez, Sebastián Ventura
Inf. Sci.3
2017 Multi-target support vector regression via correlation regressor chains
Gabriella Melki, Alberto Cano 0001, Vojislav Kecman, Sebastián Ventura
Inf. Sci.4
2016 Subgroup discovery on big data: Pruning the search space on exhaustive search algorithms
abstract
Subgroup Discovery is a broadly applicable supervised local pattern mining method to search relations between different properties with respect to a target variable. With the exponential growth in data storage, the massive data gathered has hampered the performance of current techniques. In this regard, our aim is to propose two new algorithms to discover subgroups on Big Data by using MapReduce. Apache Spark was used to tackle the Big Data requirements. The experimental study includes more than 50 large datasets and a set of efficient algorithms. Search spaces bigger than 1.276 · 1015subgroups are used. The experimental study reveals the alluring results in efficiency when optimistic estimates are considered, as well as demonstrating the usefulness of using Apache Spark to tackle Big Data.
Francisco Padillo, José María Luna, Sebastián Ventura
IEEE BigData3
2016 LAIM discretization for multi-label data
Alberto Cano 0001, José María Luna, Eva Lucrecia Gibaja Galindo, Sebastián Ventura
Inf. Sci.4
2016 Discovering useful patterns from multiple instance data
José María Luna, Alberto Cano 0001, Virgilijus Sakalauskas, Sebastián Ventura
Inf. Sci.4
2016 Effective lazy learning algorithm based on a data gravitation model for multi-label learning
Oscar Gabriel Reyes Pupo, Carlos Morell 0001, Sebastián Ventura
Inf. Sci.3
2016 Mining exceptional relationships with grammar-guided genetic programming
José María Luna, Mykola Pechenizkiy, Sebastián Ventura
Knowl. Inf. Syst.3
2015 An approach for the evolutionary discovery of software architectures
Aurora Ramírez 0001, José Raúl Romero, Sebastián Ventura
Inf. Sci.3
2015 Speeding up multiple instance learning classification rules on GPUs
Alberto Cano 0001, Amelia Zafra, Sebastián Ventura
Knowl. Inf. Syst.3
2014 On the adaptability of G3PARM to the extraction of rare association rules
José María Luna, José Raúl Romero, Sebastián Ventura
Knowl. Inf. Syst.3
2013 Grammar-based multi-objective algorithms for mining association rules
José María Luna, José Raúl Romero, Sebastián Ventura
Data Knowl. Eng.3
2013 An interpretable classification rule mining algorithm
Alberto Cano 0001, Amelia Zafra, Sebastián Ventura
Inf. Sci.3
2013 HyDR-MI: A hybrid algorithm to reduce dimensionality in multiple instance learning
Amelia Zafra, Mykola Pechenizkiy, Sebastián Ventura
Inf. Sci.3
2013 DRAL: a tool for discovering relevant e-activities for learners
Amelia Zafra, Cristóbal Romero 0001, Sebastián Ventura
Knowl. Inf. Syst.3
2012 Design and behavior study of a grammar-guided genetic programming algorithm for mining association rules
José María Luna, José Raúl Romero, Sebastián Ventura
Knowl. Inf. Syst.3
2010 G3P-MI: A genetic programming algorithm for multiple instance learning
Amelia Zafra, Sebastián Ventura
Inf. Sci.2
2007 Multi-objective Genetic Programming for Multiple Instance Learning
Amelia Zafra, Sebastián Ventura
ECML2