VLDB 2026 Research / reviewers in the wild / expert
Verónica Bolón-Canedo
dblp:43/8485
· DBLP profile ↗
98ranked-venue papers
31as first author
41since 2021 · last 2026
0000-0002-0524-6427ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 74 · 27 first-author · 29 since 2021Databases, data management, data science and information retrieval · 15 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Systems, architecture and hardware · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SMOTE k-out: Enhancing Class Separability through Outer Synthetic SamplingabstractOversampling techniques are commonly used to address class imbalance in supervised classification, with SMOTE being a popular approach.However, traditional SMOTE generates synthetic samples within the neighbourhood of minority instances, which can increase data complexity and hinder class separability.This work proposes SMOTE k-out, which creates synthetic samples outside the local neighbourhood to increase minority class sparsity.This aims to reduce overfitting and mitigate the impact of noise, thereby improving the definition of the decision boundary.Experiments on multiple imbalanced datasets demonstrate that SMOTE k-out consistently reduces complexity and achieves higher accuracy and F-measure, particularly with SVM and LDA classifiers. Verónica Bolón-Canedo, José Luis Morillo-Salas, Laura Moran-Fernandez, Amparo Alonso-Betanzos |
ESANN | 1 |
| 2026 | Information-Theoretic Unsupervised Feature Selection for High-Dimensional Spatial DataabstractHigh-dimensional unlabelled datasets present significant challenges for efficient analysis, storage and interpretation.Unsupervised feature selection offers a way to retain the most informative variables while discarding redundant or uninformative ones, enabling more scalable processing.We introduce a spatially aware, unsupervised method that uses information theoretic criteria to identify informative variables while limiting redundancy, producing compact and spatially dispersed subsets of features.Our approach avoids dependence on labelled data or modelspecific wrappers, making it suitable for large unstructured datasets.Experiments on MNIST and EMNIST datasets, including high-resolution upscaled versions, show that the selected features preserve both discriminative structure and reconstruction quality better than chosen supervised and unsupervised baselines, demonstrating the effectiveness of entropy and mutual information coupling in unlabelled high-dimensional settings. Samuel Suárez-Marcote, Abhijeet Vishwasrao, Ricardo Vinuesa, Laura Moran-Fernandez, Verónica Bolón-Canedo |
ESANN | 5 |
| 2026 | Optimising Image Feature Extraction and Selection: A Comprehensive Review With Spark Case StudiesabstractABSTRACT As benchmark image datasets expand in sample size and feature complexity, the challenge of managing increased dimensionality becomes apparent. Contrary to the expectation that more features equate to enhanced information and improved outcomes, the curse of dimensionality often hampers performance. This paper reviews existing literature on filter feature selection techniques applied to image features, highlighting their use in both classical and deep‐learning‐based feature extraction methods. Building on these findings, this study proposes a scalable approach for image feature extraction and selection using Big Data technologies, specifically Apache Spark, to efficiently process large and high‐dimensional datasets. The proposed framework integrates filter‐based feature selection methods within a distributed environment to evaluate their effectiveness in image analysis tasks. Several experiments were performed to compare the results using feature selection techniques with various reduction percentages. Results show that significant feature reduction can be achieved without compromising classification accuracy, demonstrating the potential of Spark‐based distributed processing for large‐scale image analytics. J. Guzmán Figueira-domínguez, Beatriz Remeseiro, Verónica Bolón-Canedo |
Expert Syst. J. Knowl. Eng. | 3 |
| 2026 | Enhancing Classification Performance on Imbalanced Datasets Through Complexity-Guided Oversampling With SMOTEabstractABSTRACT Improving classification performance on imbalanced datasets remains a challenging problem in machine learning. Synthetic oversampling techniques such as SMOTE are widely used to address class imbalance; however, their random interpolation strategy often ignores structural data properties, which may affect classifier generalisation. This work proposes a set of SMOTE‐based strategies that guide the generation of synthetic samples in order to produce structurally simpler training datasets that are easier for classifiers to learn. The first strategy generates more dispersed (outer) synthetic samples to increase class separability with minimal computational overhead. The second and main contribution, SMOTE‐Complex , formulates synthetic sample selection as an explicit optimisation process that minimises measurable training dataset complexity. A clustering‐based variant, SMOTE‐Complex‐Clustering , reduces computational cost by restricting optimisation to feature subspaces while preserving most structural and predictive benefits. The underlying hypothesis is that reducing the structural complexity of the training data can lead to improved predictive behaviour. Extensive experiments on binary and multiclass datasets, using multiple classifiers and complementary evaluation metrics, provide empirical support for this hypothesis across diverse structural conditions. The results indicate moderate but stable improvements in structurally favourable scenarios—particularly binary and moderately complex problems—without systematic degradation in imbalance‐sensitive metrics, while the clustering‐based refinement offers a scalable trade‐off between optimisation strength and computational efficiency. José Luis Morillo-Salas, Verónica Bolón-Canedo, Laura Moran-Fernandez, Amparo Alonso-Betanzos |
Expert Syst. J. Knowl. Eng. | 2 |
| 2026 | A hybrid metaheuristics-Bayesian optimization framework with safe transfer learning for continuous spark tuningabstractTuning configuration parameters in distributed Big Data engines such as Apache Spark is a high-dimensional, workload-dependent problem with significant impact on performance and operational cost. We address this challenge with a hybrid optimization framework that integrates Iterated Local Search, Tabu Search, and locally embedded Bayesian Optimization guided by STL-PARN (safe transfer learning with pattern-adaptive robust neighborhoods). Historical executions are partitioned into a Nucleus of reliable neighbors and a Corona of exploratory configurations, ensuring relevance while mitigating negative transfer. The surrogate within the embedded Bayesian Optimization stage decouples performance prediction from uncertainty modeling, enabling parameter-free acquisition functions that self-adapt to diverse workloads. Experiments on a modernized HiBench suite across multiple input scales show consistent gains over state-of-the-art baselines in execution time, convergence, and cost efficiency. Overall, the results demonstrate the robustness and practical value of embedding Bayesian Optimization within a global metaheuristic loop for adaptive, cost-aware Spark tuning. All source code and datasets are publicly available, supporting reproducibility and operational efficiency in large-scale data processing. Mariano Garralda-Barrio, Carlos Eiras-Franco, Verónica Bolón-Canedo |
Future Gener. Comput. Syst. | 3 |
| 2026 | Improving restaurant recommendation transparency through feature selection
Roger Bagué-Masanés, Beatriz Remeseiro, Verónica Bolón-Canedo |
Knowl. Inf. Syst. | 3 |
| 2026 | Comparison of data set sample selection algorithms for data science: a systematic reviewabstractAbstract In the era of big data, selecting representative samples has become essential to mitigate overfitting, noise, and high computational cost in machine learning. This study systematically reviews the evolution of instance selection (IS) methods, highlighting the growing importance of instance hardness (IH) as a guiding criterion to improve training efficiency and model robustness. Through a comprehensive search in Scopus and Web of Science, fifty-five studies were identified and analyzed following strict inclusion and exclusion criteria. The reviewed works were classified according to their underlying rationale–error-based, geometric, heuristic, or explainability-driven–revealing that IH principles intersect these categories as a transversal perspective on data quality. Most studies focus on enhancing predictive accuracy (56%) and computational efficiency (36%), while bias reduction and privacy preservation remain secondary. Reported outcomes show significant dataset reductions (up to 97%) with minimal accuracy loss and, in some cases, notable performance gains (+32% accuracy, +67% improvement in MSE). Despite these advances, explicit references to IH are rare, though many methods implicitly rely on related metrics such as misclassification frequency or decision-boundary proximity. Overall, IS is gaining relevance across domains such as cybersecurity, biomedicine, and computer vision, yet the field still lacks standardized methodologies and benchmarking frameworks, underscoring the need for unified, IH-informed strategies for robust and generalizable instance selection. Alberto Fernandez-Sanchez, Marcos Gestal Pose, Verónica Bolón-Canedo, Julián Dorado, Alejandro Pazos |
Neural Comput. Appl. | 3 |
| 2025 | Do not get lost in projection: finding the right distance for meaningful UMAP embeddingsabstractDimensionality reduction techniques are essential for visualizing and analyzing high-dimensional data.This study explores the impact of distance measures on the performance of Uniform Manifold Approximation and Projection (UMAP), a widely used dimensionality reduction method.We evaluate their influence on cluster separation, structure preservation, and their effectiveness when used as a preprocessing step for classification tasks on real and synthetic datasets.The results highlight the importance of tailoring distance measures to specific data contexts and provide guidance for optimizing UMAP applications. Eva Blanco-Mallo, Verónica Bolón-Canedo, Beatriz Remeseiro |
ESANN | 2 |
| 2025 | Efficient ReliefF: A Low-Power Optimization of ReliefF for Resource-Constrained Devices
Samuel Suárez-Marcote, Laura Moran-Fernandez, Verónica Bolón-Canedo |
ICANN (1) | 3 |
| 2025 | Mitigating Overfitting in Recommender Systems via Intra-domain Transfer Learning
Eva Blanco-Mallo, Pablo Pérez-Núñez, Verónica Bolón-Canedo, Beatriz Remeseiro |
IDEAL (2) | 3 |
| 2025 | Fast and Frugal Transfer Learning via Precomputed Features and Adaptive Normalization
Daniel Vila-Cruz, Verónica Bolón-Canedo, Laura Moran-Fernandez |
IDEAL (1) | 2 |
| 2025 | Optimising Resource Use Through Low-Precision Feature Selection: A Performance Analysis of Logarithmic Division and Stochastic RoundingabstractABSTRACT The growth in the number of wearable devices has increased the amount of data produced daily. Simultaneously, the limitations of such devices has also led to a growing interest in the implementation of machine learning algorithms with low‐precision computation. We propose green and efficient modifications of state‐of‐the‐art feature selection methods based on information theory and fixed‐point representation. We tested two potential improvements: stochastic rounding to prevent information loss, and logarithmic division to improve computational and energy efficiency. Experiments with several datasets showed comparable results to baseline methods, with minimal information loss in both feature selection and subsequent classification steps. Our low‐precision approach proved viable even for complex datasets like microarrays, making it suitable for energy‐efficient internet‐of‐things (IoT) devices. While further investigation into stochastic rounding did not yield significant improvements, the use of logarithmic division for probability approximation showed promising results without compromising classification performance. Our findings offer valuable insights into resource‐efficient feature selection that contribute to IoT device performance and sustainability. Samuel Suárez-Marcote, Laura Moran-Fernandez, Verónica Bolón-Canedo |
Expert Syst. J. Knowl. Eng. | 3 |
| 2025 | Adaptive incremental transfer learning for efficient performance modeling of big data workloads
Mariano Garralda-Barrio, Carlos Eiras-Franco, Verónica Bolón-Canedo |
Future Gener. Comput. Syst. | 3 |
| 2025 | Breaking boundaries: Low-precision conditional mutual information for efficient feature selectionabstract©2025 Elsevier B.V. All rights reserved. This manuscript version is made available under the CC-BY-NC-ND 4.0 license https://creativecommons.org/licenses/bync-nd/4.0/. This version of the article has been accepted for publication in Pattern Recognition. The Version of Record is available online at https://doi.org/10.1016/j.patcog.2025.111375 Laura Moran-Fernandez, Eva Blanco-Mallo, Konstantinos Sechidis, Verónica Bolón-Canedo |
Pattern Recognit. | 4 |
| 2024 | Mutual Information-Based Feature Selection for Federated Learning EnvironmentsabstractDue to the growth of Internet of Things devices, the dimensionality of the data has only increased. These devices generate a large amount of data which, after being processed, is one of the main sources of information for machine learning systems. However, only a small part of this data is truly relevant. Understanding the relevance of this data allows us to process it at the edge of the network and improve the performance of Artificial Intelligence systems. On the one hand, feature selection is one of the most common approaches to reducing irrelevant features. On the other hand, federated learning allows the exploitation of such data by machine learning models without the need to exchange raw data. Therefore, in this paper we present modifications to two widely utilised feature selection algorithms based on the Mutual Information metric to enable them to work in a federated environment. The experimental process conducted on several datasets and data distributions demonstrates the capability of our modifications to operate in such an environment without a loss of information. This enhances execution time and leverages specific particularities of these environments, such as increased security. Samuel Suárez-Marcote, Laura Moran-Fernandez, Verónica Bolón-Canedo |
ICMLA | 3 |
| 2024 | The imbalance problem: A comparison of sampling approaches using different parameters and feature selection methods in the context of classificationabstractAbstract A common situation in classification tasks is to deal with unbalanced datasets, an issue that appears when the majority class(es) has a large number of samples compared to the minority class(es). This problem is even more significant when the datasets have a large number of features but only a few samples, as is the case with microarray datasets. Traditionally, an approach to alleviate this problem has been the application of sampling methods to obtain more balanced classes, increasing the number of samples in the minority class (replicating samples or generating new synthetic samples), or decreasing the number of samples in the majority class. In this study, we have compared different balancing methods, including a novel method that applies sampling in both the minority and majority classes. The interest in applying feature selection in combination with balancing methods has also been explored. In view of the results, a recommendation of sampling method, feature selection, and classifier is proposed to improve the classification results according to the type of dataset. José Luis Morillo-Salas, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
Expert Syst. J. Knowl. Eng. | 2 |
| 2024 | Spatial-temporal feature-based End-to-end Fourier network for 3D sign language recognition
Sunusi Bala Abdullahi, Kosin Chamnongthai, Verónica Bolón-Canedo, Brais Cancela |
Expert Syst. Appl. | 3 |
| 2024 | A review of green artificial intelligence: Towards a more sustainable futureabstractGreen artificial intelligence (AI) is more environmentally friendly and inclusive than conventional AI, as it not only produces accurate results without increasing the computational cost but also ensures that any researcher with a laptop can perform high-quality research without the need for costly cloud servers. This paper discusses green AI as a pivotal approach to enhancing the environmental sustainability of AI systems. Described are AI solutions for eco-friendly practices in other fields (green-by AI), strategies for designing energy-efficient machine learning (ML) algorithms and models (green-in AI), and tools for accurately measuring and optimizing energy consumption. Also examined are the role of regulations in promoting green AI and future directions for sustainable ML. Underscored is the importance of aligning AI practices with environmental considerations, fostering a more eco-conscious and energy-efficient future for AI systems. Verónica Bolón-Canedo, Laura Moran-Fernandez, Brais Cancela, Amparo Alonso-Betanzos |
Neurocomputing | 1 |
| 2024 | Fed-mRMR: A lossless federated feature selection methodabstractFeature selection has become a mandatory task in data mining, due to the overwhelming amount of features in Big Data problems. To handle this high-dimensional data and avoid the well-known curse of dimensionality, we need to pre-select an optimal subset of features to reduce redundant computations. Federated learning is a machine learning technique based on training an algorithm over many decentralized edge devices holding local rather than global data on a centralized server. Application of this technique is extending to fields such as self-driving cars, medicine and health, and Industry 4.0, where data privacy is compulsory. Feature selection through federated learning is a complicated task since suboptimal features calculated by feature selection methods may be different in heterogeneous datasets from different nodes. In this paper, we propose a lossless federated version of the classic minimum redundancy maximum relevance (mRMR) feature selection algorithm, called federated mRMR (fed-mRMR), which, without losing any effectiveness of the original mRMR method, is applicable to federated learning approaches and capable of dealing with data that are not independent and identically distributed (non-IID data). Implementation can be found at: https://github.com/jorgehermo9/fed-mrmr Jorge Hermo, Verónica Bolón-Canedo, Susana Ladra |
Inf. Sci. | 2 |
| 2024 | Finding a needle in a haystack: insights on feature selection for classification tasksabstractAbstract The growth of Big Data has resulted in an overwhelming increase in the volume of data available, including the number of features. Feature selection, the process of selecting relevant features and discarding irrelevant ones, has been successfully used to reduce the dimensionality of datasets. However, with numerous feature selection approaches in the literature, determining the best strategy for a specific problem is not straightforward. In this study, we compare the performance of various feature selection approaches to a random selection to identify the most effective strategy for a given type of problem. We use a large number of datasets to cover a broad range of real-world challenges. We evaluate the performance of seven popular feature selection approaches and five classifiers. Our findings show that feature selection is a valuable tool in machine learning and that correlation-based feature selection is the most effective strategy regardless of the scenario. Additionally, we found that using improper thresholds with ranker approaches produces results as poor as randomly selecting a subset of features. Laura Moran-Fernandez, Verónica Bolón-Canedo |
J. Intell. Inf. Syst. | 2 |
| 2024 | CUDA acceleration of MI-based feature selection methodsabstractFeature selection algorithms are necessary nowadays for machine learning as they are capable of removing irrelevant and redundant information to reduce the dimensionality of the data and improve the quality of subsequent analyses. The problem with current feature selection approaches is that they are computationally expensive when processing large datasets. This work presents parallel implementations for Nvidia GPUs of three highly-used feature selection methods based on the Mutual Information (MI) metric: mRMR, JMI and DISR. Publicly available code includes not only CUDA implementations of the general methods, but also an adaptation of them to work with low-precision fixed point in order to further increase their performance on GPUs. The experimental evaluation was carried out on two modern Nvidia GPUs (Turing T4 and Ampere A100) with highly satisfactory results, achieving speedups of up to 283x when compared to state-of-the-art C implementations. Bieito Beceiro, Jorge González-Domínguez, Laura Moran-Fernandez, Verónica Bolón-Canedo, Juan Touriño |
J. Parallel Distributed Comput. | 4 |
| 2024 | A novel framework for generic Spark workload characterization and similar pattern recognition using machine learningabstractComprehensive workload characterization plays a pivotal role in comprehending Spark applications, as it enables the analysis of diverse aspects and behaviors. This understanding is indispensable for devising downstream tuning objectives, such as performance improvement. To address this pivotal issue, our work introduces a novel and scalable framework for generic Spark workload characterization, complemented by consistent geometric measurements. The presented approach aims to build robust workload descriptors by profiling only quantitative metrics at the application task-level, in a non-intrusive manner. We expand our framework for downstream workload pattern recognition by incorporating unsupervised machine learning techniques: clustering algorithms and feature selection. These techniques significantly improve the process of grouping similar workloads without relying on predefined labels. We effectively recognize 24 representative Spark workloads from diverse domains, including SQL, machine learning, web search, graph, and micro-benchmarks, available in HiBench. Our framework achieves a high accuracy F-Measure score of up to 90.9% and a Normalized Mutual Information of up to 94.5% in similar workload pattern recognition. These scores significantly outperform the results obtained in a comparative analysis with an established workload characterization approach in the literature. Mariano Garralda-Barrio, Carlos Eiras-Franco, Verónica Bolón-Canedo |
J. Parallel Distributed Comput. | 3 |
| 2023 | Green Machine LearningabstractGreen machine learning refers to research that is more environmentally friendly and inclusive, not only by producing novel results without increasing the computational cost, but also by ensuring that any researcher with a laptop has the opportunity to perform high-quality research without the need to use expensive cloud servers.Efficient machine learning approaches (especially deep learning) are starting to receive some attention in the research community.This tutorial is concerned with the development of machine learning algorithms that optimize efficiency rather than only accuracy.We provide an overview of this recent field, together with a review of the novel contributions to the ESANN 2023 special session on Green Machine Learning. * This work Verónica Bolón-Canedo, Laura Moran-Fernandez, Brais Cancela, Amparo Alonso-Betanzos |
ESANN | 1 |
| 2023 | Efficient feature selection for domain adaptation using Mutual Information MaximizationabstractGreen AI, an emerging research field, focuses on improving the efficiency of machine learning models.In this paper, we introduce a novel and efficient method for feature selection in domain adaptation, a type of transfer learning where the source and target domains share the feature space and task but differ in their distributions.Instead of using evolutionary algorithms, a typical approach in this field, we propose the use of filter methods, which do not require an iterative search process and are less computationally expensive.Our proposed method is Mutual Information Maximization, and our experiments show that it outperforms Particle Swarm Optimization in terms of efficiency, speed, and the ability to select a reduced subset of features while achieving competitive classification accuracy results. Guillermo Castillo García, Laura Moran-Fernandez, Verónica Bolón-Canedo |
ESANN | 3 |
| 2023 | Automated green machine learning for condition-based maintenanceabstractWithin the big data paradigm, there is an increasing demand for machine learning with automatic configuration of hyperparameters.Although several algorithms have been proposed for automatically learning time-changing concepts, they generally do not scale well to very large databases.In this context, this paper presents an automated green machine learning approach applied to condition-based maintenance with automatic data fusion and density-based anomaly detection based on locality sensitivity hashing.Experiments on numerical simulations of train-track dynamic interactions demonstrate the utility of the approach to detect railway wheel out-of-roundness.This unlocks the full potential of scalable machine learning, paving the way for environment-friendly systems and automated decision-making. Afonso Lourenço, Carolina Ferraz, Jorge Meira, Goreti Marreiros, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
ESANN | 5 |
| 2023 | Logarithmic division for green feature selection: an information-theoretic approachabstractFeature selection is a popular preprocessing step to reduce the dimensionality of the data while preserving the important information.In this paper we propose an efficient and green feature selection method based on information theory, with the novelty of using the logarithmic division and resort to fixed-point precision.The results of experiments conducted on several datasets indicate the potential of our proposal, as it does not incur in significant information loss compared to the standard method, both in the features selected and in the subsequent classification step.This finding opens up possibilities for a new family of green feature selection methods, which would help to minimize energy consumption and carbon emissions.* This work was Samuel Suárez-Marcote, Laura Moran-Fernandez, Verónica Bolón-Canedo |
ESANN | 3 |
| 2023 | A Pseudo-Label Guided Hybrid Approach for Unsupervised Domain Adaptation
Eva Blanco-Mallo, Verónica Bolón-Canedo, Beatriz Remeseiro |
IDEAL | 2 |
| 2023 | Data-driven predictive maintenance framework for railway systemsabstractThe emergence of the Industry 4.0 trend brings automation and data exchange to industrial manufacturing. Using computational systems and IoT devices allows businesses to collect and deal with vast volumes of sensorial and business process data. The growing and proliferation of big data and machine learning technologies enable strategic decisions based on the analyzed data. This study suggests a data-driven predictive maintenance framework for the air production unit (APU) system of a train of Metro do Porto. The proposed method assists in detecting failures and errors in machinery before they reach critical stages. We present an anomaly detection model following an unsupervised approach, combining the Half-Space-trees method with One Class K Nearest Neighbor, adapted to deal with data streams. We evaluate and compare our approach with the Half-Space-Trees method applied without the One Class K Nearest Neighbor combination. Our model produced few type-I errors, significantly increasing the value of precision when compared to the Half-Space-Trees model. Our proposal achieved high anomaly detection performance, predicting most of the catastrophic failures of the APU train system. Jorge Meira, Bruno M. Veloso, Verónica Bolón-Canedo, Goreti Marreiros, Amparo Alonso-Betanzos, João Gama 0001 |
Intell. Data Anal. | 3 |
| 2023 | Feature selection for domain adaptation using complexity measures and swarm intelligenceabstractParticle Swarm Optimization is an optimization algorithm that mimics the behaviour of a flock of birds, setting multiple particles that explore the search space guided by a fitness function in order to find the best possible solution.We apply the Sticky Binary Particle Swarm Optimization algorithm to perform feature selection for domain adaptation, a specific type of transfer learning in which the source and the target domain have a common feature space, a common task, but different distributions.When applying Particle Swarm Optimization, classification error is usually employed in the fitness function to evaluate the goodness of subsets of features.In this paper, we aim to compare this approach with using complexity metrics instead, under the assumption that reducing the complexity of the problem will lead to results that are independent from the classifier used for testing while being less computationally demanding.Therefore, we carried out experiments to compare the performance of both approaches in terms of classification accuracy, speed and number of features selected.We found out that our proposal, although in some cases incurs in a slight degradation of classification performance, it is indeed faster and selects fewer features, making it a feasible trade-off. Guillermo Castillo García, Laura Moran-Fernandez, Verónica Bolón-Canedo |
Neurocomputing | 3 |
| 2023 | E2E-FS: An End-to-End Feature Selection Method for Neural NetworksabstractClassic embedded feature selection algorithms are often divided in two large groups: tree-based algorithms and LASSO variants. Both approaches are focused in different aspects: while the tree-based algorithms provide a clear explanation about which variables are being used to trigger a certain output, LASSO-like approaches sacrifice a detailed explanation in favor of increasing its accuracy. In this paper, we present a novel embedded feature selection algorithm, called End-to-End Feature Selection (E2E-FS), that aims to provide both accuracy and explainability in a clever way. Despite having non-convex regularization terms, our algorithm, similar to the LASSO approach, is solved with gradient descent techniques, introducing some restrictions that force the model to specifically select a maximum number of features that are going to be used subsequently by the classifier. Although these are hard restrictions, the experimental results obtained show that this algorithm can be used with any learning model that is trained using a gradient descent algorithm. Brais Cancela, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Do all roads lead to Rome? Studying distance measures in the context of machine learningabstractMany machine learning and data mining tasks are based on distance measures, so a large amount of literature addresses this aspect somehow. Due to the broad scope of the topic, this paper aims to provide an overview of the use of these measures in the most common machine learning problems, pointing out those aspects to consider to choose the most appropriate measure for a particular task. For this purpose, the most recent works addressing the subject were reviewed and seven of the most commonly used measures were analyzed, investigating in detail their main properties and applications. Different experiments were carried out to study their relationships and compare their performance. The degradation of the results in the presence of noise was also considered, as well as the execution time required by each measure. Eva Blanco-Mallo, Laura Moran-Fernandez, Beatriz Remeseiro, Verónica Bolón-Canedo |
Pattern Recognit. | 4 |
| 2023 | Machine Learning Methods for Predicting League of Legends Game OutcomeabstractThe video gameLeague of Legendshas several professional leagues and tournaments that offer prizes reaching several million dollars, making it one of the most followed games in the Esports scene. This article addresses the prediction of the winning team in professional matches of the game, using only pregame data. We propose to improve the accuracy of the models trained with the features offered by the game application programming interface (API). To this end, new features are built to collect interesting information, such as the skills of a player handling a certain champion, the synergies between players of the same team or the ability of a player to beat another player. Then, we perform feature selection and train different classification algorithms aiming at obtaining the best model. Experimental results show classification accuracy above 0.70, which is comparable to the results of other proposals presented in the literature, but with the added benefit of using few samples and not requiring the use of external sources to collect additional statistics. Juan Agustín Hitar-García, Laura Moran-Fernandez, Verónica Bolón-Canedo |
IEEE Trans. Games | 3 |
| 2022 | The role of feature selection in personalized recommender systemsabstractRecommender systems suggest products to users, based on their popularity or the users' preferences.This paper proposes a hybrid personalized recommender system based on users' tastes and also on information available about items.We used a dataset downloaded from Tri-pAdvisor, which contains some information from restaurants (items), such as price range or special diets.Feature selection techniques are employed to analyze the impact that each variable has on personalized recommendations, allowing us to understand not only the process underlying the recommendation to favor the transparency of the system, but also what users value the most when choosing a restaurant. Roger Bagué-Masanés, Verónica Bolón-Canedo, Beatriz Remeseiro |
ESANN | 2 |
| 2022 | Feature selection for transfer learning using particle swarm optimization and complexity measuresabstractParticle Swarm Optimization is an optimization algorithm that explores a search space guided by a fitness function in order to find a good solution.We apply it to perform feature selection for domain adaptation.Usually, classification error is used in the fitness function to evaluate the goodness of subsets of features.In this paper, we propose to employ complexity metrics instead, as we assume that reducing the complexity of the problem will lead to good results while being less computationally demanding and independent from the classifier used for testing.We found out that our method is indeed faster and selects fewer features, obtaining competitive classification accuracy results. Verónica Bolón-Canedo, Guillermo Castillo García, Laura Moran-Fernandez |
ESANN | 1 |
| 2022 | When the best reviews are not placed between extremesabstractSeveral research studies have demonstrated the strong influence that online reviews exert on consumers' purchasing decisions. Specifically, those with extreme opinions, both favorable and unfavorable, are often considered more useful. This paper is focused on enhancing the detection of extreme reviews through sentiment analysis. For this purpose, a real scenario is taken into account, using the examples of all classes and dealing with the imbalance between them, which is characteristic in online reviews. The main objective is to carry out the classification with a high certainty and incurring as few errors as possible in relation to the examples belonging to the rest of the classes. Therefore, the emphasis is on the quality of the predictions rather than on the quantity. Using XLNet, we show how the transfer of knowledge extracted from the source domain (i.e., the extreme reviews) improves their detection regarding the overall number of errors made in the target domain (i.e., multi-class classification). Eva Blanco-Mallo, João Carneiro 0001, Goreti Marreiros, Beatriz Remeseiro, Verónica Bolón-Canedo |
IJCNN | 5 |
| 2022 | Less is more: Low-precision feature selection for wearablesabstractNowadays, the amount of data produced daily has significantly increased due to the growth in the number of wearable devices. Similarly, this increase is also visible in the interest of developing machine learning algorithms with reduced precision computations, due to the limitations of such devices. This work studies the effect of using low precision operations in the context of feature selection, a preprocessing step that is becoming necessary to deal with the increasing data dimensionality. This study focuses specifically on feature selection methods based on Mutual Information (one of the most popular and widely-used metrics in this area) and how low precision computations can be carried out obtaining experimental results similar to those achieved by double-precision over several low- and high-dimensional datasets. We observe that the use of 16-bit fixed-point representation makes it possible to obtain feature rankings with high similarity to those obtained in double- precision. Even the rankings obtained with 8 bits and then used in subsequent classification tasks, lead to similar accuracy (no significant difference) to the one obtained when using the 64-bit representation in certain situations. Samuel Suárez-Marcote, Laura Moran-Fernandez, Verónica Bolón-Canedo |
IJCNN | 3 |
| 2022 | Reduced precision discretization based on information theoryabstractIn recent years, new technological areas have emerged and proliferated, such as the Internet of Things or embedded systems in drones, which are usually characterized by making use of devices with strict requirements of weight, size, cost and power consumption. As a consequence, there has been a growing interest in the implementation of machine learning algorithms with reduced precision that can be embedded in these constrained devices. These algorithms cover not only learning, but they can also be applied to other stages such as feature selection or data discretization. In this work we study the behavior of the Minimum Description Length Principle (MDLP) discretizer, proposed by Fayyad and Irani, when reduced precision is used, and how much it affects to a typical machine learning pipeline. Experimental results show that the use of fixed-point format is sufficient to achieve performances similar to those obtained when using double-precision format, which opens the door to the use of reduced-precision discretizers in embedded systems, minimizing energy consumption and carbon emissions. Brais Ares, Laura Moran-Fernandez, Verónica Bolón-Canedo |
KES | 3 |
| 2022 | Machine learning techniques to predict different levels of hospital care of CoVid-19abstractIn this study, we analyze the capability of several state of the art machine learning methods to predict whether patients diagnosed with CoVid-19 (CoronaVirus disease 2019) will need different levels of hospital care assistance (regular hospital admission or intensive care unit admission), during the course of their illness, using only demographic and clinical data. For this research, a data set of 10,454 patients from 14 hospitals in Galicia (Spain) was used. Each patient is characterized by 833 variables, two of which are age and gender and the other are records of diseases or conditions in their medical history. In addition, for each patient, his/her history of hospital or intensive care unit (ICU) admissions due to CoVid-19 is available. This clinical history will serve to label each patient and thus being able to assess the predictions of the model. Our aim is to identify which model delivers the best accuracies for both hospital and ICU admissions only using demographic variables and some structured clinical data, as well as identifying which of those are more relevant in both cases. The results obtained in the experimental study show that the best models are those based on oversampling as a preprocessing phase to balance the distribution of classes. Using these models and all the available features, we achieved an area under the curve (AUC) of 76.1% and 80.4% for predicting the need of hospital and ICU admissions, respectively. Furthermore, feature selection and oversampling techniques were applied and it has been experimentally verified that the relevant variables for the classification are age and gender, since only using these two features the performance of the models is not degraded for the two mentioned prediction problems. Elena Hernández-Pereira, Oscar Fontenla-Romero, Verónica Bolón-Canedo, Brais Cancela, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos |
Appl. Intell. | 3 |
| 2022 | How important is data quality? Best classifiers vs best featuresabstractThe task of choosing the appropriate classifier for a given scenario is not an easy-to-solve question. First, there is an increasingly high number of algorithms available belonging to different families. And also there is a lack of methodologies that can help on recommending in advance a given family of algorithms for a certain type of datasets. Besides, most of these classification algorithms exhibit a degradation in the performance when faced with datasets containing irrelevant and/or redundant features. In this work we analyze the impact of feature selection in classification over several synthetic and real datasets. The experimental results obtained show that the significance of selecting a classifier decreases after applying an appropriate preprocessing step and, not only this alleviates the choice, but it also improves the results in almost all the datasets tested. Laura Moran-Fernandez, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
Neurocomputing | 2 |
| 2022 | Fast anomaly detection with locality-sensitive hashing and hyperparameter autotuningabstractThis paper presents LSHAD, an anomaly detection (AD) method based on Locality Sensitive Hashing (LSH), capable of dealing with large-scale datasets. The resulting algorithm is highly parallelizable and its implementation in Apache Spark further increases its ability to handle very large datasets. Moreover, the algorithm incorporates an automatic hyperparameter tuning mechanism so that users do not have to implement costly manual tuning. Our LSHAD method is novel as both hyperparameter automation and distributed properties are not usual in AD techniques. Our results for experiments with LSHAD across a variety of datasets point to state-of-the-art AD performance while handling much larger datasets than state-of-the-art alternatives. In addition, evaluation results for the tradeoff between AD performance and scalability show that our method offers significant advantages over competing methods. Jorge Meira, Carlos Eiras-Franco, Verónica Bolón-Canedo, Goreti Marreiros, Amparo Alonso-Betanzos |
Inf. Sci. | 3 |
| 2021 | Dealing with heterogeneity in the context of distributed feature selection for classification
José Luis Morillo-Salas, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
Knowl. Inf. Syst. | 2 |
| 2020 | Do we need hundreds of classifiers or a good feature selection?
Laura Moran-Fernandez, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
ESANN | 2 |
| 2020 | Gene Subset Selection for Transfer Learning using Bilevel Particle Swarm OptimizationabstractIn classification problems that involve multiple sources, data distributions may vary. Therefore, knowledge of dimensions that differ in the source and target data is important to reduce the distance between domains, allowing accurate transfer knowledge. Here, we present a novel method to identify (in)variant genes between source and target datasets and integrate such results to simultaneously reduce the variance between two distributions while optimizing the size and classification error of the selected subset. In particular, we use an evolutionary computation particle swarm optimization algorithm to implement such a bilevel multi-objective programming approach, allowing us to solve a gene subset selection problem. Hassen Dhrif, Verónica Bolón-Canedo, Stefan Wuchty |
ICMLA | 2 |
| 2020 | A delayed Elastic-Net approach for performing adversarial attacksabstractWith the rise of the so-called Adversarial Attacks, there is an increased concern on model security. In this paper we present two different contributions: novel measures of robustness (based on adversarial attacks) and a novel adversarial attack. The key idea behind these metrics is to obtain a measure that could compare different architectures, with independence of how the input is preprocessed (robustness against different input sizes and value ranges). To do so, a novel adversarial attack is presented, performing a delayed elastic-net adversarial attack (constraints are only used whenever a successful adversarial attack is obtained). Experimental results show that our approach obtains state-of-the-art adversarial samples, in terms of minimal perturbation distance. Finally, a benchmark of ImageNet pretrained models is used to conduct experiments aiming to shed some light about which model should be selected whenever security is a role factor. Brais Cancela, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
ICPR | 2 |
| 2020 | Can data placement be effective for Neural Networks classification tasks? Introducing the Orthogonal Loss
Brais Cancela, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
ICPR | 2 |
| 2020 | CUDA-JMI: Acceleration of feature selection on heterogeneous systems
Jorge González-Domínguez, Roberto R. Expósito, Verónica Bolón-Canedo |
Future Gener. Comput. Syst. | 3 |
| 2020 | A scalable saliency-based feature selection method with instance-level information
Brais Cancela, Verónica Bolón-Canedo, Amparo Alonso-Betanzos, João Gama 0001 |
Knowl. Based Syst. | 2 |
| 2020 | Feature selection with limited bit depth mutual information for portable embedded systems
Laura Moran-Fernandez, Konstantinos Sechidis, Verónica Bolón-Canedo, Amparo Alonso-Betanzos, Gavin Brown 0001 |
Knowl. Based Syst. | 3 |
| 2019 | Case Study of Anomaly Detection and Quality Control of Energy Efficiency and Hygrothermal Comfort in Buildingsabstract[Abstract] The aim of this work is to propose different statistical and machine learning methodologies for identifying anomalies and control the quality of energy efficiency and hygrothermal comfort in buildings. Companies focused on energy sector for buildings are interested on statistical and machine learning tools to automate the control of energy consumption and ensure quality of Heat Ventilation and Air Conditioning (HVAC) installations. Consequently, a methodology based on the application of the Local Correlation Integral (LOCI) anomaly detection technique has been proposed. In addition, the most critical variables for anomaly detection are identified by using ReliefF method. Once vectors of critical variables are obtained, multivariate and univariate control charts can be applied to control the quality of HVAC installations (consumption, thermal comfort). In order to test the proposed methodology, the companies involved in this project have provided the case study of a store of a clothing brand located in a shopping center in Panama. It is important to note that this is a controlled case study for which all the anomalies have been previously identified by maintenance personnel. Moreover, as an alternatively solution, in addition to machine learning and multivariate techniques, new nonparametric control charts for functional data based on data depth have been proposed and applied to curves of daily energy consumption in HVAC. Carlos Eiras-Franco, Miguel Flores, Verónica Bolón-Canedo, Sonia Zaragoza, Rubén Fernández-Casal, Salvador Naya, Javier Tarrío-Saavedra |
DATA | 3 |
| 2019 | Biases in feature selection with missing data
Borja Seijo-Pardo, Amparo Alonso-Betanzos, Kristin P. Bennett, Verónica Bolón-Canedo, Julie Josse, Mehreen Saeed, Isabelle Guyon |
Neurocomputing | 4 |
| 2019 | Insights into distributed feature ranking
Verónica Bolón-Canedo, Konstantinos Sechidis, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, Gavin Brown 0001 |
Inf. Sci. | 1 |
| 2019 | Parallel feature selection for distributed-memory clusters
Jorge González-Domínguez, Verónica Bolón-Canedo, Borja Freire, Juan Touriño |
Inf. Sci. | 2 |
| 2019 | Distributed classification based on distances between probability distributions in feature space
Pablo Montero-Manso, Laura Moran-Fernandez, Verónica Bolón-Canedo, José Antonio Vilar, Amparo Alonso-Betanzos |
Inf. Sci. | 3 |
| 2018 | Analysis of imputation bias for feature selection with missing data
Borja Seijo-Pardo, Amparo Alonso-Betanzos, Kristin P. Bennett, Verónica Bolón-Canedo, Isabelle Guyon, Julie Josse, Mehreen Saeed |
ESANN | 4 |
| 2018 | Predictive Maintenance in the Metallurgical Industry: Data Analysis and Feature Selection
Marta Fernandes, Alda Canito, Verónica Bolón-Canedo, Luís Conceição, Isabel Praça, Goreti Marreiros |
WorldCIST (1) | 3 |
| 2018 | On the scalability of feature selection methods on high-dimensional data
Verónica Bolón-Canedo, Diego Fernández-Francos, Diego Peteiro-Barral, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Noelia Sánchez-Maroño |
Knowl. Inf. Syst. | 1 |
| 2018 | An Information Theory-Based Feature Selection Framework for Big Data Under Apache SparkabstractWith the advent of extremely high dimensional datasets, dimensionality reduction techniques are becoming mandatory. Of the many techniques available, feature selection (FS) is of growing interest for its ability to identify both relevant features and frequently repeated instances in huge datasets. We aim to demonstrate that standard FS methods can be parallelized in big data platforms like Apache Spark so as to boost both performance and accuracy. We propose a distributed implementation of a generic FS framework that includes a broad group of well-known information theory-based methods. Experimental results for a broad set of real-world datasets show that our distributed framework is capable of rapidly dealing with ultrahigh-dimensional datasets as well as those with a huge number of samples, outperforming the sequential version in all the cases studied. Sergio Ramírez-Gallego, Héctor Mouriño-Talín, David Martínez-Rego, Verónica Bolón-Canedo, José Manuel Benítez 0001, Amparo Alonso-Betanzos, Francisco Herrera |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2017 | Algorithmic challenges in big data analytics
Verónica Bolón-Canedo, Beatriz Remeseiro, Konstantinos Sechidis, David Martínez-Rego, Amparo Alonso-Betanzos |
ESANN | 1 |
| 2017 | A distributed approach for classification using distance metrics
Laura Moran-Fernandez, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
ESANN | 2 |
| 2017 | Paving the way for providing teaching feedback in automatic evaluation of open response assignmentsabstractPeer grading has been the regular procedure to use for automatic assessment of open ended assignments in Massive Open Online Courses (MOOCs). However, and although the procedure tries to overcome the rupture of the classical teach-learn-assess/feedback cycle, it does so only in the student side, and no attempt has been made as yet in giving feedback to instructors. The work described inhere aims at filling this gap, with a proposal in which the instructors are supplied with the set of words most used by the best and worst ranked quartiles of assignments. In order to achieve this, a Gaussian Mixture Model (GMM) fed with the bag of words supplied by a previous feature selection algorithm is presented, with the goal of identifying the clusters of words related with similar grades. The results obtained over three pilot studies, containing assignments in three different disciplines, show that our model can lead to more complete information on the teacher feedback on the results of the assignments. Verónica Bolón-Canedo, Jorge Díez 0001, Oscar Luaces, Antonio Bahamonde, Amparo Alonso-Betanzos |
IJCNN | 1 |
| 2017 | Exploring the consequences of distributed feature selection in DNA microarray dataabstractMicroarray data classification has been typically seen as a difficult challenge for machine learning researchers mainly due to its high dimension in features while sample size is small. Because of this particularity, feature selection is usually applied trying to reduce its high dimensionality. However, existing algorithms may not scale well when dealing with this amount of features, and a possible solution is to distribute the features into several nodes. In this work we explore the process of distribution on microarray data - which has recently gained attention - and we evaluate to what extent it is possible to obtain similar results as those obtained with the whole dataset. We performed experiments with different aggregation methods, feature rankers and also evaluated the effect of distributing the feature ranking process in the subsequent classification performance. Verónica Bolón-Canedo, Konstantinos Sechidis, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, Gavin Brown 0001 |
IJCNN | 1 |
| 2017 | Fast-mRMR: Fast Minimum Redundancy Maximum Relevance Algorithm for High-Dimensional Big DataabstractWith the advent of large-scale problems, feature selection has become a fundamental preprocessing step to reduce input dimensionality. The minimum-redundancy-maximum-relevance (mRMR) selector is considered one of the most relevant methods for dimensionality reduction due to its high accuracy. However, it is a computationally expensive technique, sharply affected by the number of features. This paper presents fast-mRMR, an extension of mRMR, which tries to overcome this computational burden. Associated with fast-mRMR, we include a package with three implementations of this algorithm in several platforms, namely, CPU for sequential execution, GPU (graphics processing units) for parallel computing, and Apache Spark for distributed computing using big data technologies. Sergio Ramírez-Gallego, Iago Lastra, David Martínez-Rego, Verónica Bolón-Canedo, José Manuel Benítez 0001, Francisco Herrera, Amparo Alonso-Betanzos |
Int. J. Intell. Syst. | 4 |
| 2017 | Can classification performance be predicted by complexity measures? A study using microarray data
Laura Moran-Fernandez, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
Knowl. Inf. Syst. | 2 |
| 2017 | Centralized vs. distributed feature selection methods based on data complexity measures
Laura Moran-Fernandez, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
Knowl. Based Syst. | 2 |
| 2017 | Ensemble feature selection: Homogeneous and heterogeneous approaches
Borja Seijo-Pardo, Iago Porto-Díaz, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
Knowl. Based Syst. | 3 |
| 2017 | Testing Different Ensemble Configurations for Feature Selection
Borja Seijo-Pardo, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
Neural Process. Lett. | 2 |
| 2016 | Machine learning for medical applications
Verónica Bolón-Canedo, Beatriz Remeseiro, Amparo Alonso-Betanzos, Aurélio J. C. Campilho |
ESANN | 1 |
| 2016 | Data complexity measures for analyzing the effect of SMOTE over microarrays
Laura Moran-Fernandez, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
ESANN | 2 |
| 2016 | Using a feature selection ensemble on DNA microarray datasets
Borja Seijo-Pardo, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
ESANN | 2 |
| 2016 | A unified pipeline for online feature selection and classification
Verónica Bolón-Canedo, Diego Fernández-Francos, Diego Peteiro-Barral, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Noelia Sánchez-Maroño |
Expert Syst. Appl. | 1 |
| 2016 | A comparison of performance of K-complex classification methods using feature selection
Elena Hernández-Pereira, Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Diego Álvarez-Estévez, Vicente Moret-Bonillo, Amparo Alonso-Betanzos |
Inf. Sci. | 2 |
| 2015 | Feature and kernel learning
Verónica Bolón-Canedo, Michele Donini, Fabio Aiolli |
ESANN | 1 |
| 2015 | On the use of machine learning techniques for the analysis of spontaneous reactions in automated hearing assessment
Verónica Bolón-Canedo, Alba Fernández, Amparo Alonso-Betanzos, Marcos Ortega 0001, Manuel G. Penedo |
ESANN | 1 |
| 2015 | Learning features on tear film lipid layer classification
Beatriz Remeseiro, Verónica Bolón-Canedo, Amparo Alonso-Betanzos, Manuel G. Penedo |
ESANN | 2 |
| 2015 | An insight on complexity measures and classification in microarray dataabstractMicroarray data classification has been typically seen as a difficult challenge for machine learning researchers mainly due to its high dimension in feature while sample size is small. However, this type of data presents other complications such as overlapping between classes, dataset shift, class imbalance, non-linearity, or features extracted under extremely different distributions. This paper intends to analyze in depth the theoretical complexity of several popular binary datasets, by making use of complexity measures, and then connecting it with the empirical results obtained by four widely-used classifiers. Two different situations are covered: datasets with only training set and datasets originally divided into training and test sets. In both cases it is demonstrated that there exists a correlation between the complexity measures and the actual error rates, which can facilitate in the future how to deal with a given dataset. Finally, we present a case study on Prostate dataset, improving the test classification accuracy from 53% to 97%. Verónica Bolón-Canedo, Laura Moran-Fernandez, Amparo Alonso-Betanzos |
IJCNN | 1 |
| 2015 | Recent advances and emerging challenges of feature selection in the context of big data
Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
Knowl. Based Syst. | 1 |
| 2014 | Toward parallel feature selection from vertically partitioned data
Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Joana Cerviño-Rabuñal |
ESANN | 1 |
| 2014 | Learning on Vertically Partitioned Data based on Chi-square Feature Selection and Naive Bayes ClassificationabstractIn the last few years, distributed learning has been the focus of much attention due to the explosion of big databases, in some cases distributed across different nodes. However, the great majority of current selection and classification algorithms are designed for centralized learning, i.e. they use the whole dataset at once. In this paper, a new approach for learning on vertically partitioned data is presented, which covers both feature selection and classification. The approach splits the data by features, and then uses the chi-square filter and the naive Bayes classifier to learn at each node. Finally, a merging procedure is performed, which updates the learned model in an incremental fashion. The experimental results on five representative datasets show that the execution time is shortened considerably whereas the classification performance is maintained as the number of nodes increases. Verónica Bolón-Canedo, Diego Peteiro-Barral, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Noelia Sánchez-Maroño |
ICAART (1) | 1 |
| 2014 | mC-ReliefF - An Extension of ReliefF for Cost-based Feature SelectionabstractThe proliferation of high-dimensional data in the last few years has brought a necessity to use dimensionality reduction techniques, in which feature selection is arguably the most famous one. Feature selection consists of detecting relevant features and discarding the irrelevant ones. However, there are some situations where the users are not only interested in the relevance of the selected features but also in the costs that they imply (e.g. economical or computational costs). In this paper an extension of the well-known ReliefF method for feature selection is proposed, which consists of adding a new term to the function which updates the weights of the features so as to be able to reach a trade-off between the relevance of a feature and its associated cost. The behavior of the proposed method is tested on twelve heterogeneous classification datasets as well as a real application, using a support vector machine (SVM) as a classifier. The results of the experimental study show that the approach is sound, since it allows the user to reduce the cost significantly without compromising the classification error. Verónica Bolón-Canedo, Beatriz Remeseiro, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
ICAART (1) | 1 |
| 2014 | Scalability Analysis of mRMR for Microarray DataabstractLately, derived from the Big Data problem, researchers in Machine Learning became also interested not only
in accuracy, but also in scalability. Although scalability of learning methods is a trending issue, scalability of
feature selection methods has not received the same amount of attention. In this research, an attempt to study
scalability of both Feature Selection and Machine Learning on microarray datasets will be done. For this sake,
the minimum redundancy maximum relevance (mRMR) filter method has been chosen, since it claims to be
very adequate for this type of datasets. Three synthetic databases which reflect the problematics of microarray
will be evaluated with new measures, based not only in an accurate selection but also in execution time. The
results obtained are presented and discussed. Diego Rego-Fernández, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
ICAART (1) | 2 |
| 2014 | Data classification using an ensemble of filters
Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
Neurocomputing | 1 |
| 2014 | A review of microarray datasets and applied feature selection methods
Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, José Manuel Benítez 0001, Francisco Herrera |
Inf. Sci. | 1 |
| 2014 | A framework for cost-based feature selection
Verónica Bolón-Canedo, Iago Porto-Díaz, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
Pattern Recognit. | 1 |
| 2014 | A Methodology for Improving Tear Film Lipid Layer ClassificationabstractDry eye is a symptomatic disease which affects a wide range of population and has a negative impact on their daily activities. Its diagnosis can be achieved by analyzing the interference patterns of the tear film lipid layer and by classifying them into one of the Guillon categories. The manual process done by experts is not only affected by subjective factors but is also very time consuming. In this paper we propose a general methodology to the automatic classification of tear film lipid layer, using color and texture information to characterize the image and feature selection methods to reduce the processing time. The adequacy of the proposed methodology was demonstrated since it achieves classification rates over 97% while maintaining robustness and provides unbiased results. Also, it can be applied in real time, and so allows important time savings for the experts. Beatriz Remeseiro, Verónica Bolón-Canedo, Diego Peteiro-Barral, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Antonio Mosquera González, Manuel G. Penedo, Noelia Sánchez-Maroño |
IEEE J. Biomed. Health Informatics | 2 |
| 2013 | A distributed wrapper approach for feature selection
Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
ESANN | 1 |
| 2013 | Toward the scalability of neural networks through feature selection
Diego Peteiro-Barral, Verónica Bolón-Canedo, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Noelia Sánchez-Maroño |
Expert Syst. Appl. | 2 |
| 2013 | A review of feature selection methods on synthetic data
Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
Knowl. Inf. Syst. | 1 |
| 2012 | Interferential Tear Film Lipid Layer Classification: An Automatic Dry Eye TestabstractDry eye is a symptomatic disease which affects a wide range of population and has a negative impact on their daily activities, such as driving or working with computers. Its diagnosis can be achieved by several clinical tests, one of which is the analysis of the interference pattern and its classification into one of the Guillon's categories. The methodologies for automatic classification obtain promising results but at the expense of requiring a long processing time. In this research, feature selection techniques are used to reduce time whilst maintaining performance, paving the way for the development of a novel tool for automatic classification of tear film lipid layer. This tool produces significant classification rates over 96% compared with the annotations of the optometrists and provides unbiased results. Also, it works in real-time and so allows important time savings for the experts. Verónica Bolón-Canedo, Diego Peteiro-Barral, Beatriz Remeseiro, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Antonio Mosquera González, Manuel G. Penedo, Noelia Sánchez-Maroño |
ICTAI | 1 |
| 2012 | An ensemble of filters and classifiers for microarray data classification
Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
Pattern Recognit. | 1 |
| 2011 | Statistical dependence measure for feature selection in microarray datasets
Verónica Bolón-Canedo, Sohan Seth, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, José C. Príncipe |
ESANN | 1 |
| 2011 | On the behavior of feature selection methods dealing with noise and relevance over synthetic scenariosabstractAdequate identification of relevant features is fundamental in real world scenarios. The problem is specially important when the datasets have a much larger number of features than samples. However, in most cases, the relevant features in real datasets are unknown. In this paper several synthetic datasets are employed to test the effectiveness of different feature selection methods over different artificial classification scenarios, such as altered features (noise), presence of a crescent number of irrelevant features and a small ratio between number of samples and number of features. Six filters and two embedded methods are tested over five synthetic datasets, so as to be able to choose a robust and noise tolerant method, paving the way for its application to real datasets in the classification domain. Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
IJCNN | 1 |
| 2011 | Toward an ensemble of filters for classificationabstractIn this paper we propose a new framework for feature selection consisting of an ensemble of filters for classification. Five filters, based on different metrics, were involved. Two different approaches of ensembles are presented by varying the role of the classification step. The different options to build an ensemble of filters were studied in detail, dealing with issues such as the presence of redundancy when joining the features selected by different methods. The adequacy of using an ensemble of filters instead of a single filter was demonstrated over a challenging scenario such as DNA microarray data. Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
ISDA | 1 |
| 2011 | Feature selection and classification in multiple class datasets: An application to KDD Cup 99 dataset
Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
Expert Syst. Appl. | 1 |
| 2011 | A study of performance on microarray data sets for a classifier based on information theoretic learning
Iago Porto-Díaz, Verónica Bolón-Canedo, Amparo Alonso-Betanzos, Oscar Fontenla-Romero |
Neural Networks | 2 |
| 2010 | Local Modeling Classifier for Microarray Gene-Expression Data
Iago Porto-Díaz, Verónica Bolón-Canedo, Amparo Alonso-Betanzos, Oscar Fontenla-Romero |
ICANN (3) | 2 |
| 2010 | On the effectiveness of discretization on gene selection of microarray dataabstractDNA microarray data is a challenging issue for machine learning researchers due to the high number of gene expression contained and the small samples sizes. To deal with this problem, feature selection methods, such as filters and wrappers, are typically applied to reduce the dimensionality. In this work, we apply a filter method before the classification and include a discretization step. The results obtained over ten different microarray data sets confirm the adequacy of the proposed method, that achieves better performances than the classifier alone. Besides, the combination method is also compared with the approaches of other authors (using wrappers and filters), outperforming the prediction accuracy and maintaining or even decreasing the number of genes required. Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
IJCNN | 1 |
| 2010 | Multiclass classifiers vs multiple binary classifiers using filters for feature selectionabstractThere are two classical approaches for dealing with multiple class data sets: a classifier that can deal directly with them, or alternatively, dividing the problem into multiple binary sub-problems. While studies on feature selection using the first approach are relatively frequent in scientific literature, very few studies employ the latter one. Out of the four classical methods that can be employed for generating binary problems from a multiple class data set (random, exhaustive, one-vs-one and one-vs-rest), the two last were employed in this work. Besides, four different methods were used for joining the results of these binary classifiers (sum, sum with threshold, Hamming distance and loss-based function). In this paper, both approaches (multiclass and multiple binary classifiers), are carried out using a combination method composed by a discretizer (two different were employed), a filter for feature selection (two methods were chosen), and a classifier (two classifiers were tested). The different combinations of the previous methods, with and without feature selection, were tested over 21 different multiple data sets. An exhaustive study of the results and a comparison between the described methods and some others on the literature is carried out. Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, Pablo Garcia-Gonzalez, Verónica Bolón-Canedo |
IJCNN | 4 |
| 2009 | A combination of discretization and filter methods for improving classification performance in KDD Cup 99 datasetabstractKDD Cup 99 dataset is a classical challenge for computer intrusion detection as well as machine learning researchers. Due to the problematic of this dataset, several sophisticated machine learning algorithms have been tried by different authors. In this paper a new approach is proposed that consists in a combination of a discretizator, a filter method and a very simple classical classifier. The results obtained show the adequacy of the method, that achieves comparable or even better performances than those of other more complicated algorithms, but with a considerable reduction in the number of input features. The proposed method has also been tried over another two large datasets maintaining the same behavior as in the KDD Cup 99 dataset. Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
IJCNN | 1 |