Laura Moran-Fernandez

dblp:169/0933 · also Laura Morán-Fernández · DBLP profile ↗
← Back
29ranked-venue papers
9as first author
21since 2021 · last 2026
0000-0001-6703-1846ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 7 first-author · 18 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SMOTE k-out: Enhancing Class Separability through Outer Synthetic Sampling
abstract
Oversampling techniques are commonly used to address class imbalance in supervised classification, with SMOTE being a popular approach.However, traditional SMOTE generates synthetic samples within the neighbourhood of minority instances, which can increase data complexity and hinder class separability.This work proposes SMOTE k-out, which creates synthetic samples outside the local neighbourhood to increase minority class sparsity.This aims to reduce overfitting and mitigate the impact of noise, thereby improving the definition of the decision boundary.Experiments on multiple imbalanced datasets demonstrate that SMOTE k-out consistently reduces complexity and achieves higher accuracy and F-measure, particularly with SVM and LDA classifiers.
Verónica Bolón-Canedo, José Luis Morillo-Salas, Laura Moran-Fernandez, Amparo Alonso-Betanzos
ESANN3
2026 Information-Theoretic Unsupervised Feature Selection for High-Dimensional Spatial Data
abstract
High-dimensional unlabelled datasets present significant challenges for efficient analysis, storage and interpretation.Unsupervised feature selection offers a way to retain the most informative variables while discarding redundant or uninformative ones, enabling more scalable processing.We introduce a spatially aware, unsupervised method that uses information theoretic criteria to identify informative variables while limiting redundancy, producing compact and spatially dispersed subsets of features.Our approach avoids dependence on labelled data or modelspecific wrappers, making it suitable for large unstructured datasets.Experiments on MNIST and EMNIST datasets, including high-resolution upscaled versions, show that the selected features preserve both discriminative structure and reconstruction quality better than chosen supervised and unsupervised baselines, demonstrating the effectiveness of entropy and mutual information coupling in unlabelled high-dimensional settings.
Samuel Suárez-Marcote, Abhijeet Vishwasrao, Ricardo Vinuesa, Laura Moran-Fernandez, Verónica Bolón-Canedo
ESANN4
2026 Enhancing Classification Performance on Imbalanced Datasets Through Complexity-Guided Oversampling With SMOTE
abstract
ABSTRACT Improving classification performance on imbalanced datasets remains a challenging problem in machine learning. Synthetic oversampling techniques such as SMOTE are widely used to address class imbalance; however, their random interpolation strategy often ignores structural data properties, which may affect classifier generalisation. This work proposes a set of SMOTE‐based strategies that guide the generation of synthetic samples in order to produce structurally simpler training datasets that are easier for classifiers to learn. The first strategy generates more dispersed (outer) synthetic samples to increase class separability with minimal computational overhead. The second and main contribution, SMOTE‐Complex , formulates synthetic sample selection as an explicit optimisation process that minimises measurable training dataset complexity. A clustering‐based variant, SMOTE‐Complex‐Clustering , reduces computational cost by restricting optimisation to feature subspaces while preserving most structural and predictive benefits. The underlying hypothesis is that reducing the structural complexity of the training data can lead to improved predictive behaviour. Extensive experiments on binary and multiclass datasets, using multiple classifiers and complementary evaluation metrics, provide empirical support for this hypothesis across diverse structural conditions. The results indicate moderate but stable improvements in structurally favourable scenarios—particularly binary and moderately complex problems—without systematic degradation in imbalance‐sensitive metrics, while the clustering‐based refinement offers a scalable trade‐off between optimisation strength and computational efficiency.
José Luis Morillo-Salas, Verónica Bolón-Canedo, Laura Moran-Fernandez, Amparo Alonso-Betanzos
Expert Syst. J. Knowl. Eng.3
2025 Efficient ReliefF: A Low-Power Optimization of ReliefF for Resource-Constrained Devices
Samuel Suárez-Marcote, Laura Moran-Fernandez, Verónica Bolón-Canedo
ICANN (1)2
2025 Fast and Frugal Transfer Learning via Precomputed Features and Adaptive Normalization
Daniel Vila-Cruz, Verónica Bolón-Canedo, Laura Moran-Fernandez
IDEAL (1)3
2025 Optimising Resource Use Through Low-Precision Feature Selection: A Performance Analysis of Logarithmic Division and Stochastic Rounding
abstract
ABSTRACT The growth in the number of wearable devices has increased the amount of data produced daily. Simultaneously, the limitations of such devices has also led to a growing interest in the implementation of machine learning algorithms with low‐precision computation. We propose green and efficient modifications of state‐of‐the‐art feature selection methods based on information theory and fixed‐point representation. We tested two potential improvements: stochastic rounding to prevent information loss, and logarithmic division to improve computational and energy efficiency. Experiments with several datasets showed comparable results to baseline methods, with minimal information loss in both feature selection and subsequent classification steps. Our low‐precision approach proved viable even for complex datasets like microarrays, making it suitable for energy‐efficient internet‐of‐things (IoT) devices. While further investigation into stochastic rounding did not yield significant improvements, the use of logarithmic division for probability approximation showed promising results without compromising classification performance. Our findings offer valuable insights into resource‐efficient feature selection that contribute to IoT device performance and sustainability.
Samuel Suárez-Marcote, Laura Moran-Fernandez, Verónica Bolón-Canedo
Expert Syst. J. Knowl. Eng.2
2025 Breaking boundaries: Low-precision conditional mutual information for efficient feature selection
abstract
©2025 Elsevier B.V. All rights reserved. This manuscript version is made available under the CC-BY-NC-ND 4.0 license https://creativecommons.org/licenses/bync-nd/4.0/. This version of the article has been accepted for publication in Pattern Recognition. The Version of Record is available online at https://doi.org/10.1016/j.patcog.2025.111375
Laura Moran-Fernandez, Eva Blanco-Mallo, Konstantinos Sechidis, Verónica Bolón-Canedo
Pattern Recognit.1
2024 Mutual Information-Based Feature Selection for Federated Learning Environments
abstract
Due to the growth of Internet of Things devices, the dimensionality of the data has only increased. These devices generate a large amount of data which, after being processed, is one of the main sources of information for machine learning systems. However, only a small part of this data is truly relevant. Understanding the relevance of this data allows us to process it at the edge of the network and improve the performance of Artificial Intelligence systems. On the one hand, feature selection is one of the most common approaches to reducing irrelevant features. On the other hand, federated learning allows the exploitation of such data by machine learning models without the need to exchange raw data. Therefore, in this paper we present modifications to two widely utilised feature selection algorithms based on the Mutual Information metric to enable them to work in a federated environment. The experimental process conducted on several datasets and data distributions demonstrates the capability of our modifications to operate in such an environment without a loss of information. This enhances execution time and leverages specific particularities of these environments, such as increased security.
Samuel Suárez-Marcote, Laura Moran-Fernandez, Verónica Bolón-Canedo
ICMLA2
2024 A review of green artificial intelligence: Towards a more sustainable future
abstract
Green artificial intelligence (AI) is more environmentally friendly and inclusive than conventional AI, as it not only produces accurate results without increasing the computational cost but also ensures that any researcher with a laptop can perform high-quality research without the need for costly cloud servers. This paper discusses green AI as a pivotal approach to enhancing the environmental sustainability of AI systems. Described are AI solutions for eco-friendly practices in other fields (green-by AI), strategies for designing energy-efficient machine learning (ML) algorithms and models (green-in AI), and tools for accurately measuring and optimizing energy consumption. Also examined are the role of regulations in promoting green AI and future directions for sustainable ML. Underscored is the importance of aligning AI practices with environmental considerations, fostering a more eco-conscious and energy-efficient future for AI systems.
Verónica Bolón-Canedo, Laura Moran-Fernandez, Brais Cancela, Amparo Alonso-Betanzos
Neurocomputing2
2024 Finding a needle in a haystack: insights on feature selection for classification tasks
abstract
Abstract The growth of Big Data has resulted in an overwhelming increase in the volume of data available, including the number of features. Feature selection, the process of selecting relevant features and discarding irrelevant ones, has been successfully used to reduce the dimensionality of datasets. However, with numerous feature selection approaches in the literature, determining the best strategy for a specific problem is not straightforward. In this study, we compare the performance of various feature selection approaches to a random selection to identify the most effective strategy for a given type of problem. We use a large number of datasets to cover a broad range of real-world challenges. We evaluate the performance of seven popular feature selection approaches and five classifiers. Our findings show that feature selection is a valuable tool in machine learning and that correlation-based feature selection is the most effective strategy regardless of the scenario. Additionally, we found that using improper thresholds with ranker approaches produces results as poor as randomly selecting a subset of features.
Laura Moran-Fernandez, Verónica Bolón-Canedo
J. Intell. Inf. Syst.1
2024 CUDA acceleration of MI-based feature selection methods
abstract
Feature selection algorithms are necessary nowadays for machine learning as they are capable of removing irrelevant and redundant information to reduce the dimensionality of the data and improve the quality of subsequent analyses. The problem with current feature selection approaches is that they are computationally expensive when processing large datasets. This work presents parallel implementations for Nvidia GPUs of three highly-used feature selection methods based on the Mutual Information (MI) metric: mRMR, JMI and DISR. Publicly available code includes not only CUDA implementations of the general methods, but also an adaptation of them to work with low-precision fixed point in order to further increase their performance on GPUs. The experimental evaluation was carried out on two modern Nvidia GPUs (Turing T4 and Ampere A100) with highly satisfactory results, achieving speedups of up to 283x when compared to state-of-the-art C implementations.
Bieito Beceiro, Jorge González-Domínguez, Laura Moran-Fernandez, Verónica Bolón-Canedo, Juan Touriño
J. Parallel Distributed Comput.3
2023 Green Machine Learning
abstract
Green machine learning refers to research that is more environmentally friendly and inclusive, not only by producing novel results without increasing the computational cost, but also by ensuring that any researcher with a laptop has the opportunity to perform high-quality research without the need to use expensive cloud servers.Efficient machine learning approaches (especially deep learning) are starting to receive some attention in the research community.This tutorial is concerned with the development of machine learning algorithms that optimize efficiency rather than only accuracy.We provide an overview of this recent field, together with a review of the novel contributions to the ESANN 2023 special session on Green Machine Learning. * This work
Verónica Bolón-Canedo, Laura Moran-Fernandez, Brais Cancela, Amparo Alonso-Betanzos
ESANN2
2023 Efficient feature selection for domain adaptation using Mutual Information Maximization
abstract
Green AI, an emerging research field, focuses on improving the efficiency of machine learning models.In this paper, we introduce a novel and efficient method for feature selection in domain adaptation, a type of transfer learning where the source and target domains share the feature space and task but differ in their distributions.Instead of using evolutionary algorithms, a typical approach in this field, we propose the use of filter methods, which do not require an iterative search process and are less computationally expensive.Our proposed method is Mutual Information Maximization, and our experiments show that it outperforms Particle Swarm Optimization in terms of efficiency, speed, and the ability to select a reduced subset of features while achieving competitive classification accuracy results.
Guillermo Castillo García, Laura Moran-Fernandez, Verónica Bolón-Canedo
ESANN2
2023 Logarithmic division for green feature selection: an information-theoretic approach
abstract
Feature selection is a popular preprocessing step to reduce the dimensionality of the data while preserving the important information.In this paper we propose an efficient and green feature selection method based on information theory, with the novelty of using the logarithmic division and resort to fixed-point precision.The results of experiments conducted on several datasets indicate the potential of our proposal, as it does not incur in significant information loss compared to the standard method, both in the features selected and in the subsequent classification step.This finding opens up possibilities for a new family of green feature selection methods, which would help to minimize energy consumption and carbon emissions.* This work was
Samuel Suárez-Marcote, Laura Moran-Fernandez, Verónica Bolón-Canedo
ESANN2
2023 Feature selection for domain adaptation using complexity measures and swarm intelligence
abstract
Particle Swarm Optimization is an optimization algorithm that mimics the behaviour of a flock of birds, setting multiple particles that explore the search space guided by a fitness function in order to find the best possible solution.We apply the Sticky Binary Particle Swarm Optimization algorithm to perform feature selection for domain adaptation, a specific type of transfer learning in which the source and the target domain have a common feature space, a common task, but different distributions.When applying Particle Swarm Optimization, classification error is usually employed in the fitness function to evaluate the goodness of subsets of features.In this paper, we aim to compare this approach with using complexity metrics instead, under the assumption that reducing the complexity of the problem will lead to results that are independent from the classifier used for testing while being less computationally demanding.Therefore, we carried out experiments to compare the performance of both approaches in terms of classification accuracy, speed and number of features selected.We found out that our proposal, although in some cases incurs in a slight degradation of classification performance, it is indeed faster and selects fewer features, making it a feasible trade-off.
Guillermo Castillo García, Laura Moran-Fernandez, Verónica Bolón-Canedo
Neurocomputing2
2023 Do all roads lead to Rome? Studying distance measures in the context of machine learning
abstract
Many machine learning and data mining tasks are based on distance measures, so a large amount of literature addresses this aspect somehow. Due to the broad scope of the topic, this paper aims to provide an overview of the use of these measures in the most common machine learning problems, pointing out those aspects to consider to choose the most appropriate measure for a particular task. For this purpose, the most recent works addressing the subject were reviewed and seven of the most commonly used measures were analyzed, investigating in detail their main properties and applications. Different experiments were carried out to study their relationships and compare their performance. The degradation of the results in the presence of noise was also considered, as well as the execution time required by each measure.
Eva Blanco-Mallo, Laura Moran-Fernandez, Beatriz Remeseiro, Verónica Bolón-Canedo
Pattern Recognit.2
2023 Machine Learning Methods for Predicting League of Legends Game Outcome
abstract
The video gameLeague of Legendshas several professional leagues and tournaments that offer prizes reaching several million dollars, making it one of the most followed games in the Esports scene. This article addresses the prediction of the winning team in professional matches of the game, using only pregame data. We propose to improve the accuracy of the models trained with the features offered by the game application programming interface (API). To this end, new features are built to collect interesting information, such as the skills of a player handling a certain champion, the synergies between players of the same team or the ability of a player to beat another player. Then, we perform feature selection and train different classification algorithms aiming at obtaining the best model. Experimental results show classification accuracy above 0.70, which is comparable to the results of other proposals presented in the literature, but with the added benefit of using few samples and not requiring the use of external sources to collect additional statistics.
Juan Agustín Hitar-García, Laura Moran-Fernandez, Verónica Bolón-Canedo
IEEE Trans. Games2
2022 Feature selection for transfer learning using particle swarm optimization and complexity measures
abstract
Particle Swarm Optimization is an optimization algorithm that explores a search space guided by a fitness function in order to find a good solution.We apply it to perform feature selection for domain adaptation.Usually, classification error is used in the fitness function to evaluate the goodness of subsets of features.In this paper, we propose to employ complexity metrics instead, as we assume that reducing the complexity of the problem will lead to good results while being less computationally demanding and independent from the classifier used for testing.We found out that our method is indeed faster and selects fewer features, obtaining competitive classification accuracy results.
Verónica Bolón-Canedo, Guillermo Castillo García, Laura Moran-Fernandez
ESANN3
2022 Less is more: Low-precision feature selection for wearables
abstract
Nowadays, the amount of data produced daily has significantly increased due to the growth in the number of wearable devices. Similarly, this increase is also visible in the interest of developing machine learning algorithms with reduced precision computations, due to the limitations of such devices. This work studies the effect of using low precision operations in the context of feature selection, a preprocessing step that is becoming necessary to deal with the increasing data dimensionality. This study focuses specifically on feature selection methods based on Mutual Information (one of the most popular and widely-used metrics in this area) and how low precision computations can be carried out obtaining experimental results similar to those achieved by double-precision over several low- and high-dimensional datasets. We observe that the use of 16-bit fixed-point representation makes it possible to obtain feature rankings with high similarity to those obtained in double- precision. Even the rankings obtained with 8 bits and then used in subsequent classification tasks, lead to similar accuracy (no significant difference) to the one obtained when using the 64-bit representation in certain situations.
Samuel Suárez-Marcote, Laura Moran-Fernandez, Verónica Bolón-Canedo
IJCNN2
2022 Reduced precision discretization based on information theory
abstract
In recent years, new technological areas have emerged and proliferated, such as the Internet of Things or embedded systems in drones, which are usually characterized by making use of devices with strict requirements of weight, size, cost and power consumption. As a consequence, there has been a growing interest in the implementation of machine learning algorithms with reduced precision that can be embedded in these constrained devices. These algorithms cover not only learning, but they can also be applied to other stages such as feature selection or data discretization. In this work we study the behavior of the Minimum Description Length Principle (MDLP) discretizer, proposed by Fayyad and Irani, when reduced precision is used, and how much it affects to a typical machine learning pipeline. Experimental results show that the use of fixed-point format is sufficient to achieve performances similar to those obtained when using double-precision format, which opens the door to the use of reduced-precision discretizers in embedded systems, minimizing energy consumption and carbon emissions.
Brais Ares, Laura Moran-Fernandez, Verónica Bolón-Canedo
KES2
2022 How important is data quality? Best classifiers vs best features
abstract
The task of choosing the appropriate classifier for a given scenario is not an easy-to-solve question. First, there is an increasingly high number of algorithms available belonging to different families. And also there is a lack of methodologies that can help on recommending in advance a given family of algorithms for a certain type of datasets. Besides, most of these classification algorithms exhibit a degradation in the performance when faced with datasets containing irrelevant and/or redundant features. In this work we analyze the impact of feature selection in classification over several synthetic and real datasets. The experimental results obtained show that the significance of selecting a classifier decreases after applying an appropriate preprocessing step and, not only this alleviates the choice, but it also improves the results in almost all the datasets tested.
Laura Moran-Fernandez, Verónica Bolón-Canedo, Amparo Alonso-Betanzos
Neurocomputing1
2020 Do we need hundreds of classifiers or a good feature selection?
Laura Moran-Fernandez, Verónica Bolón-Canedo, Amparo Alonso-Betanzos
ESANN1
2020 Feature selection with limited bit depth mutual information for portable embedded systems
Laura Moran-Fernandez, Konstantinos Sechidis, Verónica Bolón-Canedo, Amparo Alonso-Betanzos, Gavin Brown 0001
Knowl. Based Syst.1
2019 Distributed classification based on distances between probability distributions in feature space
Pablo Montero-Manso, Laura Moran-Fernandez, Verónica Bolón-Canedo, José Antonio Vilar, Amparo Alonso-Betanzos
Inf. Sci.2
2017 A distributed approach for classification using distance metrics
Laura Moran-Fernandez, Verónica Bolón-Canedo, Amparo Alonso-Betanzos
ESANN1
2017 Can classification performance be predicted by complexity measures? A study using microarray data
Laura Moran-Fernandez, Verónica Bolón-Canedo, Amparo Alonso-Betanzos
Knowl. Inf. Syst.1
2017 Centralized vs. distributed feature selection methods based on data complexity measures
Laura Moran-Fernandez, Verónica Bolón-Canedo, Amparo Alonso-Betanzos
Knowl. Based Syst.1
2016 Data complexity measures for analyzing the effect of SMOTE over microarrays
Laura Moran-Fernandez, Verónica Bolón-Canedo, Amparo Alonso-Betanzos
ESANN1
2015 An insight on complexity measures and classification in microarray data
abstract
Microarray data classification has been typically seen as a difficult challenge for machine learning researchers mainly due to its high dimension in feature while sample size is small. However, this type of data presents other complications such as overlapping between classes, dataset shift, class imbalance, non-linearity, or features extracted under extremely different distributions. This paper intends to analyze in depth the theoretical complexity of several popular binary datasets, by making use of complexity measures, and then connecting it with the empirical results obtained by four widely-used classifiers. Two different situations are covered: datasets with only training set and datasets originally divided into training and test sets. In both cases it is demonstrated that there exists a correlation between the complexity measures and the actual error rates, which can facilitate in the future how to deal with a given dataset. Finally, we present a case study on Prostate dataset, improving the test classification accuracy from 53% to 97%.
Verónica Bolón-Canedo, Laura Moran-Fernandez, Amparo Alonso-Betanzos
IJCNN2