EDBT 2026 Demo / reviewers in the wild / expert
Verónica Bolón-Canedo
dblp:43/8485
· DBLP profile ↗
15ranked-venue papers in the field
4as first author
5since 2021 · last 2026
0000-0002-0524-6427ORCID · verified
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 8 (2 first)Data Mining & Knowledge Discovery · 5 (2 first)Database Systems & Data Management · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving restaurant recommendation transparency through feature selection
Roger Bagué-Masanés, Beatriz Remeseiro, Verónica Bolón-Canedo |
Knowl. Inf. Syst. | 3 |
| 2024 | Fed-mRMR: A lossless federated feature selection methodabstractFeature selection has become a mandatory task in data mining, due to the overwhelming amount of features in Big Data problems. To handle this high-dimensional data and avoid the well-known curse of dimensionality, we need to pre-select an optimal subset of features to reduce redundant computations. Federated learning is a machine learning technique based on training an algorithm over many decentralized edge devices holding local rather than global data on a centralized server. Application of this technique is extending to fields such as self-driving cars, medicine and health, and Industry 4.0, where data privacy is compulsory. Feature selection through federated learning is a complicated task since suboptimal features calculated by feature selection methods may be different in heterogeneous datasets from different nodes. In this paper, we propose a lossless federated version of the classic minimum redundancy maximum relevance (mRMR) feature selection algorithm, called federated mRMR (fed-mRMR), which, without losing any effectiveness of the original mRMR method, is applicable to federated learning approaches and capable of dealing with data that are not independent and identically distributed (non-IID data). Implementation can be found at: https://github.com/jorgehermo9/fed-mrmr Jorge Hermo, Verónica Bolón-Canedo, Susana Ladra |
Inf. Sci. | 2 |
| 2024 | Finding a needle in a haystack: insights on feature selection for classification tasksabstractAbstract The growth of Big Data has resulted in an overwhelming increase in the volume of data available, including the number of features. Feature selection, the process of selecting relevant features and discarding irrelevant ones, has been successfully used to reduce the dimensionality of datasets. However, with numerous feature selection approaches in the literature, determining the best strategy for a specific problem is not straightforward. In this study, we compare the performance of various feature selection approaches to a random selection to identify the most effective strategy for a given type of problem. We use a large number of datasets to cover a broad range of real-world challenges. We evaluate the performance of seven popular feature selection approaches and five classifiers. Our findings show that feature selection is a valuable tool in machine learning and that correlation-based feature selection is the most effective strategy regardless of the scenario. Additionally, we found that using improper thresholds with ranker approaches produces results as poor as randomly selecting a subset of features. Laura Moran-Fernandez, Verónica Bolón-Canedo |
J. Intell. Inf. Syst. | 2 |
| 2022 | Fast anomaly detection with locality-sensitive hashing and hyperparameter autotuningabstractThis paper presents LSHAD, an anomaly detection (AD) method based on Locality Sensitive Hashing (LSH), capable of dealing with large-scale datasets. The resulting algorithm is highly parallelizable and its implementation in Apache Spark further increases its ability to handle very large datasets. Moreover, the algorithm incorporates an automatic hyperparameter tuning mechanism so that users do not have to implement costly manual tuning. Our LSHAD method is novel as both hyperparameter automation and distributed properties are not usual in AD techniques. Our results for experiments with LSHAD across a variety of datasets point to state-of-the-art AD performance while handling much larger datasets than state-of-the-art alternatives. In addition, evaluation results for the tradeoff between AD performance and scalability show that our method offers significant advantages over competing methods. Jorge Meira, Carlos Eiras-Franco, Verónica Bolón-Canedo, Goreti Marreiros, Amparo Alonso-Betanzos |
Inf. Sci. | 3 |
| 2021 | Dealing with heterogeneity in the context of distributed feature selection for classification
José Luis Morillo-Salas, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
Knowl. Inf. Syst. | 2 |
| 2019 | Case Study of Anomaly Detection and Quality Control of Energy Efficiency and Hygrothermal Comfort in Buildingsabstract[Abstract] The aim of this work is to propose different statistical and machine learning methodologies for identifying anomalies and control the quality of energy efficiency and hygrothermal comfort in buildings. Companies focused on energy sector for buildings are interested on statistical and machine learning tools to automate the control of energy consumption and ensure quality of Heat Ventilation and Air Conditioning (HVAC) installations. Consequently, a methodology based on the application of the Local Correlation Integral (LOCI) anomaly detection technique has been proposed. In addition, the most critical variables for anomaly detection are identified by using ReliefF method. Once vectors of critical variables are obtained, multivariate and univariate control charts can be applied to control the quality of HVAC installations (consumption, thermal comfort). In order to test the proposed methodology, the companies involved in this project have provided the case study of a store of a clothing brand located in a shopping center in Panama. It is important to note that this is a controlled case study for which all the anomalies have been previously identified by maintenance personnel. Moreover, as an alternatively solution, in addition to machine learning and multivariate techniques, new nonparametric control charts for functional data based on data depth have been proposed and applied to curves of daily energy consumption in HVAC. Carlos Eiras-Franco, Miguel Flores, Verónica Bolón-Canedo, Sonia Zaragoza, Rubén Fernández-Casal, Salvador Naya, Javier Tarrío-Saavedra |
DATA | 3 |
| 2019 | Insights into distributed feature ranking
Verónica Bolón-Canedo, Konstantinos Sechidis, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, Gavin Brown 0001 |
Inf. Sci. | 1 |
| 2019 | Parallel feature selection for distributed-memory clusters
Jorge González-Domínguez, Verónica Bolón-Canedo, Borja Freire, Juan Touriño |
Inf. Sci. | 2 |
| 2019 | Distributed classification based on distances between probability distributions in feature space
Pablo Montero-Manso, Laura Moran-Fernandez, Verónica Bolón-Canedo, José Antonio Vilar, Amparo Alonso-Betanzos |
Inf. Sci. | 3 |
| 2018 | On the scalability of feature selection methods on high-dimensional data
Verónica Bolón-Canedo, Diego Fernández-Francos, Diego Peteiro-Barral, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Noelia Sánchez-Maroño |
Knowl. Inf. Syst. | 1 |
| 2017 | Fast-mRMR: Fast Minimum Redundancy Maximum Relevance Algorithm for High-Dimensional Big DataabstractWith the advent of large-scale problems, feature selection has become a fundamental preprocessing step to reduce input dimensionality. The minimum-redundancy-maximum-relevance (mRMR) selector is considered one of the most relevant methods for dimensionality reduction due to its high accuracy. However, it is a computationally expensive technique, sharply affected by the number of features. This paper presents fast-mRMR, an extension of mRMR, which tries to overcome this computational burden. Associated with fast-mRMR, we include a package with three implementations of this algorithm in several platforms, namely, CPU for sequential execution, GPU (graphics processing units) for parallel computing, and Apache Spark for distributed computing using big data technologies. Sergio Ramírez-Gallego, Iago Lastra, David Martínez-Rego, Verónica Bolón-Canedo, José Manuel Benítez 0001, Francisco Herrera, Amparo Alonso-Betanzos |
Int. J. Intell. Syst. | 4 |
| 2017 | Can classification performance be predicted by complexity measures? A study using microarray data
Laura Moran-Fernandez, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
Knowl. Inf. Syst. | 2 |
| 2016 | A comparison of performance of K-complex classification methods using feature selection
Elena Hernández-Pereira, Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Diego Álvarez-Estévez, Vicente Moret-Bonillo, Amparo Alonso-Betanzos |
Inf. Sci. | 2 |
| 2014 | A review of microarray datasets and applied feature selection methods
Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, José Manuel Benítez 0001, Francisco Herrera |
Inf. Sci. | 1 |
| 2013 | A review of feature selection methods on synthetic data
Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
Knowl. Inf. Syst. | 1 |