EDBT 2026 Demo / reviewers in the wild / expert
Amparo Alonso-Betanzos
dblp:04/4614
· DBLP profile ↗
143ranked-venue papers
11as first author
27since 2021 · last 2026
0000-0003-0950-0012ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 119 · 10 first-author · 24 since 2021Databases, data management, data science and information retrieval · 15 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SMOTE k-out: Enhancing Class Separability through Outer Synthetic SamplingabstractOversampling techniques are commonly used to address class imbalance in supervised classification, with SMOTE being a popular approach.However, traditional SMOTE generates synthetic samples within the neighbourhood of minority instances, which can increase data complexity and hinder class separability.This work proposes SMOTE k-out, which creates synthetic samples outside the local neighbourhood to increase minority class sparsity.This aims to reduce overfitting and mitigate the impact of noise, thereby improving the definition of the decision boundary.Experiments on multiple imbalanced datasets demonstrate that SMOTE k-out consistently reduces complexity and achieves higher accuracy and F-measure, particularly with SVM and LDA classifiers. Verónica Bolón-Canedo, José Luis Morillo-Salas, Laura Moran-Fernandez, Amparo Alonso-Betanzos |
ESANN | 4 |
| 2026 | FedHENet: A Frugal Federated Learning Framework for Heterogeneous EnvironmentsabstractFederated Learning (FL) enables collaborative training without centralizing data, essential for privacy compliance in real-world scenarios involving sensitive visual information.Most FL approaches rely on expensive, iterative deep network optimization, which still risks privacy via shared gradients.In this work, we propose FedHENet, extending the FedHEONN framework to image classification.By using a fixed, pretrained feature extractor and learning only a single output layer, we avoid costly local fine-tuning.This layer is learned by analytically aggregating client knowledge in a single round of communication using homomorphic encryption (HE).Experiments show that FedHENet achieves competitive accuracy compared to iterative FL baselines while demonstrating superior stability performance and up to 70% better energy efficiency.Crucially, our method is hyperparameter-free, removing the carbon footprint associated with hyperparameter tuning in standard FL. Alejandro Dopico-Castro, Oscar Fontenla-Romero, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Iván Pérez Digón |
ESANN | 3 |
| 2026 | Enhancing Classification Performance on Imbalanced Datasets Through Complexity-Guided Oversampling With SMOTEabstractABSTRACT Improving classification performance on imbalanced datasets remains a challenging problem in machine learning. Synthetic oversampling techniques such as SMOTE are widely used to address class imbalance; however, their random interpolation strategy often ignores structural data properties, which may affect classifier generalisation. This work proposes a set of SMOTE‐based strategies that guide the generation of synthetic samples in order to produce structurally simpler training datasets that are easier for classifiers to learn. The first strategy generates more dispersed (outer) synthetic samples to increase class separability with minimal computational overhead. The second and main contribution, SMOTE‐Complex , formulates synthetic sample selection as an explicit optimisation process that minimises measurable training dataset complexity. A clustering‐based variant, SMOTE‐Complex‐Clustering , reduces computational cost by restricting optimisation to feature subspaces while preserving most structural and predictive benefits. The underlying hypothesis is that reducing the structural complexity of the training data can lead to improved predictive behaviour. Extensive experiments on binary and multiclass datasets, using multiple classifiers and complementary evaluation metrics, provide empirical support for this hypothesis across diverse structural conditions. The results indicate moderate but stable improvements in structurally favourable scenarios—particularly binary and moderately complex problems—without systematic degradation in imbalance‐sensitive metrics, while the clustering‐based refinement offers a scalable trade‐off between optimisation strength and computational efficiency. José Luis Morillo-Salas, Verónica Bolón-Canedo, Laura Moran-Fernandez, Amparo Alonso-Betanzos |
Expert Syst. J. Knowl. Eng. | 4 |
| 2026 | Contrastive Learning for Explanation RankingabstractAbstract Explainable recommendation systems enhance user trust and satisfaction by revealing the reasoning behind personalized recommendations. Approaching this as a post-hoc explanation-ranking problem over a fixed pool of candidate explanations, we propose Contrastive Learning for Explanation Ranking (CLER), a model that learns user, item, and explanation representations with a Normalized Temperature-scaled Binary Cross-Entropy (NT-BXent) loss. This function specifically applies a per-row reweighting strategy, preventing the vast number of negative examples from dominating the objective. We evaluate CLER on the Amazon, TripAdvisor, and Yelp datasets from the EXTRA benchmark. Across traditional ranking metrics, CLER achieves the strongest results among the compared baselines. Miguel Escarda-Fernández, Brais Cancela, Carlos Eiras-Franco, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos |
Mach. Learn. | 5 |
| 2025 | Rethinking Efficiency in Machine LearningabstractThe success of Artificial Intelligence (AI) has so far relied on developing increasingly precise models. However, this has come at the cost of greater complexity, requiring a higher number of parameters to estimate. As a result, model transparency and explainability have diminished, while the energy demands for training and deployment have skyrocketed. It is estimated that by 2030, AI could account for more than 30% of the planet's total energy consumption. Amparo Alonso-Betanzos |
GECCO | 1 |
| 2025 | Sustainable Techniques to Improve Data Quality for Training Image-Based Explanatory Models for Recommender Systems
Jorge Paz-Ruza, David Esteban Martínez, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas |
ICANN (3) | 3 |
| 2025 | Empowering AI Through Frugality
Amparo Alonso-Betanzos |
ICPRAM | 1 |
| 2025 | Predictively Combatting Toxicity in Health-related Online Discussions through Machine LearningabstractIn health-related topics, user toxicity in online discussions frequently becomes a source of social conflict or promotion of dangerous, unscientific behaviour; common approaches for battling it include different forms of detection, flagging and/or removal of existing toxic comments, which is often counterproductive for platforms and users alike. In this work, we propose the alternative of combatting user toxicity predictively, anticipating where a user could interact toxically in health-related online discussions. Applying a Collaborative Filtering-based Machine Learning methodology, we predict the toxicity in COVID-related conversations between any user and subcommunity of Reddit, surpassing 80% predictive performance in relevant metrics, and allowing us to prevent the pairing of conflicting users and subcommunities. Jorge Paz-Ruza, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Carlos Eiras-Franco |
IJCNN | 2 |
| 2025 | An agent-based model to simulate the public acceptability of social innovationsabstractAbstract The successful adoption of social innovations, such as renewable energy systems or pollution reduction plans in cities, depends, to a large extent, on the willingness and participation of the population in their development and implementation. We present an agent‐based model (ABM) to analyze the process of citizen acceptability of a social innovation that uses a variety of agents to represent individual citizens and relevant groups of citizens. Citizen agents make use of the HUMAT cognitive decision‐making model, based on psychosocial theories, to decide on their support for the social innovation considering how their needs will be satisfied if they decide to support (or not) the innovation project, and the influence exerted by the agents in their environment. The ABM was initially developed to represent the urban and transport planning superblock project in the city of Vitoria‐Gasteiz (Spain). The ABM simulations make it possible to study the evolution of public acceptance of social innovation, with the results providing insights to the social dynamics and individual factors that affect the acceptance of the project, enabling an evaluation of how to devise new policies that increase public acceptance. Sufficiently generic to be easily adaptable to different types of social innovations, the ABM is a powerful tool to explore different scenarios and design strategies that foster the acceptance and sustainable adoption of social innovations. Alejandro Rodríguez-Arias, Noelia Sánchez-Maroño, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Isabel Lema-Blanco, Adina Dumitru |
Expert Syst. J. Knowl. Eng. | 4 |
| 2025 | Performance and sustainability of BERT derivatives in dyadic dataabstract[Abstract]: In recent years, the Natural Language Processing (NLP) field has experienced a revolution, where numerous models – based on the Transformer architecture – have emerged to process the ever-growing volume of online text-generated data. This architecture has been the basis for the rise of Large Language Models (LLMs). Enabling their application to many diverse tasks in which they excel with just a fine-tuning process that comes right after a vast pre-training phase. However, their sustainability can often be overlooked, especially regarding computational and environmental costs. Our research aims to compare various BERT derivatives in the context of a dyadic data task while also drawing attention to the growing need for sustainable AI solutions. To this end, we utilize a selection of transformer models in an explainable recommendation setting, modeled as a multi-label classification task originating from a social network context, where users, restaurants, and reviews interact. Miguel Escarda-Fernández, Carlos Eiras-Franco, Brais Cancela, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos |
Expert Syst. Appl. | 5 |
| 2025 | Beyond RMSE and MAE: Introducing EAUC to Unmask Hidden Bias and Unfairness in Dyadic Regression ModelsabstractDyadic regression models, which output real-valued predictions for pairs of entities, are fundamental in many domains [e.g., obtaining user-product ratings in recommender systems (RSs)] and promising and under exploration in others (e.g., tuning patient-drug dosages in precision pharmacology). In this work, we prove that nonuniform observed value distributions of individual entities lead to severe biases in state-of-the-art models, skewing predictions toward the average of observed past values for the entity and providing worse-than-random predictive power in eccentric yet crucial cases; we name this phenomenon eccentricity bias. We show that global error metrics like root-mean-squared error (RMSE) are insufficient to capture this bias, and we introduce eccentricity area under the curve (EAUC) as a novel metric that can quantify it in all studied domains and models. We prove the intuitive interpretation of EAUC by experimenting with naive post-training bias corrections and theorize other options to use EAUC to guide the construction of fair models. This work contributes a bias-aware evaluation of dyadic regression to prevent unfairness in critical real-world applications of such systems. Jorge Paz-Ruza, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Brais Cancela, Carlos Eiras-Franco |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | AI-based algorithm for intrusion detection on a real datasetabstractIn the realm of cybersecurity, the detection of network intrusions stands as a paramount challenge, with ever-evolving threats demanding innovative solutions.This study delves into the application of diverse machine learning algorithms on a contemporary dataset (UGR'16) comprising real-world instances of intrusion in software systems.Specifically, several Machine Learning models (Outlier Detectors, Ensemble Methods, Deep Learning, and Conventional Classifiers) were tested and compared with previously reported results using a standard methodology.The obtained results reveal that the Ensemble Methods have been capable of improving the results from prior research.Particularly, the Extreme Gradient Boosting (XGBoost) algorithm offers better results than the original solution with Random Forest, with an AUC of 0.9218 as opposed to 0.8977, and more than four times as fast for the problem to solve. David Esteban Martínez, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Elena Hernández-Pereira, Alejandro Esteban Martínez |
ESANN | 3 |
| 2024 | Agent-based model to assess public acceptance of an energy sustainability project in the island of El HierroabstractThis paper presents an agent-based model to address citizen acceptance of energy sustainability projects, using the 100% Renewable El Hierro Project as a case study. The island of El Hierro is a pioneer in the search for sustainable energy sources, and this project seeks to understand how the individual psychosocial needs of El Hierro’s citizens influence their perception and adoption of energy policies. The HUMAT architecture, inspired by psychological and sociological theories, is adapted to model the behaviour of the citizens of El Hierro in relation to the energy sustainability project. Each agent in the model represents an individual citizen and their behaviour is determined by their psychosocial needs, which include factors such as the island’s energy independence, economic sustainability or environmental quality. Simulations are used to assess the impact of different communication strategies by project stakeholders, such as the press or local government, on the evolution of public acceptance of a project extension. This study highlights the importance of integrating multidisciplinary approaches combining psychology, sociology and engineering to address energy sustainability challenges in specific contexts such as El Hierro. It also offers new perspectives on how to design and implement energy policy communication campaigns in a way that is more acceptable to the public. Alejandro Rodríguez-Arias, Noelia Sánchez-Maroño, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos |
KES | 4 |
| 2024 | The imbalance problem: A comparison of sampling approaches using different parameters and feature selection methods in the context of classificationabstractAbstract A common situation in classification tasks is to deal with unbalanced datasets, an issue that appears when the majority class(es) has a large number of samples compared to the minority class(es). This problem is even more significant when the datasets have a large number of features but only a few samples, as is the case with microarray datasets. Traditionally, an approach to alleviate this problem has been the application of sampling methods to obtain more balanced classes, increasing the number of samples in the minority class (replicating samples or generating new synthetic samples), or decreasing the number of samples in the majority class. In this study, we have compared different balancing methods, including a novel method that applies sampling in both the minority and majority classes. The interest in applying feature selection in combination with balancing methods has also been explored. In view of the results, a recommendation of sampling method, feature selection, and classifier is proposed to improve the classification results according to the type of dataset. José Luis Morillo-Salas, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
Expert Syst. J. Knowl. Eng. | 3 |
| 2024 | A review of green artificial intelligence: Towards a more sustainable futureabstractGreen artificial intelligence (AI) is more environmentally friendly and inclusive than conventional AI, as it not only produces accurate results without increasing the computational cost but also ensures that any researcher with a laptop can perform high-quality research without the need for costly cloud servers. This paper discusses green AI as a pivotal approach to enhancing the environmental sustainability of AI systems. Described are AI solutions for eco-friendly practices in other fields (green-by AI), strategies for designing energy-efficient machine learning (ML) algorithms and models (green-in AI), and tools for accurately measuring and optimizing energy consumption. Also examined are the role of regulations in promoting green AI and future directions for sustainable ML. Underscored is the importance of aligning AI practices with environmental considerations, fostering a more eco-conscious and energy-efficient future for AI systems. Verónica Bolón-Canedo, Laura Moran-Fernandez, Brais Cancela, Amparo Alonso-Betanzos |
Neurocomputing | 4 |
| 2023 | Green Machine LearningabstractGreen machine learning refers to research that is more environmentally friendly and inclusive, not only by producing novel results without increasing the computational cost, but also by ensuring that any researcher with a laptop has the opportunity to perform high-quality research without the need to use expensive cloud servers.Efficient machine learning approaches (especially deep learning) are starting to receive some attention in the research community.This tutorial is concerned with the development of machine learning algorithms that optimize efficiency rather than only accuracy.We provide an overview of this recent field, together with a review of the novel contributions to the ESANN 2023 special session on Green Machine Learning. * This work Verónica Bolón-Canedo, Laura Moran-Fernandez, Brais Cancela, Amparo Alonso-Betanzos |
ESANN | 4 |
| 2023 | Automated green machine learning for condition-based maintenanceabstractWithin the big data paradigm, there is an increasing demand for machine learning with automatic configuration of hyperparameters.Although several algorithms have been proposed for automatically learning time-changing concepts, they generally do not scale well to very large databases.In this context, this paper presents an automated green machine learning approach applied to condition-based maintenance with automatic data fusion and density-based anomaly detection based on locality sensitivity hashing.Experiments on numerical simulations of train-track dynamic interactions demonstrate the utility of the approach to detect railway wheel out-of-roundness.This unlocks the full potential of scalable machine learning, paving the way for environment-friendly systems and automated decision-making. Afonso Lourenço, Carolina Ferraz, Jorge Meira, Goreti Marreiros, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
ESANN | 6 |
| 2023 | Data-driven predictive maintenance framework for railway systemsabstractThe emergence of the Industry 4.0 trend brings automation and data exchange to industrial manufacturing. Using computational systems and IoT devices allows businesses to collect and deal with vast volumes of sensorial and business process data. The growing and proliferation of big data and machine learning technologies enable strategic decisions based on the analyzed data. This study suggests a data-driven predictive maintenance framework for the air production unit (APU) system of a train of Metro do Porto. The proposed method assists in detecting failures and errors in machinery before they reach critical stages. We present an anomaly detection model following an unsupervised approach, combining the Half-Space-trees method with One Class K Nearest Neighbor, adapted to deal with data streams. We evaluate and compare our approach with the Half-Space-Trees method applied without the One Class K Nearest Neighbor combination. Our model produced few type-I errors, significantly increasing the value of precision when compared to the Half-Space-Trees model. Our proposal achieved high anomaly detection performance, predicting most of the catastrophic failures of the APU train system. Jorge Meira, Bruno M. Veloso, Verónica Bolón-Canedo, Goreti Marreiros, Amparo Alonso-Betanzos, João Gama 0001 |
Intell. Data Anal. | 5 |
| 2023 | E2E-FS: An End-to-End Feature Selection Method for Neural NetworksabstractClassic embedded feature selection algorithms are often divided in two large groups: tree-based algorithms and LASSO variants. Both approaches are focused in different aspects: while the tree-based algorithms provide a clear explanation about which variables are being used to trigger a certain output, LASSO-like approaches sacrifice a detailed explanation in favor of increasing its accuracy. In this paper, we present a novel embedded feature selection algorithm, called End-to-End Feature Selection (E2E-FS), that aims to provide both accuracy and explainability in a clever way. Despite having non-convex regularization terms, our algorithm, similar to the LASSO approach, is solved with gradient descent techniques, introducing some restrictions that force the model to specifically select a maximum number of features that are going to be used subsequently by the classifier. Although these are hard restrictions, the experimental results obtained show that this algorithm can be used with any learning model that is trained using a gradient descent algorithm. Brais Cancela, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Sustainable Personalisation and Explainability in Dyadic Data SystemsabstractSystems that rely on dyadic data, which relate entities of two types together, have become ubiquitously used in fields such as media services, tourism business, e-commerce, and others. However, these systems have had a tendency to be black-box systems, despite their objective of influencing people's decisions. There is a lack of research on providing personalised explanations to the outputs of systems that make use of such data, that is, integrating the idea of Explainable Artificial Intelligence into the field of dyadic data. Moreover, the existing approaches rely heavily on Deep Learning models for their training, reducing their overall sustainability. In this work, we propose a computationally efficient model which provides personalisation by generating explanations based on user-created images. In the context of a particular dyadic data system, the restaurant review platform TripAdvisor, we predict, for any (user, restaurant) pair, the review of the restaurant that is most adequate to present it to the user, based on their personal preferences. This model exploits the usage of efficient Matrix Factorisation techniques combined with feature-rich embeddings of the pre-trained Image Classification models, developing a method capable of providing transparency to dyadic data systems while reducing as much as 80% the carbon emissions of training compared to alternative approaches. Jorge Paz-Ruza, Carlos Eiras-Franco, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos |
KES | 4 |
| 2022 | How Agent-based modeling can help to foster sustainability projectsabstract[Abstract] The Sustainable Development Goals (SDGs) adopted by the United Nations require relevant social changes that sometimes involve the development of innovative projects that cause rejection and confrontation. Agent-Based Models (ABM) are powerful tools to represent the behavior of systems, and they have become valuable for the social sciences as they can simulate the behavior of a society under different conditions. Superblocks are innovative city projects that reorganize urban space and minimize private motorized transport. In this paper, we present an ABM that simulates the implantation of superblocks in two Spanish cities: Vitoria-Gasteiz and Barcelona. The interest of this model is to provide policymakers with relevant scientific information that can be used to support their planning and decision-making processes by running possible alternative policy scenarios. This paper presents the details of the designed model and the simulation of different policy scenarios to increase the acceptability rates of citizens about the project, demonstrating how the model takes into account local differences and its usefulness for those political leaders from other cities interested in implementing this type of project. Noelia Sánchez-Maroño, Alejandro Rodríguez-Arias, Adina Dumitru, Isabel Lema-Blanco, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos |
KES | 6 |
| 2022 | Machine learning techniques to predict different levels of hospital care of CoVid-19abstractIn this study, we analyze the capability of several state of the art machine learning methods to predict whether patients diagnosed with CoVid-19 (CoronaVirus disease 2019) will need different levels of hospital care assistance (regular hospital admission or intensive care unit admission), during the course of their illness, using only demographic and clinical data. For this research, a data set of 10,454 patients from 14 hospitals in Galicia (Spain) was used. Each patient is characterized by 833 variables, two of which are age and gender and the other are records of diseases or conditions in their medical history. In addition, for each patient, his/her history of hospital or intensive care unit (ICU) admissions due to CoVid-19 is available. This clinical history will serve to label each patient and thus being able to assess the predictions of the model. Our aim is to identify which model delivers the best accuracies for both hospital and ICU admissions only using demographic variables and some structured clinical data, as well as identifying which of those are more relevant in both cases. The results obtained in the experimental study show that the best models are those based on oversampling as a preprocessing phase to balance the distribution of classes. Using these models and all the available features, we achieved an area under the curve (AUC) of 76.1% and 80.4% for predicting the need of hospital and ICU admissions, respectively. Furthermore, feature selection and oversampling techniques were applied and it has been experimentally verified that the relevant variables for the classification are age and gender, since only using these two features the performance of the models is not degraded for the two mentioned prediction problems. Elena Hernández-Pereira, Oscar Fontenla-Romero, Verónica Bolón-Canedo, Brais Cancela, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos |
Appl. Intell. | 6 |
| 2022 | How important is data quality? Best classifiers vs best featuresabstractThe task of choosing the appropriate classifier for a given scenario is not an easy-to-solve question. First, there is an increasingly high number of algorithms available belonging to different families. And also there is a lack of methodologies that can help on recommending in advance a given family of algorithms for a certain type of datasets. Besides, most of these classification algorithms exhibit a degradation in the performance when faced with datasets containing irrelevant and/or redundant features. In this work we analyze the impact of feature selection in classification over several synthetic and real datasets. The experimental results obtained show that the significance of selecting a classifier decreases after applying an appropriate preprocessing step and, not only this alleviates the choice, but it also improves the results in almost all the datasets tested. Laura Moran-Fernandez, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
Neurocomputing | 3 |
| 2022 | Fast anomaly detection with locality-sensitive hashing and hyperparameter autotuningabstractThis paper presents LSHAD, an anomaly detection (AD) method based on Locality Sensitive Hashing (LSH), capable of dealing with large-scale datasets. The resulting algorithm is highly parallelizable and its implementation in Apache Spark further increases its ability to handle very large datasets. Moreover, the algorithm incorporates an automatic hyperparameter tuning mechanism so that users do not have to implement costly manual tuning. Our LSHAD method is novel as both hyperparameter automation and distributed properties are not usual in AD techniques. Our results for experiments with LSHAD across a variety of datasets point to state-of-the-art AD performance while handling much larger datasets than state-of-the-art alternatives. In addition, evaluation results for the tradeoff between AD performance and scalability show that our method offers significant advantages over competing methods. Jorge Meira, Carlos Eiras-Franco, Verónica Bolón-Canedo, Goreti Marreiros, Amparo Alonso-Betanzos |
Inf. Sci. | 5 |
| 2021 | Scalable feature selection using ReliefF aided by locality-sensitive hashingabstractFeature selection algorithms, such as ReliefF, are very important for processing high-dimensionality data sets. However, widespread use of popular and effective such algorithms is limited by their computational cost. We describe an adaptation of the ReliefF algorithm that simplifies the costliest of its step by approximating the nearest neighbor graph using locality-sensitive hashing (LSH). The resulting ReliefF-LSH algorithm can process data sets that are too large for the original ReliefF, a capability further enhanced by distributed implementation in Apache Spark. Furthermore, ReliefF-LSH obtains better results and is more generally applicable than currently available alternatives to the original ReliefF, as it can handle regression and multiclass data sets. The fact that it does not require any additional hyperparameters with respect to ReliefF also avoids costly tuning. A set of experiments demonstrates the validity of this new approach and confirms its good scalability. Carlos Eiras-Franco, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Antonio Bahamonde |
Int. J. Intell. Syst. | 3 |
| 2021 | Dealing with heterogeneity in the context of distributed feature selection for classification
José Luis Morillo-Salas, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
Knowl. Inf. Syst. | 3 |
| 2021 | Wavefront Marching Methods: A Unified Algorithm to Solve Eikonal and Static Hamilton-Jacobi EquationsabstractThis paper presents a unified propagation method for dealing with both the classic Eikonal equation, where the motion direction does not affect the propagation, and the more general static Hamilton-Jacobi equations, where it does. While classic Fast Marching Method (FMM) techniques achieve the solution to the Eikonal equation with a O(M log M) (or O(M) assuming some modifications), solving the more general static Hamilton-Jacobi equation requires a higher complexity. The proposed framework maintains the O(M log M) complexity for both problems, while achieving higher accuracy than available state-of-the-art. The key idea behind the proposed method is the creation of 'mini wave-fronts', where the solution is interpolated to minimize the discretization error. Experimental results show how our algorithm can outperform the state-of-the-art both in precision and computational cost. Brais Cancela, Amparo Alonso-Betanzos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Do we need hundreds of classifiers or a good feature selection?
Laura Moran-Fernandez, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
ESANN | 3 |
| 2020 | A delayed Elastic-Net approach for performing adversarial attacksabstractWith the rise of the so-called Adversarial Attacks, there is an increased concern on model security. In this paper we present two different contributions: novel measures of robustness (based on adversarial attacks) and a novel adversarial attack. The key idea behind these metrics is to obtain a measure that could compare different architectures, with independence of how the input is preprocessed (robustness against different input sizes and value ranges). To do so, a novel adversarial attack is presented, performing a delayed elastic-net adversarial attack (constraints are only used whenever a successful adversarial attack is obtained). Experimental results show that our approach obtains state-of-the-art adversarial samples, in terms of minimal perturbation distance. Finally, a benchmark of ImageNet pretrained models is used to conduct experiments aiming to shed some light about which model should be selected whenever security is a role factor. Brais Cancela, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
ICPR | 3 |
| 2020 | Can data placement be effective for Neural Networks classification tasks? Introducing the Orthogonal Loss
Brais Cancela, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
ICPR | 3 |
| 2020 | Community detection and social network analysis based on the Italian wars of the 15th century
Javier Fumanal, Amparo Alonso-Betanzos, Oscar Cordón, Humberto Bustince, Maria Minárová |
Future Gener. Comput. Syst. | 2 |
| 2020 | A scalable saliency-based feature selection method with instance-level information
Brais Cancela, Verónica Bolón-Canedo, Amparo Alonso-Betanzos, João Gama 0001 |
Knowl. Based Syst. | 3 |
| 2020 | Feature selection with limited bit depth mutual information for portable embedded systems
Laura Moran-Fernandez, Konstantinos Sechidis, Verónica Bolón-Canedo, Amparo Alonso-Betanzos, Gavin Brown 0001 |
Knowl. Based Syst. | 4 |
| 2020 | Fast Distributed kNN Graph Construction Using Auto-tuned Locality-sensitive HashingabstractThe k -nearest-neighbors ( k NN) graph is a popular and powerful data structure that is used in various areas of Data Science, but the high computational cost of obtaining it hinders its use on large datasets. Approximate solutions have been described in the literature using diverse techniques, among which Locality-sensitive Hashing (LSH) is a promising alternative that still has unsolved problems. We present Variable Resolution Locality-sensitive Hashing, an algorithm that addresses these problems to obtain an approximate k NN graph at a significantly reduced computational cost. Its usability is greatly enhanced by its capacity to automatically find adequate hyperparameter values, a common hindrance to LSH-based methods. Moreover, we provide an implementation in the distributed computing framework Apache Spark that takes advantage of the structure of the algorithm to efficiently distribute the computational load across multiple machines, enabling practitioners to apply this solution to very large datasets. Experimental results show that our method offers significant improvements over the state-of-the-art in the field and shows very good scalability as more machines are added to the computation. Carlos Eiras-Franco, David Martínez-Rego, Leslie Kanthan, César Piñeiro, Antonio Bahamonde, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos |
ACM Trans. Intell. Syst. Technol. | 7 |
| 2020 | One-Class Convex Hull-Based Algorithm for Classification in Distributed EnvironmentsabstractIn this paper, a new one-class classification algorithm capable of working in distributed environments is presented. In it, convex hull is used to build the boundary of the target class defining the one-class problem in each of the distributed nodes. Therefore, we will consider several classifiers, each one determined using a given local data partition, and the goal is to obtain a global classification decision. In order to obtain this final decision, two different algebraic combination rules were proposed: 1) sum and 2) majority voting. Experimental results show that this method opens the possibility of tackling practical one-class classification problems in distributed big data scenarios in an efficient and accurate way. Diego Fernández-Francos, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2019 | A scalable decision-tree-based method to explain interactions in dyadic data
Carlos Eiras-Franco, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Antonio Bahamonde |
Decis. Support Syst. | 3 |
| 2019 | Biases in feature selection with missing data
Borja Seijo-Pardo, Amparo Alonso-Betanzos, Kristin P. Bennett, Verónica Bolón-Canedo, Julie Josse, Mehreen Saeed, Isabelle Guyon |
Neurocomputing | 2 |
| 2019 | Insights into distributed feature ranking
Verónica Bolón-Canedo, Konstantinos Sechidis, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, Gavin Brown 0001 |
Inf. Sci. | 4 |
| 2019 | Large scale anomaly detection in mixed numerical and categorical input spaces
Carlos Eiras-Franco, David Martínez-Rego, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Antonio Bahamonde |
Inf. Sci. | 4 |
| 2019 | Distributed classification based on distances between probability distributions in feature space
Pablo Montero-Manso, Laura Moran-Fernandez, Verónica Bolón-Canedo, José Antonio Vilar, Amparo Alonso-Betanzos |
Inf. Sci. | 5 |
| 2019 | Distributed correlation-based feature selection in spark
Raul-Jose Palma-Mendoza, Luis de-Marcos, Daniel Rodríguez-García, Amparo Alonso-Betanzos |
Inf. Sci. | 4 |
| 2018 | Analysis of imputation bias for feature selection with missing data
Borja Seijo-Pardo, Amparo Alonso-Betanzos, Kristin P. Bennett, Verónica Bolón-Canedo, Isabelle Guyon, Julie Josse, Mehreen Saeed |
ESANN | 2 |
| 2018 | On the scalability of feature selection methods on high-dimensional data
Verónica Bolón-Canedo, Diego Fernández-Francos, Diego Peteiro-Barral, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Noelia Sánchez-Maroño |
Knowl. Inf. Syst. | 4 |
| 2018 | An Information Theory-Based Feature Selection Framework for Big Data Under Apache SparkabstractWith the advent of extremely high dimensional datasets, dimensionality reduction techniques are becoming mandatory. Of the many techniques available, feature selection (FS) is of growing interest for its ability to identify both relevant features and frequently repeated instances in huge datasets. We aim to demonstrate that standard FS methods can be parallelized in big data platforms like Apache Spark so as to boost both performance and accuracy. We propose a distributed implementation of a generic FS framework that includes a broad group of well-known information theory-based methods. Experimental results for a broad set of real-world datasets show that our distributed framework is capable of rapidly dealing with ultrahigh-dimensional datasets as well as those with a huge number of samples, outperforming the sequential version in all the cases studied. Sergio Ramírez-Gallego, Héctor Mouriño-Talín, David Martínez-Rego, Verónica Bolón-Canedo, José Manuel Benítez 0001, Amparo Alonso-Betanzos, Francisco Herrera |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2017 | Algorithmic challenges in big data analytics
Verónica Bolón-Canedo, Beatriz Remeseiro, Konstantinos Sechidis, David Martínez-Rego, Amparo Alonso-Betanzos |
ESANN | 5 |
| 2017 | Scalable approximate k-NN Graph construction based on Locality Sensitive Hashing
Carlos Eiras-Franco, Leslie Kanthan, Amparo Alonso-Betanzos, David Martínez-Rego |
ESANN | 3 |
| 2017 | Mutual information for improving the efficiency of the SCH algorithm
Diego Fernández-Francos, Oscar Fontenla-Romero, Amparo Alonso-Betanzos, Gavin Brown 0001 |
ESANN | 3 |
| 2017 | A distributed approach for classification using distance metrics
Laura Moran-Fernandez, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
ESANN | 3 |
| 2017 | Paving the way for providing teaching feedback in automatic evaluation of open response assignmentsabstractPeer grading has been the regular procedure to use for automatic assessment of open ended assignments in Massive Open Online Courses (MOOCs). However, and although the procedure tries to overcome the rupture of the classical teach-learn-assess/feedback cycle, it does so only in the student side, and no attempt has been made as yet in giving feedback to instructors. The work described inhere aims at filling this gap, with a proposal in which the instructors are supplied with the set of words most used by the best and worst ranked quartiles of assignments. In order to achieve this, a Gaussian Mixture Model (GMM) fed with the bag of words supplied by a previous feature selection algorithm is presented, with the goal of identifying the clusters of words related with similar grades. The results obtained over three pilot studies, containing assignments in three different disciplines, show that our model can lead to more complete information on the teacher feedback on the results of the assignments. Verónica Bolón-Canedo, Jorge Díez 0001, Oscar Luaces, Antonio Bahamonde, Amparo Alonso-Betanzos |
IJCNN | 5 |
| 2017 | Exploring the consequences of distributed feature selection in DNA microarray dataabstractMicroarray data classification has been typically seen as a difficult challenge for machine learning researchers mainly due to its high dimension in features while sample size is small. Because of this particularity, feature selection is usually applied trying to reduce its high dimensionality. However, existing algorithms may not scale well when dealing with this amount of features, and a possible solution is to distribute the features into several nodes. In this work we explore the process of distribution on microarray data - which has recently gained attention - and we evaluate to what extent it is possible to obtain similar results as those obtained with the whole dataset. We performed experiments with different aggregation methods, feature rankers and also evaluated the effect of distributing the feature ranking process in the subsequent classification performance. Verónica Bolón-Canedo, Konstantinos Sechidis, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, Gavin Brown 0001 |
IJCNN | 4 |
| 2017 | Fast-mRMR: Fast Minimum Redundancy Maximum Relevance Algorithm for High-Dimensional Big DataabstractWith the advent of large-scale problems, feature selection has become a fundamental preprocessing step to reduce input dimensionality. The minimum-redundancy-maximum-relevance (mRMR) selector is considered one of the most relevant methods for dimensionality reduction due to its high accuracy. However, it is a computationally expensive technique, sharply affected by the number of features. This paper presents fast-mRMR, an extension of mRMR, which tries to overcome this computational burden. Associated with fast-mRMR, we include a package with three implementations of this algorithm in several platforms, namely, CPU for sequential execution, GPU (graphics processing units) for parallel computing, and Apache Spark for distributed computing using big data technologies. Sergio Ramírez-Gallego, Iago Lastra, David Martínez-Rego, Verónica Bolón-Canedo, José Manuel Benítez 0001, Francisco Herrera, Amparo Alonso-Betanzos |
Int. J. Intell. Syst. | 7 |
| 2017 | Can classification performance be predicted by complexity measures? A study using microarray data
Laura Moran-Fernandez, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
Knowl. Inf. Syst. | 3 |
| 2017 | Volume, variety and velocity in Data Science
Amparo Alonso-Betanzos, José A. Gámez 0001, Francisco Herrera, José M. Puerta, José Cristóbal Riquelme Santos |
Knowl. Based Syst. | 1 |
| 2017 | Content-based methods in peer assessment of open-response questions to grade students as authors and as graders
Oscar Luaces, Jorge Díez 0001, Amparo Alonso-Betanzos, Alicia Troncoso Lora, Antonio Bahamonde |
Knowl. Based Syst. | 3 |
| 2017 | Centralized vs. distributed feature selection methods based on data complexity measures
Laura Moran-Fernandez, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
Knowl. Based Syst. | 3 |
| 2017 | Ensemble feature selection: Homogeneous and heterogeneous approaches
Borja Seijo-Pardo, Iago Porto-Díaz, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
Knowl. Based Syst. | 4 |
| 2017 | Testing Different Ensemble Configurations for Feature Selection
Borja Seijo-Pardo, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
Neural Process. Lett. | 3 |
| 2016 | Machine learning for medical applications
Verónica Bolón-Canedo, Beatriz Remeseiro, Amparo Alonso-Betanzos, Aurélio J. C. Campilho |
ESANN | 3 |
| 2016 | One-class classification algorithm based on convex hull
Diego Fernández-Francos, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
ESANN | 3 |
| 2016 | Data complexity measures for analyzing the effect of SMOTE over microarrays
Laura Moran-Fernandez, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
ESANN | 3 |
| 2016 | Using a feature selection ensemble on DNA microarray datasets
Borja Seijo-Pardo, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
ESANN | 3 |
| 2016 | A unified pipeline for online feature selection and classification
Verónica Bolón-Canedo, Diego Fernández-Francos, Diego Peteiro-Barral, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Noelia Sánchez-Maroño |
Expert Syst. Appl. | 4 |
| 2016 | A comparison of performance of K-complex classification methods using feature selection
Elena Hernández-Pereira, Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Diego Álvarez-Estévez, Vicente Moret-Bonillo, Amparo Alonso-Betanzos |
Inf. Sci. | 6 |
| 2016 | Fault detection via recurrence time statistics and one-class classification
David Martínez-Rego, Oscar Fontenla-Romero, Amparo Alonso-Betanzos, José C. Príncipe |
Pattern Recognit. Lett. | 3 |
| 2015 | On the use of machine learning techniques for the analysis of spontaneous reactions in automated hearing assessment
Verónica Bolón-Canedo, Alba Fernández, Amparo Alonso-Betanzos, Marcos Ortega 0001, Manuel G. Penedo |
ESANN | 3 |
| 2015 | Learning features on tear film lipid layer classification
Beatriz Remeseiro, Verónica Bolón-Canedo, Amparo Alonso-Betanzos, Manuel G. Penedo |
ESANN | 3 |
| 2015 | An insight on complexity measures and classification in microarray dataabstractMicroarray data classification has been typically seen as a difficult challenge for machine learning researchers mainly due to its high dimension in feature while sample size is small. However, this type of data presents other complications such as overlapping between classes, dataset shift, class imbalance, non-linearity, or features extracted under extremely different distributions. This paper intends to analyze in depth the theoretical complexity of several popular binary datasets, by making use of complexity measures, and then connecting it with the empirical results obtained by four widely-used classifiers. Two different situations are covered: datasets with only training set and datasets originally divided into training and test sets. In both cases it is demonstrated that there exists a correlation between the complexity measures and the actual error rates, which can facilitate in the future how to deal with a given dataset. Finally, we present a case study on Prostate dataset, improving the test classification accuracy from 53% to 97%. Verónica Bolón-Canedo, Laura Moran-Fernandez, Amparo Alonso-Betanzos |
IJCNN | 3 |
| 2015 | Stream change detection via passive-aggressive classification and Bernoulli CUSUM
David Martínez-Rego, Diego Fernández-Francos, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
Inf. Sci. | 4 |
| 2015 | Recent advances and emerging challenges of feature selection in the context of big data
Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
Knowl. Based Syst. | 3 |
| 2015 | A factorization approach to evaluate open-response assignments in MOOCs using preference learning on peer assessments
Oscar Luaces, Jorge Díez 0001, Amparo Alonso-Betanzos, Alicia Troncoso Lora, Antonio Bahamonde |
Knowl. Based Syst. | 3 |
| 2015 | An Agent-Based Model for Simulating Environmental Behavior in an Educational Organization
Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, Oscar Fontenla-Romero, C. Brinquis-Núñez, J. Gareth Polhill, Tony Craig, Adina Dumitru, R. García-Mira |
Neural Process. Lett. | 2 |
| 2014 | Influence of internal values and social networks for achieving sustainable organizationsabstractThe LOw Carbon At Work (LOCAW) project has studied the everyday behavior of employees in six organizations in order to achieve a more sustainable Europe. Of these six, four organizations were involved in backcasting workshops to obtain future scenarios aimed at significantly improving engagement with pro-environmental behaviors by 2050. From these scenarios policies were extracted from the workshop participants that achieve this aim in their organization. Agent Based Models (ABM) were designed to model the organizations using actual information from the organization; the design also placed special emphasis on the representation of the social network. ABMs were then used to simulate the effects of the different policies derived from the backcasting scenarios. In this paper, the results for two organizations, UDC and Aquatim, are shown. These experimental results demonstrate the influence of different social networks and internal motivations of employees to determine the effectiveness of a given policy. Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, Oscar Fontenla-Romero, C. Brinquis-Núñez, J. Gareth Polhill, Tony Craig |
ECAI | 2 |
| 2014 | Modeling consumption of contents and advertising in online newspapers
Iago Porto-Díaz, David Martínez-Rego, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
ESANN | 4 |
| 2014 | Learning on Vertically Partitioned Data based on Chi-square Feature Selection and Naive Bayes ClassificationabstractIn the last few years, distributed learning has been the focus of much attention due to the explosion of big databases, in some cases distributed across different nodes. However, the great majority of current selection and classification algorithms are designed for centralized learning, i.e. they use the whole dataset at once. In this paper, a new approach for learning on vertically partitioned data is presented, which covers both feature selection and classification. The approach splits the data by features, and then uses the chi-square filter and the naive Bayes classifier to learn at each node. Finally, a merging procedure is performed, which updates the learned model in an incremental fashion. The experimental results on five representative datasets show that the execution time is shortened considerably whereas the classification performance is maintained as the number of nodes increases. Verónica Bolón-Canedo, Diego Peteiro-Barral, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Noelia Sánchez-Maroño |
ICAART (1) | 3 |
| 2014 | mC-ReliefF - An Extension of ReliefF for Cost-based Feature SelectionabstractThe proliferation of high-dimensional data in the last few years has brought a necessity to use dimensionality reduction techniques, in which feature selection is arguably the most famous one. Feature selection consists of detecting relevant features and discarding the irrelevant ones. However, there are some situations where the users are not only interested in the relevance of the selected features but also in the costs that they imply (e.g. economical or computational costs). In this paper an extension of the well-known ReliefF method for feature selection is proposed, which consists of adding a new term to the function which updates the weights of the features so as to be able to reach a trade-off between the relevance of a feature and its associated cost. The behavior of the proposed method is tested on twelve heterogeneous classification datasets as well as a real application, using a support vector machine (SVM) as a classifier. The results of the experimental study show that the approach is sound, since it allows the user to reduce the cost significantly without compromising the classification error. Verónica Bolón-Canedo, Beatriz Remeseiro, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
ICAART (1) | 4 |
| 2014 | Scalability Analysis of mRMR for Microarray DataabstractLately, derived from the Big Data problem, researchers in Machine Learning became also interested not only
in accuracy, but also in scalability. Although scalability of learning methods is a trending issue, scalability of
feature selection methods has not received the same amount of attention. In this research, an attempt to study
scalability of both Feature Selection and Machine Learning on microarray datasets will be done. For this sake,
the minimum redundancy maximum relevance (mRMR) filter method has been chosen, since it claims to be
very adequate for this type of datasets. Three synthetic databases which reflect the problematics of microarray
will be evaluated with new measures, based not only in an accurate selection but also in execution time. The
results obtained are presented and discussed. Diego Rego-Fernández, Verónica Bolón-Canedo, Amparo Alonso-Betanzos |
ICAART (1) | 3 |
| 2014 | Data classification using an ensemble of filters
Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
Neurocomputing | 3 |
| 2014 | A review of microarray datasets and applied feature selection methods
Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, José Manuel Benítez 0001, Francisco Herrera |
Inf. Sci. | 3 |
| 2014 | A framework for cost-based feature selection
Verónica Bolón-Canedo, Iago Porto-Díaz, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
Pattern Recognit. | 4 |
| 2014 | A Methodology for Improving Tear Film Lipid Layer ClassificationabstractDry eye is a symptomatic disease which affects a wide range of population and has a negative impact on their daily activities. Its diagnosis can be achieved by analyzing the interference patterns of the tear film lipid layer and by classifying them into one of the Guillon categories. The manual process done by experts is not only affected by subjective factors but is also very time consuming. In this paper we propose a general methodology to the automatic classification of tear film lipid layer, using color and texture information to characterize the image and feature selection methods to reduce the processing time. The adequacy of the proposed methodology was demonstrated since it achieves classification rates over 97% while maintaining robustness and provides unbiased results. Also, it can be applied in real time, and so allows important time savings for the experts. Beatriz Remeseiro, Verónica Bolón-Canedo, Diego Peteiro-Barral, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Antonio Mosquera González, Manuel G. Penedo, Noelia Sánchez-Maroño |
IEEE J. Biomed. Health Informatics | 4 |
| 2013 | A distributed wrapper approach for feature selection
Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
ESANN | 3 |
| 2013 | Toward the scalability of neural networks through feature selection
Diego Peteiro-Barral, Verónica Bolón-Canedo, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Noelia Sánchez-Maroño |
Expert Syst. Appl. | 3 |
| 2013 | A review of feature selection methods on synthetic data
Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
Knowl. Inf. Syst. | 3 |
| 2013 | A Minimum Volume Covering Approach with a Set of EllipsoidsabstractA technique for adjusting a minimum volume set of covering ellipsoids technique is elaborated. Solutions to this problem have potential application in one-class classification and clustering problems. Its main original features are: 1) It avoids the direct evaluation of determinants by using diagonalization properties of the involved matrices, 2) it identifies and removes outliers from the estimation process, 3) it avoids binary variables resulting from the combinatorial character of the assignment problem that are replaced by continuous variables in the range [0,1], 4) the problem can be solved by a bilevel algorithm that in its first level determines the ellipsoids and in its second level reassigns the data points to ellipsoids and identifies outliers based on an algorithm that forces the Karush-Kuhn-Tucker conditions to be satisfied. Two theorems provide rigorous bases for the proposed methods. Finally, a set of examples of application in different fields is given to illustrate the power of the method and its practical performance. David Martínez-Rego, Enrique F. Castillo, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2012 | One-class classifier based on extreme value statistics
David Martínez-Rego, Evan Kriminger, José C. Príncipe, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
ESANN | 5 |
| 2012 | Interferential Tear Film Lipid Layer Classification: An Automatic Dry Eye TestabstractDry eye is a symptomatic disease which affects a wide range of population and has a negative impact on their daily activities, such as driving or working with computers. Its diagnosis can be achieved by several clinical tests, one of which is the analysis of the interference pattern and its classification into one of the Guillon's categories. The methodologies for automatic classification obtain promising results but at the expense of requiring a long processing time. In this research, feature selection techniques are used to reduce time whilst maintaining performance, paving the way for the development of a novel tool for automatic classification of tear film lipid layer. This tool produces significant classification rates over 96% compared with the annotations of the optometrists and provides unbiased results. Also, it works in real-time and so allows important time savings for the experts. Verónica Bolón-Canedo, Diego Peteiro-Barral, Beatriz Remeseiro, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Antonio Mosquera González, Manuel G. Penedo, Noelia Sánchez-Maroño |
ICTAI | 4 |
| 2012 | An ensemble of filters and classifiers for microarray data classification
Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
Pattern Recognit. | 3 |
| 2012 | Nonlinear single layer neural network training algorithm for incremental, nonstationary and distributed learning scenarios
David Martínez-Rego, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
Pattern Recognit. | 3 |
| 2011 | Statistical dependence measure for feature selection in microarray datasets
Verónica Bolón-Canedo, Sohan Seth, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, José C. Príncipe |
ESANN | 4 |
| 2011 | On the behavior of feature selection methods dealing with noise and relevance over synthetic scenariosabstractAdequate identification of relevant features is fundamental in real world scenarios. The problem is specially important when the datasets have a much larger number of features than samples. However, in most cases, the relevant features in real datasets are unknown. In this paper several synthetic datasets are employed to test the effectiveness of different feature selection methods over different artificial classification scenarios, such as altered features (noise), presence of a crescent number of irrelevant features and a small ratio between number of samples and number of features. Six filters and two embedded methods are tested over five synthetic datasets, so as to be able to choose a robust and noise tolerant method, paving the way for its application to real datasets in the classification domain. Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
IJCNN | 3 |
| 2011 | Power wind mill fault detection via one-class ν-SVM vibration signal analysisabstractVibration analysis is one of the most used techniques for predictive maintenance in high-speed rotating machinery. Using the information contained in the vibration signals, a system for alarm detection and diagnosis of failures in mechanical components of power wind mills is devised. As previous failure data collection is unfeasible in real life scenarios, the method to be employed should be capable of discerning between failure and normal data, being only trained with the latter type. Other interesting capability of such a method is the possibility of measuring the evolution of the failure. Taking into account these restrictions, a method that uses the one-class-ν-SVM paradigm is employed. In order to test its adequacy, three different scenarios are tested: (a) a simulated scenario, (b) a controlled experimental scenario with real vibrational data, and (c) a real scenario using vibrational data captured from a windmill power machine installed in a wind farm in North West Spain. The results showed not only the capabilities of the method for detecting the failure in advance to the breakpoint of the component in all three scenarios, but also its capacity to present a qualitative indication on the evolution of the defect. Finally, the results of the SVM paradigm are compared to one of the most used novelty detection methods, obtaining more accurate results under noisy circumstances. David Martínez-Rego, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
IJCNN | 3 |
| 2011 | Toward an ensemble of filters for classificationabstractIn this paper we propose a new framework for feature selection consisting of an ensemble of filters for classification. Five filters, based on different metrics, were involved. Two different approaches of ensembles are presented by varying the role of the classification step. The different options to build an ensemble of filters were studied in detail, dealing with issues such as the presence of redundancy when joining the features selected by different methods. The adequacy of using an ensemble of filters instead of a single filter was demonstrated over a challenging scenario such as DNA microarray data. Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
ISDA | 3 |
| 2011 | Reducing dimensionality in a database of sleep EEG arousals
Diego Álvarez-Estévez, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, Vicente Moret-Bonillo |
Expert Syst. Appl. | 3 |
| 2011 | Feature selection and classification in multiple class datasets: An application to KDD Cup 99 dataset
Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
Expert Syst. Appl. | 3 |
| 2011 | Efficiency of local models ensembles for time series prediction
David Martínez-Rego, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
Expert Syst. Appl. | 3 |
| 2011 | Combining functional networks and sensitivity analysis as wrapper method for feature selection
Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
Expert Syst. Appl. | 2 |
| 2011 | A robust incremental learning method for non-stationary environments
David Martínez-Rego, Beatriz Pérez-Sánchez, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
Neurocomputing | 4 |
| 2011 | A study of performance on microarray data sets for a classifier based on information theoretic learning
Iago Porto-Díaz, Verónica Bolón-Canedo, Amparo Alonso-Betanzos, Oscar Fontenla-Romero |
Neural Networks | 3 |
| 2010 | Fault Prognosis of Mechanical Components Using On-Line Learning Neural Networks
David Martínez-Rego, Oscar Fontenla-Romero, Beatriz Pérez-Sánchez, Amparo Alonso-Betanzos |
ICANN (1) | 4 |
| 2010 | Local Modeling Classifier for Microarray Gene-Expression Data
Iago Porto-Díaz, Verónica Bolón-Canedo, Amparo Alonso-Betanzos, Oscar Fontenla-Romero |
ICANN (3) | 3 |
| 2010 | On the effectiveness of discretization on gene selection of microarray dataabstractDNA microarray data is a challenging issue for machine learning researchers due to the high number of gene expression contained and the small samples sizes. To deal with this problem, feature selection methods, such as filters and wrappers, are typically applied to reduce the dimensionality. In this work, we apply a filter method before the classification and include a discretization step. The results obtained over ten different microarray data sets confirm the adequacy of the proposed method, that achieves better performances than the classifier alone. Besides, the combination method is also compared with the approaches of other authors (using wrappers and filters), outperforming the prediction accuracy and maintaining or even decreasing the number of genes required. Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
IJCNN | 3 |
| 2010 | Multiclass classifiers vs multiple binary classifiers using filters for feature selectionabstractThere are two classical approaches for dealing with multiple class data sets: a classifier that can deal directly with them, or alternatively, dividing the problem into multiple binary sub-problems. While studies on feature selection using the first approach are relatively frequent in scientific literature, very few studies employ the latter one. Out of the four classical methods that can be employed for generating binary problems from a multiple class data set (random, exhaustive, one-vs-one and one-vs-rest), the two last were employed in this work. Besides, four different methods were used for joining the results of these binary classifiers (sum, sum with threshold, Hamming distance and loss-based function). In this paper, both approaches (multiclass and multiple binary classifiers), are carried out using a combination method composed by a discretizer (two different were employed), a filter for feature selection (two methods were chosen), and a classifier (two classifiers were tested). The different combinations of the previous methods, with and without feature selection, were tested over 21 different multiple data sets. An exhaustive study of the results and a comparison between the described methods and some others on the literature is carried out. Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, Pablo Garcia-Gonzalez, Verónica Bolón-Canedo |
IJCNN | 2 |
| 2010 | A Log Analyzer Agent for Intrusion Detection in a Multi-Agent System
Iago Porto-Díaz, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
KES (1) | 3 |
| 2010 | A new convex objective function for the supervised learning of single-layer neural networks
Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas, Beatriz Pérez-Sánchez, Amparo Alonso-Betanzos |
Pattern Recognit. | 4 |
| 2009 | Combining Feature Selection and Local Modelling in the KDD Cup 99 Dataset
Iago Porto-Díaz, David Martínez-Rego, Amparo Alonso-Betanzos, Oscar Fontenla-Romero |
ICANN (1) | 3 |
| 2009 | A combination of discretization and filter methods for improving classification performance in KDD Cup 99 datasetabstractKDD Cup 99 dataset is a classical challenge for computer intrusion detection as well as machine learning researchers. Due to the problematic of this dataset, several sophisticated machine learning algorithms have been tried by different authors. In this paper a new approach is proposed that consists in a combination of a discretizator, a filter method and a very simple classical classifier. The results obtained show the adequacy of the method, that achieves comparable or even better performances than those of other more complicated algorithms, but with a considerable reduction in the number of input features. The proposed method has also been tried over another two large datasets maintaining the same behavior as in the KDD Cup 99 dataset. Verónica Bolón-Canedo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos |
IJCNN | 3 |
| 2009 | A new supervised local modelling classifier based on information theoryabstractIn this paper, a novel supervised architecture for binary classification based on local modelling and information theory is described. The architecture is composed of two steps: in the first one, a separating borderline between the two classes is piecewise constructed by a set of centroids calculated by a modified clustering algorithm, based on information theory; each of these centroids define a region where, in the second step of the proposed architecture, a hyperplane is constructed and adjusted by means of one-layer neural networks. This new method allows for binary classification while maintaining adequate use of computational resources, a common problem for machine learning methods. The proposed architecture is applied over classical benchmark classification problems and data sets, and its results are compared with those obtained by other well-known statistical and machine learning classifiers. David Martínez-Rego, Oscar Fontenla-Romero, Iago Porto-Díaz, Amparo Alonso-Betanzos |
IJCNN | 4 |
| 2009 | Conversion methods for symbolic features: A comparison applied to an intrusion detection problem
Elena Hernández-Pereira, Juan A. Suárez-Romero, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
Expert Syst. Appl. | 4 |
| 2008 | A Regularized Learning Method for Neural Networks Based on Sensitivity Analysis
Bertha Guijarro-Berdiñas, Oscar Fontenla-Romero, Beatriz Pérez-Sánchez, Amparo Alonso-Betanzos |
ESANN | 4 |
| 2008 | A Method for Time Series Prediction using a Combination of Linear Models
David Martínez-Rego, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
ESANN | 3 |
| 2007 | Classification of computer intrusions using functional networks. A comparative study
Amparo Alonso-Betanzos, Noelia Sánchez-Maroño, Félix M. Carballal-Fortes, Juan A. Suárez-Romero, Beatriz Pérez-Sánchez |
ESANN | 1 |
| 2007 | An Improved Version of the Wrapper Feature Selection Method Based on Functional Decomposition
Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, Beatriz Pérez-Sánchez |
ICANN (2) | 2 |
| 2007 | A Comparative Study of Local Classifiers Based on Clustering Techniques and One-Layer Neural Networks
Yuridia Gago-Pallares, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
IDEAL | 3 |
| 2007 | Filter Methods for Feature Selection - A Comparative Study
Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, María Tombilla-Sanromán |
IDEAL | 2 |
| 2007 | A Misuse Detection Agent for Intrusion Detection in a Multi-agent Architecture
Eduardo Mosqueira-Rey, Amparo Alonso-Betanzos, Belén Baldonedo del Río, Jesús Lago Piñeiro |
KES-AMSTA | 2 |
| 2007 | Functional Network Topology Learning and Sensitivity Analysis Based on ANOVA DecompositionabstractA new methodology for learning the topology of a functional network from data, based on the ANOVA decomposition technique, is presented. The method determines sensitivity (importance) indices that allow a decision to be made as to which set of interactions among variables is relevant and which is irrelevant to the problem under study. This immediately suggests the network topology to be used in a given problem. Moreover, local sensitivities to small changes in the data can be easily calculated. In this way, the dual optimization problem gives the local sensitivities. The methods are illustrated by their application to artificial and real examples. Enrique F. Castillo, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, Carmen Castillo |
Neural Comput. | 3 |
| 2006 | A Fast Classification Algorithm Based on Local Models
Sabela Platero-Santos, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
IDEAL | 3 |
| 2006 | Functional Networks and Analysis of Variance for Feature Selection
Noelia Sánchez-Maroño, María Caamaño-Fernández, Enrique F. Castillo, Amparo Alonso-Betanzos |
IDEAL | 4 |
| 2006 | A Very Fast Learning Method for Neural Networks Based on Sensitivity AnalysisabstractThis paper introduces a learning method for two-layer feedforward neural networks based on sensitivity analysis, which uses a linear training algorithm for each of the two layers. First, random values are assigned to the outputs of the first layer; later, these initial values are updated based on sensitivity formulas, which use the weights in each of the layers; the process is repeated until convergence. Since these weights are learnt solving a linear system of equations, there is an important saving in computational time. The method also gives the local sensitivities of the least square errors with respect to input and output data, with no extra computational cost, because the necessary information becomes available without extra calculations. This method, called the Sensitivity-Based Linear Learning Method, can also be used to provide an initial set of weights, which significantly improves the behavior of other learning algorithms. The theoretical basis for the method is given and its performance is illustrated by its application to several examples in which it is compared with several learning algorithms and well known data sets. The results have shown a learning speed generally faster than other existing methods. In addition, it can be used as an initialization tool for other well known methods with significant improvements. Enrique F. Castillo, Bertha Guijarro-Berdiñas, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
J. Mach. Learn. Res. | 4 |
| 2005 | a new wrapper method for feature subset selection
Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, Enrique F. Castillo |
ESANN | 2 |
| 2005 | Modelling Engineering Problems Using Dimensional Analysis for Feature Extraction
Noelia Sánchez-Maroño, Oscar Fontenla-Romero, Enrique F. Castillo, Amparo Alonso-Betanzos |
ICANN (2) | 4 |
| 2005 | A new method for sleep apnea classification using wavelets and feedforward neural networks
Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Vicente Moret-Bonillo |
Artif. Intell. Medicine | 3 |
| 2005 | Linear-least-squares initialization of multilayer perceptrons through backpropagation of the desired responseabstractTraining multilayer neural networks is typically carried out using descent techniques such as the gradient-based backpropagation (BP) of error or the quasi-Newton approaches including the Levenberg-Marquardt algorithm. This is basically due to the fact that there are no analytical methods to find the optimal weights, so iterative local or global optimization techniques are necessary. The success of iterative optimization procedures is strictly dependent on the initial conditions, therefore, in this paper, we devise a principled novel method of backpropagating the desired response through the layers of a multilayer perceptron (MLP), which enables us to accurately initialize these neural networks in the minimum mean-square-error sense, using the analytic linear least squares solution. The generated solution can be used as an initial condition to standard iterative optimization algorithms. However, simulations demonstrate that in most cases, the performance achieved through the proposed initialization scheme leaves little room for further improvement in the mean-square-error (MSE) over the training set. In addition, the performance of the network optimized with the proposed approach also generalizes well to testing data. A rigorous derivation of the initialization algorithm is presented and its high performance is verified with a number of benchmark training problems including chaotic time-series prediction, classification, and nonlinear system identification with MLPs. Deniz Erdogmus, Oscar Fontenla-Romero, José C. Príncipe, Amparo Alonso-Betanzos, Enrique F. Castillo |
IEEE Trans. Neural Networks | 4 |
| 2004 | Shear strength prediction using dimensional analysis and functional networks
Amparo Alonso-Betanzos, Enrique F. Castillo, Oscar Fontenla-Romero, Noelia Sánchez-Maroño |
ESANN | 1 |
| 2004 | A measure of fault tolerance for functional networks
Oscar Fontenla-Romero, Enrique F. Castillo, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas |
Neurocomputing | 3 |
| 2003 | A Bayesian Neural Network Approach for Sleep Apnea Classification
Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Ana del Rocío Fraga-Iglesias, Vicente Moret-Bonillo |
AIME | 3 |
| 2003 | Recursive Least Squares for an Entropy Regularized MSE Cost Function
Deniz Erdogmus, Yadunandana N. Rao, José C. Príncipe, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
ESANN | 5 |
| 2003 | Accelerating the convergence speed of neural networks learning methods using least squares
Oscar Fontenla-Romero, Deniz Erdogmus, José C. Príncipe, Amparo Alonso-Betanzos, Enrique F. Castillo |
ESANN | 4 |
| 2003 | Self-organizing maps and functional networks for local dynamic modeling
Noelia Sánchez-Maroño, Oscar Fontenla-Romero, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas |
ESANN | 3 |
| 2003 | Linear Least-Squares Based Methods for Neural Networks Learning
Oscar Fontenla-Romero, Deniz Erdogmus, José C. Príncipe, Amparo Alonso-Betanzos, Enrique F. Castillo |
ICANN | 4 |
| 2003 | Accurate initialization of neural network weights by backpropagation of the desired responseabstractProper initialization of neural networks is critical for a successful training of its weights. Many methods have been proposed to achieve this, including heuristic least squares approaches. In this paper, inspired by these previous attempts to train (or initialize) neural networks, we formulate a mathematically sound algorithm based on backpropagating the desired output through the layers of a multilayer perceptron. The approach is accurate up to local first order approximations of the nonlinearities. It is shown to provide successful weight initialization for many data sets by Monte Carlo experiments. Deniz Erdogmus, Oscar Fontenla-Romero, José C. Príncipe, Amparo Alonso-Betanzos, Enrique F. Castillo, Robert Jenssen |
IJCNN | 4 |
| 2003 | An intelligent system for forest fire risk prediction and fire fighting management in Galicia
Amparo Alonso-Betanzos, Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas, Elena Hernández-Pereira, Maria Inmaculada Paz-Andrade, Eulogio Jimenez, Jose Luis Legido, Tarsy Carballas |
Expert Syst. Appl. | 1 |
| 2002 | A Neural Network Approach for Forestal Fire Risk Estimation
Amparo Alonso-Betanzos, Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas, Elena Hernández-Pereira, Juan Canda, Eulogio Jimenez, Jose Luis Legido, Susana Muñiz, Cristina Paz-Andrade, Maria Inmaculada Paz-Andrade |
ECAI | 1 |
| 2002 | Local Modeling Using Self-Organizing Maps and Single Layer Neural Networks
Oscar Fontenla-Romero, Amparo Alonso-Betanzos, Enrique F. Castillo, José C. Príncipe, Bertha Guijarro-Berdiñas |
ICANN | 2 |
| 2002 | Intelligent analysis and pattern recognition in cardiotocographic signals using a tightly coupled hybrid system
Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Oscar Fontenla-Romero |
Artif. Intell. | 2 |
| 2002 | Empirical evaluation of a hybrid intelligent monitoring system using different measures of effectiveness
Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos |
Artif. Intell. Medicine | 2 |
| 2002 | A Global Optimum Approach for One-Layer Neural NetworksabstractThe article presents a method for learning the weights in one-layer feedforward neural networks minimizing either the sum of squared errors or the maximum absolute error, measured in the input scale. This leads to the existence of a global optimum that can be easily obtained solving linear systems of equations or linear programming problems, using much less computational power than the one associated with the standard methods. Another version of the method allows computing a large set of estimates for the weights, providing robust, mean or median, estimates for them, and the associated standard errors, which give a good measure for the quality of the fit. Later, the standard one-layer neural network algorithms are improved by learning the neural functions instead of assuming them known. A set of examples of applications is used to illustrate the methods. Finally, a comparison with other high-performance learning algorithms shows that the proposed methods are at least 10 times faster than the fastest standard algorithm used in the comparison. Enrique F. Castillo, Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos |
Neural Comput. | 4 |
| 2001 | Adaptive pattern recognition in the analysis of cardiotocographic recordsabstractThe recognition of accelerative and decelerative patterns in the fetal heart rate (FHR) is one of the tasks carried out manually by obstetricians when they analyze cardiotocograms for information respecting the fetal state. An approach based on artificial neural networks formed by a multilayer perceptron (MLP) is developed. However, since the system utilizes the FHR signal as direct input, an anterior stage must be incorporated that applies a principal component analysis (PCA) so as to make the system independent of the signal baseline. Furthermore, the introduction of multiresolution into the PCA has resolved other problems that were detected in the application of the system. Presented in this paper are the results of validation of these systems designated the PCA-MLP and multiresolutlon principal component analysis (MR-PCA) systems against three clinical experts. Oscar Fontenla-Romero, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas |
IEEE Trans. Neural Networks | 2 |
| 2000 | Analysis and evaluation of hard and fuzzy clustering segmentation techniques in burned patient images
Amparo Alonso-Betanzos, Bernardino Arcay Varela, Alfonso Castro Martínez |
Image Vis. Comput. | 1 |
| 1999 | Applying statistical, uncertainty-based and connectionist approaches to the prediction of fetal outcome: a comparative study
Amparo Alonso-Betanzos, Eduardo Mosqueira-Rey, Vicente Moret-Bonillo, Belén Baldonedo del Río |
Artif. Intell. Medicine | 1 |
| 1997 | Information analysis and validation of intelligent monitoring systems in intensive care unitsabstractValidation of intelligent systems is an important task to perform. Typically the results of the validation analysis are used to verify whether or not the system satisfies the initial design requirements, and to acquire new knowledge and/or refine the knowledge already acquired. In practice, the validation of intelligent systems usually requires the application of several different techniques (e.g., retrospective, prospective, quantitative). In this work the authors present the methodology devised to validate PATRICIA: an intelligent monitoring system designed to advise clinicians on the management of patients dependent on mechanical ventilation. The application of this methodology requires that appropriate validation paradigms are selected, depending on both the application domain and the characteristics of the intelligent system. The article also presents and discusses validation results. Vicente Moret-Bonillo, Eduardo Mosqueira-Rey, Amparo Alonso-Betanzos |
IEEE Trans. Inf. Technol. Biomed. | 3 |
| 1995 | The NST-EXPERT project: the need to evolve
Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Vicente Moret-Bonillo, S. Lopez-Gonzalez |
Artif. Intell. Medicine | 1 |
| 1989 | FOETOS in clinical practice: A retrospective analysis of its performance
Amparo Alonso-Betanzos, Lawrence D. Devoe, Ramón A. Castillo, Vicente Moret-Bonillo, Carlos Hernández-Sande, Nancy S. Searle |
Artif. Intell. Medicine | 1 |