Bertha Guijarro-Berdiñas

dblp:00/5605 · DBLP profile ↗
← Back
68ranked-venue papers
7as first author
25since 2021 · last 2026
0000-0001-8901-5441ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 59 · 6 first-author · 23 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FedHENet: A Frugal Federated Learning Framework for Heterogeneous Environments
abstract
Federated Learning (FL) enables collaborative training without centralizing data, essential for privacy compliance in real-world scenarios involving sensitive visual information.Most FL approaches rely on expensive, iterative deep network optimization, which still risks privacy via shared gradients.In this work, we propose FedHENet, extending the FedHEONN framework to image classification.By using a fixed, pretrained feature extractor and learning only a single output layer, we avoid costly local fine-tuning.This layer is learned by analytically aggregating client knowledge in a single round of communication using homomorphic encryption (HE).Experiments show that FedHENet achieves competitive accuracy compared to iterative FL baselines while demonstrating superior stability performance and up to 70% better energy efficiency.Crucially, our method is hyperparameter-free, removing the carbon footprint associated with hyperparameter tuning in standard FL.
Alejandro Dopico-Castro, Oscar Fontenla-Romero, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Iván Pérez Digón
ESANN4
2026 Real-time analysis of indoor sports game situations through deep learning-based classification
abstract
Live indoor sports broadcasts require dynamic camera control in response to relevant game situations such as penalties or timeouts, a process that traditionally relies on human operators. This paper presents a solution for automatic real-time classification of game states in indoor invasion sports, with handball as the primary case study and basketball as a secondary validation scenario. Our approach utilizes raw video directly from cameras, enabling real-time analysis. The system continuously processes video frames, assigning each to one of seven classes: left/right attack, left/right counterattack, left/right penalty, and timeout. A synthetic representation of each frame is used to standardize the depiction of game dynamics. The proposed pipeline includes object detection with a fine-tuned You Only Look Once (YOLO) model to locate players, the ball, and referees; object tracking to compute velocity vectors; generation of a synthetic frame representing the current game state; and final classification using a custom Dense Convolutional Network (DenseNet). Using a dataset of 20 handball matches, the proposed system achieved a macro-averaged F1-score of 96.1%, with a per-image inference time below 4 milliseconds, evaluated on 118,129 images from matches unseen during training. The same pipeline was subsequently applied to basketball using only two matches, achieving an F1-score of 92.5% on 12,390 images, thereby illustrating the transferability of the proposed approach to other indoor invasion sports. The full pipeline operates in 34.04 milliseconds with GPU acceleration, processing over 25 frames per second.
Bruno Cabado, Bertha Guijarro-Berdiñas, Emilio J. Padrón 0001
Expert Syst. Appl.2
2026 Contrastive Learning for Explanation Ranking
abstract
Abstract Explainable recommendation systems enhance user trust and satisfaction by revealing the reasoning behind personalized recommendations. Approaching this as a post-hoc explanation-ranking problem over a fixed pool of candidate explanations, we propose Contrastive Learning for Explanation Ranking (CLER), a model that learns user, item, and explanation representations with a Normalized Temperature-scaled Binary Cross-Entropy (NT-BXent) loss. This function specifically applies a per-row reweighting strategy, preventing the vast number of negative examples from dominating the objective. We evaluate CLER on the Amazon, TripAdvisor, and Yelp datasets from the EXTRA benchmark. Across traditional ranking metrics, CLER achieves the strongest results among the compared baselines.
Miguel Escarda-Fernández, Brais Cancela, Carlos Eiras-Franco, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos
Mach. Learn.4
2025 From Raw Feeds to Curated Clips: A Modular AI Framework for Sports Highlight Production
abstract
In the era of social media and instantaneous content consumption, the demand for quick post-event sports highlights has emerged, requiring systems to deliver compelling summaries in a minimal turnaround time. This paper introduces an AI-driven framework designed to streamline the creation of match recaps by leveraging real-time data acquisition and efficient post-processing. The system analyzes multi-modal inputs —including live game statistics, audio-visual feeds, and contextual cues— to automatically identify key moments (e.g., goals, pivotal plays) as they occur. By integrating lightweight neural models and rule-based prioritization, it generates timestamped clips immediately after the match concludes, significantly reducing manual editing effort. The solution not only supports human editors by providing pre-curated material but also enables fully automated highlight production for platforms requiring instant content delivery. Evaluations on soccer and basketball matches demonstrate the system’s ability to cut post-event processing time by 87.5% while maintaining 90% accuracy in event selection compared to manual curation. The work underscores the potential of hybrid AI systems to bridge real-time analytics with post-production workflows, offering scalability across sports and media formats. The generated highlights are already being published on a production platform (https://tiivii.gal), demonstrating real-world applicability.
Bruno Cabado, Bertha Guijarro-Berdiñas, Emilio J. Padrón 0001
ECAI2
2025 Predicting the State of Health of Supercapacitors Using a Federated Learning Model with Homomorphic Encryption
Víctor López 0002, Oscar Fontenla-Romero, Elena Hernández-Pereira, Bertha Guijarro-Berdiñas, Carlos Blanco-Seijo, Samuel Fernández-Paz
ICAART (3)4
2025 Communication and Negotiation to Improve Agent-Based Models
Alejandro Rodríguez-Arias, Noelia Sánchez-Maroño, Bertha Guijarro-Berdiñas
ICAART (3)3
2025 Sustainable Techniques to Improve Data Quality for Training Image-Based Explanatory Models for Recommender Systems
Jorge Paz-Ruza, David Esteban Martínez, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas
ICANN (3)4
2025 Efficient and Secure Federated Learning with Ensemble of One-Layer Neural Networks
abstract
In this work, we propose a Federated Learning (FL) method combining an ensemble of one-layer neural networks whose optimal parameters can be obtained through a non-iterative procedure. Therefore, unlike most state-of-the-art methods, the collaborative global model can be obtained using a single round of communication between all clients of the federated scheme. It presents a computationally efficient and incremental batch aggregation process that suits the needs of a realistic federated scenario, simplifying the management of the federated training process. The model provides the same performance in identically and non-identically distributed data scenarios. Besides, the model implements a Fully Homomorphic Encryption (FHE) scheme to enhance robustness against privacy leaks or attacks, enabling clients to offload computational work to the coordinator, which operates entirely on encrypted data. We achieve an efficient and secure distributed model with an improved representation capacity for this type of architecture. The source code used in the study is made publicly available.
Abel Pampín-Rodríguez, Oscar Fontenla-Romero, Elena Hernández-Pereira, Bertha Guijarro-Berdiñas
IJCNN4
2025 Predictively Combatting Toxicity in Health-related Online Discussions through Machine Learning
abstract
In health-related topics, user toxicity in online discussions frequently becomes a source of social conflict or promotion of dangerous, unscientific behaviour; common approaches for battling it include different forms of detection, flagging and/or removal of existing toxic comments, which is often counterproductive for platforms and users alike. In this work, we propose the alternative of combatting user toxicity predictively, anticipating where a user could interact toxically in health-related online discussions. Applying a Collaborative Filtering-based Machine Learning methodology, we predict the toxicity in COVID-related conversations between any user and subcommunity of Reddit, surpassing 80% predictive performance in relevant metrics, and allowing us to prevent the pairing of conflicting users and subcommunities.
Jorge Paz-Ruza, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Carlos Eiras-Franco
IJCNN3
2025 An agent-based model to simulate the public acceptability of social innovations
abstract
Abstract The successful adoption of social innovations, such as renewable energy systems or pollution reduction plans in cities, depends, to a large extent, on the willingness and participation of the population in their development and implementation. We present an agent‐based model (ABM) to analyze the process of citizen acceptability of a social innovation that uses a variety of agents to represent individual citizens and relevant groups of citizens. Citizen agents make use of the HUMAT cognitive decision‐making model, based on psychosocial theories, to decide on their support for the social innovation considering how their needs will be satisfied if they decide to support (or not) the innovation project, and the influence exerted by the agents in their environment. The ABM was initially developed to represent the urban and transport planning superblock project in the city of Vitoria‐Gasteiz (Spain). The ABM simulations make it possible to study the evolution of public acceptance of social innovation, with the results providing insights to the social dynamics and individual factors that affect the acceptance of the project, enabling an evaluation of how to devise new policies that increase public acceptance. Sufficiently generic to be easily adaptable to different types of social innovations, the ABM is a powerful tool to explore different scenarios and design strategies that foster the acceptance and sustainable adoption of social innovations.
Alejandro Rodríguez-Arias, Noelia Sánchez-Maroño, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Isabel Lema-Blanco, Adina Dumitru
Expert Syst. J. Knowl. Eng.3
2025 Performance and sustainability of BERT derivatives in dyadic data
abstract
[Abstract]: In recent years, the Natural Language Processing (NLP) field has experienced a revolution, where numerous models – based on the Transformer architecture – have emerged to process the ever-growing volume of online text-generated data. This architecture has been the basis for the rise of Large Language Models (LLMs). Enabling their application to many diverse tasks in which they excel with just a fine-tuning process that comes right after a vast pre-training phase. However, their sustainability can often be overlooked, especially regarding computational and environmental costs. Our research aims to compare various BERT derivatives in the context of a dyadic data task while also drawing attention to the growing need for sustainable AI solutions. To this end, we utilize a selection of transformer models in an explainable recommendation setting, modeled as a multi-label classification task originating from a social network context, where users, restaurants, and reviews interact.
Miguel Escarda-Fernández, Carlos Eiras-Franco, Brais Cancela, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos
Expert Syst. Appl.4
2025 Beyond RMSE and MAE: Introducing EAUC to Unmask Hidden Bias and Unfairness in Dyadic Regression Models
abstract
Dyadic regression models, which output real-valued predictions for pairs of entities, are fundamental in many domains [e.g., obtaining user-product ratings in recommender systems (RSs)] and promising and under exploration in others (e.g., tuning patient-drug dosages in precision pharmacology). In this work, we prove that nonuniform observed value distributions of individual entities lead to severe biases in state-of-the-art models, skewing predictions toward the average of observed past values for the entity and providing worse-than-random predictive power in eccentric yet crucial cases; we name this phenomenon eccentricity bias. We show that global error metrics like root-mean-squared error (RMSE) are insufficient to capture this bias, and we introduce eccentricity area under the curve (EAUC) as a novel metric that can quantify it in all studied domains and models. We prove the intuitive interpretation of EAUC by experimenting with naive post-training bias corrections and theorize other options to use EAUC to guide the construction of fair models. This work contributes a bias-aware evaluation of dyadic regression to prevent unfairness in critical real-world applications of such systems.
Jorge Paz-Ruza, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Brais Cancela, Carlos Eiras-Franco
IEEE Trans. Neural Networks Learn. Syst.3
2024 AI-based algorithm for intrusion detection on a real dataset
abstract
In the realm of cybersecurity, the detection of network intrusions stands as a paramount challenge, with ever-evolving threats demanding innovative solutions.This study delves into the application of diverse machine learning algorithms on a contemporary dataset (UGR'16) comprising real-world instances of intrusion in software systems.Specifically, several Machine Learning models (Outlier Detectors, Ensemble Methods, Deep Learning, and Conventional Classifiers) were tested and compared with previously reported results using a standard methodology.The obtained results reveal that the Ensemble Methods have been capable of improving the results from prior research.Particularly, the Extreme Gradient Boosting (XGBoost) algorithm offers better results than the original solution with Random Forest, with an AUC of 0.9218 as opposed to 0.8977, and more than four times as fast for the problem to solve.
David Esteban Martínez, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Elena Hernández-Pereira, Alejandro Esteban Martínez
ESANN2
2024 Agent-based model to assess public acceptance of an energy sustainability project in the island of El Hierro
abstract
This paper presents an agent-based model to address citizen acceptance of energy sustainability projects, using the 100% Renewable El Hierro Project as a case study. The island of El Hierro is a pioneer in the search for sustainable energy sources, and this project seeks to understand how the individual psychosocial needs of El Hierro’s citizens influence their perception and adoption of energy policies. The HUMAT architecture, inspired by psychological and sociological theories, is adapted to model the behaviour of the citizens of El Hierro in relation to the energy sustainability project. Each agent in the model represents an individual citizen and their behaviour is determined by their psychosocial needs, which include factors such as the island’s energy independence, economic sustainability or environmental quality. Simulations are used to assess the impact of different communication strategies by project stakeholders, such as the press or local government, on the evolution of public acceptance of a project extension. This study highlights the importance of integrating multidisciplinary approaches combining psychology, sociology and engineering to address energy sustainability challenges in specific contexts such as El Hierro. It also offers new perspectives on how to design and implement energy policy communication campaigns in a way that is more acceptable to the public.
Alejandro Rodríguez-Arias, Noelia Sánchez-Maroño, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos
KES3
2024 Explained anomaly detection in text reviews: Can subjective scenarios be correctly evaluated?
abstract
In the current landscape, user opinions exert an unprecedented influence on the trajectory of companies. In the field of online review platforms, these opinions, transmitted through text reviews and numerical ratings, significantly shape the credibility of products and services. For this reason, detecting inappropriate reviews becomes crucial. This paper addresses the problem of automatic anomalous review detection using a novel approach based on Anomaly Detection in the field of Natural Language Processing (NLP). Unlike other NLP tasks, anomaly detection in texts is a relatively emerging area. In this paper, we present a pipeline for opinion filtering that poses the problem of discerning between normal opinions containing relevant information about an item and anomalous opinions with unrelated content. Its key functionalities include: Classifying the reviews, assigning normality scores, and generating explanations for each classification, indispensable for the human who normally moderates these platforms. To evaluate the model, several Amazon datasets were used to demonstrate that the performance obtained is robust, obtaining an average F1 score of 91.4 detecting anomalies in the most complex scenario. In addition, a comparative study of three explainability techniques was conducted with 241 participants to measure the impact on understanding the classifications of the model and to rank their perceived usefulness of explanations. As a result, we obtained a system with great potential to automate tasks related to online review platforms, offering insights into anomaly detection applications in textual data and showing the difficulties that arise when the task to be explained presents a subjectivity component.
David Novoa-Paradela, Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas
Eng. Appl. Artif. Intell.3
2023 A Federated Learning Architecture for Anomaly Detection on the Edge Using Deep Autoencoders
abstract
Autoencoder networks are widely used in anomaly detection, however, their training can be computationally expensive, limiting their use in Edge Computing and Federated Learning scenarios, where devices are usually not very powerful. In addition, there are several ways to directly or indirectly attack the privacy of the data used by these networks, which is unacceptable in this type of scenario. Unlike traditional autoencoder networks, Deep AutoEncoder for Federated learning (DAEF) does not require several rounds of learning since it is a non-iterative method. This implies greater speed, less network traffic, and lower energy consumption, as well as preventing the privacy attacks common in iterative networks. In this paper, we present an architecture designed for the use of the DAEF network in Edge Computing and Federated Learning scenarios. Unlike other collaborative machine learning approaches, it is not server based. Consequently, all the stages of the learning process, including the model aggregation, are performed on the edge devices (nodes). Each edge node trains its local DAEF network asynchronously concerning the rest. Nodes that have completed their training can voluntarily request to add their local model information to the global one. In this way, there is not a unique aggregator node, but each node in the network will be in charge of adding its local learning to the global model. The federated learning management is handled by a coordinator node using a Message Queuing Telemetry Transport (MQTT) communication protocol, which can be assumed by any device in the network in case of failures, and the information exchanged between the nodes does not compromise the privacy of the original local datasets.
David Novoa-Paradela, Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas, Diego Orellana-Cañás
WETICE3
2023 FedHEONN: Federated and homomorphically encrypted learning method for one-layer neural networks
abstract
Federated learning (FL) is a distributed approach to developing collaborative learning models from decentralized data. This is relevant to many real applications, such as in the field of the Internet of Things, since the models can be used in edge computing devices. FL approaches are motivated by and designed to protect privacy, a highly relevant issue given current data protection regulations. Although FL methods are privacy-preserving by design, recently published papers show that privacy leaks do occur, caused by attacks designed to extract private data from information interchanged during learning. In this work, we present an FL method based on a neural network without hidden layers that incorporates homomorphic encryption (HE) to enhance robustness against the above-mentioned attacks. Unlike traditional FL methods that require multiple rounds of training for convergence, our method obtains the collaborative global model in a single training round, yielding an effective and efficient model that simplifies management of the FL training process. In addition, since our method includes HE, it is also robust against model inversion attacks. In experiments with big data sets and a large number of clients in a federated scenario, we demonstrate that use of HE does not affect the accuracy of the model, whose results are competitive with state-of-the-art machine learning models. We also show that behavior in terms of accuracy is the same for identically and non-identically distributed data scenarios.
Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas, Elena Hernández-Pereira, Beatriz Pérez-Sánchez
Future Gener. Comput. Syst.2
2023 Fast deep autoencoder for federated learning
abstract
This paper presents a novel, fast and privacy preserving implementation of deep autoencoders. DAEF (Deep AutoEncoder for Federated learning), unlike traditional neural networks, trains a deep autoencoder network in a non-iterative way, which drastically reduces training time. Training can be performed incrementally, in parallel and distributed and, thanks to its mathematical formulation, the information to be exchanged does not endanger the privacy of the training data. The method has been evaluated and compared with other state-of-the-art autoencoders, showing interesting results in terms of accuracy, speed and use of available resources. This makes DAEF a valid method for edge computing and federated learning, in addition to other classic machine learning scenarios.
David Novoa-Paradela, Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas
Pattern Recognit.3
2022 Real-time classification of handball game situations
abstract
During the broadcast of sporting events, certain situations such as a penalty or a time-out occur, for which a specific action is required. In traditional broadcasting, many people are implied in making decisions based on what is happening at any given moment. To broadcast quality and entirely automatically matches it is necessary to be able to classify the important situations and then make decisions based on them. This paper presents a solution based on deep learning which is able to classify the main states of a handball match. The generated model has been trained using 127,600 images of 13 local team matches. On a test set of 118,129 images of other 7 matches, it is able to classify these situations with an accuracy of 98.6% in only 4 milliseconds, allowing to analyze the state of the game in real time. The full pipeline takes only 34.04 milliseconds using GPU acceleration, processing more than 25 frames per seconds.
Bruno Cabado, Bertha Guijarro-Berdiñas, Emilio J. Padrón 0001
ICTAI2
2022 Sustainable Personalisation and Explainability in Dyadic Data Systems
abstract
Systems that rely on dyadic data, which relate entities of two types together, have become ubiquitously used in fields such as media services, tourism business, e-commerce, and others. However, these systems have had a tendency to be black-box systems, despite their objective of influencing people's decisions. There is a lack of research on providing personalised explanations to the outputs of systems that make use of such data, that is, integrating the idea of Explainable Artificial Intelligence into the field of dyadic data. Moreover, the existing approaches rely heavily on Deep Learning models for their training, reducing their overall sustainability. In this work, we propose a computationally efficient model which provides personalisation by generating explanations based on user-created images. In the context of a particular dyadic data system, the restaurant review platform TripAdvisor, we predict, for any (user, restaurant) pair, the review of the restaurant that is most adequate to present it to the user, based on their personal preferences. This model exploits the usage of efficient Matrix Factorisation techniques combined with feature-rich embeddings of the pre-trained Image Classification models, developing a method capable of providing transparency to dyadic data systems while reducing as much as 80% the carbon emissions of training compared to alternative approaches.
Jorge Paz-Ruza, Carlos Eiras-Franco, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos
KES3
2022 How Agent-based modeling can help to foster sustainability projects
abstract
[Abstract] The Sustainable Development Goals (SDGs) adopted by the United Nations require relevant social changes that sometimes involve the development of innovative projects that cause rejection and confrontation. Agent-Based Models (ABM) are powerful tools to represent the behavior of systems, and they have become valuable for the social sciences as they can simulate the behavior of a society under different conditions. Superblocks are innovative city projects that reorganize urban space and minimize private motorized transport. In this paper, we present an ABM that simulates the implantation of superblocks in two Spanish cities: Vitoria-Gasteiz and Barcelona. The interest of this model is to provide policymakers with relevant scientific information that can be used to support their planning and decision-making processes by running possible alternative policy scenarios. This paper presents the details of the designed model and the simulation of different policy scenarios to increase the acceptability rates of citizens about the project, demonstrating how the model takes into account local differences and its usefulness for those political leaders from other cities interested in implementing this type of project.
Noelia Sánchez-Maroño, Alejandro Rodríguez-Arias, Adina Dumitru, Isabel Lema-Blanco, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos
KES5
2022 Machine learning techniques to predict different levels of hospital care of CoVid-19
abstract
In this study, we analyze the capability of several state of the art machine learning methods to predict whether patients diagnosed with CoVid-19 (CoronaVirus disease 2019) will need different levels of hospital care assistance (regular hospital admission or intensive care unit admission), during the course of their illness, using only demographic and clinical data. For this research, a data set of 10,454 patients from 14 hospitals in Galicia (Spain) was used. Each patient is characterized by 833 variables, two of which are age and gender and the other are records of diseases or conditions in their medical history. In addition, for each patient, his/her history of hospital or intensive care unit (ICU) admissions due to CoVid-19 is available. This clinical history will serve to label each patient and thus being able to assess the predictions of the model. Our aim is to identify which model delivers the best accuracies for both hospital and ICU admissions only using demographic variables and some structured clinical data, as well as identifying which of those are more relevant in both cases. The results obtained in the experimental study show that the best models are those based on oversampling as a preprocessing phase to balance the distribution of classes. Using these models and all the available features, we achieved an area under the curve (AUC) of 76.1% and 80.4% for predicting the need of hospital and ICU admissions, respectively. Furthermore, feature selection and oversampling techniques were applied and it has been experimentally verified that the relevant variables for the classification are age and gender, since only using these two features the performance of the models is not degraded for the two mentioned prediction problems.
Elena Hernández-Pereira, Oscar Fontenla-Romero, Verónica Bolón-Canedo, Brais Cancela, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos
Appl. Intell.5
2021 Federated Learning approach for SpectralClustering
abstract
Spectral clustering is a clustering paradigm that has been shown to be more effective in finding clusters with non-convex shapes than some traditional algorithms such as k-means.However, this algorithm is not directly applicable when the data is naturally distributed in different locations, as it happens in many Internet of Things scenarios.In this work, we propose a distributed spectral clustering to create a cooperative federated model to deal with those cases in which the data is distributed in different sites and with data privacy concerns.We demonstrate that sharing a minimal amount of information allows this distributed version of the spectral clustering to achieve good behavior for clustering several synthetic data sets.
Elena Hernández-Pereira, Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas, Beatriz Pérez-Sánchez
ESANN3
2021 Scalable feature selection using ReliefF aided by locality-sensitive hashing
abstract
Feature selection algorithms, such as ReliefF, are very important for processing high-dimensionality data sets. However, widespread use of popular and effective such algorithms is limited by their computational cost. We describe an adaptation of the ReliefF algorithm that simplifies the costliest of its step by approximating the nearest neighbor graph using locality-sensitive hashing (LSH). The resulting ReliefF-LSH algorithm can process data sets that are too large for the original ReliefF, a capability further enhanced by distributed implementation in Apache Spark. Furthermore, ReliefF-LSH obtains better results and is more generally applicable than currently available alternatives to the original ReliefF, as it can handle regression and multiclass data sets. The fact that it does not require any additional hyperparameters with respect to ReliefF also avoids costly tuning. A set of experiments demonstrates the validity of this new approach and confirms its good scalability.
Carlos Eiras-Franco, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Antonio Bahamonde
Int. J. Intell. Syst.2
2021 DSVD-autoencoder: A scalable distributed privacy-preserving method for one-class classification
abstract
One-class classification has gained interest as a solution to certain kinds of problems typical in a wide variety of real environments like anomaly or novelty detection. Autoencoder is the type of neural network that has been widely applied in these one-class problems. In the Big Data era, new challenges have arisen, mainly related with the data volume. Another main concern derives from Privacy issues when data is distributed and cannot be shared among locations. These two conditions make many of the classic and brilliant methods not applicable. In this paper, we present distributed singular value decomposition (DSVD-autoencoder), a method for autoencoders that allows learning in distributed scenarios without sharing raw data. Additionally, to guarantee privacy, it is noniterative and hyperparameter-free, two interesting characteristics when dealing with Big Data. In comparison with the state of the art, results demonstrate that DSVD-autoencoder provides a highly competitive solution to deal with very large data sets by reducing training from several hours to seconds while maintaining good accuracy.
Oscar Fontenla-Romero, Beatriz Pérez-Sánchez, Bertha Guijarro-Berdiñas
Int. J. Intell. Syst.3
2020 Online learning for anomaly detection via subdivisible convex hulls
abstract
Due to the frequent use of anomaly detection systems in monitoring and the lack of methods capable of learning in real time, this research presents a new method that provides such online adaptability. The method developed is called OSHULL (Online and Subdivisible Distributed Scaled Convex Hull) and bases its operation on the properties of scaled convex hulls. It begins building a convex hull, using a minimum set of data, that is adapted and subdivided along time to accurately fit the boundary of the normal class data. The method has been evaluated and compared to several main algorithms of the field using some real and artificial data sets. As a consequence, an algorithm has been obtained with online learning ability and easily configurable, all without diminishing its effectiveness in relation to other batch state-of-the-art methods. Finally, its execution can be carried out in a distributed and parallel way, which is an interesting advantage in the treatment of big data sets.
David Novoa-Paradela, Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas
IJCNN3
2020 A FIPA-ACL based communication utility for Unity
abstract
Multi-agent systems (MAS) allow facing complex, heterogeneous, distributed problems difficult to solve by only one software agent. Games often provide very adequate problems for the use of MAS. In the field of Games, Unity is one of the most used video game engines and allows the development of intelligent agents in 2D/3D virtual environments. However, although Unity allows working in multi-agent environments it does not provide functionalities to facilitate the development of genuine MAS. One crucial functionality is to allow agents to communicate and understand each other through an agent communication language (ACL). The objective of this work is to supply Unity with this capacity, adding support for the creation, sending and reception of messages. As a result, a communication system and an ACL are provided that facilitate the development of MAS in Unity, allowing to solve problems that require collaborative solutions among agents.
Alejandro Rodríguez-Arias, Bertha Guijarro-Berdiñas, Noelia Sánchez-Maroño
WETICE2
2020 Fast Distributed kNN Graph Construction Using Auto-tuned Locality-sensitive Hashing
abstract
The k -nearest-neighbors ( k NN) graph is a popular and powerful data structure that is used in various areas of Data Science, but the high computational cost of obtaining it hinders its use on large datasets. Approximate solutions have been described in the literature using diverse techniques, among which Locality-sensitive Hashing (LSH) is a promising alternative that still has unsolved problems. We present Variable Resolution Locality-sensitive Hashing, an algorithm that addresses these problems to obtain an approximate k NN graph at a significantly reduced computational cost. Its usability is greatly enhanced by its capacity to automatically find adequate hyperparameter values, a common hindrance to LSH-based methods. Moreover, we provide an implementation in the distributed computing framework Apache Spark that takes advantage of the structure of the algorithm to efficiently distribute the computational load across multiple machines, enabling practitioners to apply this solution to very large datasets. Experimental results show that our method offers significant improvements over the state-of-the-art in the field and shows very good scalability as more machines are added to the computation.
Carlos Eiras-Franco, David Martínez-Rego, Leslie Kanthan, César Piñeiro, Antonio Bahamonde, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos
ACM Trans. Intell. Syst. Technol.6
2019 A scalable decision-tree-based method to explain interactions in dyadic data
Carlos Eiras-Franco, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Antonio Bahamonde
Decis. Support Syst.2
2019 Large scale anomaly detection in mixed numerical and categorical input spaces
Carlos Eiras-Franco, David Martínez-Rego, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Antonio Bahamonde
Inf. Sci.3
2018 LANN-DSVD: A privacy-preserving distributed algorithm for machine learning
Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas, Beatriz Pérez-Sánchez, Marcelo Gómez-Casal
ESANN2
2018 On the scalability of feature selection methods on high-dimensional data
Verónica Bolón-Canedo, Diego Fernández-Francos, Diego Peteiro-Barral, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Noelia Sánchez-Maroño
Knowl. Inf. Syst.5
2018 LANN-SVD: A Non-Iterative SVD-Based Learning Algorithm for One-Layer Neural Networks
abstract
In the scope of data analytics, the volume of a data set can be defined as a product of instance size and dimensionality of the data. In many real problems, data sets are mainly large only on one of these aspects. Machine learning methods proposed in the literature are able to efficiently learn in only one of these two situations, when the number of variables is much greater than instances or vice versa. However, there is no proposal allowing to efficiently handle either circumstances in a large-scale scenario. In this brief, we present an approach to integrally address both situations, large dimensionality or large instance size, by using a singular value decomposition (SVD) within a learning algorithm for one-layer feedforward neural network. As a result, a noniterative solution is obtained, where the weights can be calculated in a closed-form manner, thereby avoiding low convergence rate and also hyperparameter tuning. The proposed learning method, LANN-SVD in short, presents a good computational efficiency for large-scale data analytic. Comprehensive comparisons were conducted to assess LANN-SVD against other state-of-the-art algorithms. The results of this brief exhibited the superior efficiency of the proposed method in any circumstance.
Oscar Fontenla-Romero, Beatriz Pérez-Sánchez, Bertha Guijarro-Berdiñas
IEEE Trans. Neural Networks Learn. Syst.3
2016 A fast learning algorithm for high dimensional problems: an application to microarrays
Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas, Beatriz Pérez-Sánchez, Diego Rego-Fernández, David Martínez-Rego
ESANN2
2016 Distributed learning algorithm for feedforward neural networks
Oscar Fontenla-Romero, Beatriz Pérez-Sánchez, Bertha Guijarro-Berdiñas, Diego Rego-Fernández
ESANN3
2016 A unified pipeline for online feature selection and classification
Verónica Bolón-Canedo, Diego Fernández-Francos, Diego Peteiro-Barral, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Noelia Sánchez-Maroño
Expert Syst. Appl.5
2015 Distributed One-Class Support Vector Machine
abstract
This paper presents a novel distributed one-class classification approach based on an extension of the ν-SVM method, thus permitting its application to Big Data data sets. In our method we will consider several one-class classifiers, each one determined using a given local data partition on a processor, and the goal is to find a global model. The cornerstone of this method is the novel mathematical formulation that makes the optimization problem separable whilst avoiding some data points considered as outliers in the final solution. This is particularly interesting and important because the decision region generated by the method will be unaffected by the position of the outliers and the form of the data will fit more precisely. Another interesting property is that, although built in parallel, the classifiers exchange data during learning in order to improve their individual specialization. Experimental results using different datasets demonstrate the good performance in accuracy of the decision regions of the proposed method in comparison with other well-known classifiers while saving training time due to its distributed nature.
Enrique F. Castillo, Diego Peteiro-Barral, Bertha Guijarro-Berdiñas, Oscar Fontenla-Romero
Int. J. Neural Syst.3
2014 Learning on Vertically Partitioned Data based on Chi-square Feature Selection and Naive Bayes Classification
abstract
In the last few years, distributed learning has been the focus of much attention due to the explosion of big databases, in some cases distributed across different nodes. However, the great majority of current selection and classification algorithms are designed for centralized learning, i.e. they use the whole dataset at once. In this paper, a new approach for learning on vertically partitioned data is presented, which covers both feature selection and classification. The approach splits the data by features, and then uses the chi-square filter and the naive Bayes classifier to learn at each node. Finally, a merging procedure is performed, which updates the learned model in an incremental fashion. The experimental results on five representative datasets show that the execution time is shortened considerably whereas the classification performance is maintained as the number of nodes increases.
Verónica Bolón-Canedo, Diego Peteiro-Barral, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Noelia Sánchez-Maroño
ICAART (1)4
2014 Self-adaptive Topology Neural Network for Online Incremental Learning
abstract
Many real problems in machine learning are of dynamic nature. In that cases, the model used for the learning process should work in real time and have the ability to act and react by itself, adjusting its controlling parameters, even its structures, depending on the requirements of the process. In a previous work, the authors proposed an online learning method for two-layer feedforward neural networks that presents two main characteristics. Firstly, it is effective in dynamic environments as well as in stationary contexts. Secondly, it allows to incorporate new hidden neurons during learning without loosing the knowledge already acquired. In this paper, we extended this previous algorithm including a mechanism to automatically adapt the network topology according with the needs of the learning process. This automatic estimation technique is based on the Vapnik-Chervonenkis dimension. The theoretical basis for the method is given and its performance is illustrated by means of its application to different system identification problems. The results confirm that the proposed method is able to check whether new hidden units should be added depending on the requirements of the online learning process.
Beatriz Pérez-Sánchez, Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas
ICAART (1)3
2014 A Methodology for Improving Tear Film Lipid Layer Classification
abstract
Dry eye is a symptomatic disease which affects a wide range of population and has a negative impact on their daily activities. Its diagnosis can be achieved by analyzing the interference patterns of the tear film lipid layer and by classifying them into one of the Guillon categories. The manual process done by experts is not only affected by subjective factors but is also very time consuming. In this paper we propose a general methodology to the automatic classification of tear film lipid layer, using color and texture information to characterize the image and feature selection methods to reduce the processing time. The adequacy of the proposed methodology was demonstrated since it achieves classification rates over 97% while maintaining robustness and provides unbiased results. Also, it can be applied in real time, and so allows important time savings for the experts.
Beatriz Remeseiro, Verónica Bolón-Canedo, Diego Peteiro-Barral, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Antonio Mosquera González, Manuel G. Penedo, Noelia Sánchez-Maroño
IEEE J. Biomed. Health Informatics5
2013 A Distributed Learning Algorithm Based on Frontier Vector Quantization and Information Theory
Diego Peteiro-Barral, Bertha Guijarro-Berdiñas
ICANN2
2013 An online learning algorithm for adaptable topologies of neural networks
Beatriz Pérez-Sánchez, Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas, David Martínez-Rego
Expert Syst. Appl.3
2013 Toward the scalability of neural networks through feature selection
Diego Peteiro-Barral, Verónica Bolón-Canedo, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Noelia Sánchez-Maroño
Expert Syst. Appl.4
2013 A comparative study of the scalability of a sensitivity-based learning algorithm for artificial neural networks
Diego Peteiro-Barral, Bertha Guijarro-Berdiñas, Beatriz Pérez-Sánchez, Oscar Fontenla-Romero
Expert Syst. Appl.2
2012 Interferential Tear Film Lipid Layer Classification: An Automatic Dry Eye Test
abstract
Dry eye is a symptomatic disease which affects a wide range of population and has a negative impact on their daily activities, such as driving or working with computers. Its diagnosis can be achieved by several clinical tests, one of which is the analysis of the interference pattern and its classification into one of the Guillon's categories. The methodologies for automatic classification obtain promising results but at the expense of requiring a long processing time. In this research, feature selection techniques are used to reduce time whilst maintaining performance, paving the way for the development of a novel tool for automatic classification of tear film lipid layer. This tool produces significant classification rates over 96% compared with the annotations of the optometrists and provides unbiased results. Also, it works in real-time and so allows important time savings for the experts.
Verónica Bolón-Canedo, Diego Peteiro-Barral, Beatriz Remeseiro, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Antonio Mosquera González, Manuel G. Penedo, Noelia Sánchez-Maroño
ICTAI5
2012 An Analysis of Clustering Approaches to Distributed Learning on Heterogeneously Distributed Datasets
abstract
Advances in communication technologies have contributed to the proliferation of distributed datasets. The most effective approach to distributed learning is to learn locally and then combine the local models. In general, distributed algorithms assume that there is a single model that could be induced from the distributed datasets. Under this view, distribution is treated exclusively as a technical issue. However, real-world distributed datasets frequently present an intrinsic data skewness among their partitions. Despite of its importance, up to the authors’ knowledge, its impact has been barely investigated in the literature. In this paper, the performance of different cluster-based distributed learning methods is analyzed over distinct scenarios by incrementing the differences in the probabilistic distribution of data among partitions. Based on these results the best approach is suggested at every scenario.
Diego Peteiro-Barral, Bertha Guijarro-Berdiñas
KES2
2012 A mixture of experts for classifying sleep apneas
Bertha Guijarro-Berdiñas, Elena Hernández-Pereira, Diego Peteiro-Barral
Expert Syst. Appl.1
2011 A distributed learning algorithm based on two-layer artificial neural networks and genetic algorithms
Diego Peteiro-Barral, Bertha Guijarro-Berdiñas, Beatriz Pérez-Sánchez, Oscar Fontenla-Romero
ESANN2
2011 Dealing with "Very Large" Datasets - An Overview of a Promising Research Line: Distributed Learning
Diego Peteiro-Barral, Bertha Guijarro-Berdiñas, Beatriz Pérez-Sánchez
ICAART (1)2
2010 A Privacy-Preserving Distributed and Incremental Learning Method for Intrusion Detection
Bertha Guijarro-Berdiñas, Santiago Fernández-Lorenzo, Noelia Sánchez-Maroño, Oscar Fontenla-Romero
ICANN (1)1
2010 An incremental learning method for neural networks in adaptive environments
abstract
Many real scenarios in machine learning are non-stationary. These challenges forces to develop new algorithms that are able to deal with changes in the underlying problem to be learnt. These changes can be gradual or abrupt. As the dynamics of the changes can be different, the existing machine learning algorithms exhibit difficulties to cope with them. In this work we propose a new method, that is based in the introduction of a forgetting function in an incremental online learning algorithm for two-layer feedforward neural networks. This forgetting function gives a monotonically crescent importance to new data. Due to this fact, the network forgets in presence of changes while maintaining a stable behavior when the context is stationary. The theoretical basis for the method is given and its performance is illustrated by evaluating its behavior. The results confirm that the proposed method is able to work in evolving environments.
Beatriz Pérez-Sánchez, Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas
IJCNN3
2010 A new convex objective function for the supervised learning of single-layer neural networks
Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas, Beatriz Pérez-Sánchez, Amparo Alonso-Betanzos
Pattern Recognit.2
2008 A Regularized Learning Method for Neural Networks Based on Sensitivity Analysis
Bertha Guijarro-Berdiñas, Oscar Fontenla-Romero, Beatriz Pérez-Sánchez, Amparo Alonso-Betanzos
ESANN1
2007 A Fast Semi-linear Backpropagation Learning Algorithm
Bertha Guijarro-Berdiñas, Oscar Fontenla-Romero, Beatriz Pérez-Sánchez, Paula Fraguela
ICANN (1)1
2007 A Linear Learning Method for Multilayer Perceptrons Using Least-Squares
Bertha Guijarro-Berdiñas, Oscar Fontenla-Romero, Beatriz Pérez-Sánchez, Paula Fraguela
IDEAL1
2006 A Very Fast Learning Method for Neural Networks Based on Sensitivity Analysis
abstract
This paper introduces a learning method for two-layer feedforward neural networks based on sensitivity analysis, which uses a linear training algorithm for each of the two layers. First, random values are assigned to the outputs of the first layer; later, these initial values are updated based on sensitivity formulas, which use the weights in each of the layers; the process is repeated until convergence. Since these weights are learnt solving a linear system of equations, there is an important saving in computational time. The method also gives the local sensitivities of the least square errors with respect to input and output data, with no extra computational cost, because the necessary information becomes available without extra calculations. This method, called the Sensitivity-Based Linear Learning Method, can also be used to provide an initial set of weights, which significantly improves the behavior of other learning algorithms. The theoretical basis for the method is given and its performance is illustrated by its application to several examples in which it is compared with several learning algorithms and well known data sets. The results have shown a learning speed generally faster than other existing methods. In addition, it can be used as an initialization tool for other well known methods with significant improvements.
Enrique F. Castillo, Bertha Guijarro-Berdiñas, Oscar Fontenla-Romero, Amparo Alonso-Betanzos
J. Mach. Learn. Res.2
2005 A new method for sleep apnea classification using wavelets and feedforward neural networks
Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Vicente Moret-Bonillo
Artif. Intell. Medicine2
2004 A measure of fault tolerance for functional networks
Oscar Fontenla-Romero, Enrique F. Castillo, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas
Neurocomputing4
2003 A Bayesian Neural Network Approach for Sleep Apnea Classification
Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Ana del Rocío Fraga-Iglesias, Vicente Moret-Bonillo
AIME2
2003 Self-organizing maps and functional networks for local dynamic modeling
Noelia Sánchez-Maroño, Oscar Fontenla-Romero, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas
ESANN4
2003 An intelligent system for forest fire risk prediction and fire fighting management in Galicia
Amparo Alonso-Betanzos, Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas, Elena Hernández-Pereira, Maria Inmaculada Paz-Andrade, Eulogio Jimenez, Jose Luis Legido, Tarsy Carballas
Expert Syst. Appl.3
2002 A Neural Network Approach for Forestal Fire Risk Estimation
Amparo Alonso-Betanzos, Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas, Elena Hernández-Pereira, Juan Canda, Eulogio Jimenez, Jose Luis Legido, Susana Muñiz, Cristina Paz-Andrade, Maria Inmaculada Paz-Andrade
ECAI3
2002 Local Modeling Using Self-Organizing Maps and Single Layer Neural Networks
Oscar Fontenla-Romero, Amparo Alonso-Betanzos, Enrique F. Castillo, José C. Príncipe, Bertha Guijarro-Berdiñas
ICANN5
2002 Intelligent analysis and pattern recognition in cardiotocographic signals using a tightly coupled hybrid system
Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Oscar Fontenla-Romero
Artif. Intell.1
2002 Empirical evaluation of a hybrid intelligent monitoring system using different measures of effectiveness
Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos
Artif. Intell. Medicine1
2002 A Global Optimum Approach for One-Layer Neural Networks
abstract
The article presents a method for learning the weights in one-layer feedforward neural networks minimizing either the sum of squared errors or the maximum absolute error, measured in the input scale. This leads to the existence of a global optimum that can be easily obtained solving linear systems of equations or linear programming problems, using much less computational power than the one associated with the standard methods. Another version of the method allows computing a large set of estimates for the weights, providing robust, mean or median, estimates for them, and the associated standard errors, which give a good measure for the quality of the fit. Later, the standard one-layer neural network algorithms are improved by learning the neural functions instead of assuming them known. A set of examples of applications is used to illustrate the methods. Finally, a comparison with other high-performance learning algorithms shows that the proposed methods are at least 10 times faster than the fastest standard algorithm used in the comparison.
Enrique F. Castillo, Oscar Fontenla-Romero, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos
Neural Comput.3
2001 Adaptive pattern recognition in the analysis of cardiotocographic records
abstract
The recognition of accelerative and decelerative patterns in the fetal heart rate (FHR) is one of the tasks carried out manually by obstetricians when they analyze cardiotocograms for information respecting the fetal state. An approach based on artificial neural networks formed by a multilayer perceptron (MLP) is developed. However, since the system utilizes the FHR signal as direct input, an anterior stage must be incorporated that applies a principal component analysis (PCA) so as to make the system independent of the signal baseline. Furthermore, the introduction of multiresolution into the PCA has resolved other problems that were detected in the application of the system. Presented in this paper are the results of validation of these systems designated the PCA-MLP and multiresolutlon principal component analysis (MR-PCA) systems against three clinical experts.
Oscar Fontenla-Romero, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas
IEEE Trans. Neural Networks3
1995 The NST-EXPERT project: the need to evolve
Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Vicente Moret-Bonillo, S. Lopez-Gonzalez
Artif. Intell. Medicine2