Carlos Eiras-Franco

dblp:192/9059 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0001-6322-7593ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-author · 2 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A hybrid metaheuristics-Bayesian optimization framework with safe transfer learning for continuous spark tuning
abstract
Tuning configuration parameters in distributed Big Data engines such as Apache Spark is a high-dimensional, workload-dependent problem with significant impact on performance and operational cost. We address this challenge with a hybrid optimization framework that integrates Iterated Local Search, Tabu Search, and locally embedded Bayesian Optimization guided by STL-PARN (safe transfer learning with pattern-adaptive robust neighborhoods). Historical executions are partitioned into a Nucleus of reliable neighbors and a Corona of exploratory configurations, ensuring relevance while mitigating negative transfer. The surrogate within the embedded Bayesian Optimization stage decouples performance prediction from uncertainty modeling, enabling parameter-free acquisition functions that self-adapt to diverse workloads. Experiments on a modernized HiBench suite across multiple input scales show consistent gains over state-of-the-art baselines in execution time, convergence, and cost efficiency. Overall, the results demonstrate the robustness and practical value of embedding Bayesian Optimization within a global metaheuristic loop for adaptive, cost-aware Spark tuning. All source code and datasets are publicly available, supporting reproducibility and operational efficiency in large-scale data processing.
Mariano Garralda-Barrio, Carlos Eiras-Franco, Verónica Bolón-Canedo
Future Gener. Comput. Syst.2
2026 Contrastive Learning for Explanation Ranking
abstract
Abstract Explainable recommendation systems enhance user trust and satisfaction by revealing the reasoning behind personalized recommendations. Approaching this as a post-hoc explanation-ranking problem over a fixed pool of candidate explanations, we propose Contrastive Learning for Explanation Ranking (CLER), a model that learns user, item, and explanation representations with a Normalized Temperature-scaled Binary Cross-Entropy (NT-BXent) loss. This function specifically applies a per-row reweighting strategy, preventing the vast number of negative examples from dominating the objective. We evaluate CLER on the Amazon, TripAdvisor, and Yelp datasets from the EXTRA benchmark. Across traditional ranking metrics, CLER achieves the strongest results among the compared baselines.
Miguel Escarda-Fernández, Brais Cancela, Carlos Eiras-Franco, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos
Mach. Learn.3
2025 Evaluate with the Inverse: Efficient Approximation of Latent Explanation Quality Distribution
abstract
Obtaining high-quality explanations of a model's output enables developers to identify and correct biases, align the system's behavior with human values, and ensure ethical compliance. Explainable Artificial Intelligence (XAI) practitioners rely on specific measures to gauge the quality of such explanations. These measures assess key attributes, such as how closely an explanation aligns with a model's decision process (faithfulness), how accurately it pinpoints the relevant input features (localization), and its consistency across different cases (robustness). Despite providing valuable information, these measures do not fully address a critical practitioner's concern: how does the quality of a given explanation compare to other potential explanations? Traditionally, the quality of an explanation has been assessed by comparing it to a randomly generated counterpart. This paper introduces an alternative: the Quality Gap Estimate (QGE). The QGE method offers a direct comparison to what can be viewed as the `inverse' explanation, one that conceptually represents the antithesis of the original explanation. Our extensive testing across multiple model architectures, datasets, and established quality metrics demonstrates that the QGE method is superior to the traditional approach. Furthermore, we show that QGE enhances the statistical reliability of these quality assessments. This advance represents a significant step toward a more insightful evaluation of explanations that enables a more effective inspection of a model's behavior.
Carlos Eiras-Franco, Anna Hedström, Marina M.-C. Höhne
AAAI1
2025 Predictively Combatting Toxicity in Health-related Online Discussions through Machine Learning
abstract
In health-related topics, user toxicity in online discussions frequently becomes a source of social conflict or promotion of dangerous, unscientific behaviour; common approaches for battling it include different forms of detection, flagging and/or removal of existing toxic comments, which is often counterproductive for platforms and users alike. In this work, we propose the alternative of combatting user toxicity predictively, anticipating where a user could interact toxically in health-related online discussions. Applying a Collaborative Filtering-based Machine Learning methodology, we predict the toxicity in COVID-related conversations between any user and subcommunity of Reddit, surpassing 80% predictive performance in relevant metrics, and allowing us to prevent the pairing of conflicting users and subcommunities.
Jorge Paz-Ruza, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Carlos Eiras-Franco
IJCNN4
2025 Performance and sustainability of BERT derivatives in dyadic data
abstract
[Abstract]: In recent years, the Natural Language Processing (NLP) field has experienced a revolution, where numerous models – based on the Transformer architecture – have emerged to process the ever-growing volume of online text-generated data. This architecture has been the basis for the rise of Large Language Models (LLMs). Enabling their application to many diverse tasks in which they excel with just a fine-tuning process that comes right after a vast pre-training phase. However, their sustainability can often be overlooked, especially regarding computational and environmental costs. Our research aims to compare various BERT derivatives in the context of a dyadic data task while also drawing attention to the growing need for sustainable AI solutions. To this end, we utilize a selection of transformer models in an explainable recommendation setting, modeled as a multi-label classification task originating from a social network context, where users, restaurants, and reviews interact.
Miguel Escarda-Fernández, Carlos Eiras-Franco, Brais Cancela, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos
Expert Syst. Appl.2
2025 Adaptive incremental transfer learning for efficient performance modeling of big data workloads
Mariano Garralda-Barrio, Carlos Eiras-Franco, Verónica Bolón-Canedo
Future Gener. Comput. Syst.2
2025 Beyond RMSE and MAE: Introducing EAUC to Unmask Hidden Bias and Unfairness in Dyadic Regression Models
abstract
Dyadic regression models, which output real-valued predictions for pairs of entities, are fundamental in many domains [e.g., obtaining user-product ratings in recommender systems (RSs)] and promising and under exploration in others (e.g., tuning patient-drug dosages in precision pharmacology). In this work, we prove that nonuniform observed value distributions of individual entities lead to severe biases in state-of-the-art models, skewing predictions toward the average of observed past values for the entity and providing worse-than-random predictive power in eccentric yet crucial cases; we name this phenomenon eccentricity bias. We show that global error metrics like root-mean-squared error (RMSE) are insufficient to capture this bias, and we introduce eccentricity area under the curve (EAUC) as a novel metric that can quantify it in all studied domains and models. We prove the intuitive interpretation of EAUC by experimenting with naive post-training bias corrections and theorize other options to use EAUC to guide the construction of fair models. This work contributes a bias-aware evaluation of dyadic regression to prevent unfairness in critical real-world applications of such systems.
Jorge Paz-Ruza, Amparo Alonso-Betanzos, Bertha Guijarro-Berdiñas, Brais Cancela, Carlos Eiras-Franco
IEEE Trans. Neural Networks Learn. Syst.5
2024 A novel framework for generic Spark workload characterization and similar pattern recognition using machine learning
abstract
Comprehensive workload characterization plays a pivotal role in comprehending Spark applications, as it enables the analysis of diverse aspects and behaviors. This understanding is indispensable for devising downstream tuning objectives, such as performance improvement. To address this pivotal issue, our work introduces a novel and scalable framework for generic Spark workload characterization, complemented by consistent geometric measurements. The presented approach aims to build robust workload descriptors by profiling only quantitative metrics at the application task-level, in a non-intrusive manner. We expand our framework for downstream workload pattern recognition by incorporating unsupervised machine learning techniques: clustering algorithms and feature selection. These techniques significantly improve the process of grouping similar workloads without relying on predefined labels. We effectively recognize 24 representative Spark workloads from diverse domains, including SQL, machine learning, web search, graph, and micro-benchmarks, available in HiBench. Our framework achieves a high accuracy F-Measure score of up to 90.9% and a Normalized Mutual Information of up to 94.5% in similar workload pattern recognition. These scores significantly outperform the results obtained in a comparative analysis with an established workload characterization approach in the literature.
Mariano Garralda-Barrio, Carlos Eiras-Franco, Verónica Bolón-Canedo
J. Parallel Distributed Comput.2
2022 Sustainable Personalisation and Explainability in Dyadic Data Systems
abstract
Systems that rely on dyadic data, which relate entities of two types together, have become ubiquitously used in fields such as media services, tourism business, e-commerce, and others. However, these systems have had a tendency to be black-box systems, despite their objective of influencing people's decisions. There is a lack of research on providing personalised explanations to the outputs of systems that make use of such data, that is, integrating the idea of Explainable Artificial Intelligence into the field of dyadic data. Moreover, the existing approaches rely heavily on Deep Learning models for their training, reducing their overall sustainability. In this work, we propose a computationally efficient model which provides personalisation by generating explanations based on user-created images. In the context of a particular dyadic data system, the restaurant review platform TripAdvisor, we predict, for any (user, restaurant) pair, the review of the restaurant that is most adequate to present it to the user, based on their personal preferences. This model exploits the usage of efficient Matrix Factorisation techniques combined with feature-rich embeddings of the pre-trained Image Classification models, developing a method capable of providing transparency to dyadic data systems while reducing as much as 80% the carbon emissions of training compared to alternative approaches.
Jorge Paz-Ruza, Carlos Eiras-Franco, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos
KES2
2022 Fast anomaly detection with locality-sensitive hashing and hyperparameter autotuning
abstract
This paper presents LSHAD, an anomaly detection (AD) method based on Locality Sensitive Hashing (LSH), capable of dealing with large-scale datasets. The resulting algorithm is highly parallelizable and its implementation in Apache Spark further increases its ability to handle very large datasets. Moreover, the algorithm incorporates an automatic hyperparameter tuning mechanism so that users do not have to implement costly manual tuning. Our LSHAD method is novel as both hyperparameter automation and distributed properties are not usual in AD techniques. Our results for experiments with LSHAD across a variety of datasets point to state-of-the-art AD performance while handling much larger datasets than state-of-the-art alternatives. In addition, evaluation results for the tradeoff between AD performance and scalability show that our method offers significant advantages over competing methods.
Jorge Meira, Carlos Eiras-Franco, Verónica Bolón-Canedo, Goreti Marreiros, Amparo Alonso-Betanzos
Inf. Sci.2
2021 Scalable feature selection using ReliefF aided by locality-sensitive hashing
abstract
Feature selection algorithms, such as ReliefF, are very important for processing high-dimensionality data sets. However, widespread use of popular and effective such algorithms is limited by their computational cost. We describe an adaptation of the ReliefF algorithm that simplifies the costliest of its step by approximating the nearest neighbor graph using locality-sensitive hashing (LSH). The resulting ReliefF-LSH algorithm can process data sets that are too large for the original ReliefF, a capability further enhanced by distributed implementation in Apache Spark. Furthermore, ReliefF-LSH obtains better results and is more generally applicable than currently available alternatives to the original ReliefF, as it can handle regression and multiclass data sets. The fact that it does not require any additional hyperparameters with respect to ReliefF also avoids costly tuning. A set of experiments demonstrates the validity of this new approach and confirms its good scalability.
Carlos Eiras-Franco, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Antonio Bahamonde
Int. J. Intell. Syst.1
2020 Fast Distributed kNN Graph Construction Using Auto-tuned Locality-sensitive Hashing
abstract
The k -nearest-neighbors ( k NN) graph is a popular and powerful data structure that is used in various areas of Data Science, but the high computational cost of obtaining it hinders its use on large datasets. Approximate solutions have been described in the literature using diverse techniques, among which Locality-sensitive Hashing (LSH) is a promising alternative that still has unsolved problems. We present Variable Resolution Locality-sensitive Hashing, an algorithm that addresses these problems to obtain an approximate k NN graph at a significantly reduced computational cost. Its usability is greatly enhanced by its capacity to automatically find adequate hyperparameter values, a common hindrance to LSH-based methods. Moreover, we provide an implementation in the distributed computing framework Apache Spark that takes advantage of the structure of the algorithm to efficiently distribute the computational load across multiple machines, enabling practitioners to apply this solution to very large datasets. Experimental results show that our method offers significant improvements over the state-of-the-art in the field and shows very good scalability as more machines are added to the computation.
Carlos Eiras-Franco, David Martínez-Rego, Leslie Kanthan, César Piñeiro, Antonio Bahamonde, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos
ACM Trans. Intell. Syst. Technol.1
2019 Case Study of Anomaly Detection and Quality Control of Energy Efficiency and Hygrothermal Comfort in Buildings
abstract
[Abstract] The aim of this work is to propose different statistical and machine learning methodologies for identifying anomalies and control the quality of energy efficiency and hygrothermal comfort in buildings. Companies focused on energy sector for buildings are interested on statistical and machine learning tools to automate the control of energy consumption and ensure quality of Heat Ventilation and Air Conditioning (HVAC) installations. Consequently, a methodology based on the application of the Local Correlation Integral (LOCI) anomaly detection technique has been proposed. In addition, the most critical variables for anomaly detection are identified by using ReliefF method. Once vectors of critical variables are obtained, multivariate and univariate control charts can be applied to control the quality of HVAC installations (consumption, thermal comfort). In order to test the proposed methodology, the companies involved in this project have provided the case study of a store of a clothing brand located in a shopping center in Panama. It is important to note that this is a controlled case study for which all the anomalies have been previously identified by maintenance personnel. Moreover, as an alternatively solution, in addition to machine learning and multivariate techniques, new nonparametric control charts for functional data based on data depth have been proposed and applied to curves of daily energy consumption in HVAC.
Carlos Eiras-Franco, Miguel Flores, Verónica Bolón-Canedo, Sonia Zaragoza, Rubén Fernández-Casal, Salvador Naya, Javier Tarrío-Saavedra
DATA1
2019 A scalable decision-tree-based method to explain interactions in dyadic data
Carlos Eiras-Franco, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Antonio Bahamonde
Decis. Support Syst.1
2019 Large scale anomaly detection in mixed numerical and categorical input spaces
Carlos Eiras-Franco, David Martínez-Rego, Bertha Guijarro-Berdiñas, Amparo Alonso-Betanzos, Antonio Bahamonde
Inf. Sci.1
2017 Scalable approximate k-NN Graph construction based on Locality Sensitive Hashing
Carlos Eiras-Franco, Leslie Kanthan, Amparo Alonso-Betanzos, David Martínez-Rego
ESANN1