VLDB 2026 Research / reviewers in the wild / expert
Yoan Gutiérrez
dblp:66/9766 · also Yoan Gutiérrez Vázquez
· DBLP profile ↗
22ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0002-4052-7427ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoMLabstractExperts in machine learning leverage domain knowledge to navigate decisions in model selection, hyperparameter optimization, and resource allocation.This is particularly critical for fine-tuning language models (LMs), where repeated trials incur substantial computational overhead and environmental impact.However, no existing automated framework simultaneously tackles the entire model selection and hyperparameter optimization (HPO) task for resource-efficient LM fine-tuning.We introduce XAutoLM, a meta-learning-augmented AutoML framework that reuses past experiences to optimize discriminative and generative LM fine-tuning pipelines efficiently.XAutoLM learns from stored successes and failures by extracting task-and system-level meta-features to bias its sampling toward valuable configurations and away from costly dead ends.On four text classification and two question-answering benchmarks, XAutoLM surpasses zero-shot optimizer's peak F 1 on five of six tasks, cuts mean evaluation time of pipelines by up to 4.5x, reduces search error ratios by up to sevenfold, and uncovers up to 50% more pipelines above the zero-shot Pareto front.In contrast, simpler memory-based baselines suffer negative transfer.We release XAutoLM and our experience store to catalyze resource-efficient, Green AI fine-tuning in the NLP community. Ernesto Luis Estevanell-Valladares, Suilan Estévez-Velarde, Yoan Gutiérrez, Andrés Montoyo, Ruslan Mitkov |
EMNLP | 3 |
| 2025 | Bias mitigation for fair automation of classification tasksabstractAbstract The incorporation of machine learning algorithms into high‐risk decision‐making tasks has raised some alarms in the scientific community. Research shows that machine learning‐based technologies can contain biases that cause unfair decisions for certain population groups. The fundamental danger of ignoring this problem is that machine learning methods can not only reflect the biases present in our society but could also amplify them. This article presents the design and validation of a technology to assist the fair automation of classification problems. In essence, the proposal is based on taking advantage of the intermediate solutions generated during the resolution of classification problems through using Auto‐ML tools, in particular, AutoGOAL, to create unbiased/fair classifiers. The technology employs a multi‐objective optimization search to find the collection of models with the best trade‐offs between performance and fairness. To solve the optimization problem, we introduce a combination of Probabilistic Grammatical Evolution Search and NSGA‐II. The technology was evaluated using the Adult dataset from the UCI repository, a common benchmark in related research. Results were compared with other published results in scenarios with single and multiple fairness definitions. Our experiments demonstrate the technology's ability to automate classification tasks while incorporating fairness constraints. Additionally, our method achieves competitive results against other bias mitigation techniques. A notable advantage of our approach is its minimal requirement for machine learning expertise, thanks to its Auto‐ML foundation. This makes the technology accessible and valuable for advancing fairness in machine learning applications. The source code is available online for the research community. Juan Pablo Consuegra-Ayala, Yoan Gutiérrez, Yudivián Almeida-Cruz, Manuel Palomar |
Expert Syst. J. Knowl. Eng. | 2 |
| 2024 | A comprehensive methodology to construct standardised datasets for Science and Technology ParksabstractThis work presents a standardised approach to create datasets for Science and Technology Parks (STPs), facilitating future analysis of STP characteristics, trends and performance. STPs are the most representative examples of innovation ecosystems. The ETL (extraction-transformation-load) structure was adapted to a global field study of STPs. A selection stage and quality check were incorporated, and the methodology was applied to Spanish STPs. This study applies diverse techniques such as expert labelling and information extraction which uses language technologies. A novel methodology for building quality and standardised STP datasets was designed and applied to a Spanish STP case study with 49 STPs. An updatable dataset and a list of the main features impacting STPs are presented. Twenty-one (n=21) core features were refined and selected, with fifteen of them (71.4%) being robust enough for developing further quality analysis. The methodology presented integrates different sources with heterogeneous information that is often decentralised, disaggregated and in different formats: excel files, and unstructured information in HTML or PDF format. The existence of this updatable dataset and the defined methodology will enable powerful AI tools to be applied that focus on more sophisticated analysis, such as taxonomy, monitoring, and predictive and prescriptive analytics in the innovation ecosystems field. Olga Francés, Javi Fernández, José Ignacio Abreu, Yoan Gutiérrez, Manuel Palomar |
Data Knowl. Eng. | 4 |
| 2024 | KD SENSO-MERGER: An architecture for semantic integration of heterogeneous dataabstractThis paper presents KD SENSO-MERGER, a novel Knowledge Discovery (KD) architecture that is capable of semantically integrating heterogeneous data from various sources of structured and unstructured data (i.e. geolocations, demographic, socio-economic, user reviews, and comments). This goal drives the main design approach of the architecture. It works by building internal representations that adapt and merge knowledge across multiple domains, ensuring that the knowledge base is continuously updated. To deal with the challenge of integrating heterogeneous data, this proposal puts forward the corresponding solutions: (i) knowledge extraction, addressed via a plugin-based architecture of knowledge sensors; (ii) data integrity, tackled by an architecture designed to deal with uncertain or noisy information; (iii) scalability, this is also supported by the plugin-based architecture as only relevant knowledge to the scenario is integrated by switching-off non-relevant sensors. Also, we minimize the expert knowledge required, which may pose a bottleneck when integrating a fast-paced stream of new sources. As proof of concept, we developed a case study that deploys the architecture to integrate population census and economic data, municipal cartography, and Google Reviews to analyze the socio-economic contexts of educational institutions. The knowledge discovered enables us to answer questions that are not possible through individual sources. Thus, companies or public entities can discover patterns of behavior or relationships that would otherwise not be visible and this would allow extracting valuable information for the decision-making process. Yoan Gutiérrez, José Ignacio Abreu, Andrés Montoyo, Rafael Muñoz 0001, Suilan Estévez-Velarde |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Automatic annotation of protected attributes to support fairness optimization
Juan Pablo Consuegra-Ayala, Yoan Gutiérrez, Yudivián Almeida-Cruz, Manuel Palomar |
Inf. Sci. | 2 |
| 2022 | Why are some social-media contents more popular than others? Opinion and association rules mining applied to virality patterns discoveryabstractDiscovering the main features of virality patterns in Twitter is the focus of this research. Five trending topics related to the COVID-19 pandemic were selected for the study, with Spanish as the target language. To carry out the discovery of virality patterns, we applied opinion mining techniques that enable us to structure the information based on the polarity of the messages and the emotions they contain. After transforming the information from an unstructured textual representation to a structured one, data mining techniques were applied, specifically association rules mining. Message patterns with the highest virality (high shares and high likes), and at the same time the most relevant characteristics of the patterns with less impact were extracted. After an exhaustive analysis of the most relevant non-redundant rules, it can be concluded that messages with a high-negative polarity and a very high emotional charge, especially emotions that have intensified with the COVID-19 pandemic, such as fear, sadness, anger and surprise are more likely to go viral in social media. By contrast, messages with little news coverage in the media, few authors, and the absence of surprise are relevant features when it comes to seeing messages with very low dissemination in social media. Estela Saquete Boró, José Jacobo Zubcoff, Yoan Gutiérrez, Patricio Martínez-Barco, Javi Fernández |
Expert Syst. Appl. | 3 |
| 2022 | Intelligent ensembling of auto-ML system outputs for solving classification problems
Juan Pablo Consuegra-Ayala, Yoan Gutiérrez, Yudivián Almeida-Cruz, Manuel Palomar |
Inf. Sci. | 2 |
| 2021 | General-purpose hierarchical optimisation of machine learning pipelines with grammatical evolution
Suilan Estévez-Velarde, Yoan Gutiérrez, Yudivián Almeida-Cruz, Andrés Montoyo |
Inf. Sci. | 2 |
| 2021 | Automatic extension of corpora from the intelligent ensembling of eHealth knowledge discovery systems outputs
Juan Pablo Consuegra-Ayala, Yoan Gutiérrez, Alejandro Piad-Morffis, Yudivián Almeida-Cruz, Manuel Palomar |
J. Biomed. Informatics | 2 |
| 2020 | Automatic Discovery of Heterogeneous Machine Learning Pipelines: An Application to Natural Language ProcessingabstractThis paper presents AutoGOAL, a system for automatic machine learning (AutoML) that uses heterogeneous techniques.In contrast with existing AutoML approaches, our contribution can automatically build machine learning pipelines that combine techniques and algorithms from different frameworks, including shallow classifiers, natural language processing tools, and neural networks.We define the heterogeneous AutoML optimization problem as the search for the best sequence of algorithms that transforms specific input data into the desired output.This provides a novel theoretical and practical approach to AutoML.Our proposal is experimentally evaluated in diverse machine learning problems and compared with alternative approaches, showing that it is competitive with other AutoML alternatives in standard benchmarks.Furthermore, it can be applied to novel scenarios, such as several NLP tasks, where existing alternatives cannot be directly deployed.The system is freely available and includes in-built compatibility with a large number of popular machine learning frameworks, which makes our approach useful for solving practical problems with relative ease and effort. Suilan Estévez-Velarde, Yoan Gutiérrez, Andrés Montoyo, Yudivián Almeida-Cruz |
COLING | 2 |
| 2020 | A computational ecosystem to support eHealth Knowledge Discovery technologies in SpanishabstractThe massive amount of biomedical information published online requires the development of automatic knowledge discovery technologies to effectively make use of this available content. To foster and support this, the research community creates linguistic resources, such as annotated corpora, and designs shared evaluation campaigns and academic competitive challenges. This work describes an ecosystem that facilitates research and development in knowledge discovery in the biomedical domain, specifically in Spanish language. To this end, several resources are developed and shared with the research community, including a novel semantic annotation model, an annotated corpus of 1045 sentences, and computational resources to build and evaluate automatic knowledge discovery techniques. Furthermore, a research task is defined with objective evaluation criteria, and an online evaluation environment is setup and maintained, enabling researchers interested in this task to obtain immediate feedback and compare their results with the state-of-the-art. As a case study, we analyze the results of a competitive challenge based on these resources and provide guidelines for future research. The constructed ecosystem provides an effective learning and evaluation environment to encourage research in knowledge discovery in Spanish biomedical documents. Alejandro Piad-Morffis, Yoan Gutiérrez, Yudivián Almeida-Cruz, Rafael Muñoz 0001 |
J. Biomed. Informatics | 2 |
| 2019 | AutoML Strategy Based on Grammatical Evolution: A Case Study about Knowledge Discovery from TextabstractThe process of extracting knowledge from natural language text poses a complex problem that requires both a combination of machine learning techniques and proper feature selection.Recent advances in Automatic Machine Learning (AutoML) provide effective tools to explore large sets of algorithms, hyperparameters and features to find out the most suitable combination of them.This paper proposes a novel AutoML strategy based on probabilistic grammatical evolution, which is evaluated on the health domain by facing the knowledge discovery challenge in Spanish text documents.Our approach achieves state-ofthe-art results and provides interesting insights into the best combination of parameters and algorithms to use when dealing with this challenge.Source code is provided for the research community. Suilan Estévez-Velarde, Yoan Gutiérrez, Andrés Montoyo, Yudivián Almeida-Cruz |
ACL (1) | 2 |
| 2019 | Optimizing Natural Language Processing Pipelines: Opinion Mining Case Study
Suilan Estévez-Velarde, Yoan Gutiérrez, Andrés Montoyo, Yudivián Almeida-Cruz |
CIARP | 2 |
| 2019 | Developing an ontology schema for enriching and linking digital media assets
Yoan Gutiérrez, David Tomás 0001, Isabel Moreno |
Future Gener. Comput. Syst. | 1 |
| 2019 | A corpus to support eHealth Knowledge Discovery technologies
Alejandro Piad-Morffis, Yoan Gutiérrez, Rafael Muñoz 0001 |
J. Biomed. Informatics | 2 |
| 2019 | Socialising around media - Improving the second screen experience through semantic analysis, context awareness and dynamic communities
David Tomás 0001, Yoan Gutiérrez, Atta Badii, Marco Tiemann, Fotis Aisopos |
Multim. Tools Appl. | 2 |
| 2017 | Spreading semantic information by Word Sense Disambiguation
Yoan Gutiérrez, Sonia Vázquez, Andrés Montoyo |
Knowl. Based Syst. | 1 |
| 2016 | A semantic framework for textual data enrichment
Yoan Gutiérrez, Sonia Vázquez, Andrés Montoyo |
Expert Syst. Appl. | 1 |
| 2015 | Developing an Ontology to Capture Documents' SemanticsabstractThis ontology aims to capture the semantics of documents through a set of key aspects in texts, such as the temporal dimension, presence of named entities, detection of opinionated information, or conceptual classifications. In addition, the ontology provides a lexical dimension, where the sentence of each document, and a possible summary derived from it, are taken into account. These are determining factors for setting up our own interpretation of possible scenarios (a meta-level specification) and vocabulary. Since our ontology aims to be reused by a large community, we tried to establish basic NLP terminology that was hierarchized by experts in this research field. Elena Lloret, Yoan Gutiérrez, José M. Gómez |
KEOD | 2 |
| 2012 | A graph-Based Approach to WSD Using Relevant Semantic Trees and N-Cliques Model
Yoan Gutiérrez, Sonia Vázquez, Andrés Montoyo |
CICLing (1) | 1 |
| 2011 | Word Sense Disambiguation: A Graph-Based Approach Using N-Cliques Partitioning Technique
Yoan Gutiérrez, Sonia Vázquez, Andrés Montoyo |
NLDB | 1 |
| 2011 | An Unsupervised Method to Improve Spanish Stemmer
Antonio Fernández Orquín, Josval Díaz, Yoan Gutiérrez, Rafael Muñoz 0001 |
NLDB | 3 |