Carlos Badenes-Olmedo

dblp:186/2838 · also Carlos Badenes · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0002-2753-9917ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Prompt-Based Syntactic Adaptation for Easy-to-Read Texts Using Large Language Models
Pablo Gómez-Álvarez, Carlos Badenes-Olmedo, Isam Diab-Lozano
ICCHP (1)2
2026 QuerIA: adaptive question generation and evaluation in higher education using large language models and contextual learning
abstract
• Presents QuerIA, an open-source system for generating and evaluating Bloom-aligned questions from academic materials. • Combines semantic chunking and prompt-based templates to control difficulty and support multiple-choice and open-ended formats. • Piloted with 341 participants from 17 faculties, generating nearly 6,000 questions and over 800 evaluations. • Results confirm consistent variation in difficulty levels and positive pedagogical perceptions across domains. • Demonstrates the feasibility of local LLM deployment on standard, low-resource hardware, ensuring scalability and data governance in academic settings. This paper presents QuerIA, a system that automates the generation and evaluation of educational questionnaires in Spanish using large language models (LLMs). QuerIA integrates semantic chunking with a simplified Bloom’s taxonomy, enabling controllable variation of cognitive difficulty across multiple-choice (MCQ) and open-ended (OEQ) question formats. Instructional documents are segmented into coherent units, and pedagogically aligned prompts guide the generation of question–answer pairs. The system was deployed as a locally hosted web service at Universidad Politécnica de Madrid, ensuring institutional data governance and low-resource feasibility. During one academic semester, 341 participants from 17 faculties created nearly 6,000 questions and provided more than 800 evaluations on clarity, alignment, and pedagogical value. Analyses show that the predefined difficulty levels correspond to distinct linguistic and psychometric patterns, and that user feedback confirms the clarity and educational usefulness of the generated items. QuerIA demonstrates that LLM-based question generation can be applied at scale in higher education, supporting formative assessment while offering open resources for future research.
Carlos Badenes-Olmedo, Paul Eyzaguirre Barreda
Expert Syst. Appl.1
2026 ESQAD: A curriculum-aligned dataset for question answer generation in Spanish
abstract
This paper presents ESQAD (Educational Spanish Question-Answer Dataset), an open-access resource for question-answer generation (QAG) in Spanish aligned with national curricula. ESQAD comprises 6,980 expert-curated exam questions, 340 literary comprehension items, 47 legal FAQs, and 913 LLM-generated pairs validated in academic settings. The dataset covers diverse linguistic and cognitive features, including subject classification, question typologies, and Bloom-level complexity. We analyze lexical diversity, cognitive intent, and structure across subsets. The automatically generated portion was created with Bloom-guided prompts and large language models (LLMs), and evaluated through both computational metrics and classroom validation. Results show that cognitive-level control is reliable for factual recall but only partial for higher-order reasoning. All resources, including dataset, code, and validation tool, are publicly available. ESQAD provides a structured benchmark for curriculum-aware QAG in under-resourced languages, with applications in educational NLP, question generation, and adaptive learning.
Carlos Badenes-Olmedo, Paul Eyzaguirre Barreda, Noa Chu-Artzt, Joaquín Gayoso-Cabada
Knowl. Based Syst.1
2023 Automatic Topic Label Generation using Conversational Models
abstract
In probabilistic topic models, a topic is characterised by a set of words, with a probability associated to each of them. Even though it is not necessary to understand the meaning of topics to perform common downstream tasks where topic models are used, such as topic inference or document similarity, there have been attempts to uncover the semantics of topics by providing labels to them, consisting in a couple of concepts. In this paper we propose a methodology, Conversational Probabilistic Topic Labelling (CPTL), to study whether conversational models can be used to generate labels that describe probabilistic topics given their most representative keywords. We evaluate and compare the performance of a selection of conversational models for the topic label generation task with the performance of a task-specific language model trained to generate topic labels.
Virginia Ramón-Ferrer, Carlos Badenes-Olmedo, Óscar Corcho
K-CAP2
2023 Lessons learned to enable question answering on knowledge graphs extracted from scientific publications: A case study on the coronavirus literature
abstract
The article presents a workflow to create a question-answering system whose knowledge base combines knowledge graphs and scientific publications on coronaviruses. It is based on the experience gained in modeling evidence from research articles to provide answers to questions in natural language. The work contains best practices for acquiring scientific publications, tuning language models to identify and normalize relevant entities, creating representational models based on probabilistic topics, and formalizing an ontology that describes the associations between domain concepts supported by the scientific literature. All the resources generated in the domain of coronavirus are available openly as part of the Drugs4COVID initiative, and can be (re)-used independently or as a whole. They can be exploited by scientific communities conducting research related to SARS-CoV-2/COVID-19 and also by therapeutic communities, laboratories, etc., wishing to find and understand relationships between symptoms, drugs, active ingredients and their documentary evidence.
Carlos Badenes-Olmedo, Óscar Corcho
J. Biomed. Informatics1
2022 EBOCA: Evidences for BiOmedical Concepts Association Ontology
Andrea Álvarez Pérez, Ana Iglesias-Molina, Lucía Prieto Santamaría, María Poveda-Villalón, Carlos Badenes-Olmedo, Alejandro Rodríguez González
EKAW5
2021 A High-Level Ontology Network for ICT Infrastructures
Óscar Corcho, David Fraga 0001, Jhon Toledo, Julián Arenas-Guerrero, Carlos Badenes-Olmedo, Mingxue Wang, Hu Peng, Nicholas Burrett, Jose Mora, Puchao Zhang
ISWC5
2020 Enhancing Public Procurement in the European Union Through Constructing and Exploiting an Integrated Knowledge Graph
Ahmet Soylu, Óscar Corcho, Brian Elvesæter, Carlos Badenes-Olmedo, Francisco Yedro Martínez, Matej Kovacic, Matej Posinkovic, Ian Makgill, Chris Taggart, Elena Simperl, Till C. Lech, Dumitru Roman
ISWC (2)4
2019 Scalable Cross-lingual Document Similarity through Language-specific Concept Hierarchies
abstract
With the ongoing growth in number of digital articles in a wider set of languages and the expanding use of different languages, we need annotation methods that enable browsing multi-lingual corpora. Multilingual probabilistic topic models have recently emerged as a group of semi-supervised machine learning models that can be used to perform thematic explorations on collections of texts in multiple languages. However, these approaches require theme-aligned training data to create a language-independent space. This constraint limits the amount of scenarios that this technique can offer solutions to train and makes it difficult to scale up to situations where a huge collection of multi-lingual documents are required during the training phase. This paper presents an unsupervised document similarity algorithm that does not require parallel or comparable corpora, or any other type of translation resource. The algorithm annotates topics automatically created from documents in a single language with cross-lingual labels and describes documents by hierarchies of multi-lingual concepts from independently-trained models. Experiments performed on the English, Spanish and French editions of JCR-Acquis corpora reveal promising results on classifying and sorting documents by similar content.
Carlos Badenes-Olmedo, José Luis Redondo García, Óscar Corcho
K-CAP1
2017 Distributing Text Mining tasks with librAIry
abstract
We present librAIry, a novel architecture to store, process and analyze large collections of textual resources, integrating existing algorithms and tools into a common, distributed, high-performance workflow. Available text mining techniques can be incorporated as independent plug&play modules working in a collaborative manner into the framework. In the absence of a pre-defined flow, librAIry leverages on the aggregation of operations executed by different components in response to an emergent chain of events. Extensive use of Linked Data (LD) and Representational State Transfer (REST) principles are made to provide individually addressable resources from textual documents. We have described the architecture design and its implementation and tested its effectiveness in real-world scenarios such as collections of research papers, patents or ICT aids, with the objective of providing solutions for decision makers and experts in those domains. Major advantages of the framework and lessons-learned from these experiments are reported.
Carlos Badenes-Olmedo, José Luis Redondo García, Óscar Corcho
DocEng1
2017 Efficient Clustering from Distributions over Topics
abstract
There are many scenarios where we may want to find pairs of textually similar documents in a large corpus (e.g. a researcher doing literature review, or an R&D project manager analyzing project proposals). To programmatically discover those connections can help experts to achieve those goals, but brute-force pairwise comparisons are not computationally adequate when the size of the document corpus is too large. Some algorithms in the literature divide the search space into regions containing potentially similar documents, which are later processed separately from the rest in order to reduce the number of pairs compared. However, this kind of unsupervised methods still incur in high temporal costs. In this paper, we present an approach that relies on the results of a topic modeling algorithm over the documents in a collection, as a means to identify smaller subsets of documents where the similarity function can then be computed. This approach has proved to obtain promising results when identifying similar documents in the domain of scientific publications. We have compared our approach against state of the art clustering techniques and with different configurations for the topic modeling algorithm. Results suggest that our approach outperforms (> 0.5) the other analyzed techniques in terms of efficiency.
Carlos Badenes-Olmedo, José Luis Redondo García, Óscar Corcho
K-CAP1