David Tomás 0001

dblp:12/2774 · also David Tomás Díaz · DBLP profile ↗
← Back
30ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0003-3287-9366ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 13 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Ethical documentation of open-source AI models: a multivocal literature review and large-scale analysis of model cards
abstract
This study investigates ethical documentation for open-source AI models through a combination of a Multivocal Literature Review (MLR) with a large-scale empirical analysis of real-world Model Cards. We aim to address the gap in how ethical considerations are documented and implemented in both research and practice. Publications from major scholarly databases and arXiv (N = 36) were systematically retrieved, screened, and analysed. In parallel, metadata and documentation from the most-downloaded models on Hugging Face Hub were processed, yielding over 60,000 valid Model Cards after filtering. The literature review shows that Model Cards are the predominant artifact in open publication settings, but ethical coverage is uneven: transparency is most frequently addressed, while societal and environmental aspects are rarely discussed. Notably, bias is often referenced generically, with little methodological consideration. Finally, the Model Card analysis reveals three issues that indicate an implementation gap: (i) only 16.8% include explicit references to ethics (e.g., bias, limitations, risks), (ii) these sections are typically brief, and (iii) their relative prevalence declines over time despite absolute growth in model releases. These findings underscore the importance of standardized documentation, clearer guidelines, and machine-readable, verifiable ethical properties to support responsible model development and reuse.
Carmen García-Barceló, David Tomás 0001, Jose-Norberto Mazón
Expert Syst. Appl.2
2025 Exploring Content-Based Catalogs for Enhanced Discovery Services in Data Spaces
Adriana Morejón, Alberto Berenguer, Lucia de Espona, David Tomás 0001, Jose-Norberto Mazón
DOLAP4
2025 Integrating advanced vision-language models for context recognition in risks assessment
abstract
This study proposes an open-environment, multi-label human risk classification framework, capable of identifying possible risks to which individuals appearing on input video data are exposed. The framework consists of an ensemble of models covering object detection, action recognition, context understanding and text classification tasks. Each model is evaluated separately in the context of home environments, with the overall framework performing well in each evaluation after fine tuning. The models were evaluated using a combination of several datasets, including Charades, ETRI-Activity3D, and custom video question answering and risk datasets. This study exploits the ability of large language models to interpret semantic visual features combined with textual input in order to understand the context in which the person is placed. The framework’s ability to output multiple risks and its cross-domain capabilities make it a powerful tool that can enhance current risk management systems in a variety of scenarios, such as homes, construction sites and industry.
Javier Rodríguez-Juan, David Ortiz-Perez, José García Rodríguez 0001, David Tomás 0001, Grzegorz J. Nalepa
Neurocomputing4
2025 CogniAlign: Word-level multimodal speech alignment with gated cross-attention for Alzheimer's detection
abstract
Early detection of cognitive disorders such as Alzheimer’s disease is critical for enabling timely clinical intervention and improving patient outcomes. In this work, we introduce CogniAlign, a multimodal architecture for Alzheimer’s detection that integrates audio and textual modalities, two non-intrusive sources of information that offer complementary insights into cognitive health. Unlike prior approaches that fuse modalities at a coarse level, CogniAlign leverages a word-level temporal alignment strategy that synchronizes audio embeddings with corresponding textual tokens based on transcription timestamps. This alignment supports the development of token-level fusion techniques, enabling more precise cross-modal interactions. To fully exploit this alignment, we propose a Gated Cross-Attention Fusion mechanism, where audio features attend over textual representations, guided by the superior unimodal performance of the text modality. In addition, we incorporate prosodic cues, specifically interword pauses, by inserting pause tokens into the text and generating audio embeddings for silent intervals, further enriching both streams. We evaluate CogniAlign on the ADReSSo dataset, where it achieves an accuracy of 87.35% over a Leave-One-Subject-Out setup and of 90.36% over a 5 fold Cross-Validation, outperforming existing state-of-the-art methods. A detailed ablation study confirms the advantages of our alignment strategy, attention-based fusion, and prosodic modeling. Finally, we perform a corpus analysis to assess the impact of the proposed prosodic features and apply Integrated Gradients to identify the most influential input segments used by the model in predicting cognitive health outcomes.
David Ortiz-Perez, Manuel Benavent-Lledó, Javier Rodríguez-Juan, José García Rodríguez 0001, David Tomás 0001
Knowl. Based Syst.5
2024 Evaluating the Impact of Content Deletion on Tabular Data Similarity and Retrieval Using Contextual Word Embeddings
Alberto Berenguer, David Tomás 0001, Jose-Norberto Mazón
ECIR (2)2
2024 Multimodal Fusion Strategies for Emotion Recognition
abstract
Emotions play a crucial role in our daily lives, influencing how we face challenges throughout the day and shaping our behavior, even when we are not consciously aware of them. Detecting emotional states in others relies on comprehending the collective impact of a variety of actions that emotions can produce, such as facial expressions, posture, tone of voice, or speech. To address this challenge, we propose a multimodal transformer-based model designed to recognize the emotional moods of individuals using video data, including audio and text transcriptions. Consequently, our model extracts the most relevant information from each modality to make a final prediction. Throughout this work, different fusion architectures have been integrated with transformer-based models to determine the optimal combination. This study examines the performance of individual modalities and their combinations using the CMU-MOSEI dataset. This dataset encompasses preprocessed video, audio, and text data. Our best model achieves a weighted accuracy of 85.59% on this dataset, surpassing previous works for this task.
David Ortiz-Perez, Manuel Benavent-Lledó, David Mulero-Pérez, David Tomás 0001, José García Rodríguez 0001
IJCNN4
2024 Cognitive Insights Across Languages: Enhancing Multimodal Interview Analysis
David Ortiz-Perez, José García Rodríguez 0001, David Tomás 0001
INTERSPEECH3
2024 Word embeddings for retrieving tabular data from research publications
abstract
Abstract Scientists face challenges when finding datasets related to their research problems due to the limitations of current dataset search engines. Existing tools for searching research datasets rely on publication content or metadata, do not considering the data contained in the publication in the form of tables. Moreover, scientists require more elaborate inputs and functionalities to retrieve different parts of an article, such as data presented in tables, based on their search purposes. Therefore, this paper proposes a novel approach to retrieve relevant tabular datasets from publications. The input of our system is a research problem stated as an abstract from a scientific paper, and the output is a set of relevant tables from publications that are related to the research problem. This approach aims to provide a better solution for scientists to find useful datasets that support them in addressing their research problems. To validate this approach, experiments were conducted using word embedding from different language models to calculate the semantic similarity between abstracts and tables. The results showed that contextual models significantly outperformed non-contextual models, especially when pre-trained with scientific data. Furthermore, the importance of context was found to be crucial for improving the results.
Alberto Berenguer, Jose-Norberto Mazón, David Tomás 0001
Mach. Learn.3
2023 Tabular Open Government Data Search for Data Spaces based on Word Embeddings
Alberto Berenguer, David Tomás 0001, Jose-Norberto Mazón
DOLAP2
2023 How Challenging is Multimodal Irony Detection?
Manuj Malik, David Tomás 0001, Paolo Rosso
NLDB2
2023 A Deep Learning-Based Multimodal Architecture to predict Signs of Dementia
abstract
This paper proposes a multimodal deep learning architecture combining text and audio information to predict dementia, a disease which affects around 55 million people all over the world and makes them in some cases dependent people. The system was evaluated on the DementiaBank Pitt Corpus dataset, which includes audio recordings as well as their transcriptions for healthy people and people with dementia. Different models have been used and tested, including Convolutional Neural Networks (CNN) for audio classification, Transformers for text classification, and a combination of both in a multimodal ensemble. These models have been evaluated on a test set, obtaining the best results by using the text modality, achieving 90.36% accuracy on the task of detecting dementia. Additionally, an analysis of the corpus has been conducted for the sake of explainability, aiming to obtain more information about how the models generate their predictions and identify patterns in the data.
David Ortiz-Perez, Pablo Ruiz-Ponce, David Tomás 0001, José García Rodríguez 0001, Maria Flores Vizcaya-Moreno, Marco Leo
Neurocomputing3
2023 Detecting and locating trending places using multimodal social network data
abstract
Abstract This paper presents a machine learning-based classifier for detecting points of interest through the combined use of images and text from social networks. This model exploits the transfer learning capabilities of the neural network architecture CLIP (Contrastive Language-Image Pre-Training) in multimodal environments using image and text. Different methodologies based on multimodal information are explored for the geolocation of the places detected. To this end, pre-trained neural network models are used for the classification of images and their associated texts. The result is a system that allows creating new synergies between images and texts in order to detect and geolocate trending places that has not been previously tagged by any other means, providing potentially relevant information for tasks such as cataloging specific types of places in a city for the tourism industry. The experiments carried out reveal that, in general, textual information is more accurate and relevant than visual cues in this multimodal setting.
Luis Lucas, David Tomás 0001, José García Rodríguez 0001
Multim. Tools Appl.2
2023 Contextual word embeddings for tabular data search and integration
abstract
Abstract This paper presents a new approach to retrieve and further integrate tabular datasets (collections of rows and columns) using union and join operations. In this work, both processes were carried out using a similarity measure based on contextual word embeddings, which allows finding semantically similar tables and overcome the recall problem of lexical approaches based on string similarity. This work is the first attempt to use contextual word embeddings in the whole pipeline of table search and integration, including for the first time their use in the join operation. A comprehensive analysis of their performance was carried out on both retrieving and integrating tabular datasets, comparing them with context-free models. Column headings and cell values were used as contextual information and their impact on each task was evaluated. The results revealed that contextual models significantly outperform context-free models and a traditional weighting schema in ad hoc table retrieval. In the data integration task, contextual models also improved the results on union operation compared to context-free approaches.
José Pilaluisa, David Tomás 0001, Borja Navarro-Colorado, Jose-Norberto Mazón
Neural Comput. Appl.2
2021 Towards a tabular open data search engine for public sector information
abstract
Public Sector Information (PSI) scenarios require tools that support retrieval of tabular open data beyond keyword-based search on metadata. This paper presents a novel interface for searching tabular open data, as well as a search engine that retrieves tabular data by considering table contents apart from metadata. Our search engine uses word embeddings to calculate the semantic similarity between tabular open data, providing a ranking of candidate tabular datasets to be integrated with an input query table according to the different intentions of the open data reuser (e.g. column or row extension as well as data completion). An initial set of experiments have been conducted, showing promising results in this task.
Alberto Berenguer, Jose-Norberto Mazón, David Tomás 0001
IEEE BigData3
2020 Fighting post-truth using natural language processing: A review and open challenges
abstract
Post-truth is a term that describes a distorting phenomenon that aims to manipulate public opinion and behavior. One of its key engines is the spread of Fake News. Nowadays most news is rapidly disseminated in written language via digital media and social networks. Therefore, to detect fake news it is becoming increasingly necessary to apply Artificial Intelligence (AI) and, more specifically Natural Language Processing (NLP). This paper presents a review of the application of AI to the complex task of automatically detecting fake news. The review begins with a definition and classification of fake news. Considering the complexity of the fake news detection task, a divide-and-conquer methodology was applied to identify a series of subtasks to tackle the problem from a computational perspective. As a result, the following subtasks were identified: deception detection; stance detection; controversy and polarization; automated fact checking; clickbait detection; and, credibility scores. From each subtask, a PRISMA compliant systematic review of the main studies was undertaken, searching Google Scholar. The various approaches and technologies are surveyed, as well as the resources and competitions that have been involved in resolving the different subtasks. The review concludes with a roadmap for addressing the future challenges that have emerged from the analysis of the state of the art, providing a rich source of potential work for the research community going forward.
Estela Saquete Boró, David Tomás 0001, Paloma Moreda, Patricio Martínez-Barco, Manuel Palomar
Expert Syst. Appl.2
2019 Developing an ontology schema for enriching and linking digital media assets
Yoan Gutiérrez, David Tomás 0001, Isabel Moreno
Future Gener. Comput. Syst.2
2019 Socialising around media - Improving the second screen experience through semantic analysis, context awareness and dynamic communities
David Tomás 0001, Yoan Gutiérrez, Atta Badii, Marco Tiemann, Fotis Aisopos
Multim. Tools Appl.1
2017 Users and uses of a global union catalog: A mixed-methods study of WorldCat.org
abstract
This paper presents the first large‐scale investigation of the users and uses of WorldCat.org , the world's largest bibliographic database and global union catalog. Using a mixed‐methods approach involving focus group interviews with 120 participants, an online survey with 2,918 responses, and an analysis of transaction logs of approximately 15 million sessions from WorldCat.org , the study provides a new understanding of the context for global union catalog use. We find that WorldCat.org is accessed by a diverse population, with the three primary user groups being librarians, students, and academics. Use of the system is found to fall within three broad types of work‐task (professional, academic, and leisure), and we also present an emergent taxonomy of search tasks that encompass known‐item, unknown‐item, and institutional information searches. Our results support the notion that union catalogs are primarily used for known‐item searches, although the volume of traffic to WorldCat.org means that unknown‐item searches nonetheless represent an estimated 250,000 sessions per month. Search engine referrals account for almost half of all traffic, but although WorldCat.org effectively connects users referred from institutional library catalogs to other libraries holding a sought item, users arriving from a search engine are less likely to connect to a library.
Simon Wakeling, Paul D. Clough, Lynn Silipigni Connaway, Barbara Anne Sen, David Tomás 0001
J. Assoc. Inf. Sci. Technol.5
2013 A portable multilingual medical directory by automatic categorization of Wikipedia articles
abstract
Wikipedia has become one of the most important sources of information available all over the world. However, the categorization of Wikipedia articles is not standardized and the searches are mainly performed on keywords rather than concepts. In this paper we present an application that builds a hierarchical structure to organize all Wikipedia entries, so that medical articles can be reached from general to particular, using the well known Medical Subject Headings (MeSH) thesaurus. Moreover, the language links between articles will allow using the directory created in different languages. The final system can be packed and ported to mobile devices as a standalone offline application.
Fernando Ruiz-Rico, María-Consuelo Rubio-Sánchez, David Tomás 0001, José Luis Vicedo González
SIGIR3
2013 A multilingual and multiplatform application for medicinal plants prescription from medical symptoms
abstract
This paper presents an application for medicinal plants prescription based on text classification techniques. The system receives as an input a free text describing the symptoms of a user, and retrieves a ranked list of medicinal plants related to those symptoms. In addition, a set of links to Wikipedia are also provided, enriching the information about every medicinal plant presented to the user. In order to improve the accessibility to the application, the input can be written in six different languages, adapting the results accordingly. The application interface can be accessed from different devices and platforms.
Fernando Ruiz-Rico, David Tomás 0001, José Luis Vicedo González, María-Consuelo Rubio-Sánchez
SIGIR2
2013 Multi-source, multilingual information extraction and summarization - Edited by Thierry Poibeau, Horacio Saggion, Jakub Piskorski and Roman Yangarber
José Luis Vicedo González, David Tomás 0001
J. Assoc. Inf. Sci. Technol.2
2013 Minimally supervised question classification on fine-grained taxonomies
David Tomás 0001, José Luis Vicedo González
Knowl. Inf. Syst.1
2012 Question Answering and Multi-search Engines in Geo-Temporal Information Retrieval
Fernando Samuel Peregrino, David Tomás 0001, Fernando Llopis
CICLing (2)2
2012 Rada Mihalcea and Dragomir Radev: Graph-based natural language processing and information retrieval - Cambridge University Press, 2011, viii + 192 pp
David Tomás 0001
Mach. Transl.1
2011 Map-Based Filters for Fuzzy Entities in Geographical Information Retrieval
Fernando Samuel Peregrino, David Tomás 0001, Fernando Llopis
NLDB2
2011 Exploiting Unlabeled Data for Question Classification
David Tomás 0001, Claudio Giuliano
NLDB1
2011 The QALL-ME Framework: A specifiable-domain multilingual Question Answering architecture
Óscar Ferrández, Christian Spurk, Milen Kouylekov, Iustin Dornescu, Sergio Ferrández, Matteo Negri, Rubén Izquierdo, David Tomás 0001, Constantin Orasan, Günter Neumann, Bernardo Magnini, José Luis Vicedo González
J. Web Semant.8
2009 A Parallel Corpus Labeled Using Open and Restricted Domain Ontologies
Ester Boldrini, Sergio Ferrández, Rubén Izquierdo, David Tomás 0001, José Luis Vicedo González
CICLing4
2009 A semi-supervised approach to question classification
David Tomás 0001, Claudio Giuliano
ESANN1
2008 The QALL-ME Benchmark: a Multilingual Resource of Annotated Spoken Requests for Question Answering
Elena Cabrio, Milen Kouylekov, Bernardo Magnini, Matteo Negri, Laura Hasler, Constantin Orasan, David Tomás 0001, José Luis Vicedo González, Günter Neumann, Corinna Weber
LREC7