Víctor Robles

dblp:r/VictorRobles · also Victor Robles, Víctor Robles Forcada · DBLP profile ↗
← Back
30ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0003-3937-2269ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-authorSystems, architecture and hardware · 5Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
YearPublicationVenuePosition
2025 ELADAIS: An Integrated Platform for High-Impact Clinical Data Extraction, Standardization and Advanced Analytics Using OMOP-CDM
abstract
Clinical data generated in healthcare systems is increasingly recognized as a key resource for biomedical research, healthcare optimization, and population health monitoring. However, its full potential remains underexploited due to fragmentation, heterogeneity, and lack of interoperability between data sources. The ELADAIS project addresses this challenge by designing, developing, and deploying a scalable, modular, and interoperable technological platform for the extraction, transformation, storage, and advanced analysis of high-impact clinical data. Grounded in the OMOP Common Data Model (OMOP-CDM), ELADAIS integrates a microservice-based architecture, analytical environments, workflow orchestration, and federated data capabilities. The platform will be deployed at two major hospitals in Madrid, Spain, with the expectation of standardizing access to over 1 million patient records and more than 19 million clinical events. This paper presents the architectural principles and expectations of ELADAIS, highlighting its potential to accelerate reproducible and collaborative clinical research.
Alejandro Rodríguez González, Víctor Robles, Juan José Cubillas Mercado, Juan Manuel Martínez Pérez, Jose Luis González Mendez, Ernestina Menasalvas Ruiz
CBMS2
2025 GPT for medical entity recognition in Spanish
abstract
Abstract In recent years, there has been a remarkable surge in the development of Natural Language Processing (NLP) models, particularly in the realm of Named Entity Recognition (NER). Models such as BERT have demonstrated exceptional performance, leveraging annotated corpora for accurate entity identification. However, the question arises: Can newer Large Language Models (LLMs) like GPT be utilized without the need for extensive annotation, thereby enabling direct entity extraction? In this study, we explore this issue, comparing the efficacy of fine-tuning techniques with prompting methods to elucidate the potential of GPT in the identification of medical entities within Spanish electronic health records (EHR). This study utilized a dataset of Spanish EHRs related to breast cancer and implemented both a traditional NER method using BERT, and a contemporary approach that combines few shot learning and integration of external knowledge, driven by LLMs using GPT, to structure the data. The analysis involved a comprehensive pipeline that included these methods. Key performance metrics, such as precision, recall, and F-score, were used to evaluate the effectiveness of each method. This comparative approach aimed to highlight the strengths and limitations of each method in the context of structuring Spanish EHRs efficiently and accurately.The comparative analysis undertaken in this article demonstrates that both the traditional BERT-based NER method and the few-shot LLM-driven approach, augmented with external knowledge, provide comparable levels of precision in metrics such as precision, recall, and F score when applied to Spanish EHR. Contrary to expectations, the LLM-driven approach, which necessitates minimal data annotation, performs on par with BERT’s capability to discern complex medical terminologies and contextual nuances within the EHRs. The results of this study highlight a notable advance in the field of NER for Spanish EHRs, with the few shot approach driven by LLM, enhanced by external knowledge, slightly edging out the traditional BERT-based method in overall effectiveness. GPT’s superiority in F-score and its minimal reliance on extensive data annotation underscore its potential in medical data processing.
Alvaro Garcia-Barragán, Alberto González Calatayud, Oswaldo Solarte Pabón, Mariano Provencio, Ernestina Menasalvas Ruiz, Víctor Robles
Multim. Tools Appl.6
2024 Named Entity Recognition in Mammography Radiology Reports using a Multilingual Transfer Learning Approach
abstract
This study explores a multilingual transfer learning strategy for Named Entity Recognition (NER) in mammography radiology reports, aiming to improve breast cancer diagnosis. By utilizing a dataset from TecSalud, which includes mammograms and Electronic Health Records (EHRs) over ten years, this study seeks to address the linguistic barriers in medical documentation through advanced Natural Language Processing (NLP) models. Our approach involves meticulously labeling twenty-four distinct entities within the predominantly Spanish dataset, covering a range of diagnostic features and interpretive findings, highlighting the challenge of linguistic diversity in medical records and the potential of NLP to bridge this gap.The results demonstrate that fine-tuning on the last layer offers a balanced approach between simplicity and accuracy, avoiding overfitting and achieving state-of-art results.
Esteban Ricardo Salazar Cabrera, Alejandro Santos-Díaz, Ernestina Menasalvas Ruiz, José G. Tamez-Peña, Víctor Robles
CBMS5
2024 Step-forward structuring disease phenotypic entities with LLMs for disease understanding
abstract
In the rapidly evolving field of biomedical text mining, the extraction of phenotypic entities from unstructured texts remains a pivotal challenge. This paper introduces a novel method that leverage Large Language Models (LLMs) to extract phenotypical entities from freely available texts such as Wikipedia. Our approach goes beyond traditional Named Entity Recognition (NER) techniques by utilizing both local and cloud-based LLMs. We present a comprehensive comparison with state-of-the-art tools. Our study confirms the significant advantages of LLMs in identifying relevant phenotypic entities, thus enhancing the ability of researchers and clinicians to understand and respond to disease dynamics more effectively. Therefore, this work underscores the potential of next-generation LLMs to redefine the standards for the extraction of phenotypic entities in biomedical research.
Alvaro Garcia-Barragán, Alberto González Calatayud, Lucía Prieto Santamaría, Víctor Robles, Ernestina Menasalvas Ruiz
CBMS4
2023 Structuring Breast Cancer Spanish Electronic Health Records Using Deep Learning
abstract
Using Natural Language Processing (NLP) in the clinical domain has increased the possibility of automatically extracting information from oncology clinical narratives. Specifically, deep learning methods have been used to extract information in the cancer domain. However, most of the above proposals have concentrated only on extracting named entities from clinical narratives, but those proposals do not include a methodology for structuring the information after an information extraction step. In this paper, we propose an automatic pipeline based on deep learning for structuring breast cancer information from clinical narratives written in Spanish. The pipeline inputs a set of clinical documents written in narrative form and automatically generates a structured JSON file that contains the information for each patient. This pipeline integrates both clinical entity extraction and negation and uncertainty detection. Obtained results have shown that deep learning methods are feasible for structuring information in the breast cancer domain.
Alvaro Garcia-Barragán, Oswaldo Solarte Pabón, Georgiy Nedostup, Mariano Provencio, Ernestina Menasalvas Ruiz, Víctor Robles
CBMS6
2023 Transformers for extracting breast cancer information from Spanish clinical narratives
abstract
The wide adoption of electronic health records (EHRs) offers immense potential as a source of support for clinical research. However, previous studies focused on extracting only a limited set of medical concepts to support information extraction in the cancer domain for the Spanish language. Building on the success of deep learning for processing natural language texts, this paper proposes a transformer-based approach to extract named entities from breast cancer clinical notes written in Spanish and compares several language models. To facilitate this approach, a schema for annotating clinical notes with breast cancer concepts is presented, and a corpus for breast cancer is developed. Results indicate that both BERT-based and RoBERTa-based language models demonstrate competitive performance in clinical Named Entity Recognition (NER). Specifically, BETO and multilingual BERT achieve F-scores of 93.71% and 94.63%, respectively. Additionally, RoBERTa Biomedical attains an F-score of 95.01%, while RoBERTa BNE achieves an F-score of 94.54%. The findings suggest that transformers can feasibly extract information in the clinical domain in the Spanish language, with the use of models trained on biomedical texts contributing to enhanced results. The proposed approach takes advantage of transfer learning techniques by fine-tuning language models to automatically represent text features and avoiding the time-consuming feature engineering process.
Oswaldo Solarte Pabón, Orlando Montenegro, Alvaro Garcia-Barragán, Maria Torrente, Mariano Provencio, Ernestina Menasalvas Ruiz, Víctor Robles
Artif. Intell. Medicine7
2022 Deep learning to extract Breast Cancer diagnosis concepts
abstract
The wide adoption of electronic health records (EHRs) provides a potential source to support clinical research. The Bidirectional Encoder Representations from Transformers (BERT) has shown promising results in extracting information in the biomedical domain, including the cancer field. However, one of the challenges in the cancer domain is annotating resources to support information extraction. In this paper, we will show how models trained in a lung cancer corpus can be used to extract cancer concepts even in other cancer types. In particular, we will show the performance of BERT models on breast cancer data that was not used to train the models. Results are very promising as they show the possibility of applying deep learning-based models to predict cancer concepts in a different dataset to the one they were trained on, representing a considerable save of time and resources.
Oswaldo Solarte Pabón, Maria Torrente, Alvaro Garcia-Barragán, Mariano Provencio, Ernestina Menasalvas Ruiz, Víctor Robles
CBMS6
2014 Semi-supervised projected model-based clustering
Luis Guerra, Concha Bielza, Víctor Robles, Pedro Larrañaga
Data Min. Knowl. Discov.3
2014 A methodology to compare Dimensionality Reduction algorithms in terms of loss of quality
Antonio Gracia Berná, Santiago González, Víctor Robles, Ernestina Menasalvas Ruiz
Inf. Sci.3
2013 Semi-supervised Projected Clustering for Classifying GABAergic Interneurons
Luis Guerra, Ruth Benavides-Piccione, Concha Bielza, Víctor Robles, Javier DeFelipe, Pedro Larrañaga
AIME4
2012 A comparison of clustering quality indices using outliers and noise
abstract
Quality indices in clustering are used not only to assess the quality of the partitions but also to determine the number of clusters in the final result. When these indices are evaluated in a case study, real data conditions or different clustering a
Luis Guerra, Víctor Robles, Concha Bielza, Pedro Larrañaga
Intell. Data Anal.2
2011 Regularized logistic regression without a penalty term: An application to cancer classification with microarray data
Concha Bielza, Víctor Robles, Pedro Larrañaga
Expert Syst. Appl.2
2010 CliDaPa: A new approach to combining clinical data with DNA microarrays
abstract
Traditionally, clinical data have been used as the only source of information to diagnose diseases. Nowadays, other types of information, such as various forms of omics data (e.g. DNA microarrays), are taken into account to improve diagnosis and even
Santiago González, Luis Guerra, Víctor Robles, José M. Peña 0002, Fazel Famili
Intell. Data Anal.3
2010 A new initialization procedure for the distributed estimation of distribution algorithms
Santiago Muelas, José M. Peña 0002, Antonio LaTorre, Víctor Robles
Soft Comput.4
2009 An agent architecture for managing data resources in a grid environment
María S. Pérez 0001, Alberto Sánchez 0001, Jemal H. Abawajy, Víctor Robles, José M. Peña 0002
Future Gener. Comput. Syst.4
2009 Feature selection for multi-label naive Bayes classification
Min-Ling Zhang, José M. Peña 0002, Víctor Robles
Inf. Sci.3
2008 Using multiple offspring sampling to guide genetic algorithms to solve permutation problems
abstract
The correct choice of an evolutionary algorithm, a genetic repre-sentation for the problem being solved (as well as their associated variation operators) and the appropriate values for the parameters of the algorithm is a hard task and it is often considered as an opti-mization problem itself. In this contribution, we propose a new theoretical formalism, called Multiple Offspring Sampling (MOS). This new technique combines different evolutionary approaches taking advantage of the benefits provided by each of them. MOS dynamically bal-ances the participation of different mechanisms to spawn the new offspring population, according to the benefits provided by each of them in previous generations. This approach evaluates multiple offspring generation methods (for example different coding strate-gies), and configures appropriate sampling sizes. This formalism has been applied to a well-known permutation problem, the traveling salesman problem (TSP). The results on sev-eral instances of this problem show that most of the combined tech-niques outperform the results obtained by single ones.
Antonio LaTorre, José M. Peña 0002, Víctor Robles, Santiago Muelas
GECCO3
2008 Voronoi-initializated island models for solving real-coded deceptive problems
abstract
Deceptive problems have always been considered difficult for Genetic Algorithms. To cope with this characteristic, the literature has proposed the use of Parallel Genetic Algorithms (PGAs), particularly multi-population island-based models. Although the existence of multiple populations encourages population diversity, these problems are still difficult to solve. This paper introduces a new initialization mechanism for each of the populations of the islands based on Voronoi cells. In order to analyze the results, a series of different experiments using several real-value deceptive problems and a set of representative parameters (migration ratio, migration frequency and connectivity) have been chosen. The results obtained suggest that the Voronoi initialization method improves considerably the performance obtained with a traditionally uniform random initialization.
Santiago Muelas, José M. Peña 0002, Víctor Robles, Antonio LaTorre
GECCO3
2007 Design and implementation of a data mining grid-aware architecture
María S. Pérez 0001, Alberto Sánchez 0001, Víctor Robles, Pilar Herrero, José M. Peña 0002
Future Gener. Comput. Syst.3
2006 Machine learning in bioinformatics
abstract
This article reviews machine learning methods for bioinformatics. It presents modelling methods, such as supervised classification, clustering and probabilistic graphical models for knowledge discovery, as well as deterministic and stochastic heuristics for optimization. Applications in genomics, proteomics, systems biology, evolution and text mining are also shown.
Pedro Larrañaga, Borja Calvo, Roberto Santana 0001, Concha Bielza, Josu Galdiano, Iñaki Inza, José Antonio Lozano 0001, Rubén Armañanzas, Guzmán Santafé, Aritz Pérez Martínez, Víctor Robles
Briefings Bioinform.11
2006 MAPFS: A flexible multiagent parallel file system for clusters
María S. Pérez 0001, Jesús Carretero 0001, Félix García Carballeira, José M. Peña 0002, Víctor Robles
Future Gener. Comput. Syst.5
2005 Using Genetic Algorithms to Improve Accuracy of Economical Indexes Prediction
Óscar Cubo, Víctor Robles, Javier Segovia, Ernestina Menasalvas Ruiz
IDA2
2005 Extending the GA-EDA Hybrid Algorithm to Study Diversification and Intensification in GAs and EDAs
Víctor Robles, José M. Peña 0002, María S. Pérez 0001, Pilar Herrero, Óscar Cubo
IDA1
2005 A new formalism for dynamic reconfiguration of data servers in a cluster
María S. Pérez 0001, Alberto Sánchez 0001, José M. Peña 0002, Víctor Robles
J. Parallel Distributed Comput.4
2004 Cooperation model of a multiagent parallel file system for clusters
abstract
MAPFS is a parallel file system integrated with a multiagent system responsible for the information retrieval. One of the fields where the agents can be very useful is precisely in the development of information recovery systems. The usage of a multiagent system implies coordination among the agents that belong to such system. The main goal of the agent cooperation is the interaction among them for achieving a common objective in a distributed system. Thus, a communication framework must be provided. This paper shows the MAPFS cooperation model and its communication framework, emphasizing its relation with the whole system.
María S. Pérez 0001, Alberto Sánchez 0001, Víctor Robles, José M. Peña 0002, Jemal H. Abawajy
CCGRID3
2004 Design and Evaluation of an Agent-Based Communication Model for a Parallel File System
María S. Pérez 0001, Alberto Sánchez 0001, Jemal H. Abawajy, Víctor Robles, José M. Peña 0002
ICCSA (2)4
2004 GA-EDA: Hybrid Evolutionary Algorithm Using Genetic and Estimation of Distribution Algorithms
José M. Peña 0002, Víctor Robles, Pedro Larrañaga, Vanessa Herves, Francisco Rosales, María S. Pérez 0001
IEA/AIE2
2004 Bayesian network multi-classifiers for protein secondary structure prediction
Víctor Robles, Pedro Larrañaga, José M. Peña 0002, Ernestina Menasalvas Ruiz, María S. Pérez 0001, Vanessa Herves, Anita Wasilewska
Artif. Intell. Medicine1
2003 Interval Estimation Naïve Bayes
Víctor Robles, Pedro Larrañaga, José M. Peña 0002, Ernestina Menasalvas Ruiz, María S. Pérez 0001
IDA1
2003 Improvement of Naïve Bayes Collaborative Filtering Using Interval Estimation
abstract
Recommender systems emerged to help users choose among the large amount of options that ecommerce sites offer. Collaborative filtering is one of the most successful recommender techniques. Here we propose an approach to collaborative filtering based on the simple Bayesian classifier. We propose a method of increasing the efficiency of naive Bayes by applying a new semi naive Bayes approach based on interval estimation. To evaluate our algorithm we use a database of Microsoft anonymous Web data from the UCl repository. Our empirical results show that our proposed Interval based naive Bayes approach outperforms typical naive Bayes.
Víctor Robles, Pedro Larrañaga, Ernestina Menasalvas Ruiz, María S. Pérez 0001, Vanessa Herves
Web Intelligence1