VLDB 2026 Research / reviewers in the wild / expert
Víctor Robles
dblp:r/VictorRobles · also Victor Robles, Víctor Robles Forcada
· DBLP profile ↗
30ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0003-3937-2269ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-authorSystems, architecture and hardware · 5Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ELADAIS: An Integrated Platform for High-Impact Clinical Data Extraction, Standardization and Advanced Analytics Using OMOP-CDMabstractClinical data generated in healthcare systems is increasingly recognized as a key resource for biomedical research, healthcare optimization, and population health monitoring. However, its full potential remains underexploited due to fragmentation, heterogeneity, and lack of interoperability between data sources. The ELADAIS project addresses this challenge by designing, developing, and deploying a scalable, modular, and interoperable technological platform for the extraction, transformation, storage, and advanced analysis of high-impact clinical data. Grounded in the OMOP Common Data Model (OMOP-CDM), ELADAIS integrates a microservice-based architecture, analytical environments, workflow orchestration, and federated data capabilities. The platform will be deployed at two major hospitals in Madrid, Spain, with the expectation of standardizing access to over 1 million patient records and more than 19 million clinical events. This paper presents the architectural principles and expectations of ELADAIS, highlighting its potential to accelerate reproducible and collaborative clinical research. Alejandro Rodríguez González, Víctor Robles, Juan José Cubillas Mercado, Juan Manuel Martínez Pérez, Jose Luis González Mendez, Ernestina Menasalvas Ruiz |
CBMS | 2 |
| 2025 | GPT for medical entity recognition in SpanishabstractAbstract In recent years, there has been a remarkable surge in the development of Natural Language Processing (NLP) models, particularly in the realm of Named Entity Recognition (NER). Models such as BERT have demonstrated exceptional performance, leveraging annotated corpora for accurate entity identification. However, the question arises: Can newer Large Language Models (LLMs) like GPT be utilized without the need for extensive annotation, thereby enabling direct entity extraction? In this study, we explore this issue, comparing the efficacy of fine-tuning techniques with prompting methods to elucidate the potential of GPT in the identification of medical entities within Spanish electronic health records (EHR). This study utilized a dataset of Spanish EHRs related to breast cancer and implemented both a traditional NER method using BERT, and a contemporary approach that combines few shot learning and integration of external knowledge, driven by LLMs using GPT, to structure the data. The analysis involved a comprehensive pipeline that included these methods. Key performance metrics, such as precision, recall, and F-score, were used to evaluate the effectiveness of each method. This comparative approach aimed to highlight the strengths and limitations of each method in the context of structuring Spanish EHRs efficiently and accurately.The comparative analysis undertaken in this article demonstrates that both the traditional BERT-based NER method and the few-shot LLM-driven approach, augmented with external knowledge, provide comparable levels of precision in metrics such as precision, recall, and F score when applied to Spanish EHR. Contrary to expectations, the LLM-driven approach, which necessitates minimal data annotation, performs on par with BERT’s capability to discern complex medical terminologies and contextual nuances within the EHRs. The results of this study highlight a notable advance in the field of NER for Spanish EHRs, with the few shot approach driven by LLM, enhanced by external knowledge, slightly edging out the traditional BERT-based method in overall effectiveness. GPT’s superiority in F-score and its minimal reliance on extensive data annotation underscore its potential in medical data processing. Alvaro Garcia-Barragán, Alberto González Calatayud, Oswaldo Solarte Pabón, Mariano Provencio, Ernestina Menasalvas Ruiz, Víctor Robles |
Multim. Tools Appl. | 6 |
| 2024 | Named Entity Recognition in Mammography Radiology Reports using a Multilingual Transfer Learning ApproachabstractThis study explores a multilingual transfer learning strategy for Named Entity Recognition (NER) in mammography radiology reports, aiming to improve breast cancer diagnosis. By utilizing a dataset from TecSalud, which includes mammograms and Electronic Health Records (EHRs) over ten years, this study seeks to address the linguistic barriers in medical documentation through advanced Natural Language Processing (NLP) models. Our approach involves meticulously labeling twenty-four distinct entities within the predominantly Spanish dataset, covering a range of diagnostic features and interpretive findings, highlighting the challenge of linguistic diversity in medical records and the potential of NLP to bridge this gap.The results demonstrate that fine-tuning on the last layer offers a balanced approach between simplicity and accuracy, avoiding overfitting and achieving state-of-art results. Esteban Ricardo Salazar Cabrera, Alejandro Santos-Díaz, Ernestina Menasalvas Ruiz, José G. Tamez-Peña, Víctor Robles |
CBMS | 5 |
| 2024 | Step-forward structuring disease phenotypic entities with LLMs for disease understandingabstractIn the rapidly evolving field of biomedical text mining, the extraction of phenotypic entities from unstructured texts remains a pivotal challenge. This paper introduces a novel method that leverage Large Language Models (LLMs) to extract phenotypical entities from freely available texts such as Wikipedia. Our approach goes beyond traditional Named Entity Recognition (NER) techniques by utilizing both local and cloud-based LLMs. We present a comprehensive comparison with state-of-the-art tools. Our study confirms the significant advantages of LLMs in identifying relevant phenotypic entities, thus enhancing the ability of researchers and clinicians to understand and respond to disease dynamics more effectively. Therefore, this work underscores the potential of next-generation LLMs to redefine the standards for the extraction of phenotypic entities in biomedical research. Alvaro Garcia-Barragán, Alberto González Calatayud, Lucía Prieto Santamaría, Víctor Robles, Ernestina Menasalvas Ruiz |
CBMS | 4 |
| 2023 | Structuring Breast Cancer Spanish Electronic Health Records Using Deep LearningabstractUsing Natural Language Processing (NLP) in the clinical domain has increased the possibility of automatically extracting information from oncology clinical narratives. Specifically, deep learning methods have been used to extract information in the cancer domain. However, most of the above proposals have concentrated only on extracting named entities from clinical narratives, but those proposals do not include a methodology for structuring the information after an information extraction step. In this paper, we propose an automatic pipeline based on deep learning for structuring breast cancer information from clinical narratives written in Spanish. The pipeline inputs a set of clinical documents written in narrative form and automatically generates a structured JSON file that contains the information for each patient. This pipeline integrates both clinical entity extraction and negation and uncertainty detection. Obtained results have shown that deep learning methods are feasible for structuring information in the breast cancer domain. Alvaro Garcia-Barragán, Oswaldo Solarte Pabón, Georgiy Nedostup, Mariano Provencio, Ernestina Menasalvas Ruiz, Víctor Robles |
CBMS | 6 |
| 2023 | Transformers for extracting breast cancer information from Spanish clinical narrativesabstractThe wide adoption of electronic health records (EHRs) offers immense potential as a source of support for clinical research. However, previous studies focused on extracting only a limited set of medical concepts to support information extraction in the cancer domain for the Spanish language. Building on the success of deep learning for processing natural language texts, this paper proposes a transformer-based approach to extract named entities from breast cancer clinical notes written in Spanish and compares several language models. To facilitate this approach, a schema for annotating clinical notes with breast cancer concepts is presented, and a corpus for breast cancer is developed. Results indicate that both BERT-based and RoBERTa-based language models demonstrate competitive performance in clinical Named Entity Recognition (NER). Specifically, BETO and multilingual BERT achieve F-scores of 93.71% and 94.63%, respectively. Additionally, RoBERTa Biomedical attains an F-score of 95.01%, while RoBERTa BNE achieves an F-score of 94.54%. The findings suggest that transformers can feasibly extract information in the clinical domain in the Spanish language, with the use of models trained on biomedical texts contributing to enhanced results. The proposed approach takes advantage of transfer learning techniques by fine-tuning language models to automatically represent text features and avoiding the time-consuming feature engineering process. Oswaldo Solarte Pabón, Orlando Montenegro, Alvaro Garcia-Barragán, Maria Torrente, Mariano Provencio, Ernestina Menasalvas Ruiz, Víctor Robles |
Artif. Intell. Medicine | 7 |
| 2022 | Deep learning to extract Breast Cancer diagnosis conceptsabstractThe wide adoption of electronic health records (EHRs) provides a potential source to support clinical research. The Bidirectional Encoder Representations from Transformers (BERT) has shown promising results in extracting information in the biomedical domain, including the cancer field. However, one of the challenges in the cancer domain is annotating resources to support information extraction. In this paper, we will show how models trained in a lung cancer corpus can be used to extract cancer concepts even in other cancer types. In particular, we will show the performance of BERT models on breast cancer data that was not used to train the models. Results are very promising as they show the possibility of applying deep learning-based models to predict cancer concepts in a different dataset to the one they were trained on, representing a considerable save of time and resources. Oswaldo Solarte Pabón, Maria Torrente, Alvaro Garcia-Barragán, Mariano Provencio, Ernestina Menasalvas Ruiz, Víctor Robles |
CBMS | 6 |
| 2014 | Semi-supervised projected model-based clustering
Luis Guerra, Concha Bielza, Víctor Robles, Pedro Larrañaga |
Data Min. Knowl. Discov. | 3 |
| 2014 | A methodology to compare Dimensionality Reduction algorithms in terms of loss of quality
Antonio Gracia Berná, Santiago González, Víctor Robles, Ernestina Menasalvas Ruiz |
Inf. Sci. | 3 |
| 2013 | Semi-supervised Projected Clustering for Classifying GABAergic Interneurons
Luis Guerra, Ruth Benavides-Piccione, Concha Bielza, Víctor Robles, Javier DeFelipe, Pedro Larrañaga |
AIME | 4 |
| 2012 | A comparison of clustering quality indices using outliers and noiseabstractQuality indices in clustering are used not only to assess the quality of the partitions but also to determine the number of clusters in the final result. When these indices are evaluated in a case study, real data conditions or different clustering a Luis Guerra, Víctor Robles, Concha Bielza, Pedro Larrañaga |
Intell. Data Anal. | 2 |
| 2011 | Regularized logistic regression without a penalty term: An application to cancer classification with microarray data
Concha Bielza, Víctor Robles, Pedro Larrañaga |
Expert Syst. Appl. | 2 |
| 2010 | CliDaPa: A new approach to combining clinical data with DNA microarraysabstractTraditionally, clinical data have been used as the only source of information to diagnose diseases. Nowadays, other types of information, such as various forms of omics data (e.g. DNA microarrays), are taken into account to improve diagnosis and even Santiago González, Luis Guerra, Víctor Robles, José M. Peña 0002, Fazel Famili |
Intell. Data Anal. | 3 |
| 2010 | A new initialization procedure for the distributed estimation of distribution algorithms
Santiago Muelas, José M. Peña 0002, Antonio LaTorre, Víctor Robles |
Soft Comput. | 4 |
| 2009 | An agent architecture for managing data resources in a grid environment
María S. Pérez 0001, Alberto Sánchez 0001, Jemal H. Abawajy, Víctor Robles, José M. Peña 0002 |
Future Gener. Comput. Syst. | 4 |
| 2009 | Feature selection for multi-label naive Bayes classification
Min-Ling Zhang, José M. Peña 0002, Víctor Robles |
Inf. Sci. | 3 |
| 2008 | Using multiple offspring sampling to guide genetic algorithms to solve permutation problemsabstractThe correct choice of an evolutionary algorithm, a genetic repre-sentation for the problem being solved (as well as their associated variation operators) and the appropriate values for the parameters of the algorithm is a hard task and it is often considered as an opti-mization problem itself. In this contribution, we propose a new theoretical formalism, called Multiple Offspring Sampling (MOS). This new technique combines different evolutionary approaches taking advantage of the benefits provided by each of them. MOS dynamically bal-ances the participation of different mechanisms to spawn the new offspring population, according to the benefits provided by each of them in previous generations. This approach evaluates multiple offspring generation methods (for example different coding strate-gies), and configures appropriate sampling sizes. This formalism has been applied to a well-known permutation problem, the traveling salesman problem (TSP). The results on sev-eral instances of this problem show that most of the combined tech-niques outperform the results obtained by single ones. Antonio LaTorre, José M. Peña 0002, Víctor Robles, Santiago Muelas |
GECCO | 3 |
| 2008 | Voronoi-initializated island models for solving real-coded deceptive problemsabstractDeceptive problems have always been considered difficult for Genetic Algorithms. To cope with this characteristic, the literature has proposed the use of Parallel Genetic Algorithms (PGAs), particularly multi-population island-based models. Although the existence of multiple populations encourages population diversity, these problems are still difficult to solve. This paper introduces a new initialization mechanism for each of the populations of the islands based on Voronoi cells. In order to analyze the results, a series of different experiments using several real-value deceptive problems and a set of representative parameters (migration ratio, migration frequency and connectivity) have been chosen. The results obtained suggest that the Voronoi initialization method improves considerably the performance obtained with a traditionally uniform random initialization. Santiago Muelas, José M. Peña 0002, Víctor Robles, Antonio LaTorre |
GECCO | 3 |
| 2007 | Design and implementation of a data mining grid-aware architecture
María S. Pérez 0001, Alberto Sánchez 0001, Víctor Robles, Pilar Herrero, José M. Peña 0002 |
Future Gener. Comput. Syst. | 3 |
| 2006 | Machine learning in bioinformaticsabstractThis article reviews machine learning methods for bioinformatics. It presents modelling methods, such as supervised classification, clustering and probabilistic graphical models for knowledge discovery, as well as deterministic and stochastic heuristics for optimization. Applications in genomics, proteomics, systems biology, evolution and text mining are also shown. Pedro Larrañaga, Borja Calvo, Roberto Santana 0001, Concha Bielza, Josu Galdiano, Iñaki Inza, José Antonio Lozano 0001, Rubén Armañanzas, Guzmán Santafé, Aritz Pérez Martínez, Víctor Robles |
Briefings Bioinform. | 11 |
| 2006 | MAPFS: A flexible multiagent parallel file system for clusters
María S. Pérez 0001, Jesús Carretero 0001, Félix García Carballeira, José M. Peña 0002, Víctor Robles |
Future Gener. Comput. Syst. | 5 |
| 2005 | Using Genetic Algorithms to Improve Accuracy of Economical Indexes Prediction
Óscar Cubo, Víctor Robles, Javier Segovia, Ernestina Menasalvas Ruiz |
IDA | 2 |
| 2005 | Extending the GA-EDA Hybrid Algorithm to Study Diversification and Intensification in GAs and EDAs
Víctor Robles, José M. Peña 0002, María S. Pérez 0001, Pilar Herrero, Óscar Cubo |
IDA | 1 |
| 2005 | A new formalism for dynamic reconfiguration of data servers in a cluster
María S. Pérez 0001, Alberto Sánchez 0001, José M. Peña 0002, Víctor Robles |
J. Parallel Distributed Comput. | 4 |
| 2004 | Cooperation model of a multiagent parallel file system for clustersabstractMAPFS is a parallel file system integrated with a multiagent system responsible for the information retrieval. One of the fields where the agents can be very useful is precisely in the development of information recovery systems. The usage of a multiagent system implies coordination among the agents that belong to such system. The main goal of the agent cooperation is the interaction among them for achieving a common objective in a distributed system. Thus, a communication framework must be provided. This paper shows the MAPFS cooperation model and its communication framework, emphasizing its relation with the whole system. María S. Pérez 0001, Alberto Sánchez 0001, Víctor Robles, José M. Peña 0002, Jemal H. Abawajy |
CCGRID | 3 |
| 2004 | Design and Evaluation of an Agent-Based Communication Model for a Parallel File System
María S. Pérez 0001, Alberto Sánchez 0001, Jemal H. Abawajy, Víctor Robles, José M. Peña 0002 |
ICCSA (2) | 4 |
| 2004 | GA-EDA: Hybrid Evolutionary Algorithm Using Genetic and Estimation of Distribution Algorithms
José M. Peña 0002, Víctor Robles, Pedro Larrañaga, Vanessa Herves, Francisco Rosales, María S. Pérez 0001 |
IEA/AIE | 2 |
| 2004 | Bayesian network multi-classifiers for protein secondary structure prediction
Víctor Robles, Pedro Larrañaga, José M. Peña 0002, Ernestina Menasalvas Ruiz, María S. Pérez 0001, Vanessa Herves, Anita Wasilewska |
Artif. Intell. Medicine | 1 |
| 2003 | Interval Estimation Naïve Bayes
Víctor Robles, Pedro Larrañaga, José M. Peña 0002, Ernestina Menasalvas Ruiz, María S. Pérez 0001 |
IDA | 1 |
| 2003 | Improvement of Naïve Bayes Collaborative Filtering Using Interval EstimationabstractRecommender systems emerged to help users choose among the large amount of options that ecommerce sites offer. Collaborative filtering is one of the most successful recommender techniques. Here we propose an approach to collaborative filtering based on the simple Bayesian classifier. We propose a method of increasing the efficiency of naive Bayes by applying a new semi naive Bayes approach based on interval estimation. To evaluate our algorithm we use a database of Microsoft anonymous Web data from the UCl repository. Our empirical results show that our proposed Interval based naive Bayes approach outperforms typical naive Bayes. Víctor Robles, Pedro Larrañaga, Ernestina Menasalvas Ruiz, María S. Pérez 0001, Vanessa Herves |
Web Intelligence | 1 |