VLDB 2026 Research / reviewers in the wild / expert
José Alberto Benítez
dblp:203/3022 · also José Alberto Benítez-Andrades
· DBLP profile ↗
19ranked-venue papers
4as first author
15since 2021 · last 2025
0000-0002-4450-349XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fine-Tuning Transformer Models for Structuring Spanish Psychiatric Clinical NotesabstractThe unstructured nature of psychiatric clinical notes poses a significant challenge for automated information extraction and data structuring. In this study, we explore the use of transformer-based language models to perform Named Entity Recognition (NER) on de-identified Spanish electronic health records (EHRs) provided by the Psychiatry Service of Complejo Asistencial Universitario de León (CAULE). A manually annotated gold standard, consisting of 200 clinical notes, was developed by domain experts to evaluate the performance of five models: BETO (cased and uncased), ALBETO, ClinicalBERT, and Bio_ClinicalBERT. Each model was fine-tuned and assessed using a strict exact matching criterion across six clinically relevant label types. Results demonstrate that ClinicalBERT, despite being pre-trained on English medical corpora, achieved the highest macro-average F1-score on the test set (80 %). However, BETO-cased outperformed ClinicalBERT in four out of six label types, being better in categories with higher syntactic variability. Lower-performing models, such as ALBETO and Bio_ClinicalBERT, struggled to generalize to Spanish psychiatric language, likely due to domain and language mismatches. This work highlights the effectiveness of transformer-based architectures for structuring psychiatric narratives in Spanish and provides a robust foundation for future clinical NLP applications in non-English contexts. Sergio Rubio-Martín, Arturo Crespo-Álvaro, María Teresa García-Ordás, Antonio Serrano-García, Clara Margarita Franch-Pato, José Alberto Benítez |
CBMS | 6 |
| 2025 | AI-Driven Survival Prediction in Pancreatic CancerabstractPancreatic cancer remains one of the most aggressive malignancies, with limited survival rates and significant variability in patient outcomes. This study evaluates the performance of three machine learning models (Random Forest, Decision Tree, and XGBoost) in predicting patient survival at 3, 12, and 18 months, using data from the Complejo Asistencial Universitario de León (CAULE) Radiology Department. To systematically analyze the impact of different features on survival prediction, the dataset was structured into seven variable groups (G1G7), incorporating demographic, clinical, and treatment-related information. To address the inherent class imbalance in survival prediction, an Autoencoder-based synthetic data generation approach was applied, ensuring a balanced distribution of survival and non-survival cases across all timeframes. Hyperparameter tuning was performed, and experimental results indicate that Random Forest and XGBoost achieved comparable performance, both obtaining an accuracy above 81 % at 3 months, 83 % at 12 months, and 88 % at 18 months when trained on Group G7. To enhance model interpretability, SHapley Additive exPlanations (SHAP) was applied to the best-performing model, identifying key factors influencing survival. Sergio Rubio-Martín, María Teresa García-Ordás, David Corral Fontecha, Laura López-González, Gonzalo Alonso-Oláiz, Arturo Crespo-Álvaro, José Alberto Benítez |
CBMS | 7 |
| 2024 | Machine Learning in Predicting the Success of Spine Surgery: A Multivariable StudyabstractThis study explores the application of Artificial Intelligence (AI) in spine surgery, with a focus on enhancing precision and accuracy in outcome prediction. Leveraging machine learning (ML) models – including GaussianNB, ComplementNB, KNN, and Decision Trees – we analyze a rich dataset derived from 244 spine surgery patients. This dataset comprises 24 diverse variables, capturing elements such as pre-surgical conditions, socioeconomic status, psychometric evaluations, and various analytical metrics. Notably, one critical variable is the surgery’s success, serving as the primary outcome for prediction.The data was meticulously categorized into seven distinct groups, reflecting various aspects of the surgical process and patient backgrounds. This structured approach enabled a targeted and nuanced analysis, deepening our understanding of the key factors instrumental in predicting surgical outcomes. We employed a stratified split methodology for our dataset, dedicating 80% to training and 20% to testing. This was supplemented by 5-fold cross-validation and an extensive grid search optimization for refining the KNN and Decision Trees models.Our results underscore the profound capability of AI in predicting the outcomes of spine surgeries. The KNN model showed remarkable proficiency, particularly in analyzing groups defined by pre-surgical and analytical variables, demonstrating its prowess in handling complex medical datasets. This study not only evidences the effectiveness of specific ML models in medical prognostics but also highlights AI’s transformative potential in healthcare. It underlines the critical role of AI in advancing medical diagnostics and decision-making for surgeries that entail multifaceted data analysis. These insights pave the way for future research into the broader application of AI in medicine, promising more personalized and effective treatment strategies and effective treatment approaches. José Alberto Benítez, Nicolás Ordás-Reyes, Antonio Serrano-García, Marta Esteban Blanco, Jesús Betegón Nicolás, José Viloria Gutiérrez, José Ángel Hernández Encinas, Ana Lozano Muñoz, Alicia Merayo-Corcoba, Camino Prada-García |
CBMS | 1 |
| 2024 | Analysis and Detection of Melanoma through Collective Intelligence with AIabstractCancer is a widespread global health problem, claiming millions of lives each year, and skin cancer represents a significant threat as it is one of the most common types. Early tumor detection via medical imaging is critical for effective treatment. Leveraging artificial intelligence, particularly novel models like Transformers, presents promising avenues for improved diagnosis. This paper explores the efficacy of a Collective Intelligence approach using AI in classifying cancerous and non-cancerous tumors, aiming to reduce classification errors and support clinical decision-making. We created five different configurations using various datasets to compare the results. The results show solid performance for the CI in the evaluated tasks, reaching up to 75.89% accuracy. The lack of images in certain classes significantly contributes to overfitting. It is suggested to explore data expansion strategies and improve consistency in image capture for future work. Enrique Fernández-Morales, Carlos Luis Sánchez-Bocanegra, Rafael Pastor 0001, José-Juan Pereyra-Rodriguez, Juan Mario Haut, José Alberto Benítez |
CBMS | 6 |
| 2024 | Determining the severity of Parkinson's disease in patients using a multi task neural networkabstractAbstract Parkinson’s disease is easy to diagnose when it is advanced, but it is very difficult to diagnose in its early stages. Early diagnosis is essential to be able to treat the symptoms. It impacts on daily activities and reduces the quality of life of both the patients and their families and it is also the second most prevalent neurodegenerative disorder after Alzheimer in people over the age of 60. Most current studies on the prediction of Parkinson’s severity are carried out in advanced stages of the disease. In this work, the study analyzes a set of variables that can be easily extracted from voice analysis, making it a very non-intrusive technique. In this paper, a method based on different deep learning techniques is proposed with two purposes. On the one hand, to find out if a person has severe or non-severe Parkinson’s disease, and on the other hand, to determine by means of regression techniques the degree of evolution of the disease in a given patient. The UPDRS (Unified Parkinson’s Disease Rating Scale) has been used by taking into account both the motor and total labels, and the best results have been obtained using a mixed multi-layer perceptron (MLP) that classifies and regresses at the same time and the most important features of the data obtained are taken as input, using an autoencoder. A success rate of 99.15% has been achieved in the problem of predicting whether a person suffers from severe Parkinson’s disease or non-severe Parkinson’s disease. In the degree of disease involvement prediction problem case, a MSE (Mean Squared Error) of 0.15 has been obtained. Using a full deep learning pipeline for data preprocessing and classification has proven to be very promising in the field Parkinson’s outperforming the state-of-the-art proposals. María Teresa García-Ordás, José Alberto Benítez, Jose Aveleira-Mata, José-Manuel Alija-Pérez, Carmen Benavides |
Multim. Tools Appl. | 2 |
| 2024 | Conditional Weighted Linear Fitting for 2D-LiDAR-Mapping of Indoor SLAMabstractThe ability to map an unknown environment is a fundamental milestone for autonomous robotic vehicles. Solutions in this field must combine efficiency, accuracy, and precision. We propose a novel methodology for map feature extraction in indoor environments. The mathematical model and its implementation are designed to operate with 2-D light detection and ranging (LiDAR) measurements. Map parameters and associated uncertainty levels are determined through bivariate linear regression. The final step is experimental validation, using a low-cost commercial LiDAR sensor. The main contributions of the proposed methodology lie in the domains of computational efficiency and uncertainty. In addition, the results prove the ability of our methodology to handle large volumes of data while maintaining restrained growth in computational time. This outcome suggests considerable potential for real-time applications with limited hardware resources. A second methodology, extracted from the current state of the art, is used in parallel for benchmarking purposes. Natalia Prieto-Fernández, Sergio Fernández-Blanco, Álvaro Fernández-Blanco, José Alberto Benítez, Francisco Carro-De-Lorenzo, Carmen Benavides |
IEEE Trans. Ind. Informatics | 4 |
| 2023 | Detection of cerebral ischaemia using transfer learning techniquesabstractCerebrovascular accident (CVA) or stroke is one of the main causes of mortality and morbidity today, causing permanent disabilities. Its early detection helps reduce its effects and its mortality: time is brain. Currently, non-contrast computed tomography (NCCT) continues to be the first-line diagnostic method in stroke emergencies because it is a fast, available, and cost-effective technique that makes it possible to rule out haemorrhage and focus attention on the ischemic origin, that is, due to obstruction to arterial flow. NCCT are quantified using a scoring system called ASPECTS (Alberta Stroke Program Early Computed Tomography Score) according to the affected brain structures. This paper aims to detect in an initial phase those CTs of patients with stroke symptoms that present early alterations in CT density using a binary classifier of CTs without and with stroke, to alert the doctor of their existence. For this, several well-known neural network architectures are implemented in the ImageNet challenges (VGG, NasNet, ResNet and DenseNet), with 3D images, covering the entire brain volume. The training results of these networks are exposed, in which different parameters are tested to obtain maximum performance, which is achieved with a DenseNet3D network that achieves an accuracy of 98% in the training set and 95% in the test set. Cristina Antón-Munárriz, Rafael Pastor 0001, Juan Mario Haut, Antonio Robles-Gómez, Mercedes Eugenia Paoletti, José Alberto Benítez |
CBMS | 6 |
| 2023 | Early Detection of Autism Spectrum Disorder through AI-Powered Analysis of Social Media TextsabstractDetecting individuals with autism spectrum disorder (ASD) remains a challenge due to the resources and specialized professionals needed for accurate diagnosis, particularly for children where time is a critical factor. Early diagnosis of ASD is crucial for improving the quality of life for affected individuals, as it allows for timely intervention and support. In this study, we underscore the importance of artificial intelligence (AI) in developing innovative diagnostic methods, with the primary objective of creating AI models that assist in identifying users who may have ASD. Although several studies have utilized traditional machine learning (ML) and deep learning (DL) techniques to detect various illnesses, few have focused on detecting ASD using text as input. We employ natural language processing (NLP) techniques combined with AI models, specifically decision trees, extreme gradient boosting (XGB), k-nearest neighbors algorithm (KNN) as ML models, and bidirectional encoder representations from transformers (BERT) as DL models. The core idea involves extracting tweets from Twitter users through the platform's API, classifying the texts as written by individuals who claim to have ASD (ASD users) or by those without ASD (non-ASD users). We generated a dataset of 404,627 tweets and used a subset of 90,000 tweets, comprising 45,000 from each classification group, for training and testing the models. The results demonstrate a predictive model with an accuracy of over 84% when classifying texts potentially originated from ASD users. This research paves the way for using DL models to enhance the accuracy of detecting and diagnosing ASD in individuals effectively, emphasizing the critical role of AI in advancing early diagnostic methods for better patient outcomes. Sergio Rubio-Martín, María Teresa García-Ordás, Martín Bayón-Gutiérrez, Natalia Prieto-Fernández, José Alberto Benítez |
CBMS | 5 |
| 2023 | A generalized decision tree ensemble based on the NeuralNetworks architecture: Distributed Gradient Boosting Forest (DGBF)
Ángel Delgado-Panadero, José Alberto Benítez, María Teresa García-Ordás |
Appl. Intell. | 2 |
| 2023 | Multispecies bird sound recognition using a fully convolutional neural network
María Teresa García-Ordás, Sergio Rubio-Martín, José Alberto Benítez, Héctor Alaiz-Moretón, Isaías García 0001 |
Appl. Intell. | 3 |
| 2023 | Clustering Techniques Selection for a Hybrid Regression Model: A Case Study Based on a Solar Thermal SystemabstractThis work addresses the performance comparison between four clustering techniques with the objective of achieving strong hybrid models in supervised learning tasks. A real dataset from a bio-climatic house named Sotavento placed on experimental wind farm and located in Xermade (Lugo) in Galicia (Spain) has been collected. Authors have chosen the thermal solar generation system in order to study how works applying several cluster methods followed by a regression technique to predict the output temperature of the system. With the objective of defining the quality of each clustering method two possible solutions have been implemented. The first one is based on three unsupervised learning metrics (Silhouette, Calinski-Harabasz and Davies-Bouldin) while the second one, employs the most common error measurements for a regression algorithm such as Multi Layer Perceptron. María Teresa García-Ordás, Héctor Alaiz-Moretón, José Luís Casteleiro-Roca, Esteban Jove, José Alberto Benítez, Isaías García 0001, Héctor Quintián, José Luís Calvo-Rolle |
Cybern. Syst. | 5 |
| 2023 | Heart disease risk prediction using deep learning techniques with feature augmentationabstractAbstract Cardiovascular diseases state as one of the greatest risks of death for the general population. Late detection in heart diseases highly conditions the chances of survival for patients. Age, sex, cholesterol level, sugar level, heart rate, among other factors, are known to have an influence on life-threatening heart problems, but, due to the high amount of variables, it is often difficult for an expert to evaluate each patient taking this information into account. In this manuscript, the authors propose using deep learning methods, combined with feature augmentation techniques for evaluating whether patients are at risk of suffering cardiovascular disease. The results of the proposed methods outperform other state of the art methods by 4.4%, leading to a precision of a 90%, which presents a significant improvement, even more so when it comes to an affliction that affects a large population. María Teresa García-Ordás, Martín Bayón-Gutiérrez, Carmen Benavides, Jose Aveleira-Mata, José Alberto Benítez |
Multim. Tools Appl. | 5 |
| 2022 | Implementing local-explainability in Gradient Boosting Trees: Feature Contribution
Ángel Delgado-Panadero, Beatriz Hernández-Lorca, María Teresa García-Ordás, José Alberto Benítez |
Inf. Sci. | 4 |
| 2021 | BERT Model-Based Approach For Detecting Categories of Tweets in the Field of Eating Disorders (ED)abstractEating disorders (ED) are among the most widespread mental illnesses in our society today. This research work presents the study of deep learning models applied to the domain of eating disorders. For this purpose, a collection of messages from the social network Twitter was compiled using web scraping techniques. After collecting a total amount of 1,085,957 tweets, a subset of 2,000 tweets was manually classified. This classification made it possible to differentiate tweets written by people who suffer or have suffered from an ED from those written by people who have not suffered from an ED. After this, 6 predictive models based on Bidirectional Encoder Representations from Transformers (BERT) were created and a comparison was made by evaluating which model scored the best. The best scoring model was RoBERTa using the pre-trained roberta-base model with an accuracy of 87.5%. José Alberto Benítez, José-Manuel Alija-Pérez, Isaías García 0001, Carmen Benavides, Héctor Alaiz-Moretón, Rafael Pastor 0001, María Teresa García-Ordás |
CBMS | 1 |
| 2021 | Enriched multi-agent middleware for building rule-based distributed security solutions for IoT environments
Francisco J. Aguayo-Canela, Héctor Alaiz-Moretón, María Teresa García-Ordás, José Alberto Benítez, Carmen Benavides, Isaías García 0001 |
J. Supercomput. | 4 |
| 2020 | Analyzing IoT-Based Botnet Malware Activity with Distributed Low Interaction Honeypots
Sergio Vidal-González, Isaías García 0001, Héctor Alaiz-Moretón, Carmen Benavides, José Alberto Benítez, María Teresa García-Ordás, Paulo Novais |
WorldCIST (2) | 5 |
| 2020 | Social network analysis for personalized characterization and risk assessment of alcohol use disorders in adolescents using semantic technologies
José Alberto Benítez, Isaías García 0001, Carmen Benavides, Héctor Alaiz-Moretón, Alejandro Rodríguez González |
Future Gener. Comput. Syst. | 1 |
| 2020 | An ontology-based multi-domain model in social network analysis: Experimental validation and case study
José Alberto Benítez, Isaías García 0001, Carmen Benavides, Héctor Alaiz-Moretón, José Emilio Labra Gayo |
Inf. Sci. | 1 |
| 2019 | A FIPA-Compliant Framework for Integrating Rule Engines into Software Agents for Supporting Communication and Collaboration in a Multiagent Platform
Francisco J. Aguayo-Canela, Héctor Alaiz-Moretón, Isaías García 0001, Carmen Benavides, José Alberto Benítez, Paulo Novais |
WorldCIST (2) | 5 |