VLDB 2026 Research / reviewers in the wild / expert
Sergio Rubio-Martín
dblp:328/9905
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0003-0475-9949ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Radiomic Feature Robustness Through Voxel Spacing-Aware Extraction in Anisotropic CT DataabstractThis study aimed to evaluate whether voxel spacing-aware radiomic feature extraction improves reproducibility, variability, and discriminative performance compared to conventional preprocessing methods in anisotropic CT data. A curated cohort of 685 pulmonary nodules from the LIDC-IDRI dataset was analyzed. Three preprocessing strategies-no resampling, isotropic resampling, and voxel spacing-aware extraction-and one postprocessing approach, voxel spacing weighting, were systematically compared. A modified version of PyRadiomics was developed to compute texture and shape features directly from native images while incorporating physical voxel dimensions without interpolation. Among the 94 extracted features, spacingaware preprocessing improved reproducibility in 58 features, reduced variability in 48, and enhanced univariate discrimination in 37. An ensemble feature selection combining six statistical and machine learning methods identified between 18 and 20 robust features per method. Logistic Regression models trained with spacing-aware features achieved the highest composite performance score (1.498), balancing discrimination and generalizability. SHAP interpretability analysis confirmed the clinical relevance of selected geometric and texture features. Overall, voxel spacing-aware preprocessing preserved native spatial structure and mitigated the effects of voxel geometry and interpolation artifacts, yielding stable and clinically interpretable radiomic features. These findings support the adoption of spacing-aware pipelines in heterogeneous and multi-center CT radiomics studies to enhance feature robustness and reproducibility. David Corral Fontecha, Pablo Menéndez Fernández-Miranda, Sergio Rubio-Martín, Alicia Merayo-Corcoba, Laura López-González, Lara Lloret Iglesias, José Antonio Vega |
CBMS | 3 |
| 2025 | Fine-Tuning Transformer Models for Structuring Spanish Psychiatric Clinical NotesabstractThe unstructured nature of psychiatric clinical notes poses a significant challenge for automated information extraction and data structuring. In this study, we explore the use of transformer-based language models to perform Named Entity Recognition (NER) on de-identified Spanish electronic health records (EHRs) provided by the Psychiatry Service of Complejo Asistencial Universitario de León (CAULE). A manually annotated gold standard, consisting of 200 clinical notes, was developed by domain experts to evaluate the performance of five models: BETO (cased and uncased), ALBETO, ClinicalBERT, and Bio_ClinicalBERT. Each model was fine-tuned and assessed using a strict exact matching criterion across six clinically relevant label types. Results demonstrate that ClinicalBERT, despite being pre-trained on English medical corpora, achieved the highest macro-average F1-score on the test set (80 %). However, BETO-cased outperformed ClinicalBERT in four out of six label types, being better in categories with higher syntactic variability. Lower-performing models, such as ALBETO and Bio_ClinicalBERT, struggled to generalize to Spanish psychiatric language, likely due to domain and language mismatches. This work highlights the effectiveness of transformer-based architectures for structuring psychiatric narratives in Spanish and provides a robust foundation for future clinical NLP applications in non-English contexts. Sergio Rubio-Martín, Arturo Crespo-Álvaro, María Teresa García-Ordás, Antonio Serrano-García, Clara Margarita Franch-Pato, José Alberto Benítez |
CBMS | 1 |
| 2025 | AI-Driven Survival Prediction in Pancreatic CancerabstractPancreatic cancer remains one of the most aggressive malignancies, with limited survival rates and significant variability in patient outcomes. This study evaluates the performance of three machine learning models (Random Forest, Decision Tree, and XGBoost) in predicting patient survival at 3, 12, and 18 months, using data from the Complejo Asistencial Universitario de León (CAULE) Radiology Department. To systematically analyze the impact of different features on survival prediction, the dataset was structured into seven variable groups (G1G7), incorporating demographic, clinical, and treatment-related information. To address the inherent class imbalance in survival prediction, an Autoencoder-based synthetic data generation approach was applied, ensuring a balanced distribution of survival and non-survival cases across all timeframes. Hyperparameter tuning was performed, and experimental results indicate that Random Forest and XGBoost achieved comparable performance, both obtaining an accuracy above 81 % at 3 months, 83 % at 12 months, and 88 % at 18 months when trained on Group G7. To enhance model interpretability, SHapley Additive exPlanations (SHAP) was applied to the best-performing model, identifying key factors influencing survival. Sergio Rubio-Martín, María Teresa García-Ordás, David Corral Fontecha, Laura López-González, Gonzalo Alonso-Oláiz, Arturo Crespo-Álvaro, José Alberto Benítez |
CBMS | 1 |
| 2023 | Early Detection of Autism Spectrum Disorder through AI-Powered Analysis of Social Media TextsabstractDetecting individuals with autism spectrum disorder (ASD) remains a challenge due to the resources and specialized professionals needed for accurate diagnosis, particularly for children where time is a critical factor. Early diagnosis of ASD is crucial for improving the quality of life for affected individuals, as it allows for timely intervention and support. In this study, we underscore the importance of artificial intelligence (AI) in developing innovative diagnostic methods, with the primary objective of creating AI models that assist in identifying users who may have ASD. Although several studies have utilized traditional machine learning (ML) and deep learning (DL) techniques to detect various illnesses, few have focused on detecting ASD using text as input. We employ natural language processing (NLP) techniques combined with AI models, specifically decision trees, extreme gradient boosting (XGB), k-nearest neighbors algorithm (KNN) as ML models, and bidirectional encoder representations from transformers (BERT) as DL models. The core idea involves extracting tweets from Twitter users through the platform's API, classifying the texts as written by individuals who claim to have ASD (ASD users) or by those without ASD (non-ASD users). We generated a dataset of 404,627 tweets and used a subset of 90,000 tweets, comprising 45,000 from each classification group, for training and testing the models. The results demonstrate a predictive model with an accuracy of over 84% when classifying texts potentially originated from ASD users. This research paves the way for using DL models to enhance the accuracy of detecting and diagnosing ASD in individuals effectively, emphasizing the critical role of AI in advancing early diagnostic methods for better patient outcomes. Sergio Rubio-Martín, María Teresa García-Ordás, Martín Bayón-Gutiérrez, Natalia Prieto-Fernández, José Alberto Benítez |
CBMS | 1 |
| 2023 | Multispecies bird sound recognition using a fully convolutional neural network
María Teresa García-Ordás, Sergio Rubio-Martín, José Alberto Benítez, Héctor Alaiz-Moretón, Isaías García 0001 |
Appl. Intell. | 2 |