Mariano Provencio

dblp:208/6893 · also Mariano Provencio Pulla · DBLP profile ↗
← Back
20ranked-venue papers
0as first author
14since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 7 since 2021Human-computer interaction and ubiquitous computing · 9 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Theory of computation · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Large language model vs. traditional machine learning: Evaluating predictive models for early detection of tumor relapse
abstract
In this study, we evaluate the effectiveness of foundational artificial intelligence (AI) models, particularly large language models (LLMs), in comparison to traditional machine learning methods for predicting tumor relapse in patients with non-small-cell lung cancer (NSCLC). With a high recurrence risk in NSCLC, early and accurate prediction is essential for improving patient outcomes and guiding treatment decisions. Our analysis utilizes a dataset of 1,348 patients, examining the performance of traditional machine learning models such as Random Forest, alongside cutting-edge LLMs like Mistral-7B, LLaMA-7B, Falcon-7B, and GPT-based models. While the Random Forest model slightly outperforms Mistral-7B in precision–recall for relapse prediction, the comparable results suggest that both approaches offer valuable insights for early relapse detection. This study underscores the potential of integrating classical machine learning with foundational AI models to enhance predictive accuracy in cancer prognosis, providing pathways for more personalized medical interventions. • Study contrasts AI and traditional methods for NSCLC relapse prediction. • Early, accurate predictions are crucial for NSCLC patient care. • Analysis includes 1,348 NSCLC patient data points. • Random Forest edges out Mistral-7B in precision–recall. • Both models show potential for early detection of lung cancer recurrence.
Mohan Timilsina, Samuele Buosi, Maria Torrente, Mariano Provencio, Manuel Cobo, Delvys Rodriguez Abreu, Rafael Castro, Enric Carcereny, Edward Curry, Vít Novácek
Expert Syst. Appl.4
2025 GPT for medical entity recognition in Spanish
abstract
Abstract In recent years, there has been a remarkable surge in the development of Natural Language Processing (NLP) models, particularly in the realm of Named Entity Recognition (NER). Models such as BERT have demonstrated exceptional performance, leveraging annotated corpora for accurate entity identification. However, the question arises: Can newer Large Language Models (LLMs) like GPT be utilized without the need for extensive annotation, thereby enabling direct entity extraction? In this study, we explore this issue, comparing the efficacy of fine-tuning techniques with prompting methods to elucidate the potential of GPT in the identification of medical entities within Spanish electronic health records (EHR). This study utilized a dataset of Spanish EHRs related to breast cancer and implemented both a traditional NER method using BERT, and a contemporary approach that combines few shot learning and integration of external knowledge, driven by LLMs using GPT, to structure the data. The analysis involved a comprehensive pipeline that included these methods. Key performance metrics, such as precision, recall, and F-score, were used to evaluate the effectiveness of each method. This comparative approach aimed to highlight the strengths and limitations of each method in the context of structuring Spanish EHRs efficiently and accurately.The comparative analysis undertaken in this article demonstrates that both the traditional BERT-based NER method and the few-shot LLM-driven approach, augmented with external knowledge, provide comparable levels of precision in metrics such as precision, recall, and F score when applied to Spanish EHR. Contrary to expectations, the LLM-driven approach, which necessitates minimal data annotation, performs on par with BERT’s capability to discern complex medical terminologies and contextual nuances within the EHRs. The results of this study highlight a notable advance in the field of NER for Spanish EHRs, with the few shot approach driven by LLM, enhanced by external knowledge, slightly edging out the traditional BERT-based method in overall effectiveness. GPT’s superiority in F-score and its minimal reliance on extensive data annotation underscore its potential in medical data processing.
Alvaro Garcia-Barragán, Alberto González Calatayud, Oswaldo Solarte Pabón, Mariano Provencio, Ernestina Menasalvas Ruiz, Víctor Robles
Multim. Tools Appl.4
2024 Machine learning estimated probability of relapse in early-stage non-small-cell lung cancer patients with aneuploidy imputation scores and knowledge graph embeddings
abstract
Low-stage lung cancer is known to recur unpredictably, and patients receiving various treatment methods like radiation, chemotherapy, and immunotherapies have been seen to respond very differently. Identifying a priori if a patient is going to relapse or not could make a difference in terms of saving lives and personalized care offered. In this work, we provide an answer to the following research question: Is it possible to enhance the machine learning (ML) of the estimated probability of relapse in early-stage non-small-cell lung cancer (NSCLC) patients with aneuploidy imputation scores? To predict recurrence in 1,348 early-stage (I-II) NSCLC patients, we train graph ML models utilizing the Spanish pulmonary cancer group knowledge graph enriched with triples from pathway imputation. ML models trained on Knowledge graph data enriched with triples from pathway score imputation present an 82% Precision and 91% Specificity in predicting relapse over 200 patients from a held-out test set. ML models trained using graphs data could prove useful supplemental tool in the TNM classification systems and improve a lung cancer patient’s prognosis.
Samuele Buosi, Mohan Timilsina, Adrianna Janik, Luca Costabello, Maria Torrente, Mariano Provencio, Dirk Fey, Vít Novácek
Expert Syst. Appl.6
2023 Structuring Breast Cancer Spanish Electronic Health Records Using Deep Learning
abstract
Using Natural Language Processing (NLP) in the clinical domain has increased the possibility of automatically extracting information from oncology clinical narratives. Specifically, deep learning methods have been used to extract information in the cancer domain. However, most of the above proposals have concentrated only on extracting named entities from clinical narratives, but those proposals do not include a methodology for structuring the information after an information extraction step. In this paper, we propose an automatic pipeline based on deep learning for structuring breast cancer information from clinical narratives written in Spanish. The pipeline inputs a set of clinical documents written in narrative form and automatically generates a structured JSON file that contains the information for each patient. This pipeline integrates both clinical entity extraction and negation and uncertainty detection. Obtained results have shown that deep learning methods are feasible for structuring information in the breast cancer domain.
Alvaro Garcia-Barragán, Oswaldo Solarte Pabón, Georgiy Nedostup, Mariano Provencio, Ernestina Menasalvas Ruiz, Víctor Robles
CBMS4
2023 Clustering-based Pattern Discovery in Lung Cancer Treatments
abstract
Lung cancer is the leading cause of cancer death. More than 238,340 new cases of lung cancer patients are expected in 2023, with an estimation of more than 127,070 deaths. Choosing the correct treatment is an important element to enhance the probability of survival and to improve patient's quality of life. Cancer treatments might provoke secondary effects. These toxicities cause different health problems that impact the patient's quality of life. Hence, reducing treatments toxicities while maintaining or improving their effectiveness is an important goal that aims to be pursued from the clinical perspective. On the other hand, clinical guidelines include general knowledge about cancer treatment recommendations to assist clinicians. Although they provide treatment recommendations based on cancer disease aspects and individual patient features, a statistical analysis taking into account treatment outcomes is not provided here. Therefore, the comparison between clinical guidelines with treatment patterns found in clinical data, would allow to validate the patterns found, as well as discovering alternative treatment patterns. In this work, we have analyzed a dataset containing lung cancer patients information including patients' data, prescribed treatments and their outcomes. Using a Chi-square test and K-Modes clustering algorithm in combination with Pattern Discovery metrics we identify patterns, within the clusters, based on cancer stage and treatment outcomes. Obtained results are analyzed based on statistical and clinical relevance and compared with lung cancer clinical guidelines. The comparison reveals that all patterns found coincide with clinical guidelines recommendations, assessing the validity of the proposed method for pattern discovery in a clinical dataset.
Daniel Gómez-Bravo, Aaron García, Guillermo Vigueras, Belén Ríos-Sánchez, Alejandra Pérez-García, Vanessa Ospina, Maria Torrente, Ernestina Menasalvas Ruiz, Mariano Provencio, Alejandro Rodríguez González
CBMS9
2023 Machine Learning Survival Models for Relapse Prediction in a Early Stage Lung Cancer Patient
abstract
Lung cancer is one of the leading health complications causing high mortality worldwide. The relapsing behavior of medically treated early-stage lung cancer makes this disease even more complicated. Thus predicting such relapse using a data-centric approach provides a complementary perspective for clinicians to understand the disease. In this preliminary work, we explored off-the-shelf survival models to predict the relapse of early-stage lung cancer patients. We analyzed the survival models on a cohort of 1348 early-stage non-small cell lung cancer (NSCLC) patients in different timestamps. Using the prediction explanation model SHAP (SHapley Additive exPlanations), we further explained the best-performing survival model's predictions. Our explainable predictive model is a potential tool for oncologists that address an unmet clinical need for post-treatment patient stratification based on the relapse hazard.
Mohan Timilsina, Samuele Buosi, Adrianna Janik, Pasquale Minervini, Luca Costabello, Maria Torrente, Mariano Provencio, Virginia Calvo, Carlos Camps, Ana L. Ortega, Bartomeu Massutí, M. Rosario Garcia Campelo, Edel del Barco, Joaquim Bosch-Barrera, Vít Novácek
IJCNN7
2023 Transformers for extracting breast cancer information from Spanish clinical narratives
abstract
The wide adoption of electronic health records (EHRs) offers immense potential as a source of support for clinical research. However, previous studies focused on extracting only a limited set of medical concepts to support information extraction in the cancer domain for the Spanish language. Building on the success of deep learning for processing natural language texts, this paper proposes a transformer-based approach to extract named entities from breast cancer clinical notes written in Spanish and compares several language models. To facilitate this approach, a schema for annotating clinical notes with breast cancer concepts is presented, and a corpus for breast cancer is developed. Results indicate that both BERT-based and RoBERTa-based language models demonstrate competitive performance in clinical Named Entity Recognition (NER). Specifically, BETO and multilingual BERT achieve F-scores of 93.71% and 94.63%, respectively. Additionally, RoBERTa Biomedical attains an F-score of 95.01%, while RoBERTa BNE achieves an F-score of 94.54%. The findings suggest that transformers can feasibly extract information in the clinical domain in the Spanish language, with the use of models trained on biomedical texts contributing to enhanced results. The proposed approach takes advantage of transfer learning techniques by fine-tuning language models to automatically represent text features and avoiding the time-consuming feature engineering process.
Oswaldo Solarte Pabón, Orlando Montenegro, Alvaro Garcia-Barragán, Maria Torrente, Mariano Provencio, Ernestina Menasalvas Ruiz, Víctor Robles
Artif. Intell. Medicine5
2023 Synergy between imputed genetic pathway and clinical information for predicting recurrence in early stage non-small cell lung cancer
abstract
OBJECTIVE: Lung cancer exhibits unpredictable recurrence in low-stage tumors and variable responses to different therapeutic interventions. Predicting relapse in early-stage lung cancer can facilitate precision medicine and improve patient survivability. While existing machine learning models rely on clinical data, incorporating genomic information could enhance their efficiency. This study aims to impute and integrate specific types of genomic data with clinical data to improve the accuracy of machine learning models for predicting relapse in early-stage, non-small cell lung cancer patients. METHODS: The study utilized a publicly available TCGA lung cancer cohort and imputed genetic pathway scores into the Spanish Lung Cancer Group (SLCG) data, specifically in 1348 early-stage patients. Initially, tumor recurrence was predicted without imputed pathway scores. Subsequently, the SLCG data were augmented with pathway scores imputed from TCGA. The integrative approach aimed to enhance relapse risk prediction performance. RESULTS: The integrative approach achieved improved relapse risk prediction with the following evaluation metrics: an area under the precision-recall curve (PR-AUC) score of 0.75, an area under the ROC (ROC-AUC) score of 0.80, an F1 score of 0.61, and a Precision of 0.80. The prediction explanation model SHAP (SHapley Additive exPlanations) was employed to explain the machine learning model's predictions. CONCLUSION: We conclude that our explainable predictive model is a promising tool for oncologists that addresses an unmet clinical need of post-treatment patient stratification based on the relapse risk while also improving the predictive power by incorporating proxy genomic data not available for specific patients.
Mohan Timilsina, Dirk Fey, Samuele Buosi, Adrianna Janik, Luca Costabello, Enric Carcereny, Delvys Rodriguez Abreu, Manuel Cobo, Rafael Castro, Reyes Bernabé, Pasquale Minervini, Maria Torrente, Mariano Provencio, Vít Novácek
J. Biomed. Informatics13
2022 Integration of Clinical Information and Imputed Aneuploidy Scores to Enhance Relapse Prediction in Early Stage Lung Cancer Patients
Mohan Timilsina, Samuele Bousi, Dirk Fey, Adrianna Janik, Maria Torrente, Mariano Provencio, Alberto Bermúdez, Enric Carcereny, Luca Costabello, Delvys Rodriguez Abreu, Manuel Cobo, Rafael Castro, Reyes Bernabé, Maria Guirado, Pasquale Minervini, Vít Novácek
AMIA6
2022 Subgroup Discovery Analysis of Treatment Patterns in Lung Cancer Patients
abstract
Lung cancer is the leading cause of cancer death. More than 236,740 new cases of lung cancer patients are expected in 2022, with an estimation of more than 130,180 deaths. Improving the survival rates or the patient's quality of life is partially covered by a common element: treatments. Cancer treatments are well known for the toxic outcomes and secondary effects on the patients. These toxicities cause different health problems that impact the patient's quality of life. Reducing toxicities without a decline on the positive survival effect is an important goal that aims to be pursued from the clinical perspective. On the other hand, clinical guidelines include general knowl-edge about cancer treatment recommendations to assist clinicians. Although they provide treatment recommendations based on cancer disease aspects and individual patient features, a statistical analysis taking into account treatment outcomes is not provided here. Therefore, the comparison between clinical guidelines with treatment patterns found in clinical data, would allow to validate the patterns found, as well as discovering alternative treatment patterns. In this work, we have analyzed a dataset containing lung cancer patients information including patients' data, prescribed treatments and outcomes obtained. Using a Subgroup Discovery method we identify patterns based on cancer stage while relying on treatment outcomes. Results are compared with clinical guide-lines and analyzed based on statistical and medical relevance using Subgroup Discovery metrics.
Daniel Gómez-Bravo, Aaron García, Guillermo Vigueras, Belén Ríos-Sánchez, Belén Otero-Carrasco, Roberto Hernández López, Maria Torrente, Ernestina Menasalvas Ruiz, Mariano Provencio, Alejandro Rodríguez González
CBMS9
2022 Deep learning to extract Breast Cancer diagnosis concepts
abstract
The wide adoption of electronic health records (EHRs) provides a potential source to support clinical research. The Bidirectional Encoder Representations from Transformers (BERT) has shown promising results in extracting information in the biomedical domain, including the cancer field. However, one of the challenges in the cancer domain is annotating resources to support information extraction. In this paper, we will show how models trained in a lung cancer corpus can be used to extract cancer concepts even in other cancer types. In particular, we will show the performance of BERT models on breast cancer data that was not used to train the models. Results are very promising as they show the possibility of applying deep learning-based models to predict cancer concepts in a different dataset to the one they were trained on, representing a considerable save of time and resources.
Oswaldo Solarte Pabón, Maria Torrente, Alvaro Garcia-Barragán, Mariano Provencio, Ernestina Menasalvas Ruiz, Víctor Robles
CBMS4
2021 On Predicting Recurrence in Early Stage Non-small Cell Lung Cancer
Sameh K. Mohamed, Brian Walsh, Mohan Timilsina, Vít Novácek, Maria Torrente, Fabio Franco, Mariano Provencio, Adrianna Janik, Luca Costabello, Pontus Stenetorp, Pasquale Minervini
AMIA7
2021 Extracting Cancer Treatments from Clinical Text written in Spanish: A Deep Learning Approach
abstract
Extracting accurate information about cancer patients' treatments is crucial to support clinical research, treatment planning, and to improve clinical care outcomes. However, treatment information resides in unstructured clinical text, making the task of data structuring especially challenging. Although several approaches have been proposed to extract treatments from clinical text, most of these proposals have focused on the English language. In this paper, we propose a deep learning-based approach to extract cancer treatments from clinical text written in Spanish. This approach uses a Bidirectional Long Short Memory (BiLSTM) neural net with a CRF layer to perform Named Entity Recognition. An annotated corpus from clinical text written about lung cancer patients is used to train the BiLSTM-based model. Performed tests have shown a performance of 90% in the F1-score, suggesting the feasibility of our approach to extract cancer treatments from clinical narratives.
Oswaldo Solarte Pabón, Alberto Blázquez-Herranz, Maria Torrente, Alejandro Rodríguez González, Mariano Provencio, Ernestina Menasalvas Ruiz
DSAA5
2021 Towards Treatment Patterns Validation in Lung Cancer Patients
abstract
Lung cancer is the leading cause of cancer death. From the estimation of cases that will be in 2021, more than 230,000 new cases are expected to be of lung cancer patients, with an estimation of more than 131,000 deaths. Improving the survival rates or the patient's quality of life is partially covered by a common element: treatments. Collective knowledge about cancer treatment recommendations is typically included in clinical guidelines, intended to optimize patient care and assist clinicians in lung cancer treatment. These guidelines define a set of treatment paths, where recommendations depend on cancer disease aspects and individual features for a concrete patient. Although oncologists are expected to follow clinical guidelines, the inter and intrapatients' variability of response to the possible treatment combinations, makes it necessary to personalize different treatment-patterns on certain cases. Additionally, clinical guidelines are not frequently updated with new findings or lack a consistent methodology when they are frequently updated. For that reason, the analysis of patterns on both patients treated following the standard of care, or outside it, would allow to validate clinical guidelines and identify potential new treatment recommendations. In this work, we have analysed whether actual treatments prescribed to lung cancer patients follow clinical guidelines or not. Using a machine learning method that provides as output association rules (Apriori), we identify patterns based on cancer stage. These preliminary results show that treatments patterns found mostly match with clinical guidelines recommendations, validating the information included in the consulted guidelines.
Arturo Redondo, Belén Ríos-Sánchez, Guillermo Vigueras, Belén Otero-Carrasco, Roberto Hernández López, Maria Torrente, Ernestina Menasalvas Ruiz, Mariano Provencio, Alejandro Rodríguez González
DSAA8
2020 Lung Cancer Diagnosis Extraction from Clinical Notes Written in Spanish
abstract
The wide adoption of electronic health records (EHRs) offers a potential source to support research. Lung cancer is one of the most common cancer in the world. Although several tools have been developed to automatically extract concepts from oncology clinical notes, still there is a gap between concept extraction and concept understanding. The high number of clinical notes for the same patient, use of negation and proper date annotations lays in the root of the problem. In this paper, we propose an approach to accurate Lung cancer diagnosis extraction from clinical notes written in Spanish. The approach deals with a disambiguation process required to extract the correct date and diagnosis of a patient having hundreds of clinical notes and consequently hundreds of annotations. Results obtained on an annotated database of 1000 patients show an F-score of 90%.
Oswaldo Solarte Pabón, Maria Torrente, Alejandro Rodríguez González, Mariano Provencio, Ernestina Menasalvas Ruiz, Juan Manuel Tuñas
CBMS4
2020 A Data-Driven Approach for Analyzing Healthcare Services Extracted from Clinical Records
abstract
Cancer remains one of the major public health challenges worldwide. After cardiovascular diseases, cancer is one of the first causes of death and morbidity in Europe, with more than 4 million new cases and 1.9 million deaths per year. The suboptimal management of cancer patients during treatment and subsequent follows up are major obstacles in achieving better outcomes of the patients and especially regarding cost and quality of life In this paper, we present an initial data-driven approach to analyze the resources and services that are used more frequently by lung-cancer patients with the aim of identifying where the care process can be improved by paying a special attention on services before diagnosis to being able to identify possible lung-cancer patients before they are diagnosed and by reducing the length of stay in the hospital. Our approach has been built by analyzing the clinical notes of those oncological patients to extract this information and their relationships with other variables of the patient. Although the approach shown in this manuscript is very preliminary, it shows that quite interesting outcomes can be derived from further analysis.
Manuel Scurti, Ernestina Menasalvas Ruiz, Maria-Esther Vidal, Maria Torrente, Dimitrios Vogiatzis, Georgios Paliouras, Mariano Provencio, Alejandro Rodríguez González
CBMS7
2020 Reconstructing the patient's natural history from electronic health records
Marjan Najafabadipour, Massimiliano Zanin, Alejandro Rodríguez González, Maria Torrente, Beatriz Nuñez, Juan Luis Cruz-Bermúdez, Mariano Provencio, Ernestina Menasalvas Ruiz
Artif. Intell. Medicine7
2019 iASiS: Towards Heterogeneous Big Data Analysis for Personalized Medicine
abstract
The vision of IASIS project is to turn the wave of big biomedical data heading our way into actionable knowledge for decision makers. This is achieved by integrating data from disparate sources, including genomics, electronic health records and bibliography, and applying advanced analytics methods to discover useful patterns. The goal is to turn large amounts of available data into actionable information to authorities for planning public health activities and policies. The integration and analysis of these heterogeneous sources of information will enable the best decisions to be made, allowing for diagnosis and treatment to be personalised to each individual. The project offers a common representation schema for the heterogeneous data sources. The iASiS infrastructure is able to convert clinical notes into usable data, combine them with genomic data, related bibliography, image data and more, and create a global knowledge base. This facilitates the use of intelligent methods in order to discover useful patterns across different resources. Using semantic integration of data gives the opportunity to generate information that is rich, auditable and reliable. This information can be used to provide better care, reduce errors and create more confidence in sharing data, thus providing more insights and opportunities. Data resources for two different disease categories are explored within the iASiS use cases, dementia and lung cancer.
Anastasia Krithara, Fotis Aisopos, Vassiliki Rentoumi, Anastasios Nentidis, Konstantinos Bougiatiotis, Maria-Esther Vidal, Ernestina Menasalvas Ruiz, Alejandro Rodríguez González, Eleftherios Samaras, Peter Garrard, Maria Torrente, Mariano Provencio, Nikos Dimakopoulos, Rui Mauricio, Jordi Rambla De Argila, Gian Gaetano Tartaglia, Georgios Paliouras
CBMS12
2019 Recognition of Time Expressions in Spanish Electronic Health Records
abstract
The widespread adoption of Electronic Health Records (EHRs) is generating an ever-increasing amount of unstructured clinical texts. Processing time expressions from these domain-specific-texts is crucial for the discovery of patterns that can help in the detection of medical events and building the patient's natural history. In medical domain, the recognition of time information from texts is challenging due to their lack of structure; usage of various formats, styles and abbreviations; their domain specific nature; writing quality; and the presence of ambiguous expressions. Furthermore, despite of Spanish occupying the second position in the world ranking of number of native speakers, to the best of our knowledge, no Natural Language Processing (NLP) tools have been introduced for the recognition of time expressions from clinical texts, written in this particular language. Therefore, in this paper, we propose a Temporal Tagger for identifying and normalizing time expressions appeared in Spanish clinical texts. We further compare our Temporal Tagger with the Spanish version of SUTime. By using a large dataset comprising EHRs of people suffering from lung cancer, we show that our developed Temporal Tagger, with an F1 score of 0.93, outperforms SUTime, with an F1 score of 0.797.
Marjan Najafabadipour, Massimiliano Zanin, Alejandro Rodríguez González, Consuelo Gonzalo-Martín, Beatriz Nuñez, Virginia Calvo, Juan Luis Cruz-Bermúdez, Mariano Provencio, Ernestina Menasalvas Ruiz
CBMS8
2017 OncoCall: Analyzing the Outcomes of the Oncology Telephone Patient Assistance
abstract
Hospital Puerta del Hierro in Madrid, Spain, implemented in November 2011 a new service that aim to aid the patients of the oncology service with their doubts during their treatments through the use of a centralized call center. This service was created with the goal of provide a more personalized patient attention as well as to try to reduce the number of re-entries in the hospital in the emergencies. The aim of this paper is to present the main result of the analysis of the data produced by their call service in order to verify if the objectives were fulfilled as well as to gather what improvements can be done.
Ernestina Menasalvas Ruiz, Consuelo Gonzalo-Martín, Juan Manuel Tuñas, Alejandro Rodríguez González, Mariano Provencio, Cristina Gonzalez de Pedro, Marta Mendez, Olga Zaretskaia, Juan Luis Cruz, Jesús Rey, Consuelo Parejo
CBMS5