Cristina Soguero-Ruíz

dblp:76/11262 · DBLP profile ↗
← Back
31ranked-venue papers
8as first author
18since 2021 · last 2025
0000-0001-5817-989XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2025 Advanced Graph-Based Approaches for Predicting Antimicrobial Resistance in Intensive Care Units
abstract
Antimicrobial Resistance (AMR) poses a significant global public health challenge, necessitating early detection strategies to enable timely clinical interventions. Electronic Health Records (EHRs) offer extensive real-world clinical data but present challenges due to their irregularly sampled, heterogeneous, and multivariate temporal structure. This paper investigates graph-based learning models to predict AMR in Intensive Care Unit patients by systematically modeling spatial and temporal dependencies within EHR data represented as Multivariate Time Series. We propose and evaluate a novel Spatio-Temporal Graph Convolutional Neural Network architecture, demonstrating its superior predictive performance by achieving a Receiver Operating Characteristic Area Under the Curve of 80.00%, surpassing baseline models by approximately 6%. Furthermore, our analysis of the learned graph structures highlights critical clinical interactions, notably emphasizing catheterrelated variables as central nodes, aligning well with established clinical knowledge. By combining high predictive performance with enhanced interpretability, our approach presents a robust and transparent framework, well-suited for clinical applications aimed at improving AMR risk assessment and patient care management.
Paula Martín-Palomeque, Óscar Escudero-Arnanz, Cristina Soguero-Ruíz, Antonio G. Marqués
CBMS3
2025 Glucostats: an efficient Python library for glucose time series feature extraction and visual analysis
abstract
BACKGROUND: The advancement of technology and continuous glucose monitoring (CGM) systems has introduced several computational and technical challenges for clinicians and researchers. The growing volume of CGM data necessitates the development of efficient computational tools capable of handling and processing this information effectively. This paper introduces GlucoStats, an open-source and multi-processing Python library designed for efficient computation and visualization of a comprehensive set of glucose metrics derived from CGM. It simplifies the traditionally time-consuming and error-prone process of manual CGM metrics calculation, making it a valuable tool for both clinical and research applications. RESULTS: Its modular design ensures easy integration into predefined workflows, while its user-friendly interface and extensive documentation make it accessible to a broad audience, including clinicians and researchers. GlucoStats offers several key features: (i) window-based time series analysis, enabling time series division into smaller 'windows' for detailed temporal analysis, particularly beneficial for CGM data; (ii) advanced visualization tools, providing intuitive, high-quality visualizations that facilitate pattern recognition, trend analysis, and anomaly detection in CGM data; (iii) parallelization, leveraging parallel computing to efficiently handle large CGM datasets by distributing computations across multiple processors; and (iv) scikit-learn compatibility, adhering to the standardized interface of scikit-learn to allow an easy integration into machine learning pipelines for end-to-end analysis. CONCLUSIONS: GlucoStats demonstrates high efficiency in processing large-scale medical datasets in minimal time. Its modular design enables easy customization and extension, making it adaptable to diverse research and clinical needs. By offering precise CGM data analysis and user-friendly visualization tools, it serves both technical researchers and non-technical users, such as physicians and patients, with practical and research-driven applications.
Pablo Peiro-Corbacho, Francisco J. Lara-Abelenda, David Chushig-Muzo, Ana M. Wägner, Conceição Granja, Cristina Soguero-Ruíz
BMC Bioinform.6
2025 Interpretable and multimodal fusion methodology to predict severe hypoglycemia in adults with type 1 diabetes
abstract
Type 1 diabetes (T1D) causes insulin deficiency and exogenous therapy is required for maintaining targeted glucose levels. Hypoglycemia is the most frequent side effect of insulin, being severe hypoglycemia (SH) one of the most critical hazards with a range of life-threatening consequences. Artificial intelligence (AI) and multimodal fusion have boosted predictive performance in different domains. This study aims to evaluate the effectiveness of early fusion (EF) and late fusion (LF) approaches for predicting SH, to create a methodology capable of achieving robust results in datasets with a low number of samples for predicting SH and to characterize the risk factors involved in the SH onset using explainable AI (XAI). Data from a case-control study comprising adults over 60 years with T1D and with diabetes duration of 20 years were used and three types of modalities were considered: (1) continuous glucose monitoring data (time series); (2) clinical codes (text); and (3) surveys related to fear, unawareness, depression, and cognitive tests (tabular data). The results revealed that EF outperformed models trained with single-modality data by 5.8%, with an area under the receiver operating characteristic curve of 0.779. XAI techniques helped to discover that features related to fear and unawareness are mainly associated with SH. Our study introduced an interpretable and multimodal methodology capable of predicting the occurrence of SH in adults with T1D in the next year. Our interpretable methodology contributes to predicting SH and identifying related key factors, thus preventing SH complications and improving patient’s quality of life. • Multimodal fusion approaches improve the prediction of severe hypoglycemia. • Interpretability techniques lead to identifying key factors for severe hypoglycemia. • People with severe hypoglycemia show a high frequency of cardiovascular diseases. • Impaired hypoglycemia awareness elevates severe hypoglycemia risk.
Francisco J. Lara-Abelenda, David Chushig-Muzo, Ana M. Wägner, Maryam Tayefi, Cristina Soguero-Ruíz
Eng. Appl. Artif. Intell.5
2025 Transfer learning for a tabular-to-image approach: A case study for cardiovascular disease prediction
abstract
OBJECTIVE: Machine learning (ML) models have been extensively used for tabular data classification but recent works have been developed to transform tabular data into images, aiming to leverage the predictive performance of convolutional neural networks (CNNs). However, most of these approaches fail to convert data with a low number of samples and mixed-type features. This study aims: to evaluate the performance of the tabular-to-image method named low mixed-image generator for tabular data (LM-IGTD); and to assess the effectiveness of transfer learning and fine-tuning for improving predictions on tabular data. METHODS: We employed two public tabular datasets with patients diagnosed with cardiovascular diseases (CVDs): Framingham and Steno. First, both datasets were transformed into images using LM-IGTD. Then, Framingham, which contains a larger set of samples than Steno, is used to train CNN-based models. Finally, we performed transfer learning and fine-tuning using the pre-trained CNN on the Steno dataset to predict CVD risk. RESULTS: The CNN-based model with transfer learning achieved the highest AUCORC in Steno (0.855), outperforming ML models such as decision trees, K-nearest neighbors, least absolute shrinkage and selection operator (LASSO) support vector machine and TabPFN. This approach improved accuracy by 2% over the best-performing traditional model, TabPFN. CONCLUSION: To the best of our knowledge, this is the first study that evaluates the effectiveness of applying transfer learning and fine-tuning to tabular data using tabular-to-image approaches. Through the use of CNNs' predictive capabilities, our work also advances the diagnosis of CVD by providing a framework for early clinical intervention and decision-making support.
Francisco J. Lara-Abelenda, David Chushig-Muzo, Pablo Peiro-Corbacho, Vanesa Gómez-Martínez, Ana M. Wägner, Conceição Granja, Cristina Soguero-Ruíz
J. Biomed. Informatics7
2024 Irregular Temporal Classification of Multidrug Resistance Development in Intensive Care Unit Patients
abstract
The rising incidence of Multidrug-Resistant (MDR) infections in Intensive Care Units (ICUs) presents a critical challenge to global healthcare, leading to adverse patient outcomes and increased costs. Early identification of MDR pathogens is crucial for timely implementation of isolation protocols and prevention of cross-infections. This study introduces a novel Deep Learning (DL) framework for temporal prediction of MDR status in ICU patients, addressing the complexities of irregular classification tasks. We utilize Electronic Health Records from 3,502 anonymized ICU patient records spanning 17 years at the University Hospital of Fuenlabrada, Madrid, Spain. Five advanced DL architectures are systematically evaluated: Vanilla Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), Bidirectional LSTM, Gated Recurrent Unit, and Transformer, each adapted for irregular Multivariate Time Series (MTS) data. Notably, this is the first study to assess the impact of temporal window size on predictive performance and to adapt different RNN models for irregular temporal classification in MTS. Our results show that the LSTM model, when applied to a 4-day temporal window, achieves the highest predictive performance, with a mean Receiver Operating Characteristic Area Under the Curve of 79.78 ± 2.78. This study highlights the importance of early temporal windows in capturing critical indicators of MDR infections and significantly advances the state-of-the-art in MDR prediction, providing a robust foundation for more effective ICU management and potential reduction of MDR infections in healthcare settings.
Paula Martín-Palomeque, Óscar Escudero-Arnanz, Joaquín Álvarez-Rodríguez, Cristina Soguero-Ruíz
BIBM4
2024 Low-Rank Tensor Completion for Heart Failure Exacerbation Detection in Multivariate Time Series with Missing Data
abstract
Heart failure exacerbations (HFE) represent a critical challenge in healthcare due to their significant role in global mortality. The rise of home and wearable devices capable of monitoring cardiac conditions provides valuable opportunities for data-driven analysis and HFE detection. However, these devices frequently generate low-quality measurements with irregular sampling frequencies and high rates of missing data. Our paper presents a methodology that processes these measurements as a three-dimensional tensor, applying a low-rank tensor completion scheme to manage missing data effectively, thus facilitating anomaly detection without necessitating data imputation. We validate our method on a dataset from 4 patients with chronic HF in the compensation phase, collected at the Hospital Fondazione Policlinico Universitario Campus Bio-Medico in Rome, Italy. Our results demonstrate the tensor-based method’s superiority over traditional techniques, highlighting its potential for detecting anomalies within complex multivariate time series data. This research emphasizes the critical role of advanced data analysis in enhancing HFE identification, which could lead to improved patient care and reduced hospitalization rates.
Óscar Escudero-Arnanz, Rosa Sicilia, Cristina Soguero-Ruíz, I. Mora-Jiménez, Diana Lelli, Claudio Pedone, Antonio G. Marqués
CBMS3
2023 Novel Approach for AI-Based Risk Calculator Development Using Transfer Learning Suitable for Embedded Systems
abstract
Noncommunicable Diseases (NCDs), like Cardiovascular Diseases (CVD) or Diabetes Mellitus (DM) are defined as chronic conditions caused by the combination of genetic, physiological, behavioral, and environmental factors that can affect an individual's health, being a major issue for the public health system globally. Sometimes, these conditions share some of their risk factors, as occurs between CVD and DM. Current clinically validated risk calculators have been developed using different regression approaches, targeting different populations and having significant differences between their outputs and the risk factors they use to compute the risk. In this work, we present a methodology for the design of risk calculator based on Machine Learning (ML), combining the knowledge of different clinically validated cardiovascular risk calculators using transfer learning for more personalized NCD risk estimation. Besides, a hardware profiling in terms of latency and model size is performed, targeting its real-time implementation in an embedded system. Results suggest that re-training an already developed ML model with a different dataset can improve its generalization capability, being a suitable way to avoid overfitting. Moreover, profiling results shown that this type of ML-based algorithms are suitable for embedded systems implementations., having model sizes lower than 1 KB and average inference times lower than$75\ \mu\mathrm{s}$.
Antonio J. Rodríguez-Almeida, Himar Fabelo, Cristina Soguero-Ruíz, Rosa María Sanchez-Hernandez, Ana M. Wägner, Gustavo M. Callicó
DSD3
2023 Dimensionality reduction and ensemble of LSTMs for antimicrobial resistance prediction
abstract
Bacterial resistance to antibiotics has been rapidly increasing, resulting in low antibiotic effectiveness even treating common infections. The presence of resistant pathogens in environments such as a hospital Intensive Care Unit (ICU) exacerbates the critical admission-acquired infections. This work focuses on the prediction of antibiotic resistance in Pseudomonas aeruginosa nosocomial infections at the ICU, using Long Short-Term Memory (LSTM) artificial neural networks as the predictive method. The analyzed data were extracted from the Electronic Health Records (EHR) of patients admitted to the University Hospital of Fuenlabrada from 2004 to 2019 and were modeled as Multivariate Time Series. A data-driven dimensionality reduction method is built by adapting three feature importance techniques from the literature to the considered data and proposing an algorithm for selecting the most appropriate number of features. This is done using LSTM sequential capabilities so that the temporal aspect of features is taken into account. Furthermore, an ensemble of LSTMs is used to reduce the variance in performance. Our results indicate that the patient's admission information, the antibiotics administered during the ICU stay, and the previous antimicrobial resistance are the most important risk factors. Compared to other conventional dimensionality reduction schemes, our approach is able to improve performance while reducing the number of features for most of the experiments. In essence, the proposed framework achieve, in a computationally cost-efficient manner, promising results for supporting decisions in this clinical task, characterized by high dimensionality, data scarcity, and concept drift.
Álvar Hernández-Carnerero, Miquel Sànchez-Marrè, I. Mora-Jiménez, Cristina Soguero-Ruíz, Sergio Martínez-Agüero, Joaquín Álvarez-Rodríguez
Artif. Intell. Medicine4
2023 A streaming data visualization framework for supporting decision-making in the Intensive Care Unit
abstract
This research was funded by the Spanish Research Agency, grant numbers PID2021-122392OB-I00, PID2019-106623RB-C41/AEI/10.13039/501100011033 and PID2019-107768RA-I00; and by Universidad Rey Juan Carlos (URJC) and Community of Madrid, Spain , grant number 2020-66.
Miguel A. Mohedano-Munoz, Cristina Soguero-Ruíz, I. Mora-Jiménez, Manuel Rubio-Sánchez, Joaquín Álvarez-Rodríguez, Alberto Sánchez 0001
Expert Syst. Appl.2
2023 Synthetic Patient Data Generation and Evaluation in Disease Prediction Using Small and Imbalanced Datasets
abstract
The increasing prevalence of chronic non-communicable diseases makes it a priority to develop tools for enhancing their management. On this matter, Artificial Intelligence algorithms have proven to be successful in early diagnosis, prediction and analysis in the medical field. Nonetheless, two main issues arise when dealing with medical data: lack of high-fidelity datasets and maintenance of patient's privacy. To face these problems, different techniques of synthetic data generation have emerged as a possible solution. In this work, a framework based on synthetic data generation algorithms was developed. Eight medical datasets containing tabular data were used to test this framework. Three different statistical metrics were used to analyze the preservation of synthetic data integrity and six different synthetic data generation sizes were tested. Besides, the generated synthetic datasets were used to train four different supervised Machine Learning classifiers alone, and also combined with the real data. F1-score was used to evaluate classification performance. The main goal of this work is to assess the feasibility of the use of synthetic data generation in medical data in two ways: preservation of data integrity and maintenance of classification performance.
Antonio J. Rodríguez-Almeida, Himar Fabelo, Samuel Ortega, Alejandro Deniz, Francisco Balea-Fernández, Eduardo Quevedo, Cristina Soguero-Ruíz, Ana M. Wägner, Gustavo M. Callicó
IEEE J. Biomed. Health Informatics7
2022 Local Naïve Bayes for Predicting Evolution of COVID-19 Patients on Self Organizing Maps
abstract
The most recent Clinical Decision Support Systems use the potential of Machine Learning techniques to target clinical problems, avoiding the use of explicit rules. In this paper, a model to monitor and predict the risk of unfavourable evolution (UE) during hospitalization of COVID-19 patients is proposed. It combines Self Organizing Maps and local Naïve Bayes (NB) classifiers because of interpretation purposes. We used the results of six blood tests (leukocytes, D-dimer, among others) provided by a Spanish hospital group. The probabilistic approach allows us to get the daily risk of UE for each patient in an interpretable way. Several variants of the NB classifiers family have been explored, mainly weighting and likelihood estimation (parametric and nonparametric). Despite the over-simplified assumptions of the NB classifiers, they provided good predictive results in terms of sensitivity and specificity. The model with nonparametric likelihood estimation provided the best risk prediction over time even when designed with a limited number of samples. Specifically, the median value and interquartil range for the risk prediction were quite reliable even 10 days before the event day for patients hospitalized longer than 7 days. The risk median values also agree with the gold-standard for patients with a hospital stay shorter than 7 days, though the interquartil range can be too wide (probably because of the variability in the inpatient days - sometimes, just 2 days). Though a deepest analysis considering more patients and features would be convenient, our results show the potential of the proposed approach, both from a technical and clinical viewpoint.
Carlos Arias-Alcaide, Cristina Soguero-Ruíz, Paloma Santos-Alvarez, José Felipe Varona Arche, I. Mora-Jiménez
BIBM2
2022 Characterizing Cardiovascular Risk Through Unsupervised and Interpretable Techniques
Hugo Calero-Díaz, David Chushig-Muzo, Cristina Soguero-Ruíz
IDEAL3
2022 Interpretable clinical time-series modeling with intelligent feature selection for early prediction of antimicrobial multidrug resistance
abstract
Electronic health records provide rich, heterogeneous data about the evolution of the patients’ health status. However, such data need to be processed carefully, with the aim of extracting meaningful information for clinical decision support. In this paper, we leverage interpretable (deep) learning and signal processing tools to deal with multivariate time-series data collected from the Intensive Care Unit (ICU) of the University Hospital of Fuenlabrada (Madrid, Spain). The presence of antimicrobial multidrug-resistant (AMR) bacteria is one of the greatest threats to the health system in general and to the ICUs in particular due to the critical health status of the patients therein. Thus, early identification of bacteria at the ICU and early prediction of their antibiotic resistance are key for the patients’ prognosis. While intelligent data-based processing and learning schemes can contribute to this early prediction, their acceptance and deployment in the ICUs require the automatic schemes to be not only accurate but also understandable by clinicians. Accordingly, we have designed trustworthy intelligent models for the early prediction of AMR based on the combination of meaningful feature selection with interpretable recurrent neural networks. These models were created using irregularly sampled clinical measurements, both considering the health status of the patient and the global ICU environment. We explored several strategies to cope with strongly imbalance data, since only a few ICU patients are infected by AMR bacteria. It is worth noting that our approach exhibits a good balance between performance and interpretability, especially when considering the difficulty of the classification task at hand. A multitude of factors are involved in the emergence of AMR (several of them not fully understood), and the records only contain a subset of them. In addition, the limited number of patients, the imbalance between classes, and the irregularity of the data render the problem harder to solve. Our models are also enriched with SHAP post-hoc interpretability and validated by clinicians who considered model understandability and trustworthiness of paramount concern for pragmatic purposes. Moreover, we use linguistic fuzzy systems to provide clinicians with explanations in natural language. Such explanations are automatically generated from a pool of interpretable rules that describe the interaction among the most relevant features identified by SHAP. Notice that clinicians were especially satisfied with new insights provided by our models. Such insights helped them to trust the automatic schemes and use them to make (better) decisions to mitigate AMR spreading in the ICU. All in all, this work paves the way towards more comprehensible time-series analysis in the context of early AMR prediction in ICUs and reduces the time of detection of infectious diseases, opening the door to better hospital care.
Sergio Martínez-Agüero, Cristina Soguero-Ruíz, Jose Maria Alonso-Moral, I. Mora-Jiménez, Joaquín Álvarez-Rodríguez, Antonio G. Marqués
Future Gener. Comput. Syst.2
2021 Mapping Health Trajectories on Self Organizing Maps using COVID-19 Patient's Blood Tests
abstract
Since COVID-19 appeared in December 2019, scientists are researching new ways to improve the management of the disease. Considering machine learning approaches have proven to be very useful tools to discover hidden patterns in data, we propose in this paper to apply a Self Organizing Map (SOM) to characterize the health-status evolution of COVID-19 patients. The SOM is a neural network whose neurons can be represented as cells in a bi-dimensional grid preserving the mapping from the original space to the map units. We consider real-world data of hospitalized COVID-19 patients in a Spanish hospital during the first wave of the pandemic. Patients are represented by six blood tests (leukocytes and D-dimer, among others) in a daily basis. Besides, each patient is associated with one of two different health-status: favorable evolution (discharged home) and unfavorable evolution (exitus or admission to the intensive care unit). We show the potential of our approach by detailing the mapping of the health trajectory associated with different particular cases and drawing their trajectory on the bi-dimensional map of the SOM.
Carlos Arias-Alcaide, Cristina Soguero-Ruíz, Paloma Santos-Alvarez, Adrián García-Romero, I. Mora-Jiménez
BIBM2
2021 Predicting Multidrug Resistance Using Temporal Clinical Data and Machine Learning Methods
abstract
Infections caused by multidrug resistant (MR) bacteria severely jeopardize the public’s health given the inefficiency of current antibiotics to treat them. This results in a major global concern that affects any hospital service and the Intensive Care Unit (ICU) in particular. This paper aims to anticipate the antibiogram outcomes associated with MR bacteria in ICU patients by applying machine learning (ML) techniques. For this purpose, multiple clinical variables obtained from the Electronic Health Record have been employed, as the own patient’s antibiotics consumption and the drugs taken by the remaining ICU patients. A collection of 3476 patients admitted to the ICU at the University Hospital of Fuenlabrada from 2004 to 2020 were considered, 628 with MR bateria. A feature engineering (FE) and feature selection (FS) process has been conducted to extract valuable statistics from the original temporal data. The highest Accuracy and Specificity results achieved were 77% and 82%, respectively, both implementing Random Forest as classifier and without considering any FS method. The highest Sensitivity (69%) and ROC-AUC (76%) were attained with the features selected using the Chi-Square test and with both Logistic Regression and XGBoost classifiers. This work provides a promising approach to support therapy decisions by the early identification of MR infections among ICU patients.
Lidia Pascual-Sánchez, I. Mora-Jiménez, Sergio Martínez-Agüero, Joaquín Álvarez-Rodríguez, Cristina Soguero-Ruíz
BIBM5
2021 Interpreting clinical latent representations using autoencoders and probabilistic models
abstract
Electronic health records (EHRs) are a valuable data source that, in conjunction with deep learning (DL) methods, have provided important outcomes in different domains, contributing to supporting decision-making. Owing to the remarkable advancements achieved by DL-based models, autoencoders (AE) are becoming extensively used in health care. Nevertheless, AE-based models are based on nonlinear transformations, resulting in black-box models leading to a lack of interpretability, which is vital in the clinical setting. To obtain insights from AE latent representations, we propose a methodology by combining probabilistic models based on Gaussian mixture models and hierarchical clustering supported by Kullback-Leibler divergence. To validate the methodology from a clinical viewpoint, we used real-world data extracted from EHRs of the University Hospital of Fuenlabrada (Spain). Records were associated with healthy and chronic hypertensive and diabetic patients. Experimental outcomes showed that our approach can find groups of patients with similar health conditions by identifying patterns associated with diagnosis and drug codes. This work opens up promising opportunities for interpreting representations obtained by the AE-based model, bringing some light to the decision-making process made by clinical experts in daily practice.
David Chushig-Muzo, Cristina Soguero-Ruíz, Pablo de Miguel-Bohoyo, I. Mora-Jiménez
Artif. Intell. Medicine2
2021 Time series cluster kernels to exploit informative missingness and incomplete label information
abstract
The time series cluster kernel (TCK) provides a powerful tool for analysing multivariate time series subject to missing data. TCK is designed using an ensemble learning approach in which Bayesian mixture models form the base models. Because of the Bayesian approach, TCK can naturally deal with missing values without resorting to imputation and the ensemble strategy ensures robustness to hyperparameters, making it particularly well suited for unsupervised learning. However, TCK assumes missing at random and that the underlying missingness mechanism is ignorable, i.e. uninformative, an assumption that does not hold in many real-world applications, such as e.g. medicine. To overcome this limitation, we present a kernel capable of exploiting the potentially rich information in the missing values and patterns, as well as the information from the observed data. In our approach, we create a representation of the missing pattern, which is incorporated into mixed mode mixture models in such a way that the information provided by the missing patterns is effectively exploited. Moreover, we also propose a semi-supervised kernel, capable of taking advantage of incomplete label information to learn more accurate similarities. Experiments on benchmark data, as well as a real-world case study of patients described by longitudinal electronic health record data who potentially suffer from hospital-acquired infections, demonstrate the effectiveness of the proposed methods.
Karl Øyvind Mikalsen, Cristina Soguero-Ruíz, Filippo Maria Bianchi, Arthur Revhaug, Robert Jenssen
Pattern Recognit.2
2021 Data and Network Analytics for COVID-19 ICU Patients: A Case Study for a Spanish Hospital
abstract
The COVID-19 pandemic presents unprecedented challenges to the healthcare systems around the world. In 2020, Spain was among the countries with the highest Intensive Care Unit (ICU) hospitalization and mortality rates. This work analyzes data of COVID-19 patients admitted to a Spanish ICU during the first wave of the pandemic. The patients in our study either died (deceased patients) or were discharged from the ICU (non-deceased patients) and underwent the following landmarks: beginning of symptoms; arrival at the emergency department; beginning of the hospital stay; and ICU admission. Our goal is to create a graph-based data-science methodology to find associations among patients' comorbidities, previous medication, symptoms, and the COVID-19 treatment, and to analyze their evolution across landmarks. Towards that end, we first perform a hypothesis test based on bootstrap to identify discriminative features among deceased and non-deceased patients. Then, we leverage graph-based representations and network analytics to determine pairwise associations and complex relations among clinical features. The descriptive statistical analysis confirms that deceased patients exhibit multiple comorbidities with stronger levels of association and are treated with a wider range of drugs during the ICU stay. We also observe that the most common treatment was the simultaneous administration of lopinavir/ritonavir with hydroxychloroquine, regardless of the patients' outcome. Our results illustrate how graph tools and representations yield insights on the relations among comorbidities, drug treatments, and patients' evolution. All in all, the approach puts forth a new data-analysis tool for clinicians that can be applied to analyze (post-COVID) symptom/patient evolution.
Sergio Martínez-Agüero, Antonio G. Marqués, I. Mora-Jiménez, Joaquín Álvarez-Rodríguez, Cristina Soguero-Ruíz
IEEE J. Biomed. Health Informatics5
2020 Finding Associations among Chronic Conditions by Bootstrap and Multiple Correspondence Analysis
abstract
Contemporary societies are suffering from negative population growth, with the consequent population aging. The prevalence of some chronic diseases, of slow progress and long duration, have become one of the main problems for healthcare systems. In particular, high blood pressure, diabetes mellitus, chronic obstructive pulmonary disease, and depression are health-status with a high economic and social burden. In collaboration with University Hospital of Fuenlabrada (Spain), we analyze in this work data (mainly diagnoses and drugs, both coded) from patients suffering from these chronic conditions. Given the high dimensionality of the data, we performed a hypothesis test with bootstrapping in order to select discriminative features that we subsequently analyzed using Multiple Correspondence Analysis (MCA). MCA allowed us to find associations among features and health-statuses, which may reveal not evident relationships. From the analysis carried out, on the one hand, some evidences are concluded, which can be used to validate the methodology followed in this work. On the other hand, we have drawn some conclusions that could assist in clinical decision-making, such as for example, offering more specialized care to patients stratified in the same health-status.
Cristina Soguero-Ruíz, Natalia Alonso-Arteaga, Sergio Muñoz-Romero, José Luis Rojo-Álvarez, Manuel Rubio-Sánchez, Isabel Caballero López-Fajardo, I. Mora-Jiménez
BIBM1
2020 Visually guided classification trees for analyzing chronic patients
abstract
BACKGROUND: Chronic diseases are becoming more widespread each year in developed countries, mainly due to increasing life expectancy. Among them, diabetes mellitus (DM) and essential hypertension (EH) are two of the most prevalent ones. Furthermore, they can be the onset of other chronic conditions such as kidney or obstructive pulmonary diseases. The need to comprehend the factors related to such complex diseases motivates the development of interpretative and visual analysis methods, such as classification trees, which not only provide predictive models for diagnosing patients, but can also help to discover new clinical insights. RESULTS: In this paper, we analyzed healthy and chronic (diabetic, hypertensive) patients associated with the University Hospital of Fuenlabrada in Spain. Each patient was classified into a single health status according to clinical risk groups (CRGs). The CRGs characterize a patient through features such as age, gender, diagnosis codes, and drug codes. Based on these features and the CRGs, we have designed classification trees to determine the most discriminative decision features among different health statuses. In particular, we propose to make use of statistical data visualizations to guide the selection of features in each node when constructing a tree. We created several classification trees to distinguish among patients with different health statuses. We analyzed their performance in terms of classification accuracy, and drew clinical conclusions regarding the decision features considered in each tree. As expected, healthy patients and patients with a single chronic condition were better classified than patients with comorbidities. The constructed classification trees also show that the use of antipsychotics and the diagnosis of chronic airway obstruction are relevant for classifying patients with more than one chronic condition, in conjunction with the usual DM and/or EH diagnoses. CONCLUSIONS: We propose a methodology for constructing classification trees in a visually guided manner. The approach allows clinicians to progressively select the decision features at each of the tree nodes. The process is guided by exploratory data analysis visualizations, which may provide new insights and unexpected clinical information.
Cristina Soguero-Ruíz, I. Mora-Jiménez, Miguel A. Mohedano-Munoz, Manuel Rubio-Sánchez, Pablo de Miguel-Bohoyo, Alberto Sánchez 0001
BMC Bioinform.1
2020 Informative variable identifier: Expanding interpretability in feature selection
Sergio Muñoz-Romero, Arantza Gorostiaga, Cristina Soguero-Ruíz, I. Mora-Jiménez, José Luis Rojo-Álvarez
Pattern Recognit.3
2019 Noisy multi-label semi-supervised dimensionality reduction
abstract
Noisy labeled data represent a rich source of information that often are easily accessible and cheap to obtain, but label noise might also have many negative consequences if not accounted for. How to fully utilize noisy labels has been studied extensively within the framework of standard supervised machine learning over a period of several decades. However, very little research has been conducted on solving the challenge posed by noisy labels in non-standard settings. This includes situations where only a fraction of the samples are labeled (semi-supervised) and each high-dimensional sample is associated with multiple labels. In this work, we present a novel semi-supervised and multi-label dimensionality reduction method that effectively utilizes information from both noisy multi-labels and unlabeled data. With the proposed Noisy multi-label semi-supervised dimensionality reduction (NMLSDR) method, the noisy multi-labels are denoised and unlabeled data are labeled simultaneously via a specially designed label propagation algorithm. NMLSDR then learns a projection matrix for reducing the dimensionality by maximizing the dependence between the enlarged and denoised multi-label space and the features in the projected space. Extensive experiments on synthetic data, benchmark datasets, as well as a real-world case study, demonstrate the effectiveness of the proposed algorithm and show that it outperforms state-of-the-art multi-label feature extraction algorithms.
Karl Øyvind Mikalsen, Cristina Soguero-Ruíz, Filippo Maria Bianchi, Robert Jenssen
Pattern Recognit.2
2018 Using multi-anchors to identify patients suffering from multimorbidities
Karl Øyvind Mikalsen, Cristina Soguero-Ruíz, I. Mora-Jiménez, Isabel Caballero-López-Fando, Robert Jenssen
BIBM2
2018 Scaled radial axes for interactive visual feature selection: A case study for analyzing chronic conditions
Alberto Sánchez 0001, Cristina Soguero-Ruíz, I. Mora-Jiménez, Francisco Javier Rivas-Flores, Dirk J. Lehmann, Manuel Rubio-Sánchez
Expert Syst. Appl.2
2018 Time series cluster kernel for learning similarities between multivariate time series with missing data
Karl Øyvind Mikalsen, Filippo Maria Bianchi, Cristina Soguero-Ruíz, Robert Jenssen
Pattern Recognit.3
2016 Predicting colorectal surgical complications using heterogeneous clinical data and kernel methods
Cristina Soguero-Ruíz, Kristian Hindberg, I. Mora-Jiménez, José Luis Rojo-Álvarez, Stein Olav Skrøvseth, Fred Godtliebsen, Kim Mortensen, Arthur Revhaug, Rolv-Ole Lindsetmo, Knut Magne Augestad, Robert Jenssen
J. Biomed. Informatics1
2016 Support Vector Feature Selection for Early Detection of Anastomosis Leakage From Bag-of-Words in Electronic Health Records
abstract
The free text in electronic health records (EHRs) conveys a huge amount of clinical information about health state and patient history. Despite a rapidly growing literature on the use of machine learning techniques for extracting this information, little effort has been invested toward feature selection and the features' corresponding medical interpretation. In this study, we focus on the task of early detection of anastomosis leakage (AL), a severe complication after elective surgery for colorectal cancer (CRC) surgery, using free text extracted from EHRs. We use a bag-of-words model to investigate the potential for feature selection strategies. The purpose is earlier detection of AL and prediction of AL with data generated in the EHR before the actual complication occur. Due to the high dimensionality of the data, we derive feature selection strategies using the robust support vector machine linear maximum margin classifier, by investigating: 1) a simple statistical criterion (leave-one-out-based test); 2) an intensive-computation statistical criterion (Bootstrap resampling); and 3) an advanced statistical criterion (kernel entropy). Results reveal a discriminatory power for early detection of complications after CRC (sensitivity 100%; specificity 72%). These results can be used to develop prediction models, based on EHR data, that can support surgeons and patients in the preoperative decision making phase.
Cristina Soguero-Ruíz, Kristian Hindberg, José Luis Rojo-Álvarez, Stein Olav Skrøvseth, Fred Godtliebsen, Kim Mortensen, Arthur Revhaug, Rolv-Ole Lindsetmo, Knut Magne Augestad, Robert Jenssen
IEEE J. Biomed. Health Informatics1
2015 Data-driven Temporal Prediction of Surgical Site Infection
Cristina Soguero-Ruíz, Fei Wang 0001, Robert Jenssen, Knut Magne Augestad, José Luis Rojo-Álvarez, I. Mora-Jiménez, Rolv-Ole Lindsetmo, Stein Olav Skrøvseth
AMIA1
2012 Deal Effect Curve and Promotional Models - Using Machine Learning and Bootstrap Resampling Test
Cristina Soguero-Ruíz, Francisco Javier Gimeno-Blanes, I. Mora-Jiménez, María Pilar Martínez-Ruiz, José Luis Rojo-Álvarez
ICPRAM (2)1
2012 On the differential benchmarking of promotional efficiency with machine learning modeling (I): Principles and statistical comparison
Cristina Soguero-Ruíz, Francisco Javier Gimeno-Blanes, I. Mora-Jiménez, María Pilar Martínez-Ruiz, José Luis Rojo-Álvarez
Expert Syst. Appl.1
2012 On the differential benchmarking of promotional efficiency with machine learning modelling (II): Practical applications
Cristina Soguero-Ruíz, Francisco Javier Gimeno-Blanes, I. Mora-Jiménez, María Pilar Martínez-Ruiz, José Luis Rojo-Álvarez
Expert Syst. Appl.1