VLDB 2026 Research / reviewers in the wild / expert
David Chushig-Muzo
dblp:272/0816
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0001-5585-2305ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AI-Based Data Augmentation for Cleft Lip Assessment: Generating Realistic Synthetic Videos to Improve Clinical Outcomes
Julián Puga, Malena Loza, David Chushig-Muzo, Luis Bote-Curiel, Felipe Grijalva |
IDEAL (1) | 3 |
| 2025 | A Comparative Study of Machine Learning Models for Two-Tier Android Malware Classification with Dynamic Behavioral Analysis
Felipe Grijalva, David Chushig-Muzo, Luis Bote-Curiel, Malena Loza |
IDEAL (1) | 3 |
| 2025 | Glucostats: an efficient Python library for glucose time series feature extraction and visual analysisabstractBACKGROUND: The advancement of technology and continuous glucose monitoring (CGM) systems has introduced several computational and technical challenges for clinicians and researchers. The growing volume of CGM data necessitates the development of efficient computational tools capable of handling and processing this information effectively. This paper introduces GlucoStats, an open-source and multi-processing Python library designed for efficient computation and visualization of a comprehensive set of glucose metrics derived from CGM. It simplifies the traditionally time-consuming and error-prone process of manual CGM metrics calculation, making it a valuable tool for both clinical and research applications. RESULTS: Its modular design ensures easy integration into predefined workflows, while its user-friendly interface and extensive documentation make it accessible to a broad audience, including clinicians and researchers. GlucoStats offers several key features: (i) window-based time series analysis, enabling time series division into smaller 'windows' for detailed temporal analysis, particularly beneficial for CGM data; (ii) advanced visualization tools, providing intuitive, high-quality visualizations that facilitate pattern recognition, trend analysis, and anomaly detection in CGM data; (iii) parallelization, leveraging parallel computing to efficiently handle large CGM datasets by distributing computations across multiple processors; and (iv) scikit-learn compatibility, adhering to the standardized interface of scikit-learn to allow an easy integration into machine learning pipelines for end-to-end analysis. CONCLUSIONS: GlucoStats demonstrates high efficiency in processing large-scale medical datasets in minimal time. Its modular design enables easy customization and extension, making it adaptable to diverse research and clinical needs. By offering precise CGM data analysis and user-friendly visualization tools, it serves both technical researchers and non-technical users, such as physicians and patients, with practical and research-driven applications. Pablo Peiro-Corbacho, Francisco J. Lara-Abelenda, David Chushig-Muzo, Ana M. Wägner, Conceição Granja, Cristina Soguero-Ruíz |
BMC Bioinform. | 3 |
| 2025 | Interpretable and multimodal fusion methodology to predict severe hypoglycemia in adults with type 1 diabetesabstractType 1 diabetes (T1D) causes insulin deficiency and exogenous therapy is required for maintaining targeted glucose levels. Hypoglycemia is the most frequent side effect of insulin, being severe hypoglycemia (SH) one of the most critical hazards with a range of life-threatening consequences. Artificial intelligence (AI) and multimodal fusion have boosted predictive performance in different domains. This study aims to evaluate the effectiveness of early fusion (EF) and late fusion (LF) approaches for predicting SH, to create a methodology capable of achieving robust results in datasets with a low number of samples for predicting SH and to characterize the risk factors involved in the SH onset using explainable AI (XAI). Data from a case-control study comprising adults over 60 years with T1D and with diabetes duration of 20 years were used and three types of modalities were considered: (1) continuous glucose monitoring data (time series); (2) clinical codes (text); and (3) surveys related to fear, unawareness, depression, and cognitive tests (tabular data). The results revealed that EF outperformed models trained with single-modality data by 5.8%, with an area under the receiver operating characteristic curve of 0.779. XAI techniques helped to discover that features related to fear and unawareness are mainly associated with SH. Our study introduced an interpretable and multimodal methodology capable of predicting the occurrence of SH in adults with T1D in the next year. Our interpretable methodology contributes to predicting SH and identifying related key factors, thus preventing SH complications and improving patient’s quality of life. • Multimodal fusion approaches improve the prediction of severe hypoglycemia. • Interpretability techniques lead to identifying key factors for severe hypoglycemia. • People with severe hypoglycemia show a high frequency of cardiovascular diseases. • Impaired hypoglycemia awareness elevates severe hypoglycemia risk. Francisco J. Lara-Abelenda, David Chushig-Muzo, Ana M. Wägner, Maryam Tayefi, Cristina Soguero-Ruíz |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Transfer learning for a tabular-to-image approach: A case study for cardiovascular disease predictionabstractOBJECTIVE: Machine learning (ML) models have been extensively used for tabular data classification but recent works have been developed to transform tabular data into images, aiming to leverage the predictive performance of convolutional neural networks (CNNs). However, most of these approaches fail to convert data with a low number of samples and mixed-type features. This study aims: to evaluate the performance of the tabular-to-image method named low mixed-image generator for tabular data (LM-IGTD); and to assess the effectiveness of transfer learning and fine-tuning for improving predictions on tabular data. METHODS: We employed two public tabular datasets with patients diagnosed with cardiovascular diseases (CVDs): Framingham and Steno. First, both datasets were transformed into images using LM-IGTD. Then, Framingham, which contains a larger set of samples than Steno, is used to train CNN-based models. Finally, we performed transfer learning and fine-tuning using the pre-trained CNN on the Steno dataset to predict CVD risk. RESULTS: The CNN-based model with transfer learning achieved the highest AUCORC in Steno (0.855), outperforming ML models such as decision trees, K-nearest neighbors, least absolute shrinkage and selection operator (LASSO) support vector machine and TabPFN. This approach improved accuracy by 2% over the best-performing traditional model, TabPFN. CONCLUSION: To the best of our knowledge, this is the first study that evaluates the effectiveness of applying transfer learning and fine-tuning to tabular data using tabular-to-image approaches. Through the use of CNNs' predictive capabilities, our work also advances the diagnosis of CVD by providing a framework for early clinical intervention and decision-making support. Francisco J. Lara-Abelenda, David Chushig-Muzo, Pablo Peiro-Corbacho, Vanesa Gómez-Martínez, Ana M. Wägner, Conceição Granja, Cristina Soguero-Ruíz |
J. Biomed. Informatics | 2 |
| 2022 | Characterizing Cardiovascular Risk Through Unsupervised and Interpretable Techniques
Hugo Calero-Díaz, David Chushig-Muzo, Cristina Soguero-Ruíz |
IDEAL | 2 |
| 2021 | Interpreting clinical latent representations using autoencoders and probabilistic modelsabstractElectronic health records (EHRs) are a valuable data source that, in conjunction with deep learning (DL) methods, have provided important outcomes in different domains, contributing to supporting decision-making. Owing to the remarkable advancements achieved by DL-based models, autoencoders (AE) are becoming extensively used in health care. Nevertheless, AE-based models are based on nonlinear transformations, resulting in black-box models leading to a lack of interpretability, which is vital in the clinical setting. To obtain insights from AE latent representations, we propose a methodology by combining probabilistic models based on Gaussian mixture models and hierarchical clustering supported by Kullback-Leibler divergence. To validate the methodology from a clinical viewpoint, we used real-world data extracted from EHRs of the University Hospital of Fuenlabrada (Spain). Records were associated with healthy and chronic hypertensive and diabetic patients. Experimental outcomes showed that our approach can find groups of patients with similar health conditions by identifying patterns associated with diagnosis and drug codes. This work opens up promising opportunities for interpreting representations obtained by the AE-based model, bringing some light to the decision-making process made by clinical experts in daily practice. David Chushig-Muzo, Cristina Soguero-Ruíz, Pablo de Miguel-Bohoyo, I. Mora-Jiménez |
Artif. Intell. Medicine | 1 |