Carlos Sáez 0001

dblp:70/960 · also Carlos J. Saez · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0003-2678-8249ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Security and privacy · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 An end-to-end solution for out-of-hospital emergency medical dispatch triage based on multimodal and continual deep learning
Pablo Ferri, Carlos Sáez 0001, Antonio Felix de Castro, Purificación Sánchez-Cuesta, Juan Miguel García-Gómez
Artif. Intell. Medicine2
2022 Multi-PheWAS intersection approach to identify sex differences across comorbidities in 59 140 pediatric patients with autism spectrum disorder
abstract
OBJECTIVE: To identify differences related to sex and define autism spectrum disorder (ASD) comorbidities female-enriched through a comprehensive multi-PheWAS intersection approach on big, real-world data. Although sex difference is a consistent and recognized feature of ASD, additional clinical correlates could help to identify potential disease subgroups, based on sex and age. MATERIALS AND METHODS: We performed a systematic comorbidity analysis on 1860 groups of comorbidities exploring all spectrum of known disease, in 59 140 individuals (11 440 females) with ASD from 4 age groups. We explored ASD sex differences in 2 independent real-world datasets, across all potential comorbidities by comparing (1) females with ASD vs males with ASD and (2) females with ASD vs females without ASD. RESULTS: We identified 27 different comorbidities that appeared significantly more frequently in females with ASD. The comorbidities were mostly neurological (eg, epilepsy, odds ratio [OR] > 1.8, 3-18 years of age), congenital (eg, chromosomal anomalies, OR > 2, 3-18 years of age), and mental disorders (eg, intellectual disability, OR > 1.7, 6-18 years of age). Novel comorbidities included endocrine metabolic diseases (eg, failure to thrive, OR = 2.5, ages 0-2), digestive disorders (gastroesophageal reflux disease: OR = 1.7, 6-11 years of age; and constipation: OR > 1.6, 3-11 years of age), and sense organs (strabismus: OR > 1.8, 3-18 years of age). DISCUSSION: A multi-PheWAS intersection approach on real-world data as presented in this study uniquely contributes to the growing body of research regarding sex-based comorbidity analysis in ASD population. CONCLUSIONS: Our findings provide insights into female-enriched ASD comorbidities that are potentially important in diagnosis, as well as the identification of distinct comorbidity patterns influencing anticipatory treatment or referrals. The code is publicly available (https://github.com/hms-dbmi/sexDifferenceInASD).
Alba Gutiérrez-Sacristán, Carlos Sáez 0001, Carlos De Niz, Niloofar Jalali, Thomas N. Desain, Ranjay Kumar, Joany M. Zachariasse, Kathe P. Fox, Nathan P. Palmer, Isaac S. Kohane, Paul Avillach
J. Am. Medical Informatics Assoc.2
2022 Multisource and temporal variability in Portuguese hospital administrative datasets: Data quality implications
abstract
BACKGROUND: Unexpected variability across healthcare datasets may indicate data quality issues and thereby affect the credibility of these data for reutilization. No gold-standard reference dataset or methods for variability assessment are usually available for these datasets. In this study, we aim to describe the process of discovering data quality implications by applying a set of methods for assessing variability between sources and over time in a large hospital database. METHODS: We described and applied a set of multisource and temporal variability assessment methods in a large Portuguese hospitalization database, in which variation in condition-specific hospitalization ratios derived from clinically coded data were assessed between hospitals (sources) and over time. We identified condition-specific admissions using the Clinical Classification Software (CCS), developed by the Agency of Health Care Research and Quality. A Statistical Process Control (SPC) approach based on funnel plots of condition-specific standardized hospitalization ratios (SHR) was used to assess multisource variability, whereas temporal heat maps and Information-Geometric Temporal (IGT) plots were used to assess temporal variability by displaying temporal abrupt changes in data distributions. Results were presented for the 15 most common inpatient conditions (CCS) in Portugal. MAIN FINDINGS: Funnel plot assessment allowed the detection of several outlying hospitals whose SHRs were much lower or higher than expected. Adjusting SHR for hospital characteristics, beyond age and sex, considerably affected the degree of multisource variability for most diseases. Overall, probability distributions changed over time for most diseases, although heterogeneously. Abrupt temporal changes in data distributions for acute myocardial infarction and congestive heart failure coincided with the periods comprising the transition to the International Classification of Diseases, 10th revision, Clinical Modification, whereas changes in the Diagnosis-Related Groups software seem to have driven changes in data distributions for both acute myocardial infarction and liveborn admissions. The analysis of heat maps also allowed the detection of several discontinuities at hospital level over time, in some cases also coinciding with the aforementioned factors. CONCLUSIONS: This paper described the successful application of a set of reproducible, generalizable and systematic methods for variability assessment, including visualization tools that can be useful for detecting abnormal patterns in healthcare data, also addressing some limitations of common approaches. The presented method for multisource variability assessment is based on SPC, which is an advantage considering the lack of gold standard for such process. Properly controlling for hospital characteristics and differences in case-mix for estimating SHR is critical for isolating data quality-related variability among data sources. The use of IGT plots provides an advantage over common methods for temporal variability assessment due its suitability for multitype and multimodal data, which are common characteristics of healthcare data. The novelty of this work is the use of a set of methods to discover new data quality insights in healthcare data.
Júlio Souza, Ismael Caballero 0001, João Vasco Santos, Mariana Lobo, Andreia Pinto, João Viana, Carlos Sáez 0001, Fernando Lopes 0003, Alberto Freitas
J. Biomed. Informatics7
2021 Measuring Variability in Acute Myocardial Infarction Coding Using a Statistical Process Control and Probabilistic Temporal Data Quality Control Approaches
Júlio Souza, Ismael Caballero 0001, João Vasco Santos, Mariana Lobo, Andreia Pinto, João Viana, Carlos Sáez 0001, Alberto Freitas
WorldCIST (2)7
2021 Deep ensemble multitask classification of emergency medical call incidents combining multimodal data improves emergency medical dispatch
Pablo Ferri, Carlos Sáez 0001, Antonio Felix de Castro, Javier Juan-Albarracín, Vicent Blanes-Selva, Purificación Sánchez-Cuesta, Juan Miguel García-Gómez
Artif. Intell. Medicine2
2021 Potential limitations in COVID-19 machine learning due to data source variability: A case study in the nCov2019 dataset
abstract
OBJECTIVE: The lack of representative coronavirus disease 2019 (COVID-19) data is a bottleneck for reliable and generalizable machine learning. Data sharing is insufficient without data quality, in which source variability plays an important role. We showcase and discuss potential biases from data source variability for COVID-19 machine learning. MATERIALS AND METHODS: We used the publicly available nCov2019 dataset, including patient-level data from several countries. We aimed to the discovery and classification of severity subgroups using symptoms and comorbidities. RESULTS: Cases from the 2 countries with the highest prevalence were divided into separate subgroups with distinct severity manifestations. This variability can reduce the representativeness of training data with respect the model target populations and increase model complexity at risk of overfitting. CONCLUSIONS: Data source variability is a potential contributor to bias in distributed research networks. We call for systematic assessment and reporting of data source variability and data quality in COVID-19 data sharing, as key information for reliable and generalizable machine learning.
Carlos Sáez 0001, Nekane Romero-Garcia, J. Alberto Conejero, Juan Miguel García-Gómez
J. Am. Medical Informatics Assoc.1
2021 Predicting morbidity by local similarities in multi-scale patient trajectories
Lucia A. Carrasco-Ribelles, Jose Ramón Pardo-Mas, Salvador Tortajada, Carlos Sáez 0001, Bernardo Valdivieso, Juan Miguel García-Gómez
J. Biomed. Informatics4
2017 Discovering Data Source Stability Patterns in Biomedical Repositories Based on Simplicial Projections from Probability Distribution Distances
abstract
The degree of homogeneity of statistical distributions among data sources is a critical issue when reusing data of Integrated Data Repositories (IDR). Evaluating this data source stability is of utmost importance in order to ensure a confident data reuse. This work tackles the task of discovering and classifying patterns among the statistical distributions of multiple sources in IDRs, by means of a novel approach based on simplicial projections from probability distribution distances, combined with Density-based spatial clustering of applications with noise (DBSCAN). The results on the evaluated 20 public repositories support the existence of four main data source stability patterns in biomedical repositories: the global stability pattern (GSP), the local stability pattern (LSP), the sparse stability pattern (SSP) and the instability pattern (IP).
Pablo Ferri, Carlos Sáez 0001, Juan Miguel García-Gómez
CBMS2
2016 Applying probabilistic temporal and multisite data quality control methods to a public health mortality registry in Spain: a systematic approach to quality control of repositories
abstract
OBJECTIVE: To assess the variability in data distributions among data sources and over time through a case study of a large multisite repository as a systematic approach to data quality (DQ). MATERIALS AND METHODS: Novel probabilistic DQ control methods based on information theory and geometry are applied to the Public Health Mortality Registry of the Region of Valencia, Spain, with 512 143 entries from 2000 to 2012, disaggregated into 24 health departments. The methods provide DQ metrics and exploratory visualizations for (1) assessing the variability among multiple sources and (2) monitoring and exploring changes with time. The methods are suited to big data and multitype, multivariate, and multimodal data. RESULTS: The repository was partitioned into 2 probabilistically separated temporal subgroups following a change in the Spanish National Death Certificate in 2009. Punctual temporal anomalies were noticed due to a punctual increment in the missing data, along with outlying and clustered health departments due to differences in populations or in practices. DISCUSSION: Changes in protocols, differences in populations, biased practices, or other systematic DQ problems affected data variability. Even if semantic and integration aspects are addressed in data sharing infrastructures, probabilistic variability may still be present. Solutions include fixing or excluding data and analyzing different sites or time periods separately. A systematic approach to assessing temporal and multisite variability is proposed. CONCLUSION: Multisite and temporal variability in data distributions affects DQ, hindering data reuse, and an assessment of such variability should be a part of systematic DQ procedures.
Carlos Sáez 0001, Oscar Zurriaga, Jordi Pérez-Panadés, Inma Melchor, Montserrat Robles, Juan Miguel García-Gómez
J. Am. Medical Informatics Assoc.1
2015 Probabilistic change detection and visualization methods for the assessment of temporal stability in biomedical data quality
Carlos Sáez 0001, Pedro Pereira Rodrigues, João Gama 0001, Montserrat Robles, Juan Miguel García-Gómez
Data Min. Knowl. Discov.1
2008 A Security Model and its Application to a Distributed Decision Support System for Healthcare
abstract
A distributed decision support system involving multiple clinical centres is crucial to the diagnosis of rare diseases. Although sharing of valid diagnosed cases can facilitate later decision making, possibly from geographically different centres, the released information could reveal patient privacy if it is not properly protected. Clinical centres may have to impose their distinct regulations and rules that govern the use of their data externally. The collaboration of centres, therefore, must respect the collective policies and ideally, serve users the most appropriate and useful resources possible in the system according to the past experience. In this way, the system’s value is entrusted and even elevated through continuous collaboration. We present in this paper a link-anonymised data scheme and in addition to that, a security model that together enforce privacy data security and secure resource access for distributed clinical centres. Our illustration of the approach involves a prototype medical decision support system, HealthAgents, for brain tumour diagnosis.
Liang Xiao 0002, Javier Vicente, Carlos Sáez 0001, Andrew Peet, Alex Gibb, Paul H. Lewis, Srinandan Dasmahapatra, Madalina Croitoru, Horacio González-Vélez, Magí Lluch i Ariet, David Dupplaw
ARES3
2007 Conceptual Graphs Based Information Retrieval in HealthAgents
abstract
This paper focuses on the problem of representing, in a meaningful way, the knowledge involved in the HealthAgents project. Our work is motivated by the complexity of representing electronic healthcare records in a consistent manner. We present HADOM (HealthAgents domain ontology) which conceptualises the required HealthAgents information and propose describing the sources knowledge by the means of conceptual graphs (CGs). This allows to build upon the existing ontology permitting for modularity and flexibility. The novelty of our approach lies in the ease with which CGs can be placed above other formalisms and their potential for optimised querying and retrieval.
Madalina Croitoru, Bo Hu 0001, Srinandan Dasmahapatra, Paul H. Lewis, David Dupplaw, Alex Gibb, Margarida Julià-Sapé, Javier Vicente, Carlos Sáez 0001, Juan Miguel García-Gómez, Roman Roset, Francesc Estanyol, Xavier Rafael Palou, Mariola Mier
CBMS9
2007 An Adaptive Security Model for Multi-agent Systems and Application to a Clinical Trials Environment
abstract
We present in this paper an adaptive security model for Multi-agent systems. A security meta-model has been developed in which the traditional role concept has been extended. The new concept incorporates the need of both security management as used by role-based access control (RBAC) and agent functional behaviour in agent-oriented Software Engineering (AOSE). Our approach avoids weaknesses of traditional RBAC approaches and provides a practically usable security model for multi-agent systems (MAS). A unified role interaction model framework has been put forward that incorporates not only functional requirements but also security constraints in MAS. A security policy rule scheme has been used to express security requirements in relation to affective roles. The major contribution of the work is that little redevelopment effort will be required when security is to be engineered into the overall MAS architecture, hence minimising the impact of the security requirements changes to the MAS architecture. We illustrate the approach through its potential application in a clinical trial setting involving a prototype medical decision support system, HealthAgents.
Liang Xiao 0002, Andrew Peet, Paul H. Lewis, Srinandan Dashmapatra, Carlos Sáez 0001, Madalina Croitoru, Javier Vicente, Horacio González-Vélez, Magí Lluch i Ariet
COMPSAC (2)5