VLDB 2026 Research / reviewers in the wild / expert
Spiros C. Denaxas
dblp:186/9095 · also Spiridon C. Denaxas
· DBLP profile ↗
16ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0001-9612-7791ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorSystems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluation of trajectory analysis for disease risk assessment: a scoping reviewabstractOBJECTIVES: Increasingly, structured longitudinal electronic health records (EHRs) are being harnessed to predict risk of having present but as yet undetected disease by analyzing "patient trajectories." Trajectory studies explore clinical event associations, characterize disease trajectories, and enhance risk prediction. This scoping review assesses study characteristics and objectives, identifies model types, and appraises model performance and reporting. MATERIALS AND METHODS: We conducted a scoping review, focused on a PubMed and Web of Science search for studies using temporal EHR sequences to identify disease signatures or predict disease presence. RESULTS: We identified 62 studies. Statistical methods, such as testing temporal associations were primarily used for clustering, while deep learning models focused on outcome prediction. Sixty-five percent of studies used secondary care data, with the most common outcomes being disease agnostic (39%) and cardiovascular disease (20%). Forty-eight studies aimed at risk prediction, with 50% comparing trajectory-based models to static baselines. Among 31 studies reporting area under the curve (AUC), temporal models showed moderate performance gains (relative/absolute AUC: median 5.7%/4.2%, range -2.6% to 58.9%/-2.3% to 33.0%). DISCUSSION: Trajectory studies are increasing in volume, but lacking in application to primary care datasets, a diverse set of diseases, external validation, and consideration of clinical applicability. CONCLUSION: While the field's nascency hinders firm conclusions, there are promising results across a range of model types and objectives. Continued research from diverse perspectives will help determine whether this growing field can deliver meaningful clinical benefits. Freya Pollington, Spiros C. Denaxas, Kezhi Li, Johan Hilge Thygesen, Georgios Lyratzopoulos, Becky White |
J. Am. Medical Informatics Assoc. | 2 |
| 2024 | Improving reporting standards for phenotyping algorithm in biomedical research: 5 fundamental dimensionsabstractINTRODUCTION: Phenotyping algorithms enable the interpretation of complex health data and definition of clinically relevant phenotypes; they have become crucial in biomedical research. However, the lack of standardization and transparency inhibits the cross-comparison of findings among different studies, limits large scale meta-analyses, confuses the research community, and prevents the reuse of algorithms, which results in duplication of efforts and the waste of valuable resources. RECOMMENDATIONS: Here, we propose five independent fundamental dimensions of phenotyping algorithms-complexity, performance, efficiency, implementability, and maintenance-through which researchers can describe, measure, and deploy any algorithms efficiently and effectively. These dimensions must be considered in the context of explicit use cases and transparent methods to ensure that they do not reflect unexpected biases or exacerbate inequities. Wei-Qi Wei, Robb Rowley, Angela M. Wood, Jacqueline Macarthur, Peter J. Embí, Spiros C. Denaxas |
J. Am. Medical Informatics Assoc. | 6 |
| 2023 | Differentiable sorting for censored time-to-event dataabstractSurvival analysis is a crucial semi-supervised task in machine learning with significant real-world applications, especially in healthcare. The most common approach to survival analysis, Cox’s partial likelihood, can be interpreted as a ranking model optimized on a lower bound of the concordance index. We follow these connections further, with listwise ranking losses that allow for a relaxation of the pairwise independence assumption. Given the inherent transitivity of ranking, we explore differentiable sorting networks as a means to introduce a stronger transitive inductive bias during optimization. Despite their potential, current differentiable sorting methods cannot account for censoring, a crucial aspect of many real-world datasets. We propose a novel method, Diffsurv, to overcome this limitation by extending differentiable sorting methods to handle censored tasks. Diffsurv predicts matrices of possible permutations that accommodate the label uncertainty introduced by censored samples. Our experiments reveal that Diffsurv outperforms established baselines in various simulated and real-world risk prediction scenarios. Furthermore, we demonstrate the algorithmic advantages of Diffsurv by presenting a novel method for top-k risk prediction that surpasses current methods. Andre Vauvelle, Benjamin Wild, Roland Eils, Spiros C. Denaxas |
NeurIPS | 4 |
| 2023 | Translating and evaluating historic phenotyping algorithms using SNOMED CTabstractOBJECTIVE: Patient phenotype definitions based on terminologies are required for the computational use of electronic health records. Within UK primary care research databases, such definitions have typically been represented as flat lists of Read terms, but Systematized Nomenclature of Medicine-Clinical Terms (SNOMED CT) (a widely employed international reference terminology) enables the use of relationships between concepts, which could facilitate the phenotyping process. We implemented SNOMED CT-based phenotyping approaches and investigated their performance in the CPRD Aurum primary care database. MATERIALS AND METHODS: We developed SNOMED CT phenotype definitions for 3 exemplar diseases: diabetes mellitus, asthma, and heart failure, using 3 methods: "primary" (primary concept and its descendants), "extended" (primary concept, descendants, and additional relations), and "value set" (based on text searches of term descriptions). We also derived SNOMED CT codelists in a semiautomated manner for 276 disease phenotypes used in a study of health across the lifecourse. Cohorts selected using each codelist were compared to "gold standard" manually curated Read codelists in a sample of 500 000 patients from CPRD Aurum. RESULTS: SNOMED CT codelists selected a similar set of patients to Read, with F1 scores exceeding 0.93, and age and sex distributions were similar. The "value set" and "extended" codelists had slightly greater recall but lower precision than "primary" codelists. We were able to represent 257 of the 276 phenotypes by a single concept hierarchy, and for 135 phenotypes, the F1 score was greater than 0.9. CONCLUSIONS: SNOMED CT provides an efficient way to define disease phenotypes, resulting in similar patient populations to manually curated codelists. Musaab Elkheder, Arturo Gonzalez-Izquierdo, Muhammad Qummer Ul Arfeen, Valerie Kuan, R. Tom Lumbers, Spiros C. Denaxas, Anoop D. Shah |
J. Am. Medical Informatics Assoc. | 6 |
| 2022 | Transforming and evaluating the UK Biobank to the OMOP Common Data Model for COVID-19 research and beyondabstractOBJECTIVE: The coronavirus disease 2019 (COVID-19) pandemic has demonstrated the value of real-world data for public health research. International federated analyses are crucial for informing policy makers. Common data models (CDMs) are critical for enabling these studies to be performed efficiently. Our objective was to convert the UK Biobank, a study of 500 000 participants with rich genetic and phenotypic data to the Observational Medical Outcomes Partnership (OMOP) CDM. MATERIALS AND METHODS: We converted UK Biobank data to OMOP CDM v. 5.3. We transformedparticipant research data on diseases collected at recruitment and electronic health records (EHRs) from primary care, hospitalizations, cancer registrations, and mortality from providers in England, Scotland, and Wales. We performed syntactic and semantic validations and compared comorbidities and risk factors between source and transformed data. RESULTS: We identified 502 505 participants (3086 with COVID-19) and transformed 690 fields (1 373 239 555 rows) to the OMOP CDM using 8 different controlled clinical terminologies and bespoke mappings. Specifically, we transformed self-reported noncancer illnesses 946 053 (83.91% of all source entries), cancers 37 802 (70.81%), medications 1 218 935 (88.25%), and prescriptions 864 788 (86.96%). In EHR, we transformed 13 028 182 (99.95%) hospital diagnoses, 6 465 399 (89.2%) procedures, 337 896 333 primary care diagnoses (CTV3, SNOMED-CT), 139 966 587 (98.74%) prescriptions (dm+d) and 77 127 (99.95%) deaths (ICD-10). We observed good concordance across demographic, risk factor, and comorbidity factors between source and transformed data. DISCUSSION AND CONCLUSION: Our study demonstrated that the OMOP CDM can be successfully leveraged to harmonize complex large-scale biobanked studies combining rich multimodal phenotypic data. Our study uncovered several challenges when transforming data from questionnaires to the OMOP CDM which require further research. The transformed UK Biobank resource is a valuable tool that can enable federated research, like COVID-19 studies. Václav Papez, Maxim Moinat, Erica A. Voss, Sofia Bazakou, Anne Van Winzum, Alessia Peviani, Stefan Payralbe, Elena Garcia Lara, Michael Kallfelz, Folkert W. Asselbergs, Daniel Prieto-Alhambra, Richard J. B. Dobson, Spiros C. Denaxas |
J. Am. Medical Informatics Assoc. | 13 |
| 2021 | Mapping the Read2/CTV3 controlled clinical terminologies to Phecodes in UK Biobank primary care electronic health records: implementation and evaluation
Spiros C. Denaxas, QiPing Feng, Ghazaleh Fatemifar, Lisa Bastarache, Vern Eric Kerchberger, Aroon D. Hingorani, R. Tom Lumbers, Josh F. Peterson, Wei-Qi Wei, Harry Hemingway |
AMIA | 1 |
| 2019 | UK phenomics platform for developing and validating electronic health record phenotypes: CALIBERabstractOBJECTIVE: Electronic health records (EHRs) are a rich source of information on human diseases, but the information is variably structured, fragmented, curated using different coding systems, and collected for purposes other than medical research. We describe an approach for developing, validating, and sharing reproducible phenotypes from national structured EHR in the United Kingdom with applications for translational research. MATERIALS AND METHODS: We implemented a rule-based phenotyping framework, with up to 6 approaches of validation. We applied our framework to a sample of 15 million individuals in a national EHR data source (population-based primary care, all ages) linked to hospitalization and death records in England. Data comprised continuous measurements (for example, blood pressure; medication information; coded diagnoses, symptoms, procedures, and referrals), recorded using 5 controlled clinical terminologies: (1) read (primary care, subset of SNOMED-CT [Systematized Nomenclature of Medicine Clinical Terms]), (2) International Classification of Diseases-Ninth Revision and Tenth Revision (secondary care diagnoses and cause of mortality), (3) Office of Population Censuses and Surveys Classification of Surgical Operations and Procedures, Fourth Revision (hospital surgical procedures), and (4) DM+D prescription codes. RESULTS: Using the CALIBER phenotyping framework, we created algorithms for 51 diseases, syndromes, biomarkers, and lifestyle risk factors and provide up to 6 validation approaches. The EHR phenotypes are curated in the open-access CALIBER Portal (https://www.caliberresearch.org/portal) and have been used by 40 national and international research groups in 60 peer-reviewed publications. CONCLUSIONS: We describe a UK EHR phenomics approach within the CALIBER EHR data platform with initial evidence of validity and use, as an important step toward international use of UK EHR data for health research. Spiros C. Denaxas, Arturo Gonzalez-Izquierdo, Kenan Direk, Natalie K. Fitzpatrick, Ghazaleh Fatemifar, Amitava Banerjee, Richard J. B. Dobson, Laurence J. Howe, Valerie Kuan, R. Tom Lumbers, Laura Pasea, Riyaz S. Patel, Anoop D. Shah, Aroon D. Hingorani, Cathie Sudlow, Harry Hemingway |
J. Am. Medical Informatics Assoc. | 1 |
| 2019 | Exploring hybrid parallel systems for probabilistic record linkage
Murilo Boratto, Pedro Alonso 0002, Clícia Pinto, Pedro Melo, Marcos E. Barreto, Spiros C. Denaxas |
J. Supercomput. | 6 |
| 2018 | On the Accuracy and Scalability of Probabilistic Data Linkage Over the Brazilian 114 Million CohortabstractData linkage refers to the process of identifying and linking records that refer to the same entity across multiple heterogeneous data sources. This method has been widely utilized across scientific domains, including public health where records from clinical, administrative, and other surveillance databases are aggregated and used for research, decision making, and assessment of public policies. When a common set of unique identifiers does not exist across sources, probabilistic linkage approaches are used to link records using a combination of attributes. These methods require a careful choice of comparison attributes as well as similarity metrics and cutoff values to decide if a given pair of records matches or not and for assessing the accuracy of the results. In large, complex datasets, linking and assessing accuracy can be challenging due to the volume and complexity of the data, the absence of a gold standard, and the challenges associated with manually reviewing a very large number of record matches. In this paper, we present AtyImo, a hybrid probabilistic linkage tool optimized for high accuracy and scalability in massive data sets. We describe the implementation details around anonymization, blocking, deterministic and probabilistic linkage, and accuracy assessment. We present results from linking a large population-based cohort of 114 million individuals in Brazil to public health and administrative databases for research. In controlled and real scenarios, we observed high accuracy of results: 93%-97% true matches. In terms of scalability, we present AtyImo's ability to link the entire cohort in less than nine days using Spark and scaling up to 20 million records in less than 12s over heterogeneous (CPU+GPU) architectures. Robespierre Pita, Clícia Pinto, Samila Sena, Rosemeire L. Fiaccone, Leila Amorim, Sandra Reis, Mauricio Barreto, Spiros C. Denaxas, Marcos E. Barreto |
IEEE J. Biomed. Health Informatics | 8 |
| 2017 | Comparing and Contrasting A Priori and A Posteriori Generalizability Assessment of Clinical Trials on Type 2 Diabetes Mellitus
Zhe He 0001, Arturo Gonzalez-Izquierdo, Spiros C. Denaxas, Andrei Sura, Yi Guo 0005, William R. Hogan, Elizabeth Shenkman, Jiang Bian 0001 |
AMIA | 3 |
| 2017 | Evaluation of Semantic Web Technologies for Storing Computable Definitions of Electronic Health Records Phenotyping Algorithms
Václav Papez, Spiros C. Denaxas, Harry Hemingway |
AMIA | 2 |
| 2017 | Methods for Enhancing the Reproducibility of Observational Research Using Electronic Health Records: Preliminary Findings from the CALIBER ResourceabstractThe ability of external investigators to reproduce published scientific findings is critical for the evaluation and validation of health research by the wider community. However, a substantial proportion of health research using electronic health records, data collected and generated during routine clinical care, potentially cannot reproduced. With the complexity, volume and variety of electronic health records made available for research steadily increasing, it is critical to ensure that findings from such data are reproducible and replicable by researchers. In this paper, we present some preliminary findings on how a series of methods and tools utilized in adjunct scientific disciplines can be used to enhance the reproducibility of research using electronic health records. Spiros C. Denaxas, Arturo Gonzalez-Izquierdo, Maria Pikoula, Kenan Direk, Natalie K. Fitzpatrick, Harry Hemingway, Liam Smeeth |
CBMS | 1 |
| 2017 | Evaluating OpenEHR for Storing Computable Representations of Electronic Health Record Phenotyping AlgorithmsabstractElectronic Health Records (EHR) are data generated during routine clinical care. EHR offer researchers unprecedented phenotypic breadth and depth and have the potential to accelerate the pace of precision medicine at scale. A main EHR use-case is creating phenotyping algorithms to define disease status, onset and severity. Currently, no common machine-readable standard exists for defining phenotyping algorithms which often are stored in human-readable formats. As a result, the translation of algorithms to implementation code is challenging and sharing across the scientific community is problematic. In this paper, we evaluate openEHR, a formal EHR data specification, for computable representations of EHR phenotyping algorithms. Václav Papez, Spiros C. Denaxas, Harry Hemingway |
CBMS | 2 |
| 2017 | Probabilistic Integration of Large Brazilian Socioeconomic and Clinical DatabasesabstractThe integration of disparate large and heterogeneous socioeconomic and clinical databases is considered essential to capture and model longitudinal and social aspects of diseases. However, such integration is challenging: databases are stored in disparate locations, make use of different identifiers, have variable data quality, record information in bespoke purpose-specific formats and have different levels of metadata. Novel computational methods are required to integrate them and enable their statistical analyses for epidemiological research purposes. In this paper, we describe a probabilistic approach for constructing a very large population-based cohort comprised of 114 million individuals using linkages between clinical databases from the National Health System and administrative databases from governmental social programmes. We present our data integration model for creating data marts (epidemiological data) and discuss our evaluation results in controlled and uncontrolled scenarios, which demonstrate that our model and tools achieve high accuracy (minimum of 91%) in different probabilistic data integration scenarios. Clícia Pinto, Robespierre Pita, George Caique Gouveia Barbosa, Bruno Rodrigues De Araújo, Juracy Bertoldo, Samila Sena, Sandra Reis, Rosemeire L. Fiaccone, Leila Amorim, Maria Yury Ichihara, Mauricio Barreto, Marcos E. Barreto, Spiros C. Denaxas |
CBMS | 13 |
| 2017 | A Machine Learning Trainable Model to Assess the Accuracy of Probabilistic Record Linkage
Robespierre Pita, Everton Mendonça, Sandra Reis, Marcos E. Barreto, Spiros C. Denaxas |
DaWaK | 5 |
| 2013 | Electronic Health Record Linkages for Translational Cardiovascular Research in Nearly 2 Million People - Clinical Disease Research Using Linked Bespoke Studies and Electronic Records (CALIBER)
Spiros C. Denaxas, Dipak Kalra, Eleni Rapsomaniki, Anoop D. Shah, Mar Pujades Rodriguez, Katherine Morley, Adam Timmis, Emily Herrett, Liam Smeeth, Harry Hemingway |
AMIA | 1 |