EDBT 2026 Demo / reviewers in the wild / expert
David J. Schlueter
dblp:250/7113
· DBLP profile ↗
9ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PheWAS analysis on large-scale biobank data with PheTKabstractSUMMARY: With the rapid growth of genetic data linked to electronic health record (EHR) data in huge cohorts, large-scale phenome-wide association study (PheWAS) have become powerful discovery tools in biomedical research. PheWAS is an analysis method to study phenotype associations utilizing longitudinal EHR data. Previous PheWAS packages were developed mostly with smaller datasets and with earlier PheWAS approaches. PheTK was designed to simplify analysis and efficiently handle biobank-scale data. PheTK uses multithreading and supports a full PheWAS workflow including extraction of data from OMOP databases and Hail matrix tables as well as PheWAS analysis for both phecode version 1.2 and phecodeX. Benchmarking results showed PheTK took 64% less time than the R PheWAS package to complete the same workflow. PheTK can be run locally or on cloud platforms such as the All of Us Researcher Workbench (All of Us) or the UK Biobank (UKB) Research Analysis Platform (RAP). AVAILABILITY AND IMPLEMENTATION: The PheTK package is freely available on the Python Package Index, on GitHub under GNU General Public License (GPL-3) at https://github.com/nhgritctran/PheTK, and on Zenodo, DOI 10.5281/zenodo.14217954, at https://doi.org/10.5281/zenodo.14217954. PheTK is implemented in Python and platform independent. Tam C. Tran, David J. Schlueter, Chenjie Zeng, Huan Mo, Robert J. Carroll, Joshua C. Denny |
Bioinform. | 2 |
| 2024 | Comparison of phenomic profiles in the All of Us Research Program against the US general population and the UK BiobankabstractIMPORTANCE: Knowledge gained from cohort studies has dramatically advanced both public and precision health. The All of Us Research Program seeks to enroll 1 million diverse participants who share multiple sources of data, providing unique opportunities for research. It is important to understand the phenomic profiles of its participants to conduct research in this cohort. OBJECTIVES: More than 280 000 participants have shared their electronic health records (EHRs) in the All of Us Research Program. We aim to understand the phenomic profiles of this cohort through comparisons with those in the US general population and a well-established nation-wide cohort, UK Biobank, and to test whether association results of selected commonly studied diseases in the All of Us cohort were comparable to those in UK Biobank. MATERIALS AND METHODS: We included participants with EHRs in All of Us and participants with health records from UK Biobank. The estimates of prevalence of diseases in the US general population were obtained from the Global Burden of Diseases (GBD) study. We conducted phenome-wide association studies (PheWAS) of 9 commonly studied diseases in both cohorts. RESULTS: This study included 287 012 participants from the All of Us EHR cohort and 502 477 participants from the UK Biobank. A total of 314 diseases curated by the GBD were evaluated in All of Us, 80.9% (N = 254) of which were more common in All of Us than in the US general population [prevalence ratio (PR) >1.1, P < 2 × 10-5]. Among 2515 diseases and phenotypes evaluated in both All of Us and UK Biobank, 85.6% (N = 2152) were more common in All of Us (PR >1.1, P < 2 × 10-5). The Pearson correlation coefficients of effect sizes from PheWAS between All of Us and UK Biobank were 0.61, 0.50, 0.60, 0.57, 0.40, 0.53, 0.46, 0.47, and 0.24 for ischemic heart diseases, lung cancer, chronic obstructive pulmonary disease, dementia, colorectal cancer, lower back pain, multiple sclerosis, lupus, and cystic fibrosis, respectively. DISCUSSION: Despite the differences in prevalence of diseases in All of Us compared to the US general population or the UK Biobank, our study supports that All of Us can facilitate rapid investigation of a broad range of diseases. CONCLUSION: Most diseases were more common in All of Us than in the general US population or the UK Biobank. Results of disease-disease association tests from All of Us are comparable to those estimated in another well-studied national cohort. Chenjie Zeng, David J. Schlueter, Tam C. Tran, Anav Babbar, Thomas Cassini, Lisa Bastarache, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 2 |
| 2023 | Systematic replication of smoking disease associations using survey responses and EHR data in the All of Us Research ProgramabstractOBJECTIVE: The All of Us Research Program (All of Us) aims to recruit over a million participants to further precision medicine. Essential to the verification of biobanks is a replication of known associations to establish validity. Here, we evaluated how well All of Us data replicated known cigarette smoking associations. MATERIALS AND METHODS: We defined smoking exposure as follows: (1) an EHR Smoking exposure that used International Classification of Disease codes; (2) participant provided information (PPI) Ever Smoking; and, (3) PPI Current Smoking, both from the lifestyle survey. We performed a phenome-wide association study (PheWAS) for each smoking exposure measurement type. For each, we compared the effect sizes derived from the PheWAS to published meta-analyses that studied cigarette smoking from PubMed. We defined two levels of replication of meta-analyses: (1) nominally replicated: which required agreement of direction of effect size, and (2) fully replicated: which required overlap of confidence intervals. RESULTS: PheWASes with EHR Smoking, PPI Ever Smoking, and PPI Current Smoking revealed 736, 492, and 639 phenome-wide significant associations, respectively. We identified 165 meta-analyses representing 99 distinct phenotypes that could be matched to EHR phenotypes. At P < .05, 74 were nominally replicated and 55 were fully replicated. At P < 2.68 × 10-5 (Bonferroni threshold), 58 were nominally replicated and 40 were fully replicated. DISCUSSION: Most phenotypes found in published meta-analyses associated with smoking were nominally replicated in All of Us. Both survey and EHR definitions for smoking produced similar results. CONCLUSION: This study demonstrated the feasibility of studying common exposures using All of Us data. David J. Schlueter, Lina M. Sulieman, Huan Mo, Jacob M. Keaton, Tracey Ferrara, Ariel Williams, Onajia J. Stubblefield, Chenjie Zeng, Tam C. Tran, Lisa Bastarache, Anav Babbar, Andrea H. Ramirez, Slavina Goleva, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 1 |
| 2022 | Impact of COVID-19 on mental health outcomes in the All of Us Research Program
Onajia J. Stubblefield, David J. Schlueter, Jacob Keaton, Ariel Williams, Slavina Goleva, Tracey Ferrara, Chenjie Zeng, Huan Mo, Joshua C. Denny |
AMIA | 2 |
| 2022 | Comparing Effect Sizes in Covid Positive Phenomic Profiles
Ariel Williams, David J. Schlueter, Jacob Keaton, Tracey Ferrara, Onajia J. Stubblefield, Kyle Webb, Slavina Goleva, Chenjie Zeng, Huan Mo, Thomas Cassini, Joshua C. Denny |
AMIA | 2 |
| 2021 | Systematic replication of smoking disease associations in the All of Us Research Program
David J. Schlueter, Lina M. Sulieman, Jacob M. Keaton, Tracey Ferrara, Kyle Webb, Ariel Williams, Francis Ratsimbazafy, Lisa Bastarache, Andrea H. Ramirez, Joshua C. Denny |
AMIA | 1 |
| 2021 | Comparing the Phenomic Profile of All of Us Research Program and National COVID Cohort Collaborative
Kyle P. Webb, David J. Schlueter, Jacob Keaton, Tracey Ferrara, Ariel Williams, Joshua C. Denny |
AMIA | 2 |
| 2020 | The All of Us Research Program Researcher Workbench Phenotype Library: Five Disease Implementations
Izabelle P. Humes, Roxana Loperena-Cortes, Melissa A. Basford, Kelsey R. Mayo, Joseph DiPaolo, David J. Schlueter, Wei-Qi Wei, Robert J. Carroll, David Glazer, Paul A. Harris, Anthony A. Philippakis, Dan M. Roden, Andrea H. Ramirez |
AMIA | 6 |
| 2019 | Detecting time-evolving phenotypic topics via tensor factorization on electronic health records: Cardiovascular disease case study
Juan Zhao 0003, David J. Schlueter, Patrick Wu, Vern Eric Kerchberger, S. Trent Rosenbloom, Quinn Stanton Wells, QiPing Feng, Joshua C. Denny, Wei-Qi Wei |
J. Biomed. Informatics | 3 |