EDBT 2026 Demo / reviewers in the wild / expert
Bhramar Mukherjee
dblp:96/8505
· DBLP profile ↗
9ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0003-0118-4561ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Generating synthetic electronic health record data: a methodological scoping review with benchmarking on phenotype data and open-source softwareabstractOBJECTIVES: To conduct a scoping review (ScR) of existing approaches for synthetic Electronic Health Records (EHR) data generation, to benchmark major methods, and to provide an open-source software and offer recommendations for practitioners. MATERIALS AND METHODS: We search three academic databases for our scoping review. Methods are benchmarked on open-source EHR datasets, Medical Information Mart for Intensive Care III and IV (MIMIC-III/IV). Seven existing methods covering major categories and two baseline methods are implemented and compared. Evaluation metrics concern data fidelity, downstream utility, privacy protection, and computational cost. RESULTS: Forty-eight studies are identified and classified into five categories. Seven open-source methods covering all categories are selected, trained on MIMIC-III, and evaluated on MIMIC-III or MIMIC-IV for transportability considerations. Among them, Generative Adversarial Network (GAN)-based methods demonstrate competitive performance in fidelity and utility on MIMIC-III, rule-based methods excel in privacy protection. Similar findings are observed on MIMIC-IV, except that GAN-based methods further outperform the baseline methods in preserving fidelity. DISCUSSION: Method choice is governed by the relative importance of the evaluation metrics in downstream use cases. We provide a decision tree to guide the choice among the benchmarked methods. An extensible Python package, "SynthEHRella", is provided to facilitate streamlined evaluations. CONCLUSION: GAN-based methods excel when distributional shifts exist between the training and testing populations. Otherwise, CorGAN and MedGAN are most suitable for association modeling and predictive modeling, respectively. Future research should prioritize enhancing fidelity of the synthetic data while controlling privacy exposure, and comprehensive benchmarking of longitudinal or conditional generation methods. Xingran Chen, Zhenke Wu, Hyunghoon Cho, Bhramar Mukherjee |
J. Am. Medical Informatics Assoc. | 5 |
| 2025 | Impacts of sample weighting on transferability of risk prediction models across EHR-Linked biobanks with different recruitment strategies
Maxwell Salvatore, Alison M. Mondul, Christopher R. Friese, David A. Hanauer, Hua Xu 0001, Celeste Leigh Pearce, Bhramar Mukherjee |
J. Biomed. Informatics | 7 |
| 2024 | Incorporating functional annotation with bilevel continuous shrinkage for polygenic risk predictionabstractBACKGROUND: Genetic variants can contribute differently to trait heritability by their functional categories, and recent studies have shown that incorporating functional annotation can improve the predictive performance of polygenic risk scores (PRSs). In addition, when only a small proportion of variants are causal variants, PRS methods that employ a Bayesian framework with shrinkage can account for such sparsity. It is possible that the annotation group level effect is also sparse. However, the number of PRS methods that incorporate both annotation information and shrinkage on effect sizes is limited. We propose a PRS method, PRSbils, which utilizes the functional annotation information with a bilevel continuous shrinkage prior to accommodate the varying genetic architectures both on the variant-specific level and on the functional annotation level. RESULTS: We conducted simulation studies and investigated the predictive performance in settings with different genetic architectures. Results indicated that when there was a relatively large variability of group-wise heritability contribution, the gain in prediction performance from the proposed method was on average 8.0% higher AUC compared to the benchmark method PRS-CS. The proposed method also yielded higher predictive performance compared to PRS-CS in settings with different overlapping patterns of annotation groups and obtained on average 6.4% higher AUC. We applied PRSbils to binary and quantitative traits in three real world data sources (the UK Biobank, the Michigan Genomics Initiative (MGI), and the Korean Genome and Epidemiology Study (KoGES)), and two sources of annotations: ANNOVAR, and pathway information from the Kyoto Encyclopedia of Genes and Genomes (KEGG), and demonstrated that the proposed method holds the potential for improving predictive performance by incorporating functional annotations. CONCLUSIONS: By utilizing a bilevel shrinkage framework, PRSbils enables the incorporation of both overlapping and non-overlapping annotations into PRS construction to improve the performance of genetic risk prediction. The software is available at https://github.com/styvon/PRSbils . Yongwen Zhuang, Na Yeon Kim, Lars G. Fritsche, Bhramar Mukherjee, Seunggeun Lee |
BMC Bioinform. | 4 |
| 2024 | To weight or not to weight? The effect of selection bias in 3 large electronic health record-linked biobanks and recommendations for practiceabstractOBJECTIVES: To develop recommendations regarding the use of weights to reduce selection bias for commonly performed analyses using electronic health record (EHR)-linked biobank data. MATERIALS AND METHODS: We mapped diagnosis (ICD code) data to standardized phecodes from 3 EHR-linked biobanks with varying recruitment strategies: All of Us (AOU; n = 244 071), Michigan Genomics Initiative (MGI; n = 81 243), and UK Biobank (UKB; n = 401 167). Using 2019 National Health Interview Survey data, we constructed selection weights for AOU and MGI to represent the US adult population more. We used weights previously developed for UKB to represent the UKB-eligible population. We conducted 4 common analyses comparing unweighted and weighted results. RESULTS: For AOU and MGI, estimated phecode prevalences decreased after weighting (weighted-unweighted median phecode prevalence ratio [MPR]: 0.82 and 0.61), while UKB estimates increased (MPR: 1.06). Weighting minimally impacted latent phenome dimensionality estimation. Comparing weighted versus unweighted phenome-wide association study for colorectal cancer, the strongest associations remained unaltered, with considerable overlap in significant hits. Weighting affected the estimated log-odds ratio for sex and colorectal cancer to align more closely with national registry-based estimates. DISCUSSION: Weighting had a limited impact on dimensionality estimation and large-scale hypothesis testing but impacted prevalence and association estimation. When interested in estimating effect size, specific signals from untargeted association analyses should be followed up by weighted analysis. CONCLUSION: EHR-linked biobanks should report recruitment and selection mechanisms and provide selection weights with defined target populations. Researchers should consider their intended estimands, specify source and target populations, and weight EHR-linked biobank analyses accordingly. Maxwell Salvatore, Ritoban Kundu, Christopher R. Friese, Seunggeun Lee, Lars G. Fritsche, Alison M. Mondul, David A. Hanauer, Celeste Leigh Pearce, Bhramar Mukherjee |
J. Am. Medical Informatics Assoc. | 10 |
| 2022 | Incorporating family disease history and controlling case-control imbalance for population-based genetic association studiesabstractMOTIVATION: In the genome-wide association analysis of population-based biobanks, most diseases have low prevalence, which results in low detection power. One approach to tackle the problem is using family disease history, yet existing methods are unable to address type I error inflation induced by increased correlation of phenotypes among closely related samples, as well as unbalanced phenotypic distribution. RESULTS: We propose a new method for genetic association test with family disease history, mixed-model-based Test with Adjusted Phenotype and Empirical saddlepoint approximation, which controls for increased phenotype correlation by adopting a two-variance-component mixed model, accounts for case-control imbalance by using empirical saddlepoint approximation, and is flexible to incorporate any existing adjusted phenotypes, such as phenotypes from the LT-FH method. We show through simulation studies and analysis of UK Biobank data of white British samples and the Korean Genome and Epidemiology Study of Korean samples that the proposed method is robust and yields better calibration compared to existing methods while gaining power for detection of variant-phenotype associations. AVAILABILITY AND IMPLEMENTATION: The summary statistics and code generated in this study are available at https://github.com/styvon/TAPE. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yongwen Zhuang, Brooke N. Wolford, Kisung Nam, Wenjian Bi, Wei Zhou 0080, Cristen J. Willer, Bhramar Mukherjee, Seunggeun Lee |
Bioinform. | 7 |
| 2022 | A Case-Crossover Phenome-wide association study (PheWAS) for understanding Post-COVID-19 diagnosis patternsabstractBACKGROUND: Post COVID-19 condition (PCC) is known to affect a large proportion of COVID-19 survivors. Robust study design and methods are needed to understand post-COVID-19 diagnosis patterns in all survivors, not just those clinically diagnosed with PCC. METHODS: We applied a case-crossover Phenome-Wide Association Study (PheWAS) in a retrospective cohort of COVID-19 survivors, comparing the occurrences of 1,671 diagnosis-based phenotype codes (PheCodes) pre- and post-COVID-19 infection periods in the same individual using a conditional logistic regression. We studied how this pattern varied by COVID-19 severity and vaccination status, and we compared to test negative and test negative but flu positive controls. RESULTS: In 44,198 SARS-CoV-2-positive patients, we foundenrichment in respiratory,circulatory, and mental health disorders post-COVID-19-infection. Top hits included anxiety disorder (p = 2.8e-109, OR = 1.7 [95 % CI: 1.6-1.8]), cardiac dysrhythmias (p = 4.9e-87, OR = 1.7 [95 % CI: 1.6-1.8]), and respiratory failure, insufficiency, arrest (p = 5.2e-75, OR = 2.9 [95 % CI: 2.6-3.3]). In severe patients, we found stronger associations with respiratory and circulatory disorders compared to mild/moderate patients. Fully vaccinated patients had mental health and chronic circulatory diseases rise to the top of the association list, similar to the mild/moderate cohort. Both control groups (test negative, test negative and flu positive) showed a different pattern of hits to SARS-CoV-2 positives. CONCLUSIONS: Patients experience myriad symptoms more than 28 days after SARS-CoV-2 infection, but especially respiratory, circulatory, and mental health disorders. Our case-crossover PheWAS approach controls for within-person confounders that are time-invariant. Comparison to test negatives and test negative but flu positive patients with a similar design helped identify enrichment specific to COVID-19. This design may be applied other emerging diseases with long-lasting effects other than a SARS-CoV-2 infection. Given the potential for bias from observational data, these results should be considered exploratory. As we look into the future, we must be aware of COVID-19 survivors' healthcare needs. Spencer R. Haupert, Lars G. Fritsche, Bhramar Mukherjee |
J. Biomed. Informatics | 5 |
| 2021 | Phenotype risk scores (PheRS) for pancreatic cancer using time-stamped electronic health record data: Discovery and validation in two large biobanksabstractBACKGROUND: Traditional methods for disease risk prediction and assessment, such as diagnostic tests using serum, urine, blood, saliva or imaging biomarkers, have been important for identifying high-risk individuals for many diseases, leading to early detection and improved survival. For pancreatic cancer, traditional methods for screening have been largely unsuccessful in identifying high-risk individuals in advance of disease progression leading to high mortality and poor survival. Electronic health records (EHR) linked to genetic profiles provide an opportunity to integrate multiple sources of patient information for risk prediction and stratification. We leverage a constellation of temporally associated diagnoses available in the EHR to construct a summary risk score, called a phenotype risk score (PheRS), for identifying individuals at high-risk for having pancreatic cancer. The proposed PheRS approach incorporates the time with respect to disease onset into the prediction framework. We combine and contrast the PheRS with more well-known measures of inherited susceptibility, namely, the polygenic risk scores (PRS) for prediction of pancreatic cancer. METHODOLOGY: We first calculated pairwise, unadjusted associations between pancreatic cancer diagnosis and all possible other diagnoses across the medical phenome. We call these pairwise associations co-occurrences. After accounting for cross-phenotype correlations, the multivariable association estimates from a subset of relatively independent diagnoses were used to create a weighted sum PheRS. We constructed time-restricted risk scores using data from 38,359 participants in the Michigan Genomics Initiative (MGI) based on the diagnoses contained in the EHR at 0, 1, 2, and 5 years prior to the target pancreatic cancer diagnosis. The PheRS was assessed for predictability in the UK Biobank (UKB). We tested the relative contribution of PheRS when added to a model containing a summary measure of inherited genetic susceptibility (PRS) plus other covariates like age, sex, smoking status, drinking status, and body mass index (BMI). RESULTS: Our exploration of co-occurrence patterns identified expected associations while also revealing unexpected relationships that may warrant closer attention. Solely using the pancreatic cancer PheRS at 5 years before the target diagnoses yielded an AUC of 0.60 (95% CI = [0.58, 0.62]) in UKB. A larger predictive model including PheRS, PRS, and the covariates at the 5-year threshold achieved an AUC of 0.74 (95% CI = [0.72, 0.76]) in UKB. We note that PheRS does contribute independently in the joint model. Finally, scores at the top percentiles of the PheRS distribution demonstrated promise in terms of risk stratification. Scores in the top 2% were 10.20 (95% CI = [9.34, 12.99]) times more likely to identify cases than those in the bottom 98% in UKB at the 5-year threshold prior to pancreatic cancer diagnosis. CONCLUSIONS: We developed a framework for creating a time-restricted PheRS from EHR data for pancreatic cancer using the rich information content of a medical phenome. In addition to identifying hypothesis-generating associations for future research, this PheRS demonstrates a potentially important contribution in identifying high-risk individuals, even after adjusting for PRS for pancreatic cancer and other traditional epidemiologic covariates. The methods are generalizable to other phenotypic traits. Maxwell Salvatore, Lauren J. Beesley, Lars G. Fritsche, David A. Hanauer, Alison M. Mondul, Celeste Leigh Pearce, Bhramar Mukherjee |
J. Biomed. Informatics | 8 |
| 2017 | Changing data practices for community health workers: Introducing digital data collection in West Bengal, IndiaabstractIn this paper, we present our findings on the experiences of West Bengal Community Health Workers (CHWs) in transitioning from paper to tablet- and mobile-based data collection. Through qualitative interviews, usability testing and timed observations, we found that efficiency and quality of data collected were comparable between the use of tablet devices and traditional paper methods, but data collection performed on smaller mobile phone interfaces was less efficient compared to paper. There was no significant difference in the quality of data collected across all three modes. In terms of work practices, we found that while initial interactions with CHWs suggested positive feelings about switching to digital devices, in their actual practices they retained and preferred the use of paper, and had workarounds to circumvent the digital data collection process. While there were foreseeable challenges around individual user experience, such as device familiarity, and application interface flexibility, the more compelling challenge in transitioning CHWs to digital data collection was organizational. The agency of CHWs within organizations, the levels of training with both data practices and devices themselves, and the sense of comfort that the data collectors felt with the overall project emerge as important factors of attention for implementers of new data management practices. Joyojeet Pal, Anjuli Dasika, Ahmad Hasan, Jackie Wolf, Nick Reid, Vaishnav Kameswaran, Purva Yardi, Allyson Mackay, Abram Wagner, Bhramar Mukherjee, Sucheta Joshi, Sujay Santra, Priyamvada Pandey |
ICTD | 10 |
| 2017 | Complete hazard ranking to analyze right-censored data: An ALS survival studyabstractSurvival analysis represents an important outcome measure in clinical research and clinical trials; further, survival ranking may offer additional advantages in clinical trials. In this study, we developed GuanRank, a non-parametric ranking-based technique to transform patients' survival data into a linear space of hazard ranks. The transformation enables the utilization of machine learning base-learners including Gaussian process regression, Lasso, and random forest on survival data. The method was submitted to the DREAM Amyotrophic Lateral Sclerosis (ALS) Stratification Challenge. Ranked first place, the model gave more accurate ranking predictions on the PRO-ACT ALS dataset in comparison to Cox proportional hazard model. By utilizing right-censored data in its training process, the method demonstrated its state-of-the-art predictive power in ALS survival ranking. Its feature selection identified multiple important factors, some of which conflicts with previous studies. Zhengnan Huang, Hongjiu Zhang, Jonathan Boss, Stephen A. Goutman, Bhramar Mukherjee, Ivo D. Dinov, Yuanfang Guan |
PLoS Comput. Biol. | 5 |