VLDB 2026 Research / reviewers in the wild / expert
Lina M. Sulieman
dblp:186/5410 · also Lina Sulieman
· DBLP profile ↗
18ranked-venue papers
8as first author
13since 2021 · last 2024
0000-0002-2680-8230ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 8 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Balancing efficacy and computational burden: weighted mean, multiple imputation, and inverse probability weighting methods for item non-response in reliable scalesabstractIMPORTANCE: Scales often arise from multi-item questionnaires, yet commonly face item non-response. Traditional solutions use weighted mean (WMean) from available responses, but potentially overlook missing data intricacies. Advanced methods like multiple imputation (MI) address broader missing data, but demand increased computational resources. Researchers frequently use survey data in the All of Us Research Program (All of Us), and it is imperative to determine if the increased computational burden of employing MI to handle non-response is justifiable. OBJECTIVES: Using the 5-item Physical Activity Neighborhood Environment Scale (PANES) in All of Us, this study assessed the tradeoff between efficacy and computational demands of WMean, MI, and inverse probability weighting (IPW) when dealing with item non-response. MATERIALS AND METHODS: Synthetic missingness, allowing 1 or more item non-response, was introduced into PANES across 3 missing mechanisms and various missing percentages (10%-50%). Each scenario compared WMean of complete questions, MI, and IPW on bias, variability, coverage probability, and computation time. RESULTS: All methods showed minimal biases (all <5.5%) for good internal consistency, with WMean suffered most with poor consistency. IPW showed considerable variability with increasing missing percentage. MI required significantly more computational resources, taking >8000 and >100 times longer than WMean and IPW in full data analysis, respectively. DISCUSSION AND CONCLUSION: The marginal performance advantages of MI for item non-response in highly reliable scales do not warrant its escalated cloud computational burden in All of Us, particularly when coupled with computationally demanding post-imputation analyses. Researchers using survey scales with low missingness could utilize WMean to reduce computing burden. Andrew Guide, Shawn Garbett, Xiaoke Feng, Brandy Mapes, Justin Cook, Lina M. Sulieman, Robert M. Cronin, Qingxia Chen |
J. Am. Medical Informatics Assoc. | 6 |
| 2024 | Illuminating the landscape of high-level clinical trial opportunities in the All of Us Research ProgramabstractOBJECTIVE: With its size and diversity, the All of Us Research Program has the potential to power and improve representation in clinical trials through ancillary studies like Nutrition for Precision Health. We sought to characterize high-level trial opportunities for the diverse participants and sponsors of future trial investment. MATERIALS AND METHODS: We matched All of Us participants with available trials on ClinicalTrials.gov based on medical conditions, age, sex, and geographic location. Based on the number of matched trials, we (1) developed the Trial Opportunities Compass (TOC) to help sponsors assess trial investment portfolios, (2) characterized the landscape of trial opportunities in a phenome-wide association study (PheWAS), and (3) assessed the relationship between trial opportunities and social determinants of health (SDoH) to identify potential barriers to trial participation. RESULTS: Our study included 181 529 All of Us participants and 18 634 trials. The TOC identified opportunities for portfolio investment and gaps in currently available trials across federal, industrial, and academic sponsors. PheWAS results revealed an emphasis on mental disorder-related trials, with anxiety disorder having the highest adjusted increase in the number of matched trials (59% [95% CI, 57-62]; P < 1e-300). Participants from certain communities underrepresented in biomedical research, including self-reported racial and ethnic minorities, had more matched trials after adjusting for other factors. Living in a nonmetropolitan area was associated with up to 13.1 times fewer matched trials. DISCUSSION AND CONCLUSION: All of Us data are a valuable resource for identifying trial opportunities to inform trial portfolio planning. Characterizing these opportunities with consideration for SDoH can provide guidance on prioritizing the most pressing barriers to trial participation. Cathy Shyr, Lina M. Sulieman, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 2 |
| 2024 | Identifying erroneous height and weight values from adult electronic health records in the All of Us research programabstractINTRODUCTION: Electronic Health Records (EHR) are a useful data source for research, but their usability is hindered by measurement errors. This study investigated an automatic error detection algorithm for adult height and weight measurements in EHR for the All of Us Research Program (All of Us). METHODS: We developed reference charts for adult heights and weights that were stratified on participant sex. Our analysis included 4,076,534 height and 5,207,328 wt measurements from ∼ 150,000 participants. Errors were identified using modified standard deviation scores, differences from their expected values, and significant changes between consecutive measurements. We evaluated our method with chart-reviewed heights (8,092) and weights (9,039) from 250 randomly selected participants and compared it with the current cleaning algorithm in All of Us. RESULTS: The proposed algorithm classified 1.4 % of height and 1.5 % of weight errors in the full cohort. Sensitivity was 90.4 % (95 % CI: 79.0-96.8 %) for heights and 65.9 % (95 % CI: 56.9-74.1 %) for weights. Precision was 73.4 % (95 % CI: 60.9-83.7 %) for heights and 62.9 (95 % CI: 54.0-71.1 %) for weights. In comparison, the current cleaning algorithm has inferior performance in sensitivity (55.8 %) and precision (16.5 %) for height errors while having higher precision (94.0 %) and lower sensitivity (61.9 %) for weight errors. DISCUSSION: Our proposed algorithm outperformed in detecting height errors compared to weights. It can serve as a valuable addition to the current All of Us cleaning algorithm for identifying erroneous height values. Andrew Guide, Lina M. Sulieman, Shawn Garbett, Robert M. Cronin, Matthew E. Spotnitz, Karthik Natarajan, Robert J. Carroll, Paul A. Harris, Qingxia Chen |
J. Biomed. Informatics | 2 |
| 2023 | Systematic replication of smoking disease associations using survey responses and EHR data in the All of Us Research ProgramabstractOBJECTIVE: The All of Us Research Program (All of Us) aims to recruit over a million participants to further precision medicine. Essential to the verification of biobanks is a replication of known associations to establish validity. Here, we evaluated how well All of Us data replicated known cigarette smoking associations. MATERIALS AND METHODS: We defined smoking exposure as follows: (1) an EHR Smoking exposure that used International Classification of Disease codes; (2) participant provided information (PPI) Ever Smoking; and, (3) PPI Current Smoking, both from the lifestyle survey. We performed a phenome-wide association study (PheWAS) for each smoking exposure measurement type. For each, we compared the effect sizes derived from the PheWAS to published meta-analyses that studied cigarette smoking from PubMed. We defined two levels of replication of meta-analyses: (1) nominally replicated: which required agreement of direction of effect size, and (2) fully replicated: which required overlap of confidence intervals. RESULTS: PheWASes with EHR Smoking, PPI Ever Smoking, and PPI Current Smoking revealed 736, 492, and 639 phenome-wide significant associations, respectively. We identified 165 meta-analyses representing 99 distinct phenotypes that could be matched to EHR phenotypes. At P < .05, 74 were nominally replicated and 55 were fully replicated. At P < 2.68 × 10-5 (Bonferroni threshold), 58 were nominally replicated and 40 were fully replicated. DISCUSSION: Most phenotypes found in published meta-analyses associated with smoking were nominally replicated in All of Us. Both survey and EHR definitions for smoking produced similar results. CONCLUSION: This study demonstrated the feasibility of studying common exposures using All of Us data. David J. Schlueter, Lina M. Sulieman, Huan Mo, Jacob M. Keaton, Tracey Ferrara, Ariel Williams, Onajia J. Stubblefield, Chenjie Zeng, Tam C. Tran, Lisa Bastarache, Anav Babbar, Andrea H. Ramirez, Slavina Goleva, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 2 |
| 2022 | Identifying Erroneous Height and Weight Values from Adult Electronic Health Records
Andrew Guide, Lina M. Sulieman, Qingxia Chen |
AMIA | 2 |
| 2022 | Self-paced Training Modality to Promote the Use of All of Us Researcher Workbench in Educational and Research Settings
Hiral Master, Lina M. Sulieman, Paul A. Harris, Karthik Natarajan, Robert J. Carroll, Kayla Marginean, Kelsey R. Mayo, Aymone Kouame |
AMIA | 2 |
| 2022 | Primum non Nocere: Challenges and Strategies for Protecting Privacy for Adolescent Patients in the 21st Century Cures Act Setting
Marianne Sharko, S. Trent Rosenbloom, Lina M. Sulieman, Jessica S. Ancker |
AMIA | 3 |
| 2022 | Predicting the Retention of Subsequent Surveys in the All of Us
Lina M. Sulieman, Xiaoke Feng, Qingxia Chen, Robert M. Cronin |
AMIA | 1 |
| 2022 | Assessing Data Quality and Diversity in the All of Us
Lina M. Sulieman, Jennifer Zhang, Kayla Marginean, Paul A. Harris, Robert J. Carroll |
AMIA | 1 |
| 2022 | Comparing medical history data derived from electronic health records and survey answers in the All of Us Research ProgramabstractOBJECTIVE: A participant's medical history is important in clinical research and can be captured from electronic health records (EHRs) and self-reported surveys. Both can be incomplete, EHR due to documentation gaps or lack of interoperability and surveys due to recall bias or limited health literacy. This analysis compares medical history collected in the All of Us Research Program through both surveys and EHRs. MATERIALS AND METHODS: The All of Us medical history survey includes self-report questionnaire that asks about diagnoses to over 150 medical conditions organized into 12 disease categories. In each category, we identified the 3 most and least frequent self-reported diagnoses and retrieved their analogues from EHRs. We calculated agreement scores and extracted participant demographic characteristics for each comparison set. RESULTS: The 4th All of Us dataset release includes data from 314 994 participants; 28.3% of whom completed medical history surveys, and 65.5% of whom had EHR data. Hearing and vision category within the survey had the highest number of responses, but the second lowest positive agreement with the EHR (0.21). The Infectious disease category had the lowest positive agreement (0.12). Cancer conditions had the highest positive agreement (0.45) between the 2 data sources. DISCUSSION AND CONCLUSION: Our study quantified the agreement of medical history between 2 sources-EHRs and self-reported surveys. Conditions that are usually undocumented in EHRs had low agreement scores, demonstrating that survey data can supplement EHR data. Disagreement between EHR and survey can help identify possible missing records and guide researchers to adjust for biases. Lina M. Sulieman, Robert M. Cronin, Robert J. Carroll, Karthik Natarajan, Kayla Marginean, Brandy Mapes, Dan M. Roden, Paul A. Harris, Andrea H. Ramirez |
J. Am. Medical Informatics Assoc. | 1 |
| 2021 | Systematic replication of smoking disease associations in the All of Us Research Program
David J. Schlueter, Lina M. Sulieman, Jacob M. Keaton, Tracey Ferrara, Kyle Webb, Ariel Williams, Francis Ratsimbazafy, Lisa Bastarache, Andrea H. Ramirez, Joshua C. Denny |
AMIA | 2 |
| 2021 | Measuring the correctness of All of Us physical measurement
Lina M. Sulieman, Karthik Natarajan, Qingxia Chen, Robert J. Carroll, Kayla Marginean, Paul A. Harris, Andrea H. Ramirez |
AMIA | 1 |
| 2021 | Comparison of family health history in surveys vs electronic health record data mapped to the observational medical outcomes partnership data model in the All of Us Research ProgramabstractOBJECTIVE: Family health history is important to clinical care and precision medicine. Prior studies show gaps in data collected from patient surveys and electronic health records (EHRs). The All of Us Research Program collects family history from participants via surveys and EHRs. This Demonstration Project aims to evaluate availability of family health history information within the publicly available data from All of Us and to characterize the data from both sources. MATERIALS AND METHODS: Surveys were completed by participants on an electronic portal. EHR data was mapped to the Observational Medical Outcomes Partnership data model. We used descriptive statistics to perform exploratory analysis of the data, including evaluating a list of medically actionable genetic disorders. We performed a subanalysis on participants who had both survey and EHR data. RESULTS: There were 54 872 participants with family history data. Of those, 26% had EHR data only, 63% had survey only, and 10.5% had data from both sources. There were 35 217 participants with reported family history of a medically actionable genetic disorder (9% from EHR only, 89% from surveys, and 2% from both). In the subanalysis, we found inconsistencies between the surveys and EHRs. More details came from surveys. When both mentioned a similar disease, the source of truth was unclear. CONCLUSIONS: Compiling data from both surveys and EHR can provide a more comprehensive source for family health history, but informatics challenges and opportunities exist. Access to more complete understanding of a person's family health history may provide opportunities for precision medicine. Robert M. Cronin, Alese E. Halvorson, Cassie Springer, Xiaoke Feng, Lina M. Sulieman, Roxana Loperena-Cortes, Kelsey R. Mayo, Robert J. Carroll, Qingxia Chen, Brian K. Ahmedani, Jason Karnes, Bruce Korf, Christopher J. O'Donnell, Andrea H. Ramirez |
J. Am. Medical Informatics Assoc. | 5 |
| 2019 | Why Patient Portal Messages Indicate Risk of Readmission for Patients with Ischemic Heart Disease
Lina M. Sulieman, Zhijun Yin, Bradley A. Malin |
AMIA | 1 |
| 2019 | A systematic literature review of machine learning in online personal health dataabstractOBJECTIVE: User-generated content (UGC) in online environments provides opportunities to learn an individual's health status outside of clinical settings. However, the nature of UGC brings challenges in both data collecting and processing. The purpose of this study is to systematically review the effectiveness of applying machine learning (ML) methodologies to UGC for personal health investigations. MATERIALS AND METHODS: We searched PubMed, Web of Science, IEEE Library, ACM library, AAAI library, and the ACL anthology. We focused on research articles that were published in English and in peer-reviewed journals or conference proceedings between 2010 and 2018. Publications that applied ML to UGC with a focus on personal health were identified for further systematic review. RESULTS: We identified 103 eligible studies which we summarized with respect to 5 research categories, 3 data collection strategies, 3 gold standard dataset creation methods, and 4 types of features applied in ML models. Popular off-the-shelf ML models were logistic regression (n = 22), support vector machines (n = 18), naive Bayes (n = 17), ensemble learning (n = 12), and deep learning (n = 11). The most investigated problems were mental health (n = 39) and cancer (n = 15). Common health-related aspects extracted from UGC were treatment experience, sentiments and emotions, coping strategies, and social support. CONCLUSIONS: The systematic review indicated that ML can be effectively applied to UGC in facilitating the description and inference of personal health. Future research needs to focus on mitigating bias introduced when building study cohorts, creating features from free text, improving clinical creditability of UGC, and model interpretability. Zhijun Yin, Lina M. Sulieman, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 2 |
| 2017 | Classifying patient portal messages using Convolutional Neural Networks
Lina M. Sulieman, David Gilmore, Christi French, Robert M. Cronin, Gretchen Purcell Jackson, Matthew Russell, Daniel Fabbri |
J. Biomed. Informatics | 1 |
| 2016 | Predicting Negative Events: Using Post-discharge Data to Detect High-Risk Patients
Lina M. Sulieman, Daniel Fabbri, Fei Wang 0001, Jianying Hu, Bradley A. Malin |
AMIA | 1 |
| 2015 | Comparison of Patient Portal Usage between Employees and Non-Employees
Lina M. Sulieman, Dara Eckerle Mize, Daniel Fabbri, S. Trent Rosenbloom |
AMIA | 1 |