Linying Zhang

dblp:181/3849 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
8since 2021 · last 2024
0000-0002-4356-4645ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 3 first-author · 8 since 2021
YearPublicationVenuePosition
2024 Causal fairness assessment of treatment allocation with electronic health records
abstract
OBJECTIVE: Healthcare continues to grapple with the persistent issue of treatment disparities, sparking concerns regarding the equitable allocation of treatments in clinical practice. While various fairness metrics have emerged to assess fairness in decision-making processes, a growing focus has been on causality-based fairness concepts due to their capacity to mitigate confounding effects and reason about bias. However, the application of causal fairness notions in evaluating the fairness of clinical decision-making with electronic health record (EHR) data remains an understudied domain. This study aims to address the methodological gap in assessing causal fairness of treatment allocation with electronic health records data. In addition, we investigate the impact of social determinants of health on the assessment of causal fairness of treatment allocation. METHODS: We propose a causal fairness algorithm to assess fairness in clinical decision-making. Our algorithm accounts for the heterogeneity of patient populations and identifies potential unfairness in treatment allocation by conditioning on patients who have the same likelihood to benefit from the treatment. We apply this framework to a patient cohort with coronary artery disease derived from an EHR database to evaluate the fairness of treatment decisions. RESULTS: Our analysis reveals notable disparities in coronary artery bypass grafting (CABG) allocation among different patient groups. Women were found to be 4.4%-7.7% less likely to receive CABG than men in two out of four treatment response strata. Similarly, Black or African American patients were 5.4%-8.7% less likely to receive CABG than others in three out of four response strata. These results were similar when social determinants of health (insurance and area deprivation index) were dropped from the algorithm. These findings highlight the presence of disparities in treatment allocation among similar patients, suggesting potential unfairness in the clinical decision-making process. CONCLUSION: This study introduces a novel approach for assessing the fairness of treatment allocation in healthcare. By incorporating responses to treatment into fairness framework, our method explores the potential of quantifying fairness from a causal perspective using EHR data. Our research advances the methodological development of fairness assessment in healthcare and highlight the importance of causality in determining treatment fairness.
Linying Zhang, Lauren R. Richter, Yixin Wang 0002, Anna Ostropolets, Noémie Elhadad, David M. Blei, George Hripcsak
J. Biomed. Informatics1
2023 Reproducible variability: assessing investigator discordance across 9 research teams attempting to reproduce the same observational study
abstract
OBJECTIVE: Observational studies can impact patient care but must be robust and reproducible. Nonreproducibility is primarily caused by unclear reporting of design choices and analytic procedures. This study aimed to: (1) assess how the study logic described in an observational study could be interpreted by independent researchers and (2) quantify the impact of interpretations' variability on patient characteristics. MATERIALS AND METHODS: Nine teams of highly qualified researchers reproduced a cohort from a study by Albogami et al. The teams were provided the clinical codes and access to the tools to create cohort definitions such that the only variable part was their logic choices. We executed teams' cohort definitions against the database and compared the number of subjects, patient overlap, and patient characteristics. RESULTS: On average, the teams' interpretations fully aligned with the master implementation in 4 out of 10 inclusion criteria with at least 4 deviations per team. Cohorts' size varied from one-third of the master cohort size to 10 times the cohort size (2159-63 619 subjects compared to 6196 subjects). Median agreement was 9.4% (interquartile range 15.3-16.2%). The teams' cohorts significantly differed from the master implementation by at least 2 baseline characteristics, and most of the teams differed by at least 5. CONCLUSIONS: Independent research teams attempting to reproduce the study based on its free-text description alone produce different implementations that vary in the population size and composition. Sharing analytical code supported by a common data model and open-source tools allows reproducing a study unambiguously thereby preserving initial design choices.
Anna Ostropolets, Yasser Albogami, Mitchell Conover, Juan M. Banda, William A. Baumgartner Jr., Clair Blacketer, Priyamvada Desai, Scott L. DuVall, Stephen P. Fortin, James P. Gilbert, Asieh Golozar, Joshua Ide, Andrew S. Kanter, David M. Kern, Chungsoo Kim, Lana Y. H. Lai, Kristine E. Lynch, Evan P. Minty, Maria Inês Neves, Ding Quan Ng, Tontel Obene, Victor Pera, Nicole Pratt, Gowtham Rao, Nadav Rappoport, Ines Reinecke, Paola Saroufim, Azza Shoaibi, Katherine Simon, Marc A. Suchard, Joel N. Swerdel, Erica A. Voss, James Weaver, Linying Zhang, George Hripcsak, Patrick B. Ryan
J. Am. Medical Informatics Assoc.36
2022 Using Data Assimilation to Predict Post-Operative Bariatric Surgery Glycemic Status in Adolescents
Lauren R. Richter, Benjamin Albert, Linying Zhang, Ilene Fennoy, David J. Albers, George Hripcsak
AMIA3
2022 Using EHR Data and Machine Learning Methods to Predict Fall Injury
Wenyu Song, Luwei Liu, Hannah Rice, Michael Sainlaire, Lillian Min, Linying Zhang, Tien Thai, Min-Jeoung Kang, Mica Curtin-Bowen, Stuart R. Lipsitz, Lipika Samal, Nancy K. Latham, Patricia C. Dykes
AMIA6
2022 Predicting hospitalization of COVID-19 positive patients using clinician-guided machine learning methods
abstract
OBJECTIVES: The coronavirus disease 2019 (COVID-19) is a resource-intensive global pandemic. It is important for healthcare systems to identify high-risk COVID-19-positive patients who need timely health care. This study was conducted to predict the hospitalization of older adults who have tested positive for COVID-19. METHODS: We screened all patients with COVID test records from 11 Mass General Brigham hospitals to identify the study population. A total of 1495 patients with age 65 and above from the outpatient setting were included in the final cohort, among which 459 patients were hospitalized. We conducted a clinician-guided, 3-stage feature selection, and phenotyping process using iterative combinations of literature review, clinician expert opinion, and electronic healthcare record data exploration. A list of 44 features, including temporal features, was generated from this process and used for model training. Four machine learning prediction models were developed, including regularized logistic regression, support vector machine, random forest, and neural network. RESULTS: All 4 models achieved area under the receiver operating characteristic curve (AUC) greater than 0.80. Random forest achieved the best predictive performance (AUC = 0.83). Albumin, an index for nutritional status, was found to have the strongest association with hospitalization among COVID positive older adults. CONCLUSIONS: In this study, we developed 4 machine learning models for predicting general hospitalization among COVID positive older adults. We identified important clinical factors associated with hospitalization and observed temporal patterns in our study cohort. Our modeling pipeline and algorithm could potentially be used to facilitate more accurate and efficient decision support for triaging COVID positive patients.
Wenyu Song, Linying Zhang, Luwei Liu, Michael Sainlaire, Mehran Karvar, Min-Jeoung Kang, Avery Pullman, Stuart R. Lipsitz, Anthony F. Massaro, Namrata Patil, Ravi Jasuja, Patricia C. Dykes
J. Am. Medical Informatics Assoc.2
2022 Adjusting for indirectly measured confounding using large-scale propensity score
abstract
Confounding remains one of the major challenges to causal inference with observational data. This problem is paramount in medicine, where we would like to answer causal questions from large observational datasets like electronic health records (EHRs) and administrative claims. Modern medical data typically contain tens of thousands of covariates. Such a large set carries hope that many of the confounders are directly measured, and further hope that others are indirectly measured through their correlation with measured covariates. How can we exploit these large sets of covariates for causal inference? To help answer this question, this paper examines the performance of the large-scale propensity score (LSPS) approach on causal analysis of medical data. We demonstrate that LSPS may adjust for indirectly measured confounders by including tens of thousands of covariates that may be correlated with them. We present conditions under which LSPS removes bias due to indirectly measured confounders, and we show that LSPS may avoid bias when inadvertently adjusting for variables (like colliders) that otherwise can induce bias. We demonstrate the performance of LSPS with both simulated medical data and real medical data.
Linying Zhang, Yixin Wang 0002, Martijn J. Schuemie, David M. Blei, George Hripcsak
J. Biomed. Informatics1
2021 Predicting Hospitalization of COVID-19 Positive Patients Using Machine Learning Methods
Wenyu Song, Linying Zhang, Michael Sainlaire, Mehran Karvar, Min-Jeoung Kang, Avery Pullman, Anthony F. Massaro, Namrata Patil, Ravi Jasuja, Patricia C. Dykes
AMIA2
2021 Predicting pressure injury using nursing assessment phenotypes and machine learning methods
abstract
OBJECTIVE: Pressure injuries are common and serious complications for hospitalized patients. The pressure injury rate is an important patient safety metric and an indicator of the quality of nursing care. Timely and accurate prediction of pressure injury risk can significantly facilitate early prevention and treatment and avoid adverse outcomes. While many pressure injury risk assessment tools exist, most were developed before there was access to large clinical datasets and advanced statistical methods, limiting their accuracy. In this paper, we describe the development of machine learning-based predictive models, using phenotypes derived from nurse-entered direct patient assessment data. METHODS: We utilized rich electronic health record data, including full assessment records entered by nurses, from 5 different hospitals affiliated with a large integrated healthcare organization to develop machine learning-based prediction models for pressure injury. Five-fold cross-validation was conducted to evaluate model performance. RESULTS: Two pressure injury phenotypes were defined for model development: nonhospital acquired pressure injury (N = 4398) and hospital acquired pressure injury (N = 1767), representing 2 distinct clinical scenarios. A total of 28 clinical features were extracted and multiple machine learning predictive models were developed for both pressure injury phenotypes. The random forest model performed best and achieved an AUC of 0.92 and 0.94 in 2 test sets, respectively. The Glasgow coma scale, a nurse-entered level of consciousness measurement, was the most important feature for both groups. CONCLUSIONS: This model accurately predicts pressure injury development and, if validated externally, may be helpful in widespread pressure injury prevention.
Wenyu Song, Min-Jeoung Kang, Linying Zhang, Wonkyung Jung, Jiyoun Song, David W. Bates, Patricia C. Dykes
J. Am. Medical Informatics Assoc.3
2020 Evaluation of Large-scale Propensity Score Modeling and Covariate Balance on Potential Unmeasured Confounding in Observational Research
Martijn J. Schuemie, Marc A. Suchard, Anna Ostropolets, Linying Zhang, Patrick B. Ryan, George Hripcsak
AMIA5
2020 Causal Inference from Observational Healthcare Data: Implications, Impacts and Innovations
George Hripcsak, David M. Blei, Elias Bareinboim, Martijn J. Schuemie, Linying Zhang
AMIA5
2020 Predicting Pressure Injury Using Nursing Assessment Phenotype and Machine Learning Methods
Wenyu Song, Min-Jeoung Kang, Linying Zhang, Jose P. Garcia, David W. Bates, Patricia C. Dykes
AMIA3
2020 The Multi-Outcome Medical Deconfounder: Assessing Treatment Effect on Multiple Renal Measures
Linying Zhang, Yixin Wang 0002, Anna Ostropolets, David M. Blei, George Hripcsak
AMIA1
2020 A scoping review of clinical decision support tools that generate new knowledge to support decision making in real time
abstract
OBJECTIVE: A growing body of observational data enabled its secondary use to facilitate clinical care for complex cases not covered by the existing evidence. We conducted a scoping review to characterize clinical decision support systems (CDSSs) that generate new knowledge to provide guidance for such cases in real time. MATERIALS AND METHODS: PubMed, Embase, ProQuest, and IEEE Xplore were searched up to May 2020. The abstracts were screened by 2 reviewers. Full texts of the relevant articles were reviewed by the first author and approved by the second reviewer, accompanied by the screening of articles' references. The details of design, implementation and evaluation of included CDSSs were extracted. RESULTS: Our search returned 3427 articles, 53 of which describing 25 CDSSs were selected. We identified 8 expert-based and 17 data-driven tools. Sixteen (64%) tools were developed in the United States, with the others mostly in Europe. Most of the tools (n = 16, 64%) were implemented in 1 site, with only 5 being actively used in clinical practice. Patient or quality outcomes were assessed for 3 (18%) CDSSs, 4 (16%) underwent user acceptance or usage testing and 7 (28%) functional testing. CONCLUSIONS: We found a number of CDSSs that generate new knowledge, although only 1 addressed confounding and bias. Overall, the tools lacked demonstration of their utility. Improvement in clinical and quality outcomes were shown only for a few CDSSs, while the benefits of the others remain unclear. This review suggests a need for a further testing of such CDSSs and, if appropriate, their dissemination.
Anna Ostropolets, Linying Zhang, George Hripcsak
J. Am. Medical Informatics Assoc.2
2019 Investigating female-male differences in risk factors for myocardial infarction using OHDSI tools
Anna Ostropolets, Linying Zhang, Jami J. Mulgrave, George Hripcsak
AMIA2
2019 Personalized treatment for type 2 diabetes using weighted k-nearest neighbors
Wenyu Song, Linying Zhang, Emily Gill, Jeremiah Z. Liu, Adam Wright
AMIA2