EDBT 2026 Demo / reviewers in the wild / expert
Yiye Zhang
dblp:127/9778
· DBLP profile ↗
30ranked-venue papers
10as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 27 · 10 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An asynchronous federated learning-assisted data sharing method for medical blockchain
Chenquan Gan, Xinghai Xiao, Yiye Zhang, Qingyi Zhu, Jichao Bi, Deepak Kumar Jain 0001, Akanksha Saini |
Appl. Intell. | 3 |
| 2025 | Extracting social support and social isolation information from clinical psychiatry notes: comparing a rule-based natural language processing system and a large language modelabstractOBJECTIVES: Social support (SS) and social isolation (SI) are social determinants of health (SDOH) associated with psychiatric outcomes. In electronic health records (EHRs), individual-level SS/SI is typically documented in narrative clinical notes rather than as structured coded data. Natural language processing (NLP) algorithms can automate the otherwise labor-intensive process of extraction of such information. MATERIALS AND METHODS: Psychiatric encounter notes from Mount Sinai Health System (MSHS, n = 300) and Weill Cornell Medicine (WCM, n = 225) were annotated to create a gold-standard corpus. A rule-based system (RBS) involving lexicons and a large language model (LLM) using FLAN-T5-XL were developed to identify mentions of SS and SI and their subcategories (eg, social network, instrumental support, and loneliness). RESULTS: For extracting SS/SI, the RBS obtained higher macroaveraged F1-scores than the LLM at both MSHS (0.89 versus 0.65) and WCM (0.85 versus 0.82). For extracting the subcategories, the RBS also outperformed the LLM at both MSHS (0.90 versus 0.62) and WCM (0.82 versus 0.81). DISCUSSION AND CONCLUSION: Unexpectedly, the RBS outperformed the LLMs across all metrics. An intensive review demonstrates that this finding is due to the divergent approach taken by the RBS and LLM. The RBS was designed and refined to follow the same specific rules as the gold-standard annotations. Conversely, the LLM was more inclusive with categorization and conformed to common English-language understanding. Both approaches offer advantages, although additional replication studies are warranted. Braja Gopal Patra, Lauren A. Lepow, Praneet Kasi Reddy Jagadeesh Kumar, Veer Vekaria, Mohit Manoj Sharma, Prakash Adekkanattu, Brian Fennessy, Gavin Hynes, Isotta Landi, Jorge A. Sanchez-Ruiz, Euijung Ryu, Joanna M. Biernacka, Girish N. Nadkarni, Ardesheer Talati, Myrna Weissman, Mark Olfson, J. John Mann, Yiye Zhang, Alexander Charney, Jyotishman Pathak |
J. Am. Medical Informatics Assoc. | 18 |
| 2025 | Machine learning applications related to suicide in military and Veterans: A scoping literature review
Yishu Wei, Yanshan Wang, Yunyu Xiao, Ronald K. Poropatich, Gretchen L. Haas, Yiye Zhang, Chunhua Weng, Jinze Liu, Lisa A. Brenner, James M. Bjork, Yifan Peng 0002 |
J. Biomed. Informatics | 7 |
| 2024 | Visualizing machine learning-based predictions of postpartum depression risk for lay audiencesabstractOBJECTIVES: To determine if different formats for conveying machine learning (ML)-derived postpartum depression risks impact patient classification of recommended actions (primary outcome) and intention to seek care, perceived risk, trust, and preferences (secondary outcomes). MATERIALS AND METHODS: We recruited English-speaking females of childbearing age (18-45 years) using an online survey platform. We created 2 exposure variables (presentation format and risk severity), each with 4 levels, manipulated within-subject. Presentation formats consisted of text only, numeric only, gradient number line, and segmented number line. For each format viewed, participants answered questions regarding each outcome. RESULTS: Five hundred four participants (mean age 31 years) completed the survey. For the risk classification question, performance was high (93%) with no significant differences between presentation formats. There were main effects of risk level (all P < .001) such that participants perceived higher risk, were more likely to agree to treatment, and more trusting in their obstetrics team as the risk level increased, but we found inconsistencies in which presentation format corresponded to the highest perceived risk, trust, or behavioral intention. The gradient number line was the most preferred format (43%). DISCUSSION AND CONCLUSION: All formats resulted high accuracy related to the classification outcome (primary), but there were nuanced differences in risk perceptions, behavioral intentions, and trust. Investigators should choose health data visualizations based on the primary goal they want lay audiences to accomplish with the ML risk score. Pooja M. Desai, Sarah Harkins, Saanjaana Rahman, Shiveen Kumar, Alison Hermann, Rochelle Joly, Yiye Zhang, Jyotishman Pathak, Jessica Kim, Deborah D'angelo, Natalie C. Benda, Meghan Reading Turchioe |
J. Am. Medical Informatics Assoc. | 7 |
| 2024 | Preparing for the bedside - optimizing a postpartum depression risk prediction model for clinical implementation in a health systemabstractOBJECTIVE: We developed and externally validated a machine-learning model to predict postpartum depression (PPD) using data from electronic health records (EHRs). Effort is under way to implement the PPD prediction model within the EHR system for clinical decision support. We describe the pre-implementation evaluation process that considered model performance, fairness, and clinical appropriateness. MATERIALS AND METHODS: We used EHR data from an academic medical center (AMC) and a clinical research network database from 2014 to 2020 to evaluate the predictive performance and net benefit of the PPD risk model. We used area under the curve and sensitivity as predictive performance and conducted a decision curve analysis. In assessing model fairness, we employed metrics such as disparate impact, equal opportunity, and predictive parity with the White race being the privileged value. The model was also reviewed by multidisciplinary experts for clinical appropriateness. Lastly, we debiased the model by comparing 5 different debiasing approaches of fairness through blindness and reweighing. RESULTS: We determined the classification threshold through a performance evaluation that prioritized sensitivity and decision curve analysis. The baseline PPD model exhibited some unfairness in the AMC data but had a fair performance in the clinical research network data. We revised the model by fairness through blindness, a debiasing approach that yielded the best overall performance and fairness, while considering clinical appropriateness suggested by the expert reviewers. DISCUSSION AND CONCLUSION: The findings emphasize the need for a thorough evaluation of intervention-specific models, considering predictive performance, fairness, and appropriateness before clinical implementation. Rochelle Joly, Meghan Reading Turchioe, Natalie C. Benda, Alison Hermann, Ashley Beecy, Jyotishman Pathak, Yiye Zhang |
J. Am. Medical Informatics Assoc. | 8 |
| 2023 | An encrypted medical blockchain data search method with access control mechanism
Chenquan Gan, Hongpeng Yang, Qingyi Zhu, Yiye Zhang, Akanksha Saini |
Inf. Process. Manag. | 4 |
| 2021 | Deep Significance Clustering (DICE) a Heterogenous Population
Yufang Huang, Peter A. D. Steel, Kelly M. Axsom, Sri Lekha Tummalapalli, Alison Hermann, Rochelle Joly, Fei Wang 0001, Jyotishman Pathak, Lakshminarayanan Subramanian, Yiye Zhang |
AMIA | 12 |
| 2021 | Impact of Social Determinants of Health on Predictive Models in 30-Day Hospital Readmission or Death for Patients with Severe Obesity
Marianne Sharko, Yongkang Zhang 0004, Yiye Zhang, Evan Sholle, Sajjad Abedian, Meghan Reading Turchioe, Jessica S. Ancker |
AMIA | 3 |
| 2021 | Prescribing Pharmacogenomics Testing: Analyzing the Acceptance of Healthcare Providers through a Survey Study
Mohit Manoj Sharma, Yonaka Harris, Yiye Zhang, Jyotishman Pathak |
AMIA | 3 |
| 2021 | MyEDCare: Evaluation of a Smartphone-based Emergency Department Discharge Process
Peter A. D. Steel, Dávid Bodnár, Maryellen Bonito, Jane Torres-Lavoro, Dona Bou Eid, Andrew Jacobowitz, Amos Shemesh, Robert Tanouye, Patrick Rumble, Daniel Dicello, Brenna Farmer, Sandra Pomerantz, Yiye Zhang |
AMIA | 15 |
| 2021 | Deep significance clustering: a novel approach for identifying risk-stratified and predictive patient subgroupsabstractOBJECTIVE: Deep significance clustering (DICE) is a self-supervised learning framework. DICE identifies clinically similar and risk-stratified subgroups that neither unsupervised clustering algorithms nor supervised risk prediction algorithms alone are guaranteed to generate. MATERIALS AND METHODS: Enabled by an optimization process that enforces statistical significance between the outcome and subgroup membership, DICE jointly trains 3 components, representation learning, clustering, and outcome prediction while providing interpretability to the deep representations. DICE also allows unseen patients to be predicted into trained subgroups for population-level risk stratification. We evaluated DICE using electronic health record datasets derived from 2 urban hospitals. Outcomes and patient cohorts used include discharge disposition to home among heart failure (HF) patients and acute kidney injury among COVID-19 (Cov-AKI) patients, respectively. RESULTS: Compared to baseline approaches including principal component analysis, DICE demonstrated superior performance in the cluster purity metrics: Silhouette score (0.48 for HF, 0.51 for Cov-AKI), Calinski-Harabasz index (212 for HF, 254 for Cov-AKI), and Davies-Bouldin index (0.86 for HF, 0.66 for Cov-AKI), and prediction metric: area under the Receiver operating characteristic (ROC) curve (0.83 for HF, 0.78 for Cov-AKI). Clinical evaluation of DICE-generated subgroups revealed more meaningful distributions of member characteristics across subgroups, and higher risk ratios between subgroups. Furthermore, DICE-generated subgroup membership alone was moderately predictive of outcomes. DISCUSSION: DICE addresses a gap in current machine learning approaches where predicted risk may not lead directly to actionable clinical steps. CONCLUSION: DICE demonstrated the potential to apply in heterogeneous populations, where having the same quantitative risk does not equate with having a similar clinical profile. Yufang Huang, Peter A. D. Steel, Kelly M. Axsom, Sri Lekha Tummalapalli, Fei Wang 0001, Jyotishman Pathak, Lakshminarayanan Subramanian, Yiye Zhang |
J. Am. Medical Informatics Assoc. | 10 |
| 2020 | Provider perspectives on the clinical utility of using a risk prediction tool for postpartum depression
Annie C. Myers, Fariha Ahsan, Rochelle Joly, Alison Hermann, Yiye Zhang, Michael Laskoff, Jyotishman Pathak, Meghan Reading Turchioe |
AMIA | 5 |
| 2020 | Identification of Alzheimer's Disease Subtypes from Electronic Health Records Using a Data-Driven Approach
Jie Xu 0012, Fei Wang 0001, Prakash Adekkanattu, Pascal S. Brandt, Guoqian Jiang, Richard C. Kiefer, Yuan Luo 0001, Chengsheng Mao, Jennifer A. Pacheco, Luke V. Rasmussen, Yiye Zhang, Richard Isaacson, Jyotishman Pathak |
AMIA | 12 |
| 2020 | Data-Driven Clinical Decision Support for Computerized Physician Order Entry: Development, Evaluation, and Implementation
Yiye Zhang, Jonathan H. Chen, Adam Wright, Jessica S. Ancker, Marc Tobias |
AMIA | 1 |
| 2020 | Cycle-Consistent Adversarial Autoencoders for Unsupervised Text Style TransferabstractUnsupervised text style transfer is full of challenges due to the lack of parallel data and difficulties in content preservation.In this paper, we propose a novel neural approach to unsupervised text style transfer which we refer to as Cycle-consistent Adversarial autoEncoders (CAE) trained from non-parallel data.CAE consists of three essential components: (1) LSTM autoencoders that encode a text in one style into its latent representation and decode an encoded representation into its original text or a transferred representation into a style-transferred text, (2) adversarial style transfer networks that use an adversarially trained generator to transform a latent representation in one style into a representation in another style, and (3) a cycle-consistent constraint that enhances the capacity of the adversarial style transfer networks in content preservation.The entire CAE with these three components can be trained end-to-end.Extensive experiments and in-depth analyses on two widely-used public datasets consistently validate the effectiveness of proposed CAE in both style transfer and content preservation against several strong baselines in terms of four automatic evaluation metrics and human evaluation. Yufang Huang, Wentao Zhu 0001, Deyi Xiong, Yiye Zhang, Changjian Hu |
COLING | 4 |
| 2020 | Sequential Pattern Mining of Longitudinal Adverse Events After Left Ventricular Assist Device ImplantabstractLeft ventricular assist devices (LVADs) are an increasingly common therapy for patients with advanced heart failure. However, implantation of the LVAD increases the risk of stroke, infection, bleeding, and other serious adverse events (AEs). Most post-LVAD AEs studies have focused on individual AEs in isolation, neglecting the possible interrelation, or causality between AEs. This study is the first to conduct an exploratory analysis to discover common sequential chains of AEs following LVAD implantation that are correlated with important clinical outcomes. This analysis was derived from 58,575 recorded AEs for 13,192 patients in International Registry for Mechanical Circulatory Support (INTERMACS) who received a continuous-flow LVAD between 2006 and 2015. The pattern mining procedure involved three main steps: (1) creating a bank of AE sequences by converting the AEs for each patient into a single, chronologically sequenced record, (2) grouping patients with similar AE sequences using hierarchical clustering, and (3) extracting temporal chains of AEs for each group of patients using Markov modeling. The mined results indicate the existence of seven groups of sequential chains of AEs, characterized by common types of AEs that occurred in a unique order. The groups were identified as: GRP1: Recurrent bleeding, GRP2: Trajectory of device malfunction & explant, GRP3: Infection, GRP4: Trajectories to transplant, GRP5: Cardiac arrhythmia, GRP6: Trajectory of neurological dysfunction & death, and GRP7: Trajectory of respiratory failure, renal dysfunction & death. These patterns of sequential post-LVAD AEs disclose potential interdependence between AEs and may aid prediction, and prevention, of subsequent AEs in future studies. Faezeh Movahedi, Robert L. Kormos, Lisa C. Lohmueller, Laura Seese, Manreet K. Kanwar, Srinivas Murali, Yiye Zhang, Rema Padman, James F. Antaki |
IEEE J. Biomed. Health Informatics | 7 |
| 2019 | A Survey of Clinicians' Perception of Pharmacogenomics Testing for Antidepressants
Yiye Zhang, Yonaka Harris, George Alexopoulos, Adam Stracher, Elizabeth Ross, Rainu Kaushal, Jyotishman Pathak |
AMIA | 1 |
| 2019 | Informatics approaches to collecting, analyzing, and addressing social determinants of health in healthcare
Yiye Zhang, Evan Sholle, Marianne Sharko, Yongkang Zhang 0004, Jessica S. Ancker |
AMIA | 1 |
| 2018 | Augmenting community-level social determinants of health data with individual-level survey data
Min-hyung Kim, Yiye Zhang, Jessica S. Ancker |
AMIA | 2 |
| 2018 | Data-driven Optimization of Order Sets Within an Electronic Health Record System
Tahir Rizvi, Soyeon Yoon, Hyun Nam Su, Huaizhu O. Gao, Victoria Tiase, Robert Leviton, Yiye Zhang |
AMIA | 8 |
| 2018 | The potential value of social determinants of health in predicting health outcomesabstractDear Dr Ohno-Machado, As Kasthurirathne and colleagues1 point out in their recent JAMIA paper, it is well established that population health is affected by socioeconomic status and other social determinants of health (SDH). It is therefore reasonable to hypothesize that SDH data should have predictive power for clinical and healthcare utilization outcomes for individual patients, yet the authors found that adding SDH data to clinical data produced no significant improvement in the performance of algorithms predicting the need for social service referrals.1 It is important to be reminded that not all available data have utility for all purposes, and that many plausible hypotheses do not survive rigorous analysis. However, we would like to make sure that the informatics community does not interpret these findings more broadly to indicate that SDH data have no value. We would like to propose several possible explanations for these interesting and surprising findings, each of which might suggest future avenues of exploration. Correlations between predictors: It is known that SDH are correlated with the risk of many clinical conditions, and it seems possible that in this particular study, the SDH variables were strongly correlated with the clinical diagnoses. For example, social determinants (eg socioeconomic status, race, social support, etc.) are well established as strong predictors of cardiovascular disease.2 If, in the current study, the information contained in the SDH was already present in the clinical variables, then adding SDH would not improve model performance. Future work in different domains might show SDH to have more predictive power. For example, with certain congenital conditions, SDH might have little influence on the risk of disease but instead could be related to access to care or quality of life for people with the condition, and thus would contribute additional information to a model containing clinical diagnoses. In addition, it is likely that many of the SDH are correlated with each other. Community-level research, which often deals with collinear predictors, often addresses this problem by collapsing correlated measures into scales (such as the Centers for Disease Control’s Social Vulnerability Index [https://svi.cdc.gov]), which are then utilized as single metrics. Choice of outcome variables: A closely related potential explanation is that the specific social service referrals chosen in the current study were not well predicted by SDH variables because they were already too well predicted by the clinical ones. For example, referrals to dietitians (one of the study’s outcome variables) might be so strongly predicted by diagnosis of diabetes, hypertension, or congestive heart failure that additional data are not helpful. Similarly, referrals to mental health services might be sufficiently strongly predicted by the presence of psychiatric diagnoses. It is possible that SDH data may have predictive power for other sorts of healthcare utilization and health outcomes, just not for these particular social services. Variability in predictors: As the authors suggest in their discussion, it is possible that the patients of the safety net health system had insufficient variability in SDH. For example, if household income did not range much above or below the poverty line in the entire population of interest, this variable would provide little discriminatory power. It is possible that models built from a more socioeconomically diverse population data set might provide different results. Comprehensiveness of the clinical data: The authors had access to unusually comprehensive clinical data through the Indiana Network for Patient Care, a well-established health information exchange organization, which allowed the researchers to leverage data such as emergency and hospital admissions from other organizations. Healthcare organizations with access to only their own local clinical data may find that SDH variables contribute more to predictive models, precisely because the SDH data might serve as proxies for some of the missing clinical data. Ecological inferences: In the absence of individual-level SDH data, the authors used community-level (ZIP code and census tract level) measures as proxies. It is possible that for certain SDH variables, community estimates are either imprecise or biased. For example, if the within-community variance in education is extremely high, then individual educational attainment will not be predicted well by the neighborhood average, creating lack of precision. Alternately, if the patients who seek care at a safety net hospital tend to be less well educated than their close neighbors, then individual educational attainment will be systematically overestimated by the neighborhood average, creating bias.3 Future work might explore whether community-level data have more utility when they describe neighborhood characteristics (such as, in this study, data about local availability of well-lit walkways or grocery stores) rather than being used to infer individual-level characteristics (such as education). A useful family of methodological approaches to account for these ecological relationships is hierarchical or multilevel models, which explicitly account for the nested structure of the data (individuals within neighborhoods, in this case). Choice of predictor variables: In addition to the rich set of social and environmental factors used by Kasthurirathne and colleagues, it is possible that others not available to the researchers might have predictive power, such as social support and social capital, or detailed employment type.4 Inspecting variable importance measures5 in the models might suggest new hypotheses about which types of SDH have the most predictive utility, and whether these point to other SDH data to collect or obtain from other sources. Given the extensive public health literature on the relationship between social determinants and health, it is exciting to see the health informatics community begin to explore mergers of clinical and SDH data sets. We appreciate the contribution of Kasthurirathne et al. to this emerging literature and welcome additional exploration of the potential utility of social determinants of health in clinical care and predictive analytics. Conflict of interest statement. None declared. Jessica S. Ancker, Min-hyung Kim, Yiye Zhang, Yongkang Zhang 0004, Jyotishman Pathak |
J. Am. Medical Informatics Assoc. | 3 |
| 2018 | Developing and maintaining clinical decision support using clinical knowledge and machine learning: the case of order setsabstractDevelopment and maintenance of order sets is a knowledge-intensive task for off-the-shelf machine-learning algorithms alone. We hypothesize that integrating clinical knowledge with machine learning can facilitate effective development and maintenance of order sets while promoting best practices in ordering. To this end, we simulated the revision of an "AM Lab Order Set" under 6 revision approaches. Revisions included changes in the order set content or default settings through 1) population statistics, 2) individualized prediction using machine learning, and 3) clinical knowledge. Revision criteria were determined using electronic health record (EHR) data from 2014 to 2015. Each revision's clinical appropriateness, workload from using the order set, and generalizability across time were evaluated using EHR data from 2016 and 2017. Our results suggest a potential order set revision approach that jointly leverages clinical knowledge and machine learning to improve usability while updating contents based on latest clinical knowledge and best practices. Yiye Zhang, Richard Trepp, Weiguang Wang 0001, Jorge M. Luna, David K. Vawdrey, Victoria Tiase |
J. Am. Medical Informatics Assoc. | 1 |
| 2018 | Association networks in a matched case-control design - Co-occurrence patterns of preexisting chronic medical conditions in patients with major depression versus their matched controls
Min-hyung Kim, Samprit Banerjee, Yize Zhao, Fei Wang 0001, Yiye Zhang, Yongjun Zhu 0001, Joseph DeFerio, Lauren Evans, Sang Min Park, Jyotishman Pathak |
J. Biomed. Informatics | 5 |
| 2017 | Redesigning the "Choice Architecture" of the EHR to Reduce Clinical Decision Support Burden
Jessica S. Ancker, Sameer Malhotra, Yiye Zhang, Adam D. Cheriff |
AMIA | 3 |
| 2017 | Revision of Order Sets Using Experts' Knowledge and Data-Driven Evidence
Yiye Zhang, Richard Trepp, Jorge M. Luna, David K. Vawdrey, Victoria Tiase |
AMIA | 1 |
| 2016 | Data-driven Clinical Pathway Learning and Outcomes Prediction for Chronic Kidney Disease
Yiye Zhang, Rema Padman |
AMIA | 1 |
| 2015 | Data Driven Order Set Development Using Metaheuristic Optimization
Yiye Zhang, Rema Padman |
AIME | 1 |
| 2015 | Paving the COWpath: Learning and visualizing clinical pathways from electronic health record data
Yiye Zhang, Rema Padman, Nirav Patel |
J. Biomed. Informatics | 1 |
| 2014 | On Learning and Visualizing Practice-based Clinical Pathways for Chronic Kidney Disease
Yiye Zhang, Rema Padman, Larry A. Wasserman |
AMIA | 1 |
| 2012 | Data-driven Order Set Generation and Evaluation in the Pediatric Environment
Yiye Zhang, James E. Levin, Rema Padman |
AMIA | 1 |