EDBT 2026 Demo / reviewers in the wild / expert
Sunyang Fu
dblp:160/1539
· DBLP profile ↗
26ranked-venue papers
5as first author
17since 2021 · last 2024
0000-0003-1691-5179ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 26 · 5 first-author · 17 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Stratifying heart failure patients with graph neural network and transformer using Electronic Health Records to optimize drug response predictionabstractOBJECTIVES: Heart failure (HF) impacts millions of patients worldwide, yet the variability in treatment responses remains a major challenge for healthcare professionals. The current treatment strategies, largely derived from population based evidence, often fail to consider the unique characteristics of individual patients, resulting in suboptimal outcomes. This study aims to develop computational models that are patient-specific in predicting treatment outcomes, by utilizing a large Electronic Health Records (EHR) database. The goal is to improve drug response predictions by identifying specific HF patient subgroups that are likely to benefit from existing HF medications. MATERIALS AND METHODS: A novel, graph-based model capable of predicting treatment responses, combining Graph Neural Network and Transformer was developed. This method differs from conventional approaches by transforming a patient's EHR data into a graph structure. By defining patient subgroups based on this representation via K-Means Clustering, we were able to enhance the performance of drug response predictions. RESULTS: Leveraging EHR data from 11 627 Mayo Clinic HF patients, our model significantly outperformed traditional models in predicting drug response using NT-proBNP as a HF biomarker across five medication categories (best RMSE of 0.0043). Four distinct patient subgroups were identified with differential characteristics and outcomes, demonstrating superior predictive capabilities over existing HF subtypes (best mean RMSE of 0.0032). DISCUSSION: These results highlight the power of graph-based modeling of EHR in improving HF treatment strategies. The stratification of patients sheds light on particular patient segments that could benefit more significantly from tailored response predictions. CONCLUSIONS: Longitudinal EHR data have the potential to enhance personalized prognostic predictions through the application of graph-based AI techniques. Shaika Chowdhury, Yongbin Chen, Pengyang Li, Sivaraman Rajaganapathy, Andrew Wen, Xiao Ma 0019, Qiying Dai, Yue Yu 0012, Sunyang Fu, Xiaoqian Jiang, Zhe He 0001, Sunghwan Sohn, Xiaoke Liu, Suzette J. Bielinski, Alanna M. Chamberlain, James R. Cerhan, Nansu Zong |
J. Am. Medical Informatics Assoc. | 9 |
| 2024 | A taxonomy for advancing systematic error analysis in multi-site electronic health record-based clinical concept extractionabstractBACKGROUND: Error analysis plays a crucial role in clinical concept extraction, a fundamental subtask within clinical natural language processing (NLP). The process typically involves a manual review of error types, such as contextual and linguistic factors contributing to their occurrence, and the identification of underlying causes to refine the NLP model and improve its performance. Conducting error analysis can be complex, requiring a combination of NLP expertise and domain-specific knowledge. Due to the high heterogeneity of electronic health record (EHR) settings across different institutions, challenges may arise when attempting to standardize and reproduce the error analysis process. OBJECTIVES: This study aims to facilitate a collaborative effort to establish common definitions and taxonomies for capturing diverse error types, fostering community consensus on error analysis for clinical concept extraction tasks. MATERIALS AND METHODS: We iteratively developed and evaluated an error taxonomy based on existing literature, standards, real-world data, multisite case evaluations, and community feedback. The finalized taxonomy was released in both .dtd and .owl formats at the Open Health Natural Language Processing Consortium. The taxonomy is compatible with several different open-source annotation tools, including MAE, Brat, and MedTator. RESULTS: The resulting error taxonomy comprises 43 distinct error classes, organized into 6 error dimensions and 4 properties, including model type (symbolic and statistical machine learning), evaluation subject (model and human), evaluation level (patient, document, sentence, and concept), and annotation examples. Internal and external evaluations revealed strong variations in error types across methodological approaches, tasks, and EHR settings. Key points emerged from community feedback, including the need to enhancing clarity, generalizability, and usability of the taxonomy, along with dissemination strategies. CONCLUSION: The proposed taxonomy can facilitate the acceleration and standardization of the error analysis process in multi-site settings, thus improving the provenance, interpretability, and portability of NLP models. Future researchers could explore the potential direction of developing automated or semi-automated methods to assist in the classification and standardization of error analysis. Sunyang Fu, Liwei Wang 0010, Andrew Wen, Nansu Zong, Anamika Kumari, Rui Zhang 0028, Yanshan Wang, Jennifer L. St. Sauver, Sunghwan Sohn |
J. Am. Medical Informatics Assoc. | 1 |
| 2024 | Automatic uncovering of patient primary concerns in portal messages using a fusion framework of pretrained language modelsabstractOBJECTIVES: The surge in patient portal messages (PPMs) with increasing needs and workloads for efficient PPM triage in healthcare settings has spurred the exploration of AI-driven solutions to streamline the healthcare workflow processes, ensuring timely responses to patients to satisfy their healthcare needs. However, there has been less focus on isolating and understanding patient primary concerns in PPMs-a practice which holds the potential to yield more nuanced insights and enhances the quality of healthcare delivery and patient-centered care. MATERIALS AND METHODS: We propose a fusion framework to leverage pretrained language models (LMs) with different language advantages via a Convolution Neural Network for precise identification of patient primary concerns via multi-class classification. We examined 3 traditional machine learning models, 9 BERT-based language models, 6 fusion models, and 2 ensemble models. RESULTS: The outcomes of our experimentation underscore the superior performance achieved by BERT-based models in comparison to traditional machine learning models. Remarkably, our fusion model emerges as the top-performing solution, delivering a notably improved accuracy score of 77.67 ± 2.74% and an F1 score of 74.37 ± 3.70% in macro-average. DISCUSSION: This study highlights the feasibility and effectiveness of multi-class classification for patient primary concern detection and the proposed fusion framework for enhancing primary concern detection. CONCLUSIONS: The use of multi-class classification enhanced by a fusion of multiple pretrained LMs not only improves the accuracy and efficiency of patient primary concern identification in PPMs but also aids in managing the rising volume of PPMs in healthcare, ensuring critical patient communications are addressed promptly and accurately. Jungwei Fan 0001, Aditya Khurana, Sunyang Fu, Dezhi Wu, Ming Huang 0006 |
J. Am. Medical Informatics Assoc. | 5 |
| 2024 | FedFSA: Hybrid and federated framework for functional status ascertainment across institutions
Sunyang Fu, Heling Jia, Maria Vassilaki, Vipina Kuttichi Keloth, Yifang Dang, Yujia Zhou 0003, Muskan Garg, Ronald C. Petersen, Jennifer L. St. Sauver, Sungrim Moon, Liwei Wang 0010, Andrew Wen, Fang Li 0011, Hua Xu 0001, Cui Tao, Jungwei Fan 0001, Sunghwan Sohn |
J. Biomed. Informatics | 1 |
| 2023 | Harnessing Transfer Learning for Dementia Prediction: Leveraging Sex-Different Mild Cognitive Impairment PrognosisabstractThis paper presents a machine learning-based prediction for dementia, leveraging transfer learning to reuse the knowledge learned from prediction of mild cognitive impairment, a precursor of dementia. We also examine the impacts of temporal aspects of longitudinal data and sex differences. The methodology encompasses key components such as setting the duration window, comparing different modeling strategies, conducting comprehensive evaluations, and examining the sex-specific impacts of simulated scenarios. The findings reveal that cognitive deficits in females, once detected at the mild cognitive impairment stage, tend to deteriorate over time, while males exhibit more diverse decline across various characteristics without highlighting specific ones. However, the underlying reasons for these sex differences remain unknown and warrant further investigation. Ziming Liu 0002, Muskan Garg, Sunyang Fu, Surjodeep Sarkar, Maria Vassilaki, Ronald C. Petersen, Jennifer L. St. Sauver, Sunghwan Sohn |
BIBM | 3 |
| 2023 | Systematic design and data-driven evaluation of social determinants of health ontology (SDoHO)abstractOBJECTIVE: Social determinants of health (SDoH) play critical roles in health outcomes and well-being. Understanding the interplay of SDoH and health outcomes is critical to reducing healthcare inequalities and transforming a "sick care" system into a "health-promoting" system. To address the SDOH terminology gap and better embed relevant elements in advanced biomedical informatics, we propose an SDoH ontology (SDoHO), which represents fundamental SDoH factors and their relationships in a standardized and measurable way. MATERIAL AND METHODS: Drawing on the content of existing ontologies relevant to certain aspects of SDoH, we used a top-down approach to formally model classes, relationships, and constraints based on multiple SDoH-related resources. Expert review and coverage evaluation, using a bottom-up approach employing clinical notes data and a national survey, were performed. RESULTS: We constructed the SDoHO with 708 classes, 106 object properties, and 20 data properties, with 1,561 logical axioms and 976 declaration axioms in the current version. Three experts achieved 0.967 agreement in the semantic evaluation of the ontology. A comparison between the coverage of the ontology and SDOH concepts in 2 sets of clinical notes and a national survey instrument also showed satisfactory results. DISCUSSION: SDoHO could potentially play an essential role in providing a foundation for a comprehensive understanding of the associations between SDoH and health outcomes and paving the way for health equity across populations. CONCLUSION: SDoHO has well-designed hierarchies, practical objective properties, and versatile functionalities, and the comprehensive semantic and coverage evaluation achieved promising performance compared to the existing ontologies relevant to SDoH. Yifang Dang, Fang Li 0011, Xinyue Hu 0002, Vipina Kuttichi Keloth, Sunyang Fu, Muhammad Amith, J. Wilfred Fan, Jingcheng Du, Evan Yu, Xiaoqian Jiang, Hua Xu 0001, Cui Tao |
J. Am. Medical Informatics Assoc. | 6 |
| 2023 | An open natural language processing (NLP) framework for EHR-based clinical research: a case demonstration using the National COVID Cohort Collaborative (N3C)abstractDespite recent methodology advancements in clinical natural language processing (NLP), the adoption of clinical NLP models within the translational research community remains hindered by process heterogeneity and human factor variations. Concurrently, these factors also dramatically increase the difficulty in developing NLP models in multi-site settings, which is necessary for algorithm robustness and generalizability. Here, we reported on our experience developing an NLP solution for Coronavirus Disease 2019 (COVID-19) signs and symptom extraction in an open NLP framework from a subset of sites participating in the National COVID Cohort (N3C). We then empirically highlight the benefits of multi-site data for both symbolic and statistical methods, as well as highlight the need for federated annotation and evaluation to resolve several pitfalls encountered in the course of these efforts. Sijia Liu 0002, Andrew Wen, Liwei Wang 0010, Sunyang Fu, Robert T. Miller, Andrew E. Williams, Daniel R. Harris, Ramakanth Kavuluru, Noor Abu-El-Rub, Dalton Schutte, Rui Zhang 0028, Masoud Rouhizadeh, John D. Osborne, Yongqun He, Umit Topaloglu, Stephanie S. Hong, Joel H. Saltz, Thomas Schaffter, Emily R. Pfaff, Christopher G. Chute, Tim Duong, Melissa A. Haendel, Rafael Fuentes, Peter Szolovits, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 5 |
| 2022 | Quality Assessment of Functional Status Documentation in EHR Across Institutions
Sunyang Fu, Maria Vassilaki, Omar A. Ibrahim, Ronald C. Petersen, Jennifer L. St. Sauver, Liwei Wang 0010, Jungwei Fan 0001, Sunghwan Sohn |
AMIA | 1 |
| 2022 | Towards User-centered Corpus Development: Lessons Learnt from Designing and Developing MedTator
Sunyang Fu, Liwei Wang 0010, Andrew Wen, Sijia Liu 0002, Sungrim Moon, Kurt Miller |
AMIA | 2 |
| 2022 | Probing Radiology Patient Experience Feedbacks with Aspect-based Sentiment Analysis
Kurt Miller, Ming Huang 0006, Sunyang Fu, Kris Abah, Andrea Maraboto Escarria, Kevin J. Peterson, Lacey Hart, Nelly Tan |
AMIA | 3 |
| 2022 | BETA: a comprehensive benchmark for computational drug-target predictionabstractInternal validation is the most popular evaluation strategy used for drug-target predictive models. The simple random shuffling in the cross-validation, however, is not always ideal to handle large, diverse and copious datasets as it could potentially introduce bias. Hence, these predictive models cannot be comprehensively evaluated to provide insight into their general performance on a variety of use-cases (e.g. permutations of different levels of connectiveness and categories in drug and target space, as well as validations based on different data sources). In this work, we introduce a benchmark, BETA, that aims to address this gap by (i) providing an extensive multipartite network consisting of 0.97 million biomedical concepts and 8.5 million associations, in addition to 62 million drug-drug and protein-protein similarities and (ii) presenting evaluation strategies that reflect seven cases (i.e. general, screening with different connectivity, target and drug screening based on categories, searching for specific drugs and targets and drug repurposing for specific diseases), a total of seven Tests (consisting of 344 Tasks in total) across multiple sampling and validation strategies. Six state-of-the-art methods covering two broad input data types (chemical structure- and gene sequence-based and network-based) were tested across all the developed Tasks. The best-worst performing cases have been analyzed to demonstrate the ability of the proposed benchmark to identify limitations of the tested methods for running over the benchmark tasks. The results highlight BETA as a benchmark in the selection of computational strategies for drug repurposing and target discovery. Nansu Zong, Ning Li 0045, Andrew Wen, Victoria Ngo, Yue Yu 0012, Ming Huang 0006, Shaika Chowdhury, Chao Jiang 0002, Sunyang Fu, Richard Weinshilboum, Guoqian Jiang, Lawrence Hunter |
Briefings Bioinform. | 9 |
| 2022 | MedTator: a serverless annotation tool for corpus developmentabstractSUMMARY: Building a high-quality annotation corpus requires expenditure of considerable time and expertise, particularly for biomedical and clinical research applications. Most existing annotation tools provide many advanced features to cover a variety of needs where the installation, integration and difficulty of use present a significant burden for actual annotation tasks. Here, we present MedTator, a serverless annotation tool, aiming to provide an intuitive and interactive user interface that focuses on the core steps related to corpus annotation, such as document annotation, corpus summarization, annotation export and annotation adjudication. AVAILABILITY AND IMPLEMENTATION: MedTator and its tutorial are freely available from https://ohnlp.github.io/MedTator. MedTator source code is available under the Apache 2.0 license: https://github.com/OHNLP/MedTator. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sunyang Fu, Liwei Wang 0010, Sijia Liu 0002, Andrew Wen |
Bioinform. | 2 |
| 2022 | Real-time risk prediction of colorectal surgery-related post-surgical complications using GRU-D model
Xiaoyang Ruan, Sunyang Fu, Curtis B. Storlie, Kellie L. Mathis, David W. Larson |
J. Biomed. Informatics | 2 |
| 2021 | Detecting Major Depressive Disorder from Clinical Notes using Neural Language Models with Distant Supervision
Bhavani Singh Agnikula Kshatriya, Nicolas A. Nunez, Manuel Gardea-Resendez, Euijung Ryu, Brandon J. Coombes, Sunyang Fu, Mark A. Frye, Joanna M. Biernacka, Yanshan Wang |
AMIA | 6 |
| 2021 | Deep Learning Approaches for Breast Cancer Characteristics Extraction from Electronic Health Records
Liwei Wang 0010, Sunyang Fu, Chetan Shenoy, Anne H. Blaes, Rui Zhang 0028 |
AMIA | 4 |
| 2021 | Early Alert of Elderly Cognitive Impairment using Temporal Streaming Clusteringabstractmore than 44 million people have been diagnosed with dementia worldwide, and this number is estimated to triple by next three decades. Given this increasing trend of older adults with cognitive impairment (CI; dementia and mild cognitive impairment) and its significant underdiagnosis, early identification of CI and understanding its progression is a critical step towards a better quality of life for the aging population. Early alert of individual health changes could facilitate better ways for clinicians to diagnose CI in its early stages and come up with more effective treatment plans. However, there is a lack of approaches to characterize patient health conditions accounting for temporal information in an unsupervised manner. Limited CI cases and its costly ascertainment in clinical settings also make unsupervised learning more promising in CI research. In this paper, a streaming clustering model was used to determine distinct patterns of older adults' health changes from their clinical visits in Mayo Clinic Study of Aging. The streaming clustering was also examined to study its ability to generate early alerts for potential incidents of CI. Our analysis demonstrated that temporal characteristics incorporated in a streaming clustering model has a promising potential to increase power in predicting CI. Omar A. Ibrahim, Sunyang Fu, Maria Vassilaki, Ronald C. Petersen, Michelle M. Mielke, Jennifer L. St. Sauver, Sunghwan Sohn |
BIBM | 2 |
| 2021 | An aberration detection-based approach for sentinel syndromic surveillance of COVID-19 and other novel influenza-like illnesses
Andrew Wen, Liwei Wang 0010, Sijia Liu 0002, Sunyang Fu, Sunghwan Sohn, Jacob A. Kugel, Vinod Kaggal, Ming Huang 0006, Yanshan Wang, Feichen Shen, Jungwei Fan 0001 |
J. Biomed. Informatics | 5 |
| 2020 | Predicting Section Location of Clinical Sentences using BERT Encoder - A Pilot Study
Sijia Liu 0002, Sunyang Fu, Sungrim Moon, Andrew Wen |
AMIA | 2 |
| 2020 | Annotating Chronic Pain Episodes in EHR Text: Guideline Development and Corpus Analysis
Luke A. Carlson, Molly M. Jeffery, Sunyang Fu, Rozalina G. McCoy, Yanshan Wang, W. M. Hooten, Jennifer L. St. Sauver, Jungwei Fan 0001 |
AMIA | 3 |
| 2020 | Big Impact from Small Data: Unsupervised Machine Learning Approaches for Chronic Pain Patient Subgrouping
Luke A. Carlson, Jennifer L. St. Sauver, Sunyang Fu, Ahmad P. Tafti, Jungwei Fan 0001, Molly M. Jeffery, Rozalina G. McCoy, Yanshan Wang |
AMIA | 3 |
| 2020 | Accelerating Development of Learning Healthcare Systems via Distantly Supervised Knowledge Discovery
Andrew Wen, Sunyang Fu, Feichen Shen |
AMIA | 2 |
| 2020 | Extracting Patient Genetic Information from Unstructured Clinical Notes
Hanzhong Yu, Luke A. Carlson, Feichen Shen, Sunyang Fu, Chen Wang 0001 |
AMIA | 5 |
| 2020 | Clinical concept extraction: A methodology review
Sunyang Fu, David Chen 0003, Sijia Liu 0002, Sungrim Moon, Kevin J. Peterson, Feichen Shen, Liwei Wang 0010, Yanshan Wang, Andrew Wen, Sunghwan Sohn |
J. Biomed. Informatics | 1 |
| 2019 | Artificial intelligence to organize patient portal messages: a journey from an ensemble deep learning text classification to rule-based named entity recognitionabstractSince the turn of the millennium, numerous healthcare venues all over the world have made a standard of communication transition from the classic telephone call to a sophisticated online patient portal system. More and more, a majority of patients prefer using portal-style communication for clinician contact, checking lab results, and other informational transactions, in which hundreds of thousands of patient portal messages (PPMs) are daily generated as free-text data with multiple requests often buried in one single message. Thus, there is a pressing need to design and implement artificial intelligence (AI) algorithms to accurately organize this wealth of data in a timely fashion. With the present contribution, an attempt was made to first develop an ensemble deep learning text classification component and then integrate it with rule-based named entity recognition to categorize free-text PPMs submitted under the “Non-Urgent Medical Question” subject in the patient portal as either containing active symptom descriptions or logistical requests (e.g., appointment rescheduling). Ahmad P. Tafti, Sunyang Fu, Aditya Khurana, George M. Mastorakos, Kenneth G. Poole, Stephen J. Traub, James A. Yiannias |
BIBM | 2 |
| 2018 | Natural Language Processing for the Identification of Silent Brain Infractions from Neuroimaging Reports
Sunyang Fu, Lester Y. Leung, Anne-Olivia Raulli, David F. Kallmes, Kristin A. Kinsman, Kristoff B. Nelson, Michael S. Clark, Patrick H. Luetmer, Paul R. Kingsbury, David M. Kent |
AMIA | 1 |
| 2018 | Supervised Learning Approach to Link Prediction in FDA Adverse Event Reporting System (FAERS) Database Network
Andy Jinseok Lee, Sunyang Fu, V. G. Vinod Vydiswaran |
AMIA | 2 |