VLDB 2026 Research / reviewers in the wild / expert
Yuan Luo 0001
dblp:90/6959-1
· DBLP profile ↗
92ranked-venue papers
18as first author
45since 2021 · last 2026
0000-0003-0195-7456ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 78 · 16 first-author · 36 since 2021Artificial intelligence and machine learning · 15 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-site analysis of COVID-19 and new-onset diabetes reveals need for improved sensitivity of EHR-based COVID-19 phenotypes - a DiCAYA Network analysisabstractOBJECTIVE: We discuss implications of potential ascertainment biases for studies examining diabetes risk following SARS-CoV-2 infection using electronic health records (EHRs). We quantitatively explore sensitivity of results to misclassification of COVID-19 status using data from the U.S.-based Diabetes in Children, Adolescents and Young Adults (DiCAYA) Network on children (≤17 years) and young adults (18-44 years). MATERIALS AND METHODS: In our retrospective case study from the DiCAYA Network, SARS-CoV-2 was identified using labs and diagnoses from June 1, 2020 to December 31, 2021. Patients were followed through December 31, 2022 for new diabetes diagnoses. Sites examined incident diabetes by COVID-19 status using Cox proportional hazards models. Results were pooled in meta-analyses. A bias analysis examined potential impact of COVID-19 misclassification scenarios on results, guided by hypotheses that sensitivity would be <50% and would be higher among those who developed diabetes. RESULTS: Prevalence of documented COVID-19 was low overall and variable across sites (children: 4.4%-7.7%, young adults: 6.2%-22.7%). Individuals with documented COVID-19 were at higher risk of incident diabetes compared to those with no documented infection, but results were heterogeneous across sites. Findings were highly sensitive to COVID-19 misclassification assumptions. Observed results could be biased away from the null under several differential misclassification scenarios. DISCUSSION: Although EHR-based documentation of COVID-19 was associated with incident diabetes, COVID-19 phenotypes likely had low sensitivity, with considerable variation across sites. Misclassification assumptions strongly impacted interpretation of results. CONCLUSION: Given the potential for low phenotype sensitivity and misclassification, caution is warranted when interpreting analyses of COVID-19 and incident diabetes using clinical or administrative databases. Lorna E. Thorpe, Jasmin Divers, Annemarie Hirsch, Brian S. Schwartz, Jihad S. Obeid, Angela Liese, Tessa L. Crume, Anna Bellatorre, Jiang Bian 0001, Yi Guo 0005, Sarah Bost, Tianchen Lyu, Matthew T. Mefford, Matt Zhou, Eva Lustigova, Levon Utidjian, Mitchell Maltenfort, Patrick Hanley, Meda E. Pavkov, Marc B. Rosenman, Andrea R. Titus, L. Charles Bailey, Christopher B. Forrest, Mitch Maltenfort, Amy Shah, Eneida A. Mendonça, G. Todd Alonso, Sara J. Deakyne Davies, H. Timothy Bunnell, Anne Kazak, Melody Kitzmiller, Manmohan Kamboj, Dimitri A. Christakis, Daksha Ranade, Annemarie G. Hirsch, Joseph J. Dewalle, H. Lester Kirchner, Meredith Lewis, Dione G. Mercer, Cara M. Nordberg, Amy Poissant, Brian E. Dixon, Shaun J. Grannis, Katie Allen, Anna Roberts, Nimish Valvi, Jeff Warvel, Ashley Wiensch, Tamara S. Hannon, Kristi Reynolds, John Chang, Don McCarthy, Rong Wei, Marc Rosenman, George Lales, Anthony Wong, Allison Zelinski, Yuan Luo 0001, Mark Weiner, Pedro Rivera, Thomas Carton, Elizabeth Nauman, Harold P. Lehmann, Meredith Akerman, Rebecca Anthopolos, Stefanie Bendik, Sarah Conderino, Andrew Fair, Jessica Guillaume, Shahidul Islam, Alan Jacobson, David C. Lee, Chinyere Okpara, Anand Rajan, Andrea Titus, Dana Dabelea, Theresa Anderson, Rebecca Conway, Toan Ong, Jack Pattee, Shawna Burgett, Elizabeth Shenkman, William T. Donahoo, William R. Hogan, Piaopiao Li, Mattia Prosperi, Yonghui Wu 0001, Angela D. Liese, Lisa Knight, Caroline Rudisill, Jessica Stucker, Deborah Bowlby, Elaine Apperson, Alex Ewing, Giuseppina Imperatore, Deborah Rolka, Ibrahim Zaganjor |
J. Am. Medical Informatics Assoc. | 62 |
| 2025 | Exploring Large Language Models for Knowledge Graph CompletionabstractKnowledge graphs play a vital role in numerous artificial intelligence tasks, yet they frequently face the issue of incompleteness. In this study, we explore utilizing Large Language Models (LLM) for knowledge graph completion. We consider triples in knowledge graphs as text sequences and introduce an innovative framework called Knowledge Graph LLM (KG-LLM) to model these triples. Our technique employs entity and relation descriptions of a triple as prompts and utilizes the response for predictions. Experiments on various benchmark knowledge graphs demonstrate that our method attains state-of-the-art performance in tasks such as triple classification and relation prediction. We also find that fine-tuning relatively smaller models (e.g., LLaMA-7B, ChatGLM-6B) outperforms recent ChatGPT and GPT-4. Jiazhen Peng, Chengsheng Mao, Yuan Luo 0001 |
ICASSP | 4 |
| 2025 | Deep Reinforcement Learning for Efficient and Fair Allocation of Healthcare ResourcesabstractThe scarcity of health care resources, such as ventilators, often leads to the unavoidable consequence of rationing, particularly during public health emergencies or in resource-constrained settings like pandemics. The absence of a universally accepted standard for resource allocation protocols results in governments relying on varying criteria and heuristic-based approaches, often yielding suboptimal and inequitable outcomes. This study addresses the societal challenge of fair and effective critical care resource allocation by leveraging deep reinforcement learning to optimize policy decisions. We propose a transformer-based deep Q-network that integrates individual patient disease progression and interaction effects among patients to enhance allocation decisions. Our method aims to improve both fairness and overall patient outcomes. Experiments using metrics such as normalized survival rates and interracial allocation rate differences demonstrate that our approach significantly reduces excess deaths and achieves more equitable resource allocation compared to severity- and comorbidity-based protocols currently in use. Our findings highlight the potential of deep reinforcement learning to address critical health care challenges. Yikuan Li, Chengsheng Mao, Kaixuan Huang, Hanyin Wang, Mengdi Wang 0001, Yuan Luo 0001 |
IJCAI | 7 |
| 2025 | Large language models accurately identify immunosuppression in intensive care unit patientsabstractOBJECTIVE: Rule-based structured data algorithms and natural language processing (NLP) approaches applied to unstructured clinical notes have limited accuracy and poor generalizability for identifying immunosuppression. Large language models (LLMs) may effectively identify patients with heterogenous types of immunosuppression from unstructured clinical notes. We compared the performance of LLMs applied to unstructured notes for identifying patients with immunosuppressive conditions or immunosuppressive medication use against 2 baselines: (1) structured data algorithms using diagnosis codes and medication orders and (2) NLP approaches applied to unstructured notes. MATERIALS AND METHODS: We used hospital admission notes from a primary cohort of 827 intensive care unit (ICU) patients at Northwestern Memorial Hospital and a validation cohort of 200 ICU patients at Beth Israel Deaconess Medical Center, along with diagnosis codes and medication orders from the primary cohort. We evaluated the performance of structured data algorithms, NLP approaches, and LLMs in identifying 7 immunosuppressive conditions and 6 immunosuppressive medications. RESULTS: In the primary cohort, structured data algorithms achieved peak F1 scores ranging from 0.30 to 0.97 for identifying immunosuppressive conditions and medications. NLP approaches achieved peak F1 scores ranging from 0 to 1. GPT-4o outperformed or matched structured data algorithms and NLP approaches across all conditions and medications, with F1 scores ranging from 0.51 to 1. GPT-4o also performed impressively in our validation cohort (F1 = 1 for 8/13 variables). DISCUSSION: LLMs, particularly GPT-4o, outperformed structured data algorithms and NLP approaches in identifying immunosuppressive conditions and medications with robust external validation. CONCLUSION: LLMs can be applied for improved cohort identification for research purposes. Vijeeth Guggilla, Mengjia Kang, Melissa J. Bak, Steven D. Tran, Anna Pawlowski, Prasanth Nannapaneni, Luke V. Rasmussen, Helen K. Donnelly, Ankit Agrawal 0001, David M. Liebovitz, Alexander V. Misharin, G. R. Scott Budinger, Richard G. Wunderink, Theresa Walunas, Catherine A. Gao, Alan R. Hauser, Alec Peltekian, Alexis Rose Wolfe, Alison L. Szabo, Alok N. Choudhary, Amy Ludwig, Anahid Amani Moghadam, Anjana V. Yeldandi, Ankit Bharat, Anna E. Pawlowski, Anthony M. Joudi, Arjun Prakash Tambe, Ashley J. Smith-Nunez, Benjamin D. Singer, Benjamin J. Ulrich, Betty Tran, Cara J. Gottardi, Chiagozie O. Pickens, Clara J. Schroedl, Daniel Meza, Dulce Sarai Garcia, Egon A. Ozer, Elen Gusman, Elisheva D. Shanes, Emily Mower Provost, Emily M. Olson, Erica Marie Hartmann, Erin A. Korth, Estefani Diaz, Estefany R. Guzman, Francisco J. Martinez, Gabrielle Matias, Hiam Abdala-Valencia, Jack T. Sumner, Jacob I Sznajder, Jacqueline M. Kruser, Jakub Glowala, James M. Walter, Jamie H. Rowell, Jason M. Arnold, John Coleman, Jon W. Lomasney, Joseph Isaac Bailey, Judd F. Hultquist, Justin A. Fiala, Justin Starren, Karen M. Ridge, Karolina Senkow, Kathryn A. Helmin, Khalilah L. Gates, Lacy Simmons, Lesley Pinzon, Lindsey D. Gradone, Lisa F. Wolfe, Lucy Luo, Luisa Morales-Nebreda, Manu Jain, Marc Sala, Maxwell Schleck, Melissa H. Ross, Melissa Querrey, Michael J. Cuttica, Michelle Hinsch Prickett, Nandita R. Nadig, Nathaniel Rhodes, Navdeep S. Chandel, Nikolay S. Markov, Peter H. S. Sporn, Qianli Liu, Rachel B. Kadar, Rachel L. Medernach, Ramon Lorenzo-Redondo, Ravi Kalhan, Rebecca K. Clepp, Richard I. Morimoto, Rogan A. Grant, Ruben J. Mylvaganam, Samuel Fenske, Scott A. Laurenzo, Seung Hye Han, Sophia Nozick, Srinivas Panchamukhi, Stephanie C. Eisenbarth, Suchitra Swaminathan, Susan R. Russell, Taylor A. Poor, Thaddeus Cybulski, Theresa A. Lombardo, Thomas Bolig, Thomas Stoeger, Tien Doan, Timothy Rowe, Wan-Ting Liao, Yuan Luo 0001, Yuliana Sokolenko, Ziyan Lu |
J. Am. Medical Informatics Assoc. | 112 |
| 2024 | Towards Expressive Graph Representations for Graph Neural NetworksabstractGraph Neural Network (GNN) aggregates the neighborhood information into the node embedding and shows its powerful capability for graph representation learning in various application areas. However, most existing GNN variants aggregate the neighborhood information in a fixed non-injective fashion, which may map different graphs or nodes to the same embedding, detrimental to the model expressiveness. In this paper, we present a theoretical framework to improve the expressive power of GNN by taking both injectivity and continuity into account. Based on the framework, we develop injective and continuous expressive Graph Neural Network (iceGNN) that learns the graph and node representations in an injective and continuous fashion, so that it can map similar nodes or graphs to similar embeddings, and non-equivalent nodes or non-isomorphic graphs to different embeddings. We validate the proposed iceGNN model for graph classification and node classification on multiple benchmark datasets. The experimental results demonstrate that our model achieves state-of-the-art performances on most of the benchmarks. Chengsheng Mao, Yuan Luo 0001 |
ICDM | 3 |
| 2024 | Longitudinal clustering of Life's Essential 8 health metrics: application of a novel unsupervised learning method in the CARDIA studyabstractOBJECTIVE: Changes in cardiovascular health (CVH) during the life course are associated with future cardiovascular disease (CVD). Longitudinal clustering analysis using subgraph augmented non-negative matrix factorization (SANMF) could create phenotypic risk profiles of clustered CVH metrics. MATERIALS AND METHODS: Life's Essential 8 (LE8) variables, demographics, and CVD events were queried over 15 ears in 5060 CARDIA participants with 18 years of subsequent follow-up. LE8 subgraphs were mined and a SANMF algorithm was applied to cluster frequently occurring subgraphs. K-fold cross-validation and diagnostics were performed to determine cluster assignment. Cox proportional hazard models were fit for future CV event risk and logistic regression was performed for cluster phenotyping. RESULTS: The cohort (54.6% female, 48.7% White) produced 3 clusters of CVH metrics: Healthy & Late Obesity (HLO) (29.0%), Healthy & Intermediate Sleep (HIS) (43.2%), and Unhealthy (27.8%). HLO had 5 ideal LE8 metrics between ages 18 and 39 years, until BMI increased at 40. HIS had 7 ideal LE8 metrics, except sleep. Unhealthy had poor levels of sleep, smoking, and diet but ideal glucose. Race and employment were significantly different by cluster (P < .001) but not sex (P = .734). For 301 incident CV events, multivariable hazard ratios (HRs) for HIS and Unhealthy were 0.73 (0.53-1.00, P = .052) and 2.00 (1.50-2.68, P < .001), respectively versus HLO. A 15-year event survival was 97.0% (HIS), 96.3% (HLO), and 90.4% (Unhealthy, P < .001). DISCUSSION AND CONCLUSION: SANMF of LE8 metrics identified 3 unique clusters of CVH behavior patterns. Clustering of longitudinal LE8 variables via SANMF is a robust tool for phenotypic risk assessment for future adverse cardiovascular events. Peter Graffy, Lindsay P. Zimmerman, Yuan Luo 0001, Jingzhi Yu, Yuni Choi, Rachel Zmora, Donald Lloyd-Jones, Norrina B. Allen |
J. Am. Medical Informatics Assoc. | 3 |
| 2024 | Using machine learning to develop smart reflex testing protocolsabstractOBJECTIVE: Reflex testing protocols allow clinical laboratories to perform second line diagnostic tests on existing specimens based on the results of initially ordered tests. Reflex testing can support optimal clinical laboratory test ordering and diagnosis. In current clinical practice, reflex testing typically relies on simple "if-then" rules; however, this limits the opportunities for reflex testing since most test ordering decisions involve more complexity than traditional rule-based approaches would allow. Here, using the analyte ferritin as an example, we propose an alternative machine learning-based approach to "smart" reflex testing. METHODS: Using deidentified patient data, we developed a machine learning model to predict whether a patient getting CBC testing will also have ferritin testing ordered. We evaluate applications of this model to reflex testing by assessing its performance in comparison to possible rule-based approaches. RESULTS: Our underlying machine learning models performed moderately well in predicting ferritin test ordering (AUC=0.731 in reference to actual ordering) and demonstrated promising potential to underlie key clinical applications. In contrast, none of the many traditionally framed, rule-based, hypothetical reflex protocols we evaluated offered sufficient agreement with actual ordering to be clinically feasible. Using chart review, we further demonstrated that the strategic deployment of our model could avoid important ferritin test ordering errors. CONCLUSIONS: Machine learning may provide a foundation for new types of reflex testing with enhanced benefits for clinical diagnosis. Matthew McDermott, Anand Dighe, Peter Szolovits, Yuan Luo 0001, Jason Baron |
J. Am. Medical Informatics Assoc. | 4 |
| 2024 | Fairness and inclusion methods for biomedical informatics research
Shyam Visweswaran, Yuan Luo 0001, Mor Peleg |
J. Biomed. Informatics | 2 |
| 2023 | Pediatric Sepsis Phenotyping Using Vital Sign TrajectoriesabstractSepsis can be life-threatening, which highlights the need to understand the condition's diverse phenotypes to enhance treatment effectiveness. Sepsis phenotypes are derived from 12-hour vital sign trajectories of children (N=12,824) with multiple organ dysfunction syndrome from 13 U.S. hospitals. Survival analysis of the two subgroups produced by hierarchical clustering (HAC) on pairwise trajectory similarity matrix from dynamic time warping (DTW) showed a hazards ratio of 4.7 for 30-day mortality, which was better than the stratification of subgroups from group-based trajectory modeling. The higher mortality subgroup from HAC on DTW displayed higher blood pressure and pulse but lower temperature, in addition to acidosis. This comprehensive analysis of phenotypes can greatly aid in early risk evaluation, tailored treatment approaches, and improved outcomes for pediatric sepsis patients. Yanyi Jenny Ding, Zhidi Luo, Mindy Szeto, Yuan Luo 0001, L. Nelson Sanchez-Pinto |
BIBM | 4 |
| 2023 | Deep Reinforcement Learning for Cost-Effective Medical Diagnosis
Yikuan Li, Joseph C. Kim, Kaixuan Huang, Yuan Luo 0001, Mengdi Wang 0001 |
ICLR | 5 |
| 2023 | Open-set recognition of breast cancer treatments
Alexander Cao, Diego Klabjan, Yuan Luo 0001 |
Artif. Intell. Medicine | 3 |
| 2023 | Characterizing variability of electronic health record-driven phenotype definitionsabstractOBJECTIVE: The aim of this study was to analyze a publicly available sample of rule-based phenotype definitions to characterize and evaluate the variability of logical constructs used. MATERIALS AND METHODS: A sample of 33 preexisting phenotype definitions used in research that are represented using Fast Healthcare Interoperability Resources and Clinical Quality Language (CQL) was analyzed using automated analysis of the computable representation of the CQL libraries. RESULTS: Most of the phenotype definitions include narrative descriptions and flowcharts, while few provide pseudocode or executable artifacts. Most use 4 or fewer medical terminologies. The number of codes used ranges from 5 to 6865, and value sets from 1 to 19. We found that the most common expressions used were literal, data, and logical expressions. Aggregate and arithmetic expressions are the least common. Expression depth ranges from 4 to 27. DISCUSSION: Despite the range of conditions, we found that all of the phenotype definitions consisted of logical criteria, representing both clinical and operational logic, and tabular data, consisting of codes from standard terminologies and keywords for natural language processing. The total number and variety of expressions are low, which may be to simplify implementation, or authors may limit complexity due to data availability constraints. CONCLUSIONS: The phenotype definitions analyzed show significant variation in specific logical, arithmetic, and other operators but are all composed of the same high-level components, namely tabular data and logical expressions. A standard representation for phenotype definitions should support these formats and be modular to support localization and shared logic. Pascal S. Brandt, Abel N. Kho, Yuan Luo 0001, Jennifer A. Pacheco, Theresa Walunas, Hakon Hakonarson, George Hripcsak, Cong Liu 0020, Ning Shang 0004, Chunhua Weng, Nephi Walton, David Carrell, Paul K. Crane, Eric B. Larson, Christopher G. Chute, Iftikhar J. Kullo, Robert J. Carroll, Joshua C. Denny, Andrea H. Ramirez, Wei-Qi Wei, Jyotishman Pathak, Laura K. Wiley, Rachel L. Richesson, Justin Starren, Luke V. Rasmussen |
J. Am. Medical Informatics Assoc. | 3 |
| 2023 | Transportability of bacterial infection prediction models for critically ill patientsabstractOBJECTIVE: Bacterial infections (BIs) are common, costly, and potentially life-threatening in critically ill patients. Patients with suspected BIs may require empiric multidrug antibiotic regimens and therefore potentially be exposed to prolonged and unnecessary antibiotics. We previously developed a BI risk model to augment practices and help shorten the duration of unnecessary antibiotics to improve patient outcomes. Here, we have performed a transportability assessment of this BI risk model in 2 tertiary intensive care unit (ICU) settings and a community ICU setting. We additionally explored how simple multisite learning techniques impacted model transportability. METHODS: Patients suspected of having a community-acquired BI were identified in 3 datasets: Medical Information Mart for Intensive Care III (MIMIC), Northwestern Medicine Tertiary (NM-T) ICUs, and NM "community-based" ICUs. ICU encounters from MIMIC and NM-T datasets were split into 70/30 train and test sets. Models developed on training data were evaluated against the NM-T and MIMIC test sets, as well as NM community validation data. RESULTS: During internal validations, models achieved AUROCs of 0.78 (MIMIC) and 0.81 (NM-T) and were well calibrated. In the external community ICU validation, the NM-T model had robust transportability (AUROC 0.81) while the MIMIC model transported less favorably (AUROC 0.74), likely due to case-mix differences. Multisite learning provided no significant discrimination benefit in internal validation studies but offered more stability during transport across all evaluation datasets. DISCUSSION: These results suggest that our BI risk models maintain predictive utility when transported to external cohorts. CONCLUSION: Our findings highlight the importance of performing external model validation on myriad clinically relevant populations prior to implementation. Garrett Eickelberg, L. Nelson Sanchez-Pinto, Adrienne S. Kline, Yuan Luo 0001 |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | A comparative study of pretrained language models for long clinical textabstractOBJECTIVE: Clinical knowledge-enriched transformer models (eg, ClinicalBERT) have state-of-the-art results on clinical natural language processing (NLP) tasks. One of the core limitations of these transformer models is the substantial memory consumption due to their full self-attention mechanism, which leads to the performance degradation in long clinical texts. To overcome this, we propose to leverage long-sequence transformer models (eg, Longformer and BigBird), which extend the maximum input sequence length from 512 to 4096, to enhance the ability to model long-term dependencies in long clinical texts. MATERIALS AND METHODS: Inspired by the success of long-sequence transformer models and the fact that clinical notes are mostly long, we introduce 2 domain-enriched language models, Clinical-Longformer and Clinical-BigBird, which are pretrained on a large-scale clinical corpus. We evaluate both language models using 10 baseline tasks including named entity recognition, question answering, natural language inference, and document classification tasks. RESULTS: The results demonstrate that Clinical-Longformer and Clinical-BigBird consistently and significantly outperform ClinicalBERT and other short-sequence transformers in all 10 downstream tasks and achieve new state-of-the-art results. DISCUSSION: Our pretrained language models provide the bedrock for clinical NLP using long texts. We have made our source code available at https://github.com/luoyuanlab/Clinical-Longformer, and the pretrained models available for public download at: https://huggingface.co/yikuan8/Clinical-Longformer. CONCLUSION: This study demonstrates that clinical knowledge-enriched long-sequence transformers are able to learn long-term dependencies in long clinical text. Our methods can also inspire the development of other domain-enriched long-sequence transformers. Yikuan Li, Ramsey M. Wehbe, Faraz S. Ahmad, Hanyin Wang, Yuan Luo 0001 |
J. Am. Medical Informatics Assoc. | 5 |
| 2023 | Patterns of diverse and changing sentiments towards COVID-19 vaccines: a sentiment analysis study integrating 11 million tweets and surveillance data across over 180 countriesabstractOBJECTIVES: Vaccines are crucial components of pandemic responses. Over 12 billion coronavirus disease 2019 (COVID-19) vaccines were administered at the time of writing. However, public perceptions of vaccines have been complex. We integrated social media and surveillance data to unravel the evolving perceptions of COVID-19 vaccines. MATERIALS AND METHODS: Applying human-in-the-loop deep learning models, we analyzed sentiments towards COVID-19 vaccines in 11 211 672 tweets of 2 203 681 users from 2020 to 2022. The diverse sentiment patterns were juxtaposed against user demographics, public health surveillance data of over 180 countries, and worldwide event timelines. A subanalysis was performed targeting the subpopulation of pregnant people. Additional feature analyses based on user-generated content suggested possible sources of vaccine hesitancy. RESULTS: Our trained deep learning model demonstrated performances comparable to educated humans, yielding an accuracy of 0.92 in sentiment analysis against our manually curated dataset. Albeit fluctuations, sentiments were found more positive over time, followed by a subsequence upswing in population-level vaccine uptake. Distinguishable patterns were revealed among subgroups stratified by demographic variables. Encouraging news or events were detected surrounding positive sentiments crests. Sentiments in pregnancy-related tweets demonstrated a lagged pattern compared with the general population, with delayed vaccine uptake trends. Feature analysis detected hesitancies stemmed from clinical trial logics, risks and complications, and urgency of scientific evidence. DISCUSSION: Integrating social media and public health surveillance data, we associated the sentiments at individual level with observed populational-level vaccination patterns. By unraveling the distinctive patterns across subpopulations, the findings provided evidence-based strategies for improving vaccine promotion during pandemics. Hanyin Wang, Yikuan Li, Meghan Hutch, Adrienne S. Kline, Sebastian Otero, Leena B. Mithal, Emily S. Miller, Andrew Naidech, Yuan Luo 0001 |
J. Am. Medical Informatics Assoc. | 9 |
| 2023 | AD-BERT: Using pre-trained language model to predict the progression from mild cognitive impairment to Alzheimer's disease
Chengsheng Mao, Jie Xu 0012, Luke V. Rasmussen, Yikuan Li, Prakash Adekkanattu, Jennifer A. Pacheco, Borna Bonakdarpour, Robert Vassar, Li Shen 0001, Guoqian Jiang, Fei Wang 0001, Jyotishman Pathak, Yuan Luo 0001 |
J. Biomed. Informatics | 13 |
| 2023 | Informative missingness: What can we learn from patterns in missing laboratory data in the electronic health record?
Amelia L. M. Tan, Emily J. Getzen, Meghan Hutch, Zachary H. Strasser, Alba Gutiérrez-Sacristán, Trang T. Le, Arianna Dagliati, Michele Morris, David A. Hanauer, Bertrand Moal, Clara-Lea Bonzel, William Yuan, Lorenzo Chiudinelli, Priyam Das, Harrison G. Zhang, Bruce J. Aronow, Paul Avillach, Gabriel A. Brat, Tianxi Cai, Chuan Hong, William G. La Cava, He Hooi Will Loh, Yuan Luo 0001, Shawn N. Murphy, Kee Yuan Hgiam, Gilbert S. Omenn, Lav P. Patel, Malarkodi J. Samayamuthu, Emily R. Shriver, Zahra Shakeri Hossein Abad, Byorn W. L. Tan, Shyam Visweswaran, Griffin M. Weber, Zongqi Xia, Bertrand Verdy, Qi Long, Danielle L. Mowery, John H. Holmes |
J. Biomed. Informatics | 23 |
| 2023 | Special issue on fairness and inclusion in biomedical informatics research: technical and social perspectives
Shyam Visweswaran, Yuan Luo 0001, Mor Peleg |
J. Biomed. Informatics | 2 |
| 2022 | Evaluation of ICU Severity Scores and Mortality Prediction by Race and Ethnicity: A Machine Learning Approach Using MIMIC-IV
Catherine A. Gao, Plamena Powla, Meghan Hutch, Oluwatosin Akinsola, Yuan Luo 0001 |
AMIA | 5 |
| 2022 | Aggregation Delayed Federated LearningabstractFederated learning is a distributed machine learning paradigm where multiple data owners (clients) collaboratively train one machine learning model while keeping data on their own devices. The heterogeneity of client datasets is one of the most important challenges of federated learning algorithms. Studies have found performance reduction with standard federated algorithms, such as FedAvg, on non-IID data. Many existing works on handling non-IID data adopt the same aggregation framework as FedAvg and focus on improving model updates either on the server side or on clients. In this work, we tackle this challenge in a different view by introducing redistribution rounds that delay the aggregation. With delayed aggregations, local models are trained on data that are more representative to the global distribution. The proposed algorithm can also be used as a federated learning paradigm, as an alternative to FedAvg, where other methods can be plugged in. We perform experiments on multiple tasks and show that the proposed framework significantly improves the performance on non-IID data. Ye Xue, Diego Klabjan, Yuan Luo 0001 |
IEEE Big Data | 3 |
| 2022 | Improving Graph Representation Learning with Distribution PreservingabstractGraph neural network (GNN) is effective to model graphs for distributed representations of nodes and an entire graph. Recently, research on the expressive power of GNN attracted growing attention. A highly expressive GNN has the ability to generate discriminative graph representations. However, in the end-to-end training process for a certain graph learning task, an expressive GNN could generate graph representations overfitting the training data for the target task but losing information important for the model generalization, thus reducing the generalizability. In this paper, we propose Distribution Preserving GNN (DP-GNN), a GNN framework that can improve the generalizability of expressive GNN models by preserving several kinds of distribution information in graph representations and node representations. Besides the generalizability, by applying an expressive GNN backbone, DP-GNN can also have high expressive power. We evaluate the proposed DP-GNN framework on multiple benchmark datasets for graph classification tasks. The experimental results demonstrate that our model achieves state-of-the-art performances. Chengsheng Mao, Yuan Luo 0001 |
ICDM | 2 |
| 2022 | Evaluating the state of the art in missing data imputation for clinical dataabstractClinical data are increasingly being mined to derive new medical knowledge with a goal of enabling greater diagnostic precision, better-personalized therapeutic regimens, improved clinical outcomes and more efficient utilization of health-care resources. However, clinical data are often only available at irregular intervals that vary between patients and type of data, with entries often being unmeasured or unknown. As a result, missing data often represent one of the major impediments to optimal knowledge derivation from clinical data. The Data Analytics Challenge on Missing data Imputation (DACMI) presented a shared clinical dataset with ground truth for evaluating and advancing the state of the art in imputing missing data for clinical time series. We extracted 13 commonly measured blood laboratory tests. To evaluate the imputation performance, we randomly removed one recorded result per laboratory test per patient admission and used them as the ground truth. DACMI is the first shared-task challenge on clinical time series imputation to our best knowledge. The challenge attracted 12 international teams spanning three continents across multiple industries and academia. The evaluation outcome suggests that competitive machine learning and statistical models (e.g. LightGBM, MICE and XGBoost) coupled with carefully engineered temporal and cross-sectional features can achieve strong imputation performance. However, care needs to be taken to prevent overblown model complexity. The challenge participating systems collectively experimented with a wide range of machine learning and probabilistic algorithms to combine temporal imputation and cross-sectional imputation, and their design principles will inform future efforts to better model clinical missing data. Yuan Luo 0001 |
Briefings Bioinform. | 1 |
| 2022 | Design and validation of a FHIR-based EHR-driven phenotyping toolboxabstractOBJECTIVES: To develop and validate a standards-based phenotyping tool to author electronic health record (EHR)-based phenotype definitions and demonstrate execution of the definitions against heterogeneous clinical research data platforms. MATERIALS AND METHODS: We developed an open-source, standards-compliant phenotyping tool known as the PhEMA Workbench that enables a phenotype representation using the Fast Healthcare Interoperability Resources (FHIR) and Clinical Quality Language (CQL) standards. We then demonstrated how this tool can be used to conduct EHR-based phenotyping, including phenotype authoring, execution, and validation. We validated the performance of the tool by executing a thrombotic event phenotype definition at 3 sites, Mayo Clinic (MC), Northwestern Medicine (NM), and Weill Cornell Medicine (WCM), and used manual review to determine precision and recall. RESULTS: An initial version of the PhEMA Workbench has been released, which supports phenotype authoring, execution, and publishing to a shared phenotype definition repository. The resulting thrombotic event phenotype definition consisted of 11 CQL statements, and 24 value sets containing a total of 834 codes. Technical validation showed satisfactory performance (both NM and MC had 100% precision and recall and WCM had a precision of 95% and a recall of 84%). CONCLUSIONS: We demonstrate that the PhEMA Workbench can facilitate EHR-driven phenotype definition, execution, and phenotype sharing in heterogeneous clinical research data environments. A phenotype definition that integrates with existing standards-compliant systems, and the use of a formal representation facilitates automation and can decrease potential for human error. Pascal S. Brandt, Jennifer A. Pacheco, Prakash Adekkanattu, Evan Sholle, Sajjad Abedian, Daniel J. Stone, David Knaack, Jie Xu 0012, Yifan Peng 0002, Natalie C. Benda, Fei Wang 0001, Yuan Luo 0001, Guoqian Jiang, Jyotishman Pathak, Luke V. Rasmussen |
J. Am. Medical Informatics Assoc. | 13 |
| 2022 | MedGCN: Medication recommendation and lab test imputation via graph convolutional networks
Chengsheng Mao, Yuan Luo 0001 |
J. Biomed. Informatics | 3 |
| 2022 | SurvMaximin: Robust federated approach to transporting survival risk prediction models
Harrison G. Zhang, Xin Xiong 0006, Chuan Hong, Griffin M. Weber, Gabriel A. Brat, Clara-Lea Bonzel, Yuan Luo 0001, Rui Duan 0004, Nathan P. Palmer, Meghan Hutch, Alba Gutiérrez-Sacristán, Riccardo Bellazzi, Luca Chiovato, Kelly Cho, Arianna Dagliati, Hossein Estiri, Noelia García-Barrio, Romain Griffier, David A. Hanauer, Yuk-Lam Ho, John H. Holmes, Mark S. Keller, Jeffrey G. Klann, Sehi L'Yi, Sara Lozano-Zahonero, Sarah E. Maidlow, Adeline Makoudjou, Alberto Malovini, Bertrand Moal, Jason H. Moore, Michele Morris, Danielle L. Mowery, Shawn N. Murphy, Antoine Neuraz, Kee Yuan Ngiam, Gilbert S. Omenn, Lav P. Patel, Miguel Pedrera-Jiménez, Andrea Prunotto, Malarkodi J. Samayamuthu, Fernando J. Sanz Vidorreta, Emily Schriver, Petra Schubert, Pablo Serrano-Balazote, Andrew M. South, Amelia L. M. Tan, Byorn W. L. Tan, Valentina Tibollo, Patric Tippmann, Shyam Visweswaran, Zongqi Xia, William Yuan, Daniela Zöller, Isaac S. Kohane, Paul Avillach, Zijian Guo 0003, Tianxi Cai |
J. Biomed. Informatics | 8 |
| 2022 | ImageGCN: Multi-Relational Image Graph Convolutional Networks for Disease Identification With Chest X-RaysabstractImage representation is a fundamental task in computer vision. However, most of the existing approaches for image representation ignore the relations between images and consider each input image independently. Intuitively, relations between images can help to understand the images and maintain model consistency over related images, leading to better explainability. In this paper, we consider modeling the image-level relations to generate more informative image representations, and propose ImageGCN, an end-to-end graph convolutional network framework for inductive multi-relational image modeling. We apply ImageGCN to chest X-ray images where rich relational information is available for disease identification. Unlike previous image representation models, ImageGCN learns the representation of an image using both its original pixel features and its relationship with other images. Besides learning informative representations for images, ImageGCN can also be used for object detection in a weakly supervised manner. The experimental results on 3 open-source x-ray datasets, ChestX-ray14, CheXpert and MIMIC-CXR demonstrate that ImageGCN can outperform respective baselines in both disease identification and localization tasks and can achieve comparable and often better results than the state-of-the-art methods. Chengsheng Mao, Yuan Luo 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2021 | PANTHER: Pathway Augmented Nonnegative Tensor Factorization for HighER-order Feature LearningabstractGenetic pathways usually encode molecular mechanisms that can inform targeted interventions. It is often challenging for existing machine learning approaches to jointly model genetic pathways (higher-order features) and variants (atomic features), and present to clinicians interpretable models. In order to build more accurate and better interpretable machine learning models for genetic medicine, we introduce Pathway Augmented Nonnegative Tensor factorization for HighER-order feature learning (PANTHER). PANTHER selects informative genetic pathways that directly encode molecular mechanisms. We apply genetically motivated constrained tensor factorization to group pathways in a way that reflects molecular mechanism interactions. We then train a softmax classifier for disease types using the identified pathway groups. We evaluated PANTHER against multiple state-of-the-art constrained tensor/matrix factorization models, as well as group guided and Bayesian hierarchical models. PANTHER outperforms all state-of-the-art comparison models significantly (p Yuan Luo 0001, Chengsheng Mao |
AAAI | 1 |
| 2021 | Open-Set Recognition with Gaussian Mixture Variational AutoencodersabstractIn inference, open-set classification is to either classify a sample into a known class from training or reject it as an unknown class. Existing deep open-set classifiers train explicit closed-set classifiers, in some cases disjointly utilizing reconstruction, which we find dilutes the latent representation's ability to distinguish unknown classes. In contrast, we train our model to cooperatively learn reconstruction and perform class-based clustering in the latent space. With this, our Gaussian mixture variational autoencoder (GMVAE) achieves more accurate and robust open-set classification results, with an average F1 increase of 0.26, through extensive experiments aided by analytical results. Alexander Cao, Yuan Luo 0001, Diego Klabjan |
AAAI | 2 |
| 2021 | Multi-site Evaluation of Longitudinal Changes in Ejection Fraction in Heart Failure Patients Through Data-driven Phenotyping
Prakash Adekkanattu, Jennifer A. Pacheco, Joseph Kabariti, Daniel J. Stone, Yue Yu 0012, Parag Goyal, Faraz S. Ahmad, Guoqian Jiang, Yuan Luo 0001, Luke V. Rasmussen, Pascal S. Brandt, Jie Xu 0012, Fei Wang 0001, Natalie C. Benda, Thomas R. Campion Jr., Jyotishman Pathak |
AMIA | 9 |
| 2021 | Addressing Bias in the Application of Machine Learning on Real-World Data
Hossein Estiri, Yuan Luo 0001, Suzanne Tamang, Harold P. Lehmann |
AMIA | 2 |
| 2021 | Multi-Modal Data Science for Healthcare: State of the Art, Challenges, and Opportunities
Yuan Luo 0001, Fei Wang 0001, Benjamin S. Glicksberg, Jessilyn Dunn, Nigam H. Shah |
AMIA | 1 |
| 2021 | Graph Based Machine Learning for Healthcare: State of the Art, Challenges, and Opportunities
Yuan Luo 0001, Fei Wang 0001, Marinka Zitnik, Shuiwang Ji |
AMIA | 1 |
| 2021 | A Deep Learning Framework Using a Pre-trained BERT Model to Predict the Risk of Progression from Mild Cognitive Impairment to Alzheimer's Disease
Chengsheng Mao, Jie Xu 0012, Luke V. Rasmussen, Jennifer A. Pacheco, Guoqian Jiang, Fei Wang 0001, Richard Isaacson, Jyotishman Pathak, Yuan Luo 0001 |
AMIA | 9 |
| 2021 | Evaluation of the Portability of Natural Language Processing-based Computable Phenotypes in the eMERGE Network
Jennifer A. Pacheco, Luke V. Rasmussen, Ken Wiley, Thomas N. Person, David J. Cronkite, Sunghwan Sohn, Shawn N. Murphy, Justin H. Gundelach, Vivian S. Gainer, Victor M. Castro, Cong Liu 0020, Todd Lingren, Frank D. Mentch, Agnes S. Sundaresan, Garrett Eickelberg, Valerie Willis, Al'ona Furmanchuk, Roshan Patel, David Carrell, Marc S. Williams, Elizabeth W. Karlson, Jodell E. Linder, Yuan Luo 0001, Chunhua Weng, Wei-Qi Wei |
AMIA | 23 |
| 2021 | FHIRTime: Standardizing Temporal Patterns Identified from Clinical Narratives Using HL7 FHIR
Daniel J. Stone, Sijia Liu 0002, Yuan Luo 0001, Andrew Wen, Nansu Zong, Luke V. Rasmussen, Prakash Adekkanattu, Pascal S. Brandt, Jennifer A. Pacheco, Fei Wang 0001, Cui Tao, Jyotishman Pathak, Guoqian Jiang |
AMIA | 3 |
| 2021 | On Constraints and Considerations for Extending Support for Natural Language Processing-Based FHIR Resource Generation
Andrew Wen, Luke V. Rasmussen, Daniel J. Stone, Sijia Liu 0002, Prakash Adekkanattu, Pascal S. Brandt, Jennifer A. Pacheco, Yuan Luo 0001, Fei Wang 0001, Jyotishman Pathak, Guoqian Jiang |
AMIA | 8 |
| 2021 | Unsupervised clustering analysis of SARS-Cov-2 population structure reveals six major subtypes at early stage across the worldabstractthe population structure of the newly emerged coronavirus SARS-CoV-2 has significant potential to inform public health management and diagnosis. As SARS-CoV-2 sequencing data accrued, grouping them into clusters is important for organizing the landscape of the population structure of the virus. Due to the limited prior information on the newly emerged coronavirus, we utilized four different clustering algorithms to group 16, S73 SARS-CoV-2 strains, which automatically enables the identification of spatial structure for SARS-CoV-2. A total of six distinct genomic clusters were identified using mutation profiles as input features. Comparison of the clustering results reveals that the four algorithms produced highly consistent results, but the state-of-the-art unsupervised deep learning clustering algorithm performed best and produced the smallest intra-cluster pairwise genetic distances. The varied proportions of the six clusters within different continents revealed specific geographical distributions. In particular, our analysis found that Oceania was the only continent on which the strains were dispersively distributed into six clusters. In summary, this study provides a concrete framework for the use of clustering methods to study the global population structure of SARS-CoV-2. In addition, clustering methods can be used for future studies of variant population structures in specific regions of these fast-growing viruses. Yawei Li 0002, Qingyun Liu 0009, Zexian Zeng, Yuan Luo 0001 |
BIBM | 4 |
| 2021 | SNPs Filtered by Allele Frequency Improve the Prediction of Hypertension SubtypesabstractHypertension is the leading global cause of cardiovascular disease and premature death. Distinct hypertension subtypes may vary in their prognoses and require different treatments. An individual’s risk for hypertension is determined by genetic and environmental factors as well as their interactions. In this work, we studied 911 African Americans and 1,171 European Americans in the Hypertension Genetic Epidemiology Network (HyperGEN) cohort. We built hypertension subtype classification models using both environmental variables and sets of genetic features selected based on different criteria. The fitted prediction models provided insights into the genetic landscape of hypertension subtypes, which may aid personalized diagnosis and treatment of hypertension in the future. Sanjiv J. Shah, Donna Arnett, Ryan Irvin, Yuan Luo 0001 |
BIBM | 5 |
| 2021 | Early Prediction of Mortality in Critical Care Setting in Sepsis Patients Using Structured Features and Unstructured Clinical NotesabstractSepsis is an important cause of mortality, especially in intensive care unit(ICU) patients. Developing novel methods to identify early mortality is critical for improving survival outcomes in sepsis patients. Using the MIMIC-III database, we integrated demographic data, physiological measurements and clinical notes. We built and applied several machine learning models to predict the risk of hospital mortality and 30-day mortality in sepsis patients. From the clinical notes, we generated clinically meaningful word representations and embeddings. Supervised learning classifiers and a deep learning architecture were used to construct prediction models. The configurations that utilized both structured and unstructured clinical features yielded competitive F-measure of 0.512. Our results showed that the approaches integrating both structured and unstructured clinical features can be effectively applied to assist clinicians in identifying the risk of mortality in sepsis patients upon admission to the ICU. Jiyoung Shin, Yikuan Li, Yuan Luo 0001 |
BIBM | 3 |
| 2021 | COVID Vaccine and Cardiovascular Risks: A Natural Language Analysis of Vaccine Adverse Event ReportsabstractAdverse events (AEs) following COVID vaccination have been intensely monitored. In our study, we developed a natural language processing system to analyze data from a spontaneous reporting system - Vaccine Adverse Event Reporting System and detect signals of AEs following administration of COVID vaccines. Our system included several components to magnify novel and rare AEs, including 1) excluding COVID positive patients, 2) excluding sentences discussing disease history or family history, 3) standardizing symptom concepts into 30 major AEs, 4) using influenza vaccine recipients as control group when calculating reporting odds ratio. We identified several cardiovascular and inflammatory-related AEs that demonstrated high odds ratio. We demonstrated our system can serve as a complementary system to identify and monitor AEs outside of pre-defined outcomes routinely monitored by existing databases or projects. Michael G. Ison, Yuan Luo 0001 |
BIBM | 3 |
| 2021 | Unsupervised Learning to Subphenotype Delirium Patients from Electronic Health RecordsabstractDelirium is a common acute onset brain dysfunction in the emergency setting and is associated with higher mortality. It is difficult to detect and monitor since its presentations and risk factors can be different depending on the underlying medical condition of patients. In our study, we aimed to identify subtypes within the delirium population and build subgroup-specific predictive models to detect delirium using Medical Information Mart for Intensive Care IV (MIMIC-IV) data. We showed that clusters exist within the delirium population. Differences in feature importance were also observed for subgroup-specific predictive models. Our work could recalibrate existing delirium prediction models for each delirium subgroup and improve the precision of delirium detection and monitoring for ICU or emergency department patients who had highly heterogeneous medical conditions. Yuan Luo 0001 |
BIBM | 2 |
| 2021 | A deep-learning-based unsupervised model on esophageal manometry using variational autoencoder
Wenjun Kou, Dustin A. Carlson, Alexandra J. Baumann, Erica Donnan, Yuan Luo 0001, John E. Pandolfino, Mozziyar Etemadi |
Artif. Intell. Medicine | 5 |
| 2021 | Deep learning for cancer type classification and driver gene identificationabstractBACKGROUND: Genetic information is becoming more readily available and is increasingly being used to predict patient cancer types as well as their subtypes. Most classification methods thus far utilize somatic mutations as independent features for classification and are limited by study power. We aim to develop a novel method to effectively explore the landscape of genetic variants, including germline variants, and small insertions and deletions for cancer type prediction. RESULTS: We proposed DeepCues, a deep learning model that utilizes convolutional neural networks to unbiasedly derive features from raw cancer DNA sequencing data for disease classification and relevant gene discovery. Using raw whole-exome sequencing as features, germline variants and somatic mutations, including insertions and deletions, were interactively amalgamated for feature generation and cancer prediction. We applied DeepCues to a dataset from TCGA to classify seven different types of major cancers and obtained an overall accuracy of 77.6%. We compared DeepCues to conventional methods and demonstrated a significant overall improvement (p < 0.001). Strikingly, using DeepCues, the top 20 breast cancer relevant genes we have identified, had a 40% overlap with the top 20 known breast cancer driver genes. CONCLUSION: Our results support DeepCues as a novel method to improve the representational resolution of DNA sequencings and its power in deriving features from raw sequences for cancer type prediction, as well as discovering new cancer relevant genes. Zexian Zeng, Chengsheng Mao, Andy H. Vo, Xiaoyu Li 0006, Janna Ore Nugent, Seema A. Khan, Susan E. Clare, Yuan Luo 0001 |
BMC Bioinform. | 8 |
| 2021 | Validation of an internationally derived patient severity phenotype to support COVID-19 analytics from electronic health record dataabstractOBJECTIVE: The Consortium for Clinical Characterization of COVID-19 by EHR (4CE) is an international collaboration addressing coronavirus disease 2019 (COVID-19) with federated analyses of electronic health record (EHR) data. We sought to develop and validate a computable phenotype for COVID-19 severity. MATERIALS AND METHODS: Twelve 4CE sites participated. First, we developed an EHR-based severity phenotype consisting of 6 code classes, and we validated it on patient hospitalization data from the 12 4CE clinical sites against the outcomes of intensive care unit (ICU) admission and/or death. We also piloted an alternative machine learning approach and compared selected predictors of severity with the 4CE phenotype at 1 site. RESULTS: The full 4CE severity phenotype had pooled sensitivity of 0.73 and specificity 0.83 for the combined outcome of ICU admission and/or death. The sensitivity of individual code categories for acuity had high variability-up to 0.65 across sites. At one pilot site, the expert-derived phenotype had mean area under the curve of 0.903 (95% confidence interval, 0.886-0.921), compared with an area under the curve of 0.956 (95% confidence interval, 0.952-0.959) for the machine learning approach. Billing codes were poor proxies of ICU admission, with as low as 49% precision and recall compared with chart review. DISCUSSION: We developed a severity phenotype using 6 code classes that proved resilient to coding variability across international institutions. In contrast, machine learning approaches may overfit hospital-specific orders. Manual chart review revealed discrepancies even in the gold-standard outcomes, possibly owing to heterogeneous pandemic conditions. CONCLUSIONS: We developed an EHR-based severity phenotype for COVID-19 in hospitalized patients and validated it at 12 international sites. Jeffrey G. Klann, Hossein Estiri, Griffin M. Weber, Bertrand Moal, Paul Avillach, Chuan Hong, Amelia L. M. Tan, Brett K. Beaulieu-Jones, Victor M. Castro, Thomas Maulhardt, Alon Geva, Alberto Malovini, Andrew M. South, Shyam Visweswaran, Michele Morris, Malarkodi J. Samayamuthu, Gilbert S. Omenn, Kee Yuan Ngiam, Kenneth D. Mandl, Martin Boeker, Karen L. Olson, Danielle L. Mowery, Robert W. Follett, David A. Hanauer, Riccardo Bellazzi, Jason H. Moore, Ne-Hooi Will Loh, Douglas S. Bell, Kavishwar B. Wagholikar, Luca Chiovato, Valentina Tibollo, Siegbert Rieg, Anthony L. L. J. Li, Vianney Jouhet, Emily Schriver, Zongqi Xia, Meghan Hutch, Yuan Luo 0001, Isaac S. Kohane, Gabriel A. Brat, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 38 |
| 2021 | Characterizing phenotypic abnormalities associated with high-risk individuals developing lung cancer using electronic health records from the All of Us researcher workbenchabstractOBJECTIVE: The study sought to test the feasibility of conducting a phenome-wide association study to characterize phenotypic abnormalities associated with individuals at high risk for lung cancer using electronic health records. MATERIALS AND METHODS: We used the beta release of the All of Us Researcher Workbench with clinical and survey data from a population of 225 000 subjects. We identified 3 cohorts of individuals at high risk to develop lung cancer based on (1) the 2013 U.S. Preventive Services Task Force criteria, (2) the long-term quitters of cigarette smoking criteria, and (3) the younger age of onset criteria. We applied the logistic regression analysis to identify the significant associations between individuals' phenotypes and their risk categories. We validated our findings against a lung cancer cohort from the same population and conducted an expert review to understand whether these associations are known or potentially novel. RESULTS: We found a total of 214 statistically significant associations (P < .05 with a Bonferroni correction and odds ratio > 1.5) enriched in the high-risk individuals from 3 cohorts, and 15 enriched in the low-risk individuals. Forty significant associations enriched in the high-risk individuals and 13 enriched in the low-risk individuals were validated in the cancer cohort. Expert review identified 15 potentially new associations enriched in the high-risk individuals. CONCLUSIONS: It is feasible to conduct a phenome-wide association study to characterize phenotypic abnormalities associated in high-risk individuals developing lung cancer using electronic health records. The All of Us Research Workbench is a promising resource for the research studies to evaluate and optimize lung cancer screening criteria. Jie Na, Nansu Zong, David E. Midthun, Yuan Luo 0001, Guoqian Jiang |
J. Am. Medical Informatics Assoc. | 5 |
| 2020 | Feasibility of Cross-Platform EHR-Driven Phenotyping Using Clinical Quality Language
Pascal S. Brandt, Richard C. Kiefer, Jennifer A. Pacheco, Prakash Adekkanattu, Evan Sholle, Faraz S. Ahmad, Jie Xu 0012, Jessica S. Ancker, Fei Wang 0001, Yuan Luo 0001, Guoqian Jiang, Jyotishman Pathak, Luke V. Rasmussen |
AMIA | 11 |
| 2020 | Unsupervised learning for systemic lupus erythematosus subtype identification: electronic health record vs registry data
Anika S. Ghosh, Anh H. Chung, Yacob Tedla, Abel N. Kho, Rosalind Ramsey-Goldman, Yuan Luo 0001, Theresa Walunas |
AMIA | 7 |
| 2020 | Identification of Alzheimer's Disease Subtypes from Electronic Health Records Using a Data-Driven Approach
Jie Xu 0012, Fei Wang 0001, Prakash Adekkanattu, Pascal S. Brandt, Guoqian Jiang, Richard C. Kiefer, Yuan Luo 0001, Chengsheng Mao, Jennifer A. Pacheco, Luke V. Rasmussen, Yiye Zhang, Richard Isaacson, Jyotishman Pathak |
AMIA | 8 |
| 2020 | A Predictive Model for Parkinson's Disease Reveals Candidate Gene Sets for Progression SubtypeabstractParkinson's Disease (PD) is the second most common neurodegenerative disease in the United States, and is characterized by the progressive decline of motor and non-motor symptoms. The progression rate and manifestation of PD is highly heterogenous, and the underlying etiology for this heterogeneity remains elusive. Although some studies have identified risk genes associated with the development of PD, it is unknown whether the patient genome influences the progression pattern of PD in any way. In this study, we used the whole-exome sequencing data of PD patients from the Parkinson's Disease Progression Marker Initiative (PPMI) to examine whether an individual's genetic profile is associated with their progression pattern of PD. We used the three distinct progression subtypes defined by Zhang et al. as the outcome variable, and trained logistic regression and support vector classifiers. Our best performing model achieved an area under the receiver operating characteristic curve of 0.69 on the test set, indicating that the genetic profile of a PD patient appears to have some relationship with their likely disease course. We then interpreted our trained model by performing Gene Set Enrichment Analysis on the sets of genes with high model coefficients for each progression subtype. The results showed enrichment for genes related to olfactory signaling in Subtype I and III, and an enrichment for genes related to protein glycosylation and glycation for Subtype I and II. Overall, our findings suggests a connection between an individual's genetic profile and PD progression subtype, and calls for controlled research studies to further examine the relationship between the implicated gene sets and each progression subtype. Saya R. Dennis, Tanya Simuni, Yuan Luo 0001 |
BIBM | 3 |
| 2020 | A Comparison of Pre-trained Vision-and-Language Models for Multimodal Representation Learning across Medical Images and ReportsabstractJoint image-text embedding extracted from medical images and associated contextual reports is the bedrock for most biomedical vision-and-language (V+L) tasks, including medical visual question answering, clinical image-text retrieval, clinical report auto-generation. In this study, we adopt four pre-trained V+L models: LXMERT, VisualBERT, UNIER and PixelBERT to learn multimodal representation from MIMIC-CXR images and associated reports. External evaluation using the OpenI dataset shows that the joint embedding learned by pre-trained V+L models demonstrates performance improvement of 1.4% in thoracic finding classification tasks compared to a pioneering CNN+RNN model. Ablation studies are conducted to further analyze the contribution of certain model components and validate the advantage of joint embedding over text-only embedding. Attention maps are also visualized to illustrate the attention mechanism of V+L models. Yikuan Li, Hanyin Wang, Yuan Luo 0001 |
BIBM | 3 |
| 2020 | Prediction of breast cancer distant recurrence using natural language processing and knowledge-guided convolutional neural network
Hanyin Wang, Yikuan Li, Seema A. Khan, Yuan Luo 0001 |
Artif. Intell. Medicine | 4 |
| 2020 | A novel normalization and differential abundance test framework for microbiome dataabstractMOTIVATION: Microbial communities have been proved to have close relationship with many diseases. The identification of differentially abundant microbial species is clinically meaningful for finding disease-related pathogenic or probiotic bacteria. However, certain characteristics of microbiome data have hurdled the accuracy and effectiveness of differential abundance analysis. The abundances or counts of microbiome species are usually on different scales and exhibit zero-inflation and over-dispersion. Normalization is a crucial step before the differential abundance test. However, existing normalization methods typically try to adjust counts on different scales to a common scale by constructing size factors with the assumption that count distributions across samples are equivalent up to a certain percentile. These methods often yield undesirable results when differentially abundant species are of low to medium abundance level. For differential abundance analysis, existing methods often use a single distribution to model the dispersion of species which lacks flexibility to catch a single species' distinctiveness. These methods tend to detect a lot of false positives and often lack of power when the effect size is small. RESULTS: We develop a novel framework for differential abundance analysis on sparse high-dimensional marker gene microbiome data. Our methodology relies on a novel network-based normalization technique and a two-stage zero-inflated mixture count regression model (RioNorm2). Our normalization method aims to find a group of relatively invariant microbiome species across samples and conditions in order to construct the size factor. Another contribution of the paper is that our testing approach can take under-sampling and over-dispersion into consideration by separating microbiome species into two groups and model them separately. Through comprehensive simulation studies, the performance of our method is consistently powerful and robust across different settings with different sample size, library size and effect size. We also demonstrate the effectiveness of our novel framework using a published dataset of metastatic melanoma and find biological insights from the results. AVAILABILITY AND IMPLEMENTATION: The R package 'RioNorm2' can be installed from Github athttps://github.com/yuanjing-ma/RioNorm2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yuanjing Ma, Yuan Luo 0001, Hongmei Jiang |
Bioinform. | 2 |
| 2020 | Predictive modeling of bacterial infections and antibiotic therapy needs in critically ill adults
Garrett Eickelberg, L. Nelson Sanchez-Pinto, Yuan Luo 0001 |
J. Biomed. Informatics | 3 |
| 2020 | Identifying sub-phenotypes of acute kidney injury using structured and unstructured electronic health record data with memory networks
Jingyuan Chou, Xi Sheryl Zhang, Yuan Luo 0001, Tamara Isakova, Prakash Adekkanattu, Jessica S. Ancker, Guoqian Jiang, Richard C. Kiefer, Jennifer A. Pacheco, Luke V. Rasmussen, Jyotishman Pathak, Fei Wang 0001 |
J. Biomed. Informatics | 4 |
| 2019 | Graph Convolutional Networks for Text ClassificationabstractText classification is an important and classical problem in natural language processing. There have been a number of studies that applied convolutional neural networks (convolution on regular grid, e.g., sequence) to classification. However, only a limited number of studies have explored the more flexible graph convolutional neural networks (convolution on non-grid, e.g., arbitrary graph) for the task. In this work, we propose to use graph convolutional networks for text classification. We build a single text graph for a corpus based on word co-occurrence and document word relations, then learn a Text Graph Convolutional Network (Text GCN) for the corpus. Our Text GCN is initialized with one-hot representation for word and document, it then jointly learns the embeddings for both words and documents, as supervised by the known class labels for documents. Our experimental results on multiple benchmark datasets demonstrate that a vanilla Text GCN without any external word embeddings or knowledge outperforms state-of-the-art methods for text classification. On the other hand, Text GCN also learns predictive word and document embeddings. In addition, experimental results show that the improvement of Text GCN over state-of-the-art comparison methods become more prominent as we lower the percentage of training data, suggesting the robustness of Text GCN to less training data in text classification. Chengsheng Mao, Yuan Luo 0001 |
AAAI | 3 |
| 2019 | Evaluating the Portability of an NLP System for Processing Echocardiograms: A Retrospective, Multi-site Observational Study
Prakash Adekkanattu, Guoqian Jiang, Yuan Luo 0001, Paul R. Kingsbury, Luke V. Rasmussen, Jennifer A. Pacheco, Richard C. Kiefer, Daniel J. Stone, Pascal S. Brandt, Yizhen Zhong, Fei Wang 0001, Jessica S. Ancker, Thomas R. Campion Jr., Jyotishman Pathak |
AMIA | 3 |
| 2019 | Considerations for Improving the Portability of Electronic Health Record-Based Phenotype Algorithms
Luke V. Rasmussen, Pascal S. Brandt, Guoqian Jiang, Richard C. Kiefer, Jennifer A. Pacheco, Prakash Adekkanattu, Jessica S. Ancker, Fei Wang 0001, Jyotishman Pathak, Yuan Luo 0001 |
AMIA | 11 |
| 2019 | Phenotyping Multiple Organ Dysfunction Syndrome Using Temporal Trends in Critically Ill ChildrenabstractMultiple organ dysfunction syndrome (MODS) is one of the most common causes of death in critically ill children. However, despite decades of clinical trials, there are no comprehensive approaches to the management of MODS or effective targeted therapies that have consistently improved outcomes. Better understanding the heterogeneity of MODS and characterizing subgroups of MODS patients could improve our understanding of the syndrome and help us develop new management strategies. We analyzed a cohort of 5,297 children with MODS from two children's hospitals and used subgraph-augmented non-negative matrix factorization (SANMF) to identify unique temporal patterns in organ dysfunction across four novel subgroups. We demonstrate that these subgroups are composed of patients with distinct clinical characteristics and are independently predictive of clinical outcomes. Our work suggests that these subgroups represent four relevant phenotypes of pediatric MODS that could be used to identify novel management strategies. Emily Kunce Stroup, Yuan Luo 0001, L. Nelson Sanchez-Pinto |
BIBM | 2 |
| 2019 | Using Machine Learning to Predict Hyperchloremia in Critically Ill PatientsabstractElevated serum chloride levels (hyperchloremia) and the administration of intravenous (IV) fluids with high chloride content have both been associated with increased morbidity and mortality in certain subgroups of critically ill patients, such as those with sepsis. Here, we demonstrate this association in a general intensive care unit (ICU) population using data from the Medical Information Mart for Intensive Care III (MIMIC-III) database and propose the use of supervised learning to predict hyperchloremia in critically ill patients. Clinical variables from records of the first 24h of adult ICU stays were represented as features for four predictive supervised learning classifiers. The best performing model was able to predict second-day hyperchloremia with an AUC of 0.80 and a ratio of 5 false alerts for every true alert, which is a clinically-actionable rate. Our results suggest that clinicians can be effectively alerted to patients at risk of developing hyperchloremia, providing an opportunity to mitigate this risk and potentially improve outcomes. Pete Yeh, Yiheng Pan, L. Nelson Sanchez-Pinto, Yuan Luo 0001 |
BIBM | 4 |
| 2019 | Mixture-based Multiple Imputation Model for Clinical Data with a Temporal DimensionabstractThe problem of missing values in multivariable time series is a key challenge in many applications such as clinical data mining. Although many imputation methods show their effectiveness in many applications, few of them are designed to accommodate clinical multivariable time series. In this work, we propose a multiple imputation model that capture both cross-sectional information and temporal correlations. We integrate Gaussian processes with mixture models and introduce individualized mixing weights to handle the variance of predictive confidence of Gaussian process models. The proposed model is compared with several state-of-the-art imputation algorithms on both real-world and synthetic datasets. Experiments show that our best model can provide more accurate imputation than the benchmarks on all of our datasets. Ye Xue, Diego Klabjan, Yuan Luo 0001 |
IEEE BigData | 3 |
| 2019 | Predicting ICU readmission using grouped physiological and medication trends
Ye Xue, Diego Klabjan, Yuan Luo 0001 |
Artif. Intell. Medicine | 3 |
| 2019 | Integrating hypertension phenotype and genotype with hybrid non-negative matrix factorizationabstractMOTIVATION: Hypertension is a heterogeneous syndrome in need of improved subtyping using phenotypic and genetic measurements with the goal of identifying subtypes of patients who share similar pathophysiologic mechanisms and may respond more uniformly to targeted treatments. Existing machine learning approaches often face challenges in integrating phenotype and genotype information and presenting to clinicians an interpretable model. We aim to provide informed patient stratification based on phenotype and genotype features. RESULTS: In this article, we present a hybrid non-negative matrix factorization (HNMF) method to integrate phenotype and genotype information for patient stratification. HNMF simultaneously approximates the phenotypic and genetic feature matrices using different appropriate loss functions, and generates patient subtypes, phenotypic groups and genetic groups. Unlike previous methods, HNMF approximates phenotypic matrix under Frobenius loss, and genetic matrix under Kullback-Leibler (KL) loss. We propose an alternating projected gradient method to solve the approximation problem. Simulation shows HNMF converges fast and accurately to the true factor matrices. On a real-world clinical dataset, we used the patient factor matrix as features and examined the association of these features with indices of cardiac mechanics. We compared HNMF with six different models using phenotype or genotype features alone, with or without NMF, or using joint NMF with only one type of loss We also compared HNMF with 3 recently published methods for integrative clustering analysis, including iClusterBayes, Bayesian joint analysis and JIVE. HNMF significantly outperforms all comparison models. HNMF also reveals intuitive phenotype-genotype interactions that characterize cardiac abnormalities. AVAILABILITY AND IMPLEMENTATION: Our code is publicly available on github at https://github.com/yuanluo/hnmf. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yuan Luo 0001, Chengsheng Mao, Yiben Yang, Fei Wang 0001, Faraz S. Ahmad, Donna Arnett, Marguerite R. Irvin, Sanjiv J. Shah |
Bioinform. | 1 |
| 2019 | Integrating hypertension phenotype and genotype with hybrid non-negative matrix factorizationabstractBioinformatics (2018) doi: 10.1093/bioinformatics/bty804 In the abstract, the availability and implementation section has been updated to include the link for the code, as follows: Our code is publicly available on github at https://github.com/yuanluo/hnmf. Yuan Luo 0001, Chengsheng Mao, Yiben Yang, Fei Wang 0001, Faraz S. Ahmad, Donna Arnett, Marguerite R. Irvin, Sanjiv J. Shah |
Bioinform. | 1 |
| 2019 | Classifying relations in clinical narratives using segment graph convolutional and recurrent neural networks (Seg-GCRNs)abstractWe propose to use segment graph convolutional and recurrent neural networks (Seg-GCRNs), which use only word embedding and sentence syntactic dependencies, to classify relations from clinical notes without manual feature engineering. In this study, the relations between 2 medical concepts are classified by simultaneously learning representations of text segments in the context of sentence syntactic dependency: preceding, concept1, middle, concept2, and succeeding segments. Seg-GCRN was systematically evaluated on the i2b2/VA relation classification challenge datasets. Experiments show that Seg-GCRN attains state-of-the-art micro-averaged F-measure for all 3 relation categories: 0.692 for classifying medical treatment-problem relations, 0.827 for medical test-problem relations, and 0.741 for medical problem-medical problem relations. Comparison with the previous state-of-the-art segment convolutional neural network (Seg-CNN) suggests that adding syntactic dependency information helps refine medical word embedding and improves concept relation classification without manual feature engineering. Seg-GCRN can be trained efficiently for the i2b2/VA dataset on a GPU platform. Yuan Luo 0001 |
J. Am. Medical Informatics Assoc. | 3 |
| 2019 | Traditional Chinese medicine clinical records classification with BERT and domain specific corporaabstractTraditional Chinese Medicine (TCM) has been developed for several thousand years and plays a significant role in health care for Chinese people. This paper studies the problem of classifying TCM clinical records into 5 main disease categories in TCM. We explored a number of state-of-the-art deep learning models and found that the recent Bidirectional Encoder Representations from Transformers can achieve better results than other deep learning models and other state-of-the-art methods. We further utilized an unlabeled clinical corpus to fine-tune the BERT language model before training the text classifier. The method only uses Chinese characters in clinical text as input without preprocessing or feature engineering. We evaluated deep learning models and traditional text classifiers on a benchmark data set. Our method achieves a state-of-the-art accuracy 89.39% ± 0.35%, Macro F1 score 88.64% ± 0.40% and Micro F1 score 89.39% ± 0.35%. We also visualized attention weights in our method, which can reveal indicative characters in clinical text. Zhe Jin 0003, Chengsheng Mao, Yin Zhang 0006, Yuan Luo 0001 |
J. Am. Medical Informatics Assoc. | 5 |
| 2019 | Developing a FHIR-based EHR phenotyping framework: A case study for identification of patients with obesity and multiple comorbidities from discharge summaries
Na Hong, Andrew Wen, Daniel J. Stone, Shintaro Tsuji, Paul R. Kingsbury, Luke V. Rasmussen, Jennifer A. Pacheco, Prakash Adekkanattu, Fei Wang 0001, Yuan Luo 0001, Jyotishman Pathak, Guoqian Jiang |
J. Biomed. Informatics | 10 |
| 2019 | Making work visible for electronic phenotype implementation: Lessons learned from the eMERGE network
Ning Shang 0004, Cong Liu 0020, Luke V. Rasmussen, Casey N. Ta, Robert J. Carroll, Barbara Benoit, Todd Lingren, Ozan Dikilitas, Frank D. Mentch, David Carrell, Wei-Qi Wei, Yuan Luo 0001, Vivian S. Gainer, Iftikhar J. Kullo, Jennifer A. Pacheco, Hakon Hakonarson, Theresa Walunas, Joshua C. Denny, Chunhua Weng |
J. Biomed. Informatics | 12 |
| 2019 | Are My EHRs Private Enough? Event-Level Privacy ProtectionabstractPrivacy is a major concern in sharing human subject data to researchers for secondary analyses. A simple binary consent (opt-in or not) may significantly reduce the amount of sharable data, since many patients might only be concerned about a few sensitive medical conditions rather than the entire medical records. We propose event-level privacy protection, and develop a feature ablation method to protect event-level privacy in electronic medical records. Using a list of 13 sensitive diagnoses, we evaluate the feasibility and the efficacy of the proposed method. As feature ablation progresses, the identifiability of a sensitive medical condition decreases with varying speeds on different diseases. We find that these sensitive diagnoses can be divided into three categories: (1) five diseases have fast declining identifiability (AUC below 0.6 with less than 400 features excluded); (2) seven diseases with progressively declining identifiability (AUC below 0.7 with between 200 and 700 features excluded); and (3) one disease with slowly declining identifiability (AUC above 0.7 with 1,000 features excluded). The fact that the majority (12 out of 13) of the sensitive diseases fall into the first two categories suggests the potential of the proposed feature ablation method as a solution for event-level record privacy protection. Chengsheng Mao, Mengxin Sun, Yuan Luo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2019 | Natural Language Processing for EHR-Based Computational PhenotypingabstractThis article reviews recent advances in applying natural language processing (NLP) to Electronic Health Records (EHRs) for computational phenotyping. NLP-based computational phenotyping has numerous applications including diagnosis categorization, novel phenotype discovery, clinical trial screening, pharmacogenomics, drug-drug interaction (DDI), and adverse drug event (ADE) detection, as well as genome-wide and phenome-wide association studies. Significant progress has been made in algorithm development and resource construction for computational phenotyping. Among the surveyed methods, well-designed keyword search and rule-based systems often achieve good performance. However, the construction of keyword and rule lists requires significant manual effort, which is difficult to scale. Supervised machine learning models have been favored because they are capable of acquiring both classification patterns and structures from data. Recently, deep learning and unsupervised learning have received growing attention, with the former favored for its performance and the latter for its ability to find novel phenotypes. Integrating heterogeneous data sources have become increasingly important and have shown promise in improving model performance. Often, better performance is achieved by combining multiple modalities of information. Despite these many advances, challenges and opportunities remain for NLP-based computational phenotyping, including better model interpretability and generalizability, and proper characterization of feature relations in clinical narratives. Zexian Zeng, Xiaoyu Li 0006, Tristan Naumann, Yuan Luo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2018 | Multi-View Graph Convolutional Network and Its Applications on Neuroimage Analysis for Parkinson's Disease
Lifang He 0001, Kun Chen 0002, Yuan Luo 0001, Fei Wang 0001 |
AMIA | 4 |
| 2018 | Supervised Nonnegative Matrix Factorization to Predict ICU Mortality Risk
Guoqing Chao, Chengsheng Mao, Fei Wang 0001, Yuan Luo 0001 |
BIBM | 5 |
| 2018 | Early Prediction of Acute Kidney Injury in Critical Care Setting Using Clinical Notes
Yikuan Li, Chengsheng Mao, Anand Srivastava, Xiaoqian Jiang, Yuan Luo 0001 |
BIBM | 6 |
| 2018 | Implementing a Portable Clinical NLP System with a Common Data Model - a Lisp Perspective
Yuan Luo 0001, Peter Szolovits |
BIBM | 1 |
| 2018 | Deep Generative Classifiers for Thoracic Disease Diagnosis with Chest X-ray Images
Chengsheng Mao, Yiheng Pan, Yuan Luo 0001, Zexian Zeng |
BIBM | 4 |
| 2018 | Characterizing Design Patterns of EHR-Driven Phenotype Extraction Algorithms
Yizhen Zhong, Luke V. Rasmussen, Jennifer A. Pacheco, Maureen E. Smith, Justin Starren, Wei-Qi Wei, Peter Speltz, Joshua C. Denny, Nephi Walton, George Hripcsak, Christopher G. Chute, Yuan Luo 0001 |
BIBM | 13 |
| 2018 | Using natural language processing and machine learning to identify breast cancer local recurrenceabstractBACKGROUND: Identifying local recurrences in breast cancer from patient data sets is important for clinical research and practice. Developing a model using natural language processing and machine learning to identify local recurrences in breast cancer patients can reduce the time-consuming work of a manual chart review. METHODS: We design a novel concept-based filter and a prediction model to detect local recurrences using EHRs. In the training dataset, we manually review a development corpus of 50 progress notes and extract partial sentences that indicate breast cancer local recurrence. We process these partial sentences to obtain a set of Unified Medical Language System (UMLS) concepts using MetaMap, and we call it positive concept set. We apply MetaMap on patients' progress notes and retain only the concepts that fall within the positive concept set. These features combined with the number of pathology reports recorded for each patient are used to train a support vector machine to identify local recurrences. RESULTS: We compared our model with three baseline classifiers using either full MetaMap concepts, filtered MetaMap concepts, or bag of words. Our model achieved the best AUC (0.93 in cross-validation, 0.87 in held-out testing). CONCLUSIONS: Compared to a labor-intensive chart review, our model provides an automated way to identify breast cancer local recurrences. We expect that by minimally adapting the positive concept set, this study has the potential to be replicated at other institutions with a moderately sized training dataset. Zexian Zeng, Sasa Espino, Ankita Roy, Xiaoyu Li 0006, Seema A. Khan, Susan E. Clare, Xia Jiang, Richard E. Neapolitan, Yuan Luo 0001 |
BMC Bioinform. | 9 |
| 2018 | Segment convolutional neural networks (Seg-CNNs) for classifying relations in clinical notesabstractWe propose Segment Convolutional Neural Networks (Seg-CNNs) for classifying relations from clinical notes. Seg-CNNs use only word-embedding features without manual feature engineering. Unlike typical CNN models, relations between 2 concepts are identified by simultaneously learning separate representations for text segments in a sentence: preceding, concept1, middle, concept2, and succeeding. We evaluate Seg-CNN on the i2b2/VA relation classification challenge dataset. We show that Seg-CNN achieves a state-of-the-art micro-average F-measure of 0.742 for overall evaluation, 0.686 for classifying medical problem-treatment relations, 0.820 for medical problem-test relations, and 0.702 for medical problem-medical problem relations. We demonstrate the benefits of learning segment-level representations. We show that medical domain word embeddings help improve relation classification. Seg-CNNs can be trained quickly for the i2b2/VA dataset on a graphics processing unit (GPU) platform. These results support the use of CNNs computed over segments of text for classifying medical relations, as they show state-of-the-art performance while requiring no manual feature engineering. Yuan Luo 0001, Özlem Uzuner, Peter Szolovits, Justin Starren |
J. Am. Medical Informatics Assoc. | 1 |
| 2018 | 3D-MICE: integration of cross-sectional and longitudinal imputation for multi-analyte longitudinal clinical dataabstractObjective: A key challenge in clinical data mining is that most clinical datasets contain missing data. Since many commonly used machine learning algorithms require complete datasets (no missing data), clinical analytic approaches often entail an imputation procedure to "fill in" missing data. However, although most clinical datasets contain a temporal component, most commonly used imputation methods do not adequately accommodate longitudinal time-based data. We sought to develop a new imputation algorithm, 3-dimensional multiple imputation with chained equations (3D-MICE), that can perform accurate imputation of missing clinical time series data. Methods: We extracted clinical laboratory test results for 13 commonly measured analytes (clinical laboratory tests). We imputed missing test results for the 13 analytes using 3 imputation methods: multiple imputation with chained equations (MICE), Gaussian process (GP), and 3D-MICE. 3D-MICE utilizes both MICE and GP imputation to integrate cross-sectional and longitudinal information. To evaluate imputation method performance, we randomly masked selected test results and imputed these masked results alongside results missing from our original data. We compared predicted results to measured results for masked data points. Results: 3D-MICE performed significantly better than MICE and GP-based imputation in a composite of all 13 analytes, predicting missing results with a normalized root-mean-square error of 0.342, compared to 0.373 for MICE alone and 0.358 for GP alone. Conclusions: 3D-MICE offers a novel and practical approach to imputing clinical laboratory time series data. 3D-MICE may provide an additional tool for use as a foundation in clinical predictive analytics and intelligent clinical decision support. Yuan Luo 0001, Peter Szolovits, Anand Dighe, Jason Baron |
J. Am. Medical Informatics Assoc. | 1 |
| 2017 | Contralateral Breast Cancer Event Detection Using Nature Language Processing
Zexian Zeng, Xiaoyu Li 0006, Sasa Espino, Ankita Roy, Kristen Kitsch, Susan E. Clare, Seema A. Khan, Yuan Luo 0001 |
AMIA | 8 |
| 2017 | Bridging semantics and syntax with graph algorithms - state-of-the-art of extracting biomedical relationsabstractResearch on extracting biomedical relations has received growing attention recently, with numerous biological and clinical applications including those in pharmacogenomics, clinical trial screening and adverse drug reaction detection. The ability to accurately capture both semantic and syntactic structures in text expressing these relations becomes increasingly critical to enable deep understanding of scientific papers and clinical narratives. Shared task challenges have been organized by both bioinformatics and clinical informatics communities to assess and advance the state-of-the-art research. Significant progress has been made in algorithm development and resource construction. In particular, graph-based approaches bridge semantics and syntax, often achieving the best performance in shared tasks. However, a number of problems at the frontiers of biomedical relation extraction continue to pose interesting challenges and present opportunities for great improvement and fruitful research. In this article, we place biomedical relation extraction against the backdrop of its versatile applications, present a gentle introduction to its general pipeline and shared resources, review the current state-of-the-art in methodology advancement, discuss limitations and point out several promising future directions. Yuan Luo 0001, Özlem Uzuner, Peter Szolovits |
Briefings Bioinform. | 1 |
| 2017 | Bridging semantics and syntax with graph algorithms - state-of-the-art of extracting biomedical relationsabstractBriefings in Bioinformatics (2017) 18(1), 2017, 160–178, doi: 10.1093/bib/bbw001 In the above article, the sentence ‘Wang et al. [84] used Latent Dirichl et al. location to create a semantic representation of biomedical named entities and used Kullback-Leibler (KL) divergence to calculate the association distance between pairs of entities in the Chem2Bio2RDF [149] semantic network’ has been corrected to ‘Wang et al. [84] used Latent Dirichlet Allocation to create a semantic representation of biomedical named entities and used Kullback-Leibler (KL) divergence to calculate the association distance between pairs of entities in the Chem2Bio2RDF [149] semantic network’. The text has been corrected online. The publisher apologizes for this error. Yuan Luo 0001, Özlem Uzuner, Peter Szolovits |
Briefings Bioinform. | 1 |
| 2017 | Tensor factorization toward precision medicineabstractPrecision medicine initiatives come amid the rapid growth in quantity and variety of biomedical data, which exceeds the capacity of matrix-oriented data representations and many current analysis algorithms. Tensor factorizations extend the matrix view to multiple modalities and support dimensionality reduction methods that identify latent groups of data for meaningful summarization of both features and instances. In this opinion article, we analyze the modest literature on applying tensor factorization to various biomedical fields including genotyping and phenotyping. Based on the cited work including work of our own, we suggest that tensor applications could serve as an effective tool to enable frequent updating of medical knowledge based on the continually growing scientific and clinical evidence. We encourage extensive experimental studies to tackle challenges including design choice of factorizations, integrating temporality and algorithm scalability. Yuan Luo 0001, Fei Wang 0001, Peter Szolovits |
Briefings Bioinform. | 1 |
| 2017 | Recurrent neural networks for classifying relations in clinical notesabstractWe proposed the first models based on recurrent neural networks (more specifically Long Short-Term Memory - LSTM) for classifying relations from clinical notes. We tested our models on the i2b2/VA relation classification challenge dataset. We showed that our segment LSTM model, with only word embedding feature and no manual feature engineering, achieved a micro-averaged f-measure of 0.661 for classifying medical problem-treatment relations, 0.800 for medical problem-test relations, and 0.683 for medical problem-medical problem relations. These results are comparable to those of the state-of-the-art systems on the i2b2/VA relation classification challenge. We compared the segment LSTM model with the sentence LSTM model, and demonstrated the benefits of exploring the difference between concept text and context text, and between different contextual parts in the sentence. We also evaluated the impact of word embedding on the performance of LSTM models and showed that medical domain word embedding help improve the relation classification. These results support the use of LSTM models for classifying relations between medical concepts, as they show comparable performance to previously published systems while requiring no manual feature engineering. Yuan Luo 0001 |
J. Biomed. Informatics | 1 |
| 2016 | Predicting ICU Mortality Risk by Grouping Temporal Trends from a Multivariate Panel of Physiologic MeasurementsabstractICU mortality risk prediction may help clinicians take effective interventions to improve patient outcome. Existing machine learning approaches often face challenges in integrating a comprehensive panel of physiologic variables and presenting to clinicians interpretable models. We aim to improve both accuracy and interpretability of prediction models by introducing Subgraph Augmented Non-negative Matrix Factorization (SANMF) on ICU physiologic time series. SANMF converts time series into a graph representation and applies frequent subgraph mining to automatically extract temporal trends. We then apply non-negative matrix factorization to group trends in a way that approximates patient pathophysiologic states. Trend groups are then used as features in training a logistic regression model for mortality risk prediction, and are also ranked according to their contribution to mortality risk. We evaluated SANMF against four empirical models on the task of predicting mortality or survival 30 days after discharge from ICU using the observed physiologic measurements between 12 and 24 hours after admission. SANMF outperforms all comparison models, and in particular, demonstrates an improvement in AUC (0.848 vs. 0.827, p<0.002) compared to a state-of-the-art machine learning method that uses manual feature engineering. Feature analysis was performed to illuminate insights and benefits of subgraph groups in mortality risk prediction. Yuan Luo 0001, Rohit Joshi, Leo A. Celi, Peter Szolovits |
AAAI | 1 |
| 2016 | Computational Phenotyping Methods
Yuan Luo 0001, Jimeng Sun 0001, Xiaoqian Jiang, Fei Wang 0001 |
AMIA | 1 |
| 2016 | Automatic identification and extraction of design patterns of EHR-driven phenotyping algorithms
Yizhen Zhong, Luke V. Rasmussen, Justin Starren, Yuan Luo 0001 |
AMIA | 4 |
| 2015 | Subgraph augmented non-negative tensor factorization (SANTF) for modeling clinical narrative textabstractOBJECTIVE: Extracting medical knowledge from electronic medical records requires automated approaches to combat scalability limitations and selection biases. However, existing machine learning approaches are often regarded by clinicians as black boxes. Moreover, training data for these automated approaches at often sparsely annotated at best. The authors target unsupervised learning for modeling clinical narrative text, aiming at improving both accuracy and interpretability. METHODS: The authors introduce a novel framework named subgraph augmented non-negative tensor factorization (SANTF). In addition to relying on atomic features (e.g., words in clinical narrative text), SANTF automatically mines higher-order features (e.g., relations of lymphoid cells expressing antigens) from clinical narrative text by converting sentences into a graph representation and identifying important subgraphs. The authors compose a tensor using patients, higher-order features, and atomic features as its respective modes. We then apply non-negative tensor factorization to cluster patients, and simultaneously identify latent groups of higher-order features that link to patient clusters, as in clinical guidelines where a panel of immunophenotypic features and laboratory results are used to specify diagnostic criteria. RESULTS AND CONCLUSION: SANTF demonstrated over 10% improvement in averaged F-measure on patient clustering compared to widely used non-negative matrix factorization (NMF) and k-means clustering methods. Multiple baselines were established by modeling patient data using patient-by-features matrices with different feature configurations and then performing NMF or k-means to cluster patients. Feature analysis identified latent groups of higher-order features that lead to medical insights. We also found that the latent groups of atomic features help to better correlate the latent groups of higher-order features. Yuan Luo 0001, Ephraim P. Hochberg, Rohit Joshi, Özlem Uzuner, Peter Szolovits |
J. Am. Medical Informatics Assoc. | 1 |
| 2014 | Quantifying Information Redundancy in Common Laboratory Tests
Yuan Luo 0001, Jason Baron, Peter Szolovits, Anand Dighe |
AMIA | 1 |
| 2014 | Research and applications: Automatic lymphoma classification with sentence subgraph mining from pathology reportsabstractOBJECTIVE: Pathology reports are rich in narrative statements that encode a complex web of relations among medical concepts. These relations are routinely used by doctors to reason on diagnoses, but often require hand-crafted rules or supervised learning to extract into prespecified forms for computational disease modeling. We aim to automatically capture relations from narrative text without supervision. METHODS: We design a novel framework that translates sentences into graph representations, automatically mines sentence subgraphs, reduces redundancy in mined subgraphs, and automatically generates subgraph features for subsequent classification tasks. To ensure meaningful interpretations over the sentence graphs, we use the Unified Medical Language System Metathesaurus to map token subsequences to concepts, and in turn sentence graph nodes. We test our system with multiple lymphoma classification tasks that together mimic the differential diagnosis by a pathologist. To this end, we prevent our classifiers from looking at explicit mentions or synonyms of lymphomas in the text. RESULTS AND CONCLUSIONS: We compare our system with three baseline classifiers using standard n-grams, full MetaMap concepts, and filtered MetaMap concepts. Our system achieves high F-measures on multiple binary classifications of lymphoma (Burkitt lymphoma, 0.8; diffuse large B-cell lymphoma, 0.909; follicular lymphoma, 0.84; Hodgkin lymphoma, 0.912). Significance tests show that our system outperforms all three baselines. Moreover, feature analysis identifies subgraph features that contribute to improved performance; these features agree with the state-of-the-art knowledge about lymphoma classification. We also highlight how these unsupervised relation features may provide meaningful insights into lymphoma classification. Yuan Luo 0001, Aliyah R. Sohani, Ephraim P. Hochberg, Peter Szolovits |
J. Am. Medical Informatics Assoc. | 1 |
| 2008 | A de-identifier for medical discharge summaries
Özlem Uzuner, Tawanda C. Sibanda, Yuan Luo 0001, Peter Szolovits |
Artif. Intell. Medicine | 3 |
| 2008 | Viewpoint Paper: Identifying Patient Smoking Status from Medical Discharge RecordsabstractClinical narrative records contain much useful information. However, most clinical narratives are in the form of fragmented English free text, showing the characteristics of a clinical sublanguage. This makes their linguistic processing, search, and retrieval challenging.1 Traditional natural language processing (NLP) tools are not designed for the fragmented free text found in narrative clinical records; therefore, they do not perform well on this type of data.2 Limited access to clinical records has been a barrier to the widespread development of medical language processing (MLP) technologies. In the absence of a standardized, publicly available ground truth that encourages the development of MLP systems and allows their head-to-head comparison, successful MLP efforts have been limited, e.g., MedLEE3 and Symtxt.4 A few MLP systems have been developed,5 and such efforts have successfully shown the usefulness of MLP in clinical settings.6–8 To improve the availability of clinical records and to contribute to the advancement of the state of the art in MLP, within the i2b2 (Informatics for Integrating Biology to the Bedside) project, the authors de-identified and released a set of clinical records from Partners HealthCare. These records provided the basis for the development of ground truth for two challenge questions: Automatic de-identification of clinical data, i.e., de-identification challenge. Automatic evaluation of the smoking status of patients based on medical records, i.e., smoking challenge. Representative teams from the MLP community participated in the two challenges and met at a workshop organized by the authors to discuss the results of the challenges. The workshop was co-sponsored by the American Medical Informatics Association and met in conjunction with its Fall Symposium in November 2006. This article provides an overview of the smoking challenge and the findings of the workshop. An overview of the de-identification challenge can be found in Uzuner et al.9 The smoking challenge continues the tradition of attempting to identify the state of the art in automatic language processing. Outside of the medical domain, there have been many efforts in this direction. Most of these efforts have been led by Message Understanding Conferences (MUC)10 and the National Institute of Standards and Technology (NIST).11 MUC organized shared tasks on named entity recognition. NIST organized a series of Text Retrieval Evaluation Conferences (TREC) on various domains including blogs and legal documents; they also organized a series of shared tasks on topic detection and tracking, speaker recognition, language recognition, spoken document retrieval, machine translation, and entity extraction. In the biomedical domain, three such prominent efforts were BioCreAtIvE12 for information extraction, ImageCLEF13,14 for image retrieval, and TREC Genomics15 for question answering and information retrieval. Keeping the goals of TREC,16 MUC,17 BioCreAtIvE,18 etc., in mind for the smoking challenge, we created a collection of actual medical discharge records. We invited the development of systems that can predict the smoking status of patients based on the narratives in these medical discharge records. We limited the scope of this task to understanding only the explicitly reported smoking information. In other words, information that implicitly reveals the smoking status was excluded from this study. Our smoking challenge continued the work on the application of classification techniques to the medical domain,19–22 and extended the MLP studies on medical discharge records.8,23–29 Information on the smoking status of patients is important for many health studies, e.g., studies on asthma; however, before this challenge, the only system for the automatic evaluation of the smoking status of patients from their records was the HITEx system.30 The data for the smoking challenge consisted exclusively of discharge summaries from Partners HealthCare. We preprocessed these records so that they were de-identified, tokenized, broken into sentences, converted into XML format, and separated into training and test sets. Institutional review boards of Partners HealthCare, Massachusetts Institute of Technology, and the State University of New York at Albany approved the challenge and the data preparation process. The data for the challenge were annotated by pulmonologists. The pulmonologists were asked to classify patient records into five possible smoking status categories. For the purposes of this challenge, we defined these categories as follows: A Past Smoker is a patient whose discharge summary asserts explicitly that the patient was a smoker one year or more ago but who has not smoked for at least one year. The assertion “past smoker” without any temporal qualifications means Past Smoker unless there is text that says that the patient stopped smoking less than one year ago. A Current Smoker is a patient whose discharge summary asserts explicitly that the patient was a smoker within the past year. The assertion “current smoker” without any temporal qualifications means Current Smoker unless there is text that says that the patient stopped smoking more than a year ago. A Smoker is a patient who is either a Current or a Past Smoker but whose medical record does not provide enough information to classify the patient as either. A Non-Smoker's discharge summary indicates that they never smoked. An Unknown is a patient whose discharge summary does not mention anything about smoking. Indecision between Current Smoker and Past Smoker does not belong to this category. Second-hand smokers are considered Non-Smokers for the purposes of this study, unless there is evidence in their record that they actively smoked. Similarly, as we are only concerned with tobacco, marijuana smoking should not affect the patients' smoking status. In addition to being provided with the above definitions, the annotators were trained on 55 sample sentences30 (see Table 1 for a subset) and 10 sample records. Annotator Training Samples Annotator Training Samples Two pulmonologists annotated each record with the smoking status of patients based strictly on the explicitly stated smoking-related facts in the records. These annotations constitute the textual judgments of the annotators. The same two pulmonologists also marked the smoking status of the patients using their medical intuitions on all information in the records. These annotations constitute the intuitive judgments of the annotators. In all, 928 records were annotated. The interannotator agreement on the textual judgments on these records, as measured by Cohen's kappa (κ),31,32 was 0.84; observed agreement on textual judgments was 0.93; specific agreement per category on textual judgments ranged from 0.4 to 0.98 (Table 2; also see the Methods section for definitions of Cohen's kappa, observed agreement, and specific agreement). The interannotator agreement on the intuitive judgments was 0.45; observed agreement on intuitive judgments was 0.73; specific agreement per category on intuitive judgments ranged from 0.3 to 0.84 (Table 2). Observed and Specific Agreement Observed and Specific Agreement Guidelines for κ are subject to interpretation and depend on parameters such as the task and categories involved.33 However, κ of 0.8 is widely used as the threshold for strong agreement.31–35 On our data, we observed strong agreement only on the textual judgments. We further observed that the intra-annotator agreement between a given doctor's intuitive and textual judgments varied from 0.62 to 0.99. This indicates that the reliance of the intuitive judgments on the explicit textual information varies considerably from doctor to doctor. Given these observations, we limited the challenge task to the identification of the smoking status based on information that is explicitly mentioned in the records. To generate the ground truth, we resolved the disagreements (on textual judgments) between the annotators by obtaining judgments from two other pulmonologists. We omitted the records that the annotators disagreed on from the challenge, unless a majority vote could identify a clear textual judgment for them. In all, 63 records were omitted from the challenge for lack of a clear textual judgment. In addition, annotation results showed heavy bias in the data for Unknown records. This is the least interesting category for the purposes of the smoking challenge as Unknown records do not contain any smoking-related information. To focus the smoking challenge less on the Unknown category and more on the other four categories, we omitted a portion of the Unknown records (363 records) from the challenge. A total of 502 de-identified medical discharge records were used for the smoking challenge. Table 3 shows the distribution of annotated records into training and test sets, and into Past Smoker, Current Smoker, Smoker, Non-Smoker, and Unknown categories. The training and test sets show similar distribution of records into the five smoking categories; however, these distributions are far from uniform. This reflects the realities of real-world data; our records were drawn at random from the Partners' database, in which some smoking categories are better represented than others. In our test set, the smallest smoking category is Smokers, with only three records. The training and test data can be obtained from i2b2.org. Smoking Status Training and Test Data Distribution Smoking Status Training and Test Data Distribution We evaluated system performances using microaveraged and macroaveraged precision, recall, and F-measure, as well as Cohen's kappa. Precision, recall, and F-measure are performance metrics frequently used in NLP.36,37 These metrics are easily derived from a binary confusion matrix. In a binary decision problem, a classifier labels entities as either positive or negative (where positive and negative represent two generic categories) and produces a confusion matrix. This matrix contains four entities: true positive (TP), true negative (TN), false positive (FP), and false negative (FN). Given such a matrix, precision is the percentage of entities classified correctly to be in a given category in relation to the total number of entities classified for the given category (Equation 1). Recall is the percentage of entities classified correctly in a given category in relation to the actual number of items in the given category (Equation 2). F-measure is the harmonic mean of precision and recall (Equation 3). β enables F-measure to favor either precision or recall. We give equal weight to precision and recall by setting β = 1. Cohen's kappa (κ) (Equation 6) is a measure of agreement31 between pairs of annotators who classify items into a set number of mutually exclusive categories. κ depends on observed agreement (Ao in Equation 7) and the agreement expected due to chance (Ae in Equation 8). A κ value of 0.8 is widely used as the threshold for strong agreement,31–35 whereas a κ of 0 indicates that the observed agreement is due to chance.33 We used κ as a measure of inter-annotator and intra-annotator agreement (see Annotations section) as a measure of agreement between two automatic systems and as a measure of agreement between a system and the ground truth (see Results and Discussion). Equation 6 through Equation 8 collectively describe κ between an automatic system and the ground truth. κ for inter-annotator and intra-annotator agreement and for inter-system agreement can be computed analogously. According to Hripcsak and Rothschild,38 there exists a correspondence between F-measure and κ. However, κ provides clearer insights into the relative strengths of the systems (see Intersystem Agreement section). We evaluated systems using F-measure but compared them using both F-measure and κ. Specific agreement (Asp)33 measures the degree of agreement (Equation 9) on each category and is not adjusted by chance. We used specific agreement to get a sense of the level of agreement between annotators without taking chance into consideration. We tested the significance of the differences of the systems using a randomization technique that is frequently utilized in NLP.39 The null hypothesis is that the absolute value of the difference in performances, e.g., F-measures, of two systems is approximately equal to zero. The randomization technique does not assume a particular distribution of the differences. Instead, it empirically generates the distribution. Given two actual systems, it randomly shuffles (at each iteration, we simulated a coin flip to decide whether the answers should be swapped) their responses to the records in the test set N times (e.g., N = 9,999), and thus creates N pairs of pseudosystems. It counts the number of times that the difference between the performances of pairs of pseudosystems is greater than the difference between the two actual systems' performances. Let this count be equal to n and compute . If s is greater than a predetermined cutoff α, then the difference of the performances of the two actual systems can be explained by chance; otherwise, the difference is significant at level α. Following MUC's example, we set α to 0.1. A total of 11 teams participated in the smoking challenge. The training data for the challenge were released in July 2006, and the test data were released for only three days in September 2006. Each team was permitted to submit up to three system on the test A total of were count only one of the three of and their other two were evaluated from the (see and are not in this In this we describe each et a this classifier that to the smoking status of patients and then and to the information. The of the that only a few in a record to the smoking status and that these could be easily by their (e.g., If more than one in a record smoking then only the was If such were then the record was classified as To predict the smoking status of a record from the test set, each from this record was compared with from the training The of the measures between each and the most similar in the training set the smoking status of the et various for text with and found that the performance from the of and that and classification more than et also from lack of explicit smoking information in the Unknown and these to further two to processing the In the they classified based on to smoking using In the they classified each explicit to smoking using and then to categories from judgments of to smoking the smoking category for the collection the category in the Given the data of their et the i2b2 data set with records and their smoking categories. that a based on the data set perform better than trained on each of the data sets This hypothesis is with a the of a system is as the sample et smoking status evaluation system was with and from the Medical This system document (e.g., medical and the status of these entities (e.g., smoking-related such as and It thus the set of and of explicit to smoking. et found that the using on the to smoking better than the classification using (see and in Table and Table this system are by the of the and for example, to the system by et The and of explicit to their For example, the does not in the section could both to a Past Smoker and to a Non-Smoker, on However, such and information was to and for Precision, and by and for Precision, and by on and by not in macroaveraged not in microaveraged the is on and by not in macroaveraged not in microaveraged the is smoking status evaluation as a task and a process. The marked smoking-related In Cohen's a smoking-related was a of specific (e.g., and The the with the from the The without any specific and them The the and classified them using created of system by a system of for the smoking status of This system utilized both and such as of and of three one and two for the smoking status of patients of smoking challenge systems as a to the and can be at For with from the and found decision the most evaluated using on the training this classifier its performance trained using only the and their as For with and were with up to five words, one of the two in the with the These two systems to most of the test data to the Unknown category. found that of as well as Table shows that from in both microaveraged and macroaveraged F-measures, at α = 0.1. the Medical to the challenge. is a system for medical text processing in and records by For the smoking challenge, this system was to English medical discharge et the smoking challenge as a classification system marked the smoking status category of each in a record and to classify the document The judgments of the training set were derived Most of the text processing was through the of the Information system of and classification was in three and by using with an In the the Unknown category was a of that excluded this category. In the the category was through the of For with was In the the Smoker, and Past Smoker categories were by a of to Current and Past This was as a temporal and was using an the in this the were not so as to information of such as the of the and such as and between Current and Past the categories were The document categories were from to as follows: Current Smoker, Past Smoker, Smoker, Non-Smoker, and Each document was the of the category for which it could provide For example, a Current Smoker was to a document of any as Current et and the system to the smoking challenge. is designed for and clinical information from free text, and contains a patient smoking status This system of four document and The and the of a record based on the of the section The to them with in a The the text and is to The to labels to each The set of is then using provides based on of patient This system for a given patient and a to the with the to each To the differences of discharge summaries from patient et to the smoking challenge. this they into the differences in the labels of the i2b2 smoking challenge and the smoking they the system to a specific than a set of and for each The reported that they found the temporal of smoking categories to be a significant challenge. et of the explicit to smoking status. and using the results of these and records using a In addition to they used and information about and used a system to predict the smoking status based on explicit of smoking. that the system they for information about the smoking status of patients the records were of explicit smoking-related For they all explicit of smoking from the records and created a data set that only the trained two on the showed that annotators precision, recall, and F-measure on the The two systems the performance of annotators with of on the test Table and 1 show the precision, recall, and F-measure, both macroaveraged and for each of the system to the smoking challenge. Table shows the results of the significance on microaveraged and macroaveraged of the this reveals that the differences in the microaveraged of the systems are not significant at α = of these systems are not from each other in their macroaveraged at this α. Results from Table by microaveraged Table shows that systems that used similar machine e.g., the systems that used not all perform This that other such as the and the used with the also to the Table also shows that a majority of the performances from systems that of the lack of explicit smoking information in the Unknown records and to further processing or classification (see the results of et and et These systems microaveraged to that were on explicit to smoking status and et of the systems that the smoking classification was Unknown unless some in the document showed it to be this in an F-measure of in the Unknown category for and et and et Table 6 shows the precision, recall, and F-measure of each system on each of the five smoking status categories. In the systems successfully the Unknown and categories; some Current and Past with the of the systems of et systems correctly classified one of the three all in Precision, and for in Precision, and for in In addition to system performance on the test set and on categories in the test set, we the performance and agreement of systems on data i.e., records, in the test For we used κ. Table shows the level of κ agreement of each system with the ground truth as well as the level of κ agreement of pairs of Most is the level of agreement between systems by the same but only for some of the This level of agreement is not these from of the same The systems of et in their to and and only in their training and used a that the from and Agreement of and the Agreement between 0.8 and is in and agreement is in Agreement of and the Agreement between 0.8 and is in and agreement is in Table also shows in the systems showed a level of agreement with each For example, each of the systems of et with each of Cohen's systems showed κ between and systems showed strong agreement with each other the differences in their to smoking status For example, and system showed κ of to compared with et Similarly, the systems of et a κ of 0.8 with that of and the classified the records at the level whereas the was at the The level of agreement between to the smoking challenge, by the of these on the ground truth, indicates the of the set of to smoking status On the other the disagreements the systems their relative For example, and disagreed with each other on records = These systems showed relative strengths in in and Current Smoker, and in Unknown and Past were only three records that system marked The state of the art in smoking status evaluation be by the strengths of such of the teams with the annotation of a few of the in the challenge mentioned in the Annotations our medical records were annotated by pulmonologists. they provide a ground truth, judgments can However, given the medical of the annotators and the agreement on the labels of these records the we the annotations as the English of the records is for all of the a majority of the systems found the classification disagreed with the judgments of we that of should be by an to measure the of the that the in the of the ground truth only on our findings on the smoking challenge, we to our in two we are by the strengths of the systems and to of these to improve the we and with data to be We that a system provide insights into the intuitive judgments on the smoking status of In this we the i2b2 smoking challenge, the data and the data preparation the evaluation and each of the We the evaluation and of the system and provided a of the of our findings for For this challenge, we a document collection derived from actual medical discharge records, this collection in to a real-world medical classification and system performances. We showed that asked to a decision on the smoking status of patients based on the explicitly stated information in medical discharge annotators with each other more than of the The systems that participated in the smoking challenge represented various from and to the differences in their to smoking status many of these systems In there were system with microaveraged above A majority of these systems of the of the challenge data, e.g., lack of to smoking in records marked of explicit to smoking. the systems in the smoking challenge showed that discharge summaries smoking status using a limited number of textual (e.g., of the smoking status from these The authors all teams for their to the challenge, for their in the of the workshop that the challenge, and and the of for their on this Özlem Uzuner, Ira Goldstein, Yuan Luo 0001, Isaac S. Kohane |
J. Am. Medical Informatics Assoc. | 3 |
| 2007 | Viewpoint Paper: Evaluating the State-of-the-Art in Automatic De-identificationabstractTo facilitate and survey studies in automatic de-identification, as a part of the i2b2 (Informatics for Integrating Biology to the Bedside) project, authors organized a Natural Language Processing (NLP) challenge on automatically removing private health information (PHI) from medical discharge records. This manuscript provides an overview of this de-identification challenge, describes the data and the annotation process, explains the evaluation metrics, discusses the nature of the systems that addressed the challenge, analyzes the results of received system runs, and identifies directions for future research. The de-indentification challenge data consisted of discharge summaries drawn from the Partners Healthcare system. Authors prepared this data for the challenge by replacing authentic PHI with synthesized surrogates. To focus the challenge on non-dictionary-based de-identification methods, the data was enriched with out-of-vocabulary PHI surrogates, i.e., made up names. The data also included some PHI surrogates that were ambiguous with medical non-PHI terms. A total of seven teams participated in the challenge. Each team submitted up to three system runs, for a total of sixteen submissions. The authors used precision, recall, and F-measure to evaluate the submitted system runs based on their token-level and instance-level performance on the ground truth. The systems with the best performance scored above 98% in F-measure for all categories of PHI. Most out-of-vocabulary PHI could be identified accurately. However, identifying ambiguous PHI proved challenging. The performance of systems on the test data set is encouraging. Future evaluations of these systems will involve larger data sets from more heterogeneous sources. Özlem Uzuner, Yuan Luo 0001, Peter Szolovits |
J. Am. Medical Informatics Assoc. | 2 |