Yifan Peng 0002

dblp:150/7295-2 · DBLP profile ↗
← Back
57ranked-venue papers
5as first author
46since 2021 · last 2026
0000-0001-9309-8331ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 46 · 5 first-author · 38 since 2021Artificial intelligence and machine learning · 10 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 A Disease-Aware Dual-Stage Framework for Chest X-ray Report Generation
abstract
Radiology report generation from chest X-rays is an important task in artificial intelligence with the potential to greatly reduce radiologists' workload and shorten patient wait times. Despite recent advances, existing approaches often lack sufficient disease-awareness in visual representations and adequate vision-language alignment to meet the specialized requirements of medical image analysis. As a result, these models usually overlook critical pathological features on chest X-rays and struggle to generate clinically accurate reports. To address these limitations, we propose a novel dual-stage disease-aware framework for chest X-ray report generation. In Stage 1, our model learns Disease-Aware Semantic Tokens (DASTs) corresponding to specific pathology categories through cross-attention mechanisms and multi-label classification, while simultaneously aligning vision and language representations via contrastive learning. In Stage 2, we introduce a Disease-Visual Attention Fusion (DVAF) module to integrate disease-aware representations with visual features, along with a Dual-Modal Similarity Retrieval (DMSR) mechanism that combines visual and disease-specific similarities to retrieve relevant exemplars, providing contextual guidance during report generation. Extensive experiments on benchmark datasets (i.e., CheXpert Plus, IU X-ray, and MIMIC-CXR) demonstrate that our disease-aware framework achieves state-of-the-art performance in chest X-ray report generation, with significant improvements in clinical accuracy and linguistic quality.
Puzhen Wu, Hexin Dong, Yi Lin 0009, Yihao Ding, Yifan Peng 0002
AAAI5
2026 SDoH-GPT: using large language models to extract social determinants of health
abstract
OBJECTIVE: Extracting social determinants of health (SDoHs) from medical notes depends heavily on labor-intensive annotations, which are typically task-specific, hampering reusability and limiting sharing. Here, we introduce SDoH-GPT, a novel framework leveraging few-shot learning large language models (LLMs) to automate the extraction of SDoH from unstructured text, aiming to improve both efficiency and generalizability. MATERIALS AND METHODS: SDoH-GPT is a framework including the few-shot learning LLM methods to extract the SDoH from medical notes and the XGBoost classifiers which continue to classify SDoH using the annotations generated by the few-shot learning LLM methods as training datasets. The unique combination of the few-shot learning LLM methods with XGBoost utilizes the strength of LLMs as great few shot learners and the efficiency of XGBoost when the training dataset is sufficient. Therefore, SDoH-GPT can extract SDoH without relying on extensive medical annotations or costly human intervention. RESULTS: Our approach achieved tenfold and twentyfold reductions in time and cost, respectively, and superior consistency with human annotators measured by Cohen's kappa of up to 0.92. The innovative combination of LLM and XGBoost can ensure high accuracy and computational efficiency while consistently maintaining 0.90+ AUROC scores. DISCUSSION: This study has verified SDoH-GPT on three datasets and highlights the potential of leveraging LLM and XGBoost to revolutionize medical note classification, demonstrating its capability to achieve highly accurate classifications with significantly reduced time and cost. CONCLUSION: The key contribution of this study is the integration of LLM with XGBoost, which enables cost-effective and high quality annotations of SDoH. This research sets the stage for SDoH can be more accessible, scalable, and impactful in driving future healthcare solutions.
Bernardo Scapini Consoli, Xizhi Wu, Song Wang 0026, Yanshan Wang, Justin F. Rousseau, Thomas Hartvigsen, Li Shen 0001, Huanmei Wu, Yifan Peng 0002, Qi Long, Tianlong Chen 0001, Ying Ding 0001
J. Am. Medical Informatics Assoc.11
2026 CPGPrompt: translating clinical guidelines into large language model-executable decision support
abstract
OBJECTIVE: Clinical practice guidelines (CPGs) provide evidence-based recommendations for patient care; however, integrating them into artificial intelligence (AI) remains challenging. Previous approaches, such as rule-based systems or black-box AI models, face significant limitations, including poor interpretability, inconsistent adherence to guidelines, and narrow domain applicability. To address this, we develop and validate CPGPrompt, an auto-prompting system that converts narrative clinical guidelines into large language models (LLMs). MATERIALS AND METHODS: Our framework translates CPGs into structured decision trees and utilizes an LLM to dynamically navigate them for patient case evaluation. Synthetic vignettes were generated across 3 domains-headache, lower back pain, and prostate cancer-and distributed into 4 categories to test different decision scenarios. System performance was assessed on both binary specialty referral decisions and fine-grained pathway classification tasks. RESULTS: The binary specialty referral classification achieved consistently strong performance across all domains (F1: 0.85-1.00), with high recall (1.00 ± 0.00). In contrast, multiclass pathway assignment showed reduced performance, with domain-specific variations: headache (F1: 0.47), lower back pain (F1: 0.72), and prostate cancer (F1: 0.77). DISCUSSION: Domain-specific performance differences reflected the structure of each guideline. The headache guideline highlighted challenges with negation handling. The lower back pain guideline required temporal reasoning. In contrast, prostate cancer pathways benefited from quantifiable laboratory tests, resulting in more reliable decision-making. CONCLUSION: CPGPrompt demonstrates generalizability across diverse clinical domains while maintaining high sensitivity for referral decisions. Its transparent, auditable framework enables the systematic identification of failure modes and provides advantages over black-box AI approaches. However, persistent challenges with subjective clinical assessments indicate a need for targeted improvements and greater clinical robustness.
Ruiqi Deng, Geoffrey Martin, Tony Wang, Yi Liu 0059, Chunhua Weng, Yanshan Wang, Justin F. Rousseau, Yifan Peng 0002
J. Am. Medical Informatics Assoc.9
2026 A critical evaluation of generative query expansion on biomedical literature retrieval
abstract
OBJECTIVE: To evaluate the effectiveness of generative query expansion for biomedical literature retrieval. MATERIALS AND METHODS: We thoroughly examined eight generative query expansion methods using three large language models across five datasets for biomedical literature retrieval. We further performed a quantitative analysis, including performance comparisons, rank transition analysis, and article-type effect analysis. We also conducted a qualitative examination of representative cases, from which we derived an error taxonomy. RESULTS: On BioASQ-Y/N, GPT-4o-based query expansion shifts Recall@10 to 0.417-0.512 and nDCG@10 to 0.358-0.479, relative to a baseline of 0.491 and 0.456. For PubMedQA, Precision@1 ranges from 0.764 to 0.876 and nDCG@10 from 0.847 to 0.931, compared with baseline values of 0.893 and 0.935. For 2019-Trec-PM, query expansion yields Recall@100 of 0.217-0.256 and nDCG@100 of 0.272-0.312, versus a baseline of 0.227 and 0.274. Similarly, for 2018-TREC-PM, Recall@100 spans 0.169-0.227 and nDCG@100 spans 0.195-0.250, relative to baseline scores of 0.164 and 0.191. For 2017-TREC-PM, Recall@100 and nDCG@100 fall within 0.111-0.139 and 0.154-0.191 under query expansion, compared with baseline metrics of 0.102 and 0.147. Both general-purpose and domain-specific Llama-based models demonstrate similar performance to GPT-4o. DISCUSSION AND CONCLUSION: The impact of query expansion varies significantly by the expansion methods and type of evidence, but is relatively agnostic to backbone model choice. Notably, query expansion primarily affects article ranking but has a limited impact on the screening stage. Our findings underscore the unique challenges of biomedical literature retrieval and highlight the need to develop domain-specific information retrieval techniques.
Yilu Fang, Fangyi Chen, Yifan Peng 0002, Chunhua Weng
J. Am. Medical Informatics Assoc.4
2025 Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
abstract
Recent advancements in multimodal large language models (MLLMs) have significantly improved performance in visual question answering. However, they often suffer from hallucinations. In this work, hallucinations are categorized into two main types: initial hallucinations and snowball hallucinations. We argue that adequate contextual information can be extracted directly from the token interaction process. Inspired by causal inference in the decoding strategy, we propose to leverage causal masks to establish information propagation between multimodal tokens. The hypothesis is that insufficient interaction between those tokens may lead the model to rely on outlier tokens, overlooking dense and rich contextual cues. Therefore, we propose to intervene in the propagation process by tackling outlier tokens to enhance in-context inference. With this goal, we present FarSight, a versatile plug-and-play decoding strategy to reduce attention interference from outlier tokens merely by optimizing the causal mask. The heart of our method is effective token propagation. We design an attention register structure within the upper triangular matrix of the causal mask, dynamically allocating attention to capture attention diverted to outlier tokens. Moreover, a positional awareness encoding method with a diminishing masking rate is proposed, allowing the model to attend to further preceding tokens, especially for video sequence tasks. With extensive experiments, FarSight demonstrates significant hallucination-mitigating performance across different MLLMs on both image and video benchmarks, proving its effectiveness.
Zhongxing Xu, Zile Huang, Haochen Xue, Ziyang Chen 0003, Zelin Peng, Sijin Zhou, Wenxue Li 0003, Yulong Li 0002, Wenxuan Song, Shiyan Su, Wei Feng 0015, Jionglong Su, Mingquan Lin, Yifan Peng 0002, Xuelian Cheng, Muhammad Imran Razzak, ZongYuan Ge
CVPR18
2025 Semi-supervised learning from small annotated data and large unlabeled data for fine-grained Participants, Intervention, Comparison, and Outcomes entity recognition
abstract
OBJECTIVE: Extracting PICO elements-Participants, Intervention, Comparison, and Outcomes-from clinical trial literature is essential for clinical evidence retrieval, appraisal, and synthesis. Existing approaches do not distinguish the attributes of PICO entities. This study aims to develop a named entity recognition (NER) model to extract PICO entities with fine granularities. MATERIALS AND METHODS: Using a corpus of 2511 abstracts with PICO mentions from 4 public datasets, we developed a semi-supervised method to facilitate the training of a NER model, FinePICO, by combining limited annotated data of PICO entities and abundant unlabeled data. For evaluation, we divided the entire dataset into 2 subsets: a smaller group with annotations and a larger group without annotations. We then established the theoretical lower and upper performance bounds based on the performance of supervised learning models trained solely on the small, annotated subset and on the entire set with complete annotations, respectively. Finally, we evaluated FinePICO on both the smaller annotated subset and the larger, initially unannotated subset. We measured the performance of FinePICO using precision, recall, and F1. RESULTS: Our method achieved precision/recall/F1 of 0.567/0.636/0.60, respectively, using a small set of annotated samples, outperforming the baseline model (F1: 0.437) by more than 16%. The model demonstrates generalizability to a different PICO framework and to another corpus, which consistently outperforms the benchmark in diverse experimental settings (P-value < .001). DISCUSSION: We developed FinePICO to recognize fine-grained PICO entities from text and validated its performance across diverse experimental settings, highlighting the feasibility of using semi-supervised learning (SSL) techniques to enhance PICO entities extraction. Future work can focus on optimizing SSL algorithms to improve efficiency and reduce computational costs. CONCLUSION: This study contributes a generalizable and effective semi-supervised approach leveraging large unlabeled data together with small, annotated data for fine-grained PICO extraction.
Fangyi Chen, Yilu Fang, Yifan Peng 0002, Chunhua Weng
J. Am. Medical Informatics Assoc.4
2025 Deciphering genomic codes using advanced natural language processing techniques: a scoping review
abstract
Objectives: The vast and complex nature of human genomic sequencing data presents challenges for effective analysis. This review aims to investigate the application of Natural Language Processing (NLP) techniques, particularly Large Language Models (LLMs) and transformer architectures, in deciphering genomic codes, focusing on tokenization, transformer models, and regulatory annotation prediction. This review aims to assess data and model accessibility in the most recent literature, gaining a better understanding of the existing capabilities and constraints of these tools in processing genomic sequencing data. Methods: Following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, our scoping review was conducted across PubMed, Medline, Scopus, Web of Science, Embase, and ACM Digital Library. Studies were included if they focused on NLP methodologies applied to genomic sequencing data analysis, without restrictions on publication date or article type. Results: A total of 26 studies published between 2021 and April 2024 were selected for review. The review highlights that tokenization and transformer models enhance the processing and understanding of genomic data, with applications in predicting regulatory annotations like transcription-factor binding sites and chromatin accessibility. Discussion: The application of NLP and LLMs to genomic sequencing data interpretation is a promising field that can help streamline the processing of large-scale genomic data while providing a better understanding of its complex structures. It can potentially drive advancements in personalized medicine by offering more efficient and scalable solutions for genomic analysis. Further research is needed to discuss and overcome limitations, enhancing model transparency and applicability.
Shuyan Cheng, Yishu Wei, Yiliang Zhou, Drew N. Wright, Jinze Liu, Yifan Peng 0002
J. Am. Medical Informatics Assoc.7
2025 A multimodal approach for few-shot biomedical named entity recognition in low-resource languages
Leilei Su, Mingquan Lin, Yifan Peng 0002, Cong Sun 0004
J. Biomed. Informatics5
2025 CLEAR: A vision to support clinical evidence lifecycle with continuous learning
abstract
Human knowledge of diseases, treatments, and prevention techniques is constantly evolving. The generation of clinical evidence using randomized controlled trials on human subjects occurs notably slowly and inefficiently. The Learning Health System (LHS) has been proposed to facilitate the continuous improvement of individual and population health through a cycle of knowledge, practice, and data. However, the gap between the demand for high-quality evidence to support clinical decisions and the available evidence continues to enlarge. While the current LHS vision articulates the integration of Real-World Data (RWD), the rapid generation of RWD often outpaces the rate of effective evidence synthesis and implementation. Considering this, we propose a new framework that more effectively leverages RWD to support the entire clinical evidence lifecycle through a continuous learning mechanism. This framework, powered by modern data science and informatics, offers enhanced scalability and efficiency. In this vision, specifically, RWD is integrated into the clinical evidence lifecycle via four closed feedback loops: 1) guiding research prioritization and study design, 2) facilitating clinical guideline development, 3) assisting guideline evaluation, and 4) supporting shared decision-making. Our framework enables rapid responsiveness to emerging health data and evolving healthcare needs, timely development of clinical guidelines to optimize clinical recommendations, and sustained improvements in clinical practice and patient outcomes. This vision calls for informatics support for an efficient, scalable, and stakeholder-aware clinical evidence lifecycle.
Yilu Fang, Fangyi Chen, George Hripcsak, Yifan Peng 0002, Patrick B. Ryan, Chunhua Weng
J. Biomed. Informatics5
2025 Scalable scientific interest profiling using large language models
Yilun Liang, Edward Sun, Betina Ross S. Idnay, Yilu Fang, Fangyi Chen, Casey N. Ta, Yifan Peng 0002, Chunhua Weng
J. Biomed. Informatics8
2025 Machine learning applications related to suicide in military and Veterans: A scoping literature review
Yishu Wei, Yanshan Wang, Yunyu Xiao, Ronald K. Poropatich, Gretchen L. Haas, Yiye Zhang, Chunhua Weng, Jinze Liu, Lisa A. Brenner, James M. Bjork, Yifan Peng 0002
J. Biomed. Informatics12
2025 CXR-LT 2024: A MICCAI challenge on long-tailed, multi-label, and zero-shot disease classification from chest X-ray
Mingquan Lin, Gregory Holste, Song Wang 0026, Yiliang Zhou, Yishu Wei, Imon Banerjee, Pengyi Chen, Tianjie Dai, Yuexi Du, Nicha C. Dvornek, Yuyan Ge, Zuwei Guo, Shohei Hanaoka, Dongkyun Kim, Pablo Messina, Yang Lu 0009, Denis Parra, Donghyun Son, Alvaro Soto, Aisha Urooj Khan, René Vidal, Yosuke Yamagishi, Pingkun Yan, Zefan Yang, Ruichi Zhang, Yang Zhou 0019, Leo A. Celi, Ronald M. Summers, Zhiyong Lu, Hao Chen 0011, Adam E. Flanders, George Shih, Zhangyang Wang, Yifan Peng 0002
Medical Image Anal.34
2025 An Interpretable Complex Knowledge Multi-Hop Reasoning Model for Predicting Synthetic Lethality in Human Cancers
abstract
Synthetic lethality (SL) has emerged as a promising strategy in cancer medicine. However, complex biomolecular interactions make wet lab methods time-consuming and expensive. Machine learning methods have gained widespread adoption for SL prediction in recent years. Although these methods have demonstrated particular effectiveness, they suffer from weak interpretability, making it difficult for users to understand the specific reasoning processes of the models. Also, they typically focus on a simple gene pair, thus struggling with more meaningful reasoning tasks involving other medical factors as in real life. To address these gaps, we propose an explainable multi-hop reasoning model EFOL-SL based on first-order logic queries. We first construct query graphs with triplet transformations for different tasks. Node embeddings are then fed into a sparse Transformer encoder and a visualized graph attention decoder to generate comprehensive multi-hop logical reasoning chains. By masking nodes in intermediate reasoning steps, our model can explicitly predict each node, allowing observation of its exact reasoning process. Additionally, we conduct extensive experiments on two widely used benchmarks with complex SL prediction tasks involving diverse medical entities. Evaluations demonstrate superior performance of our model over state-of-the-art methods on various tasks. Notably, EFOL-SL provides specific multi-hop logical reasoning chains behind its predictions, offering meaningful insights into the model's reasoning process.
Junkai Cheng, Yi-Jia Zhang 0001, Hengyi Zhang, Yifan Peng 0002, Mingyu Lu
IEEE Trans. Comput. Biol. Bioinform.4
2024 Suicide Risk Assessment on Social Media with Semi-Supervised Learning
abstract
With social media communities increasingly becoming places where suicidal individuals post and congregate, natural language processing presents an exciting avenue for the development of automated suicide risk assessment systems. However, past efforts suffer from a lack of labeled data and class imbalances within the available labeled data. To accommodate this task's imperfect data landscape, we propose a semi-supervised framework that leverages labeled (n=500) and unlabeled (n=1,500) data and expands upon the self-training algorithm with a novel pseudo-label acquisition process designed to handle imbalanced datasets. To further ensure pseudo-label quality, we manually verify a subset of the pseudo-labeled data that was not predicted unanimously across multiple trials of pseudo-label generation. We test various models to serve as the backbone for this framework, ultimately deciding that RoBERTa performs the best. Ultimately, by leveraging partially validated pseudo-labeled data in addition to ground-truth labeled data, we substantially improve our model's ability to assess suicide risk from social media posts.
Max Lovitt, Song Wang 0026, Yifan Peng 0002
IEEE Big Data4
2024 Large language models in biomedicine and health: current research landscape and future directions
abstract
Large language models in biomedicine and health: current research landscape and future directionsLarge language models (LLMs) are a specialized type of generative artificial intelligence (AI) focused on generating natural language text.These models are developed through extensive training on massive amounts of text data and use deep learning algorithms to generate new text that closely resembles human-generated text.Generative AI methods, including LLMs, are rapidly transforming various domains, including biomedicine and healthcare.[1][2][3][4][5][6] They have already demonstrated remarkable potential as a means to process and analyze large amounts of text, interpret natural language, and generate new content in these domains.For example, Nori et al reported that GPT-4 is able to correctly answer the majority of questions from medical practice licensing exams, comfortably obtaining a passing grade.7 Similarly, Stribling et al found that this model exceeded the average performance of students in the graduate medical sciences on the majority of examinations, including strong performance on short answer and essay questions.8 Even though passing the exam is not the same as applying the knowledge in a real-world setting, these results demonstrate that LLMs can generate appropriate multiple-choice and narrative responses to questions framed in natural language.ChatGPT, first released in November 2022, has garnered phenomenal attention from both the scientific community and a broader society.A keyword search of "large language models" OR "ChatGPT" in PubMed returned over 4500 articles that discuss the technology and its implications for various topics, including medical informatics, by the end of June 2024.In addition, LLM-based technologies have already been deployed in several healthcare systems and are offered as integrated products for use in the clinic within vendor electronic health record systems (for thoughts on initial evaluations of an early product, see Garcia et al 9 and Tai-Seale et al 10 ).This rapid adoption of LLMs like ChatGPT brings an unprecedented opportunity to use this novel AI technology to transform healthcare and medicine.Despite their potential benefits, LLMs can sometimes produce invalid and unsubstantiated responses, a phenomenon known as the "hallucination and confabulation issue" in the literature, or biased responses, due to the biases inherent in their training data.[11][12][13][14][15][16][17] With this great potential also comes the need for trustworthy and responsible development and use of technology.As we continue to explore the capabilities of ChatGPT and other LLMs, it is critical to address related ethical, legal, and social issues to ensure that the technology is used in ways that are safe, fair, trustworthy, and beneficial for all.In the context of biomedicine and healthcare, it is particularly important to engage stakeholders, such as AI researchers, developers of data-driven clinical decision support, care providers, and system implementers from both academic medical centers and industry, to ensure responsible use of LLMs for good.To accelerate research and development in this area, we issued a call for submissions in Summer 2023, specifically focusing on the intersection of biomedicine/health and LLMs, and invited contributions on all related aspects.We invited submissions that report on innovative informatics methods development and evaluation, as well as studies that demonstrate the effectiveness/limitations of LLMs methodologies in healthcare.We particularly encouraged submissions that address the challenges and opportunities of this intersection and offer new insights into how these fields can work together to advance healthcare.This editorial provides an overview of the papers accepted in this Focus Issue.We highlight major themes and unique aspects of the research papers in medical LLMs, discuss ongoing challenges, and recommend future research directions.Box 1 lists the relevant large language model terms and abbreviations used in this editorial. Overall statistics of the Focus IssueThis JAMIA Focus Issue on LLMs in biomedicine and health has drawn enthusiasm from many researchers across different research disciplines.In total, we received over 150 submissions from authors in 25 countries and regions across 6 continents worldwide.The rigorous JAMIA peer review process was applied to all submissions, 41 of which were ultimately accepted for publication in the Focus Issue (Table 1).The majority of the accepted papers were authored by authors in North America, followed by those in Asia and Europe (Figure 1A).The Focus Issue highlights the nature of multi-disciplinary collaboration in medical informatics research across the broad JAMIA community.The number of authors per paper varies from 1 to 23, with an average of 7.3.Many papers feature authors with diverse expertise from different departments and organizations.The authors' expertise spans a wide range of fields, including computer science, data science, informatics, statistics, medicine, nursing, clinical services, public health policies, and more.Several papers also demonstrate scientific collaborations across different sectors, including academia, government labs, research institutes, hospitals, and industry.Additionally, a few papers showcase international collaborations among authors.
Zhiyong Lu, Yifan Peng 0002, Trevor Cohen, Marzyeh Ghassemi, Chunhua Weng, Shubo Tian
J. Am. Medical Informatics Assoc.2
2024 Confidence score: a data-driven measure for inclusive systematic reviews considering unpublished preprints
abstract
OBJECTIVES: COVID-19, since its emergence in December 2019, has globally impacted research. Over 360 000 COVID-19-related manuscripts have been published on PubMed and preprint servers like medRxiv and bioRxiv, with preprints comprising about 15% of all manuscripts. Yet, the role and impact of preprints on COVID-19 research and evidence synthesis remain uncertain. MATERIALS AND METHODS: We propose a novel data-driven method for assigning weights to individual preprints in systematic reviews and meta-analyses. This weight termed the "confidence score" is obtained using the survival cure model, also known as the survival mixture model, which takes into account the time elapsed between posting and publication of a preprint, as well as metadata such as the number of first 2-week citations, sample size, and study type. RESULTS: Using 146 preprints on COVID-19 therapeutics posted from the beginning of the pandemic through April 30, 2021, we validated the confidence scores, showing an area under the curve of 0.95 (95% CI, 0.92-0.98). Through a use case on the effectiveness of hydroxychloroquine, we demonstrated how these scores can be incorporated practically into meta-analyses to properly weigh preprints. DISCUSSION: It is important to note that our method does not aim to replace existing measures of study quality but rather serves as a supplementary measure that overcomes some limitations of current approaches. CONCLUSION: Our proposed confidence score has the potential to improve systematic reviews of evidence related to COVID-19 and other clinical conditions by providing a data-driven approach to including unpublished manuscripts.
Jiayi Tong, Chongliang Luo, Yifei Sun 0007, Rui Duan 0004, M. Elle Saine, Yifan Peng 0002, Anchita Batra, Anni Pan, Olivia Wang, Ruowang Li, Arielle Marks-Anglin, Xu Zuo, Yulun Liu 0004, Jiang Bian 0001, Stephen E. Kimmel, Keith Hamilton, Adam Cuker, Rebecca A. Hubbard, Hua Xu 0001, Yong Chen 0016
J. Am. Medical Informatics Assoc.7
2024 Deep learning with noisy labels in medical prediction problems: a scoping review
abstract
OBJECTIVES: Medical research faces substantial challenges from noisy labels attributed to factors like inter-expert variability and machine-extracted labels. Despite this, the adoption of label noise management remains limited, and label noise is largely ignored. To this end, there is a critical need to conduct a scoping review focusing on the problem space. This scoping review aims to comprehensively review label noise management in deep learning-based medical prediction problems, which includes label noise detection, label noise handling, and evaluation. Research involving label uncertainty is also included. METHODS: Our scoping review follows the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. We searched 4 databases, including PubMed, IEEE Xplore, Google Scholar, and Semantic Scholar. Our search terms include "noisy label AND medical/healthcare/clinical," "uncertainty AND medical/healthcare/clinical," and "noise AND medical/healthcare/clinical." RESULTS: A total of 60 papers met inclusion criteria between 2016 and 2023. A series of practical questions in medical research are investigated. These include the sources of label noise, the impact of label noise, the detection of label noise, label noise handling techniques, and their evaluation. Categorization of both label noise detection methods and handling techniques are provided. DISCUSSION: From a methodological perspective, we observe that the medical community has been up to date with the broader deep-learning community, given that most techniques have been evaluated on medical data. We recommend considering label noise as a standard element in medical research, even if it is not dedicated to handling noisy labels. Initial experiments can start with easy-to-implement methods, such as noise-robust loss functions, weighting, and curriculum learning.
Yishu Wei, Cong Sun 0004, Mingquan Lin, Hongmei Jiang, Yifan Peng 0002
J. Am. Medical Informatics Assoc.6
2024 A span-based model for extracting overlapping PICO entities from randomized controlled trial publications
abstract
OBJECTIVES: Extracting PICO (Populations, Interventions, Comparison, and Outcomes) entities is fundamental to evidence retrieval. We present a novel method, PICOX, to extract overlapping PICO entities. MATERIALS AND METHODS: PICOX first identifies entities by assessing whether a word marks the beginning or conclusion of an entity. Then, it uses a multi-label classifier to assign one or more PICO labels to a span candidate. PICOX was evaluated using 1 of the best-performing baselines, EBM-NLP, and 3 more datasets, ie, PICO-Corpus and randomized controlled trial publications on Alzheimer's Disease (AD) or COVID-19, using entity-level precision, recall, and F1 scores. RESULTS: PICOX achieved superior precision, recall, and F1 scores across the board, with the micro F1 score improving from 45.05 to 50.87 (P ≪.01). On the PICO-Corpus, PICOX obtained higher recall and F1 scores than the baseline and improved the micro recall score from 56.66 to 67.33. On the COVID-19 dataset, PICOX also outperformed the baseline and improved the micro F1 score from 77.10 to 80.32. On the AD dataset, PICOX demonstrated comparable F1 scores with higher precision when compared to the baseline. CONCLUSION: PICOX excels in identifying overlapping entities and consistently surpasses a leading baseline across multiple datasets. Ablation studies reveal that its data augmentation strategy effectively minimizes false positives and improves precision.
Yiliang Zhou, Hua Xu 0001, Chunhua Weng, Yifan Peng 0002
J. Am. Medical Informatics Assoc.6
2024 Call for papers: Special issue on biomedical multimodal large language models - novel approaches and applications
Jiang Bian 0001, Yifan Peng 0002, Eneida A. Mendonça, Imon Banerjee, Hua Xu 0001, Casey Overby Taylor, Anália Maria Garcia Lourenço, Alejandro Rodríguez González, Elena Tutubalina
J. Biomed. Informatics2
2024 Demonstration-based learning for few-shot biomedical named entity recognition under machine reading comprehension
Leilei Su, Yifan Peng 0002, Cong Sun 0004
J. Biomed. Informatics3
2024 A survey of recent methods for addressing AI fairness and bias in biomedicine
abstract
OBJECTIVES: Artificial intelligence (AI) systems have the potential to revolutionize clinical practices, including improving diagnostic accuracy and surgical decision-making, while also reducing costs and manpower. However, it is important to recognize that these systems may perpetuate social inequities or demonstrate biases, such as those based on race or gender. Such biases can occur before, during, or after the development of AI models, making it critical to understand and address potential biases to enable the accurate and reliable application of AI models in clinical settings. To mitigate bias concerns during model development, we surveyed recent publications on different debiasing methods in the fields of biomedical natural language processing (NLP) or computer vision (CV). Then we discussed the methods, such as data perturbation and adversarial learning, that have been applied in the biomedical domain to address bias. METHODS: We performed our literature search on PubMed, ACM digital library, and IEEE Xplore of relevant articles published between January 2018 and December 2023 using multiple combinations of keywords. We then filtered the result of 10,041 articles automatically with loose constraints, and manually inspected the abstracts of the remaining 890 articles to identify the 55 articles included in this review. Additional articles in the references are also included in this review. We discuss each method and compare its strengths and weaknesses. Finally, we review other potential methods from the general domain that could be applied to biomedicine to address bias and improve fairness. RESULTS: The bias of AIs in biomedicine can originate from multiple sources such as insufficient data, sampling bias and the use of health-irrelevant features or race-adjusted algorithms. Existing debiasing methods that focus on algorithms can be categorized into distributional or algorithmic. Distributional methods include data augmentation, data perturbation, data reweighting methods, and federated learning. Algorithmic approaches include unsupervised representation learning, adversarial learning, disentangled representation learning, loss-based methods and causality-based methods.
Yifan Yang 0006, Mingquan Lin, Han Zhao 0002, Yifan Peng 0002, Furong Huang, Zhiyong Lu
J. Biomed. Informatics4
2024 Identifying social determinants of health from clinical narratives: A study of performance, documentation ratio, and potential bias
Zehao Yu 0001, Cheng Peng 0009, Xi Yang 0015, Chong Dang, Prakash Adekkanattu, Braja Gopal Patra, Yifan Peng 0002, Jyotishman Pathak, Debbie L. Wilson, Ching-Yuan Chang, Wei-Hsuan Lo-Ciganic, Thomas J. George, William R. Hogan, Yi Guo 0005, Jiang Bian 0001, Yonghui Wu 0001
J. Biomed. Informatics7
2024 Leveraging generative AI for clinical evidence synthesis needs to ensure trustworthiness
Qiao Jin 0001, Denis Jered McInerney, Yong Chen 0016, Fei Wang 0001, Curtis L. Cole, Qian Yang 0004, Yanshan Wang, Bradley A. Malin, Mor Peleg, Byron C. Wallace, Zhiyong Lu, Chunhua Weng, Yifan Peng 0002
J. Biomed. Informatics14
2024 Towards long-tailed, multi-label disease classification from chest X-ray: Overview of the CXR-LT challenge
Gregory Holste, Yiliang Zhou, Song Wang 0026, Ajay Jaiswal, Mingquan Lin, Sherry Zhuge, Yuzhe Yang 0003, Dongkyun Kim, Trong-Hieu Nguyen Mau, Minh-Triet Tran, Jaehyup Jeong, Wongi Park, Jong Bin Ryu, Feng Hong 0004, Arsh Verma, Yosuke Yamagishi, Hyeryeong Seo, Myungjoo Kang, Leo A. Celi, Zhiyong Lu, Ronald M. Summers, George Shih, Zhangyang Wang, Yifan Peng 0002
Medical Image Anal.25
2023 How Does Pruning Impact Long-Tailed Multi-label Medical Image Classifiers?
Gregory Holste, Ziyu Jiang, Ajay Jaiswal, Maria Hanna, Shlomo Minkowitz, Alan C. Legasto, Joanna G. Escalon, Sharon Steinberger, Mark Bittman, Thomas C. Shen, Ying Ding 0001, Ronald M. Summers, George Shih, Yifan Peng 0002, Zhangyang Wang
MICCAI (5)14
2023 Utilizing Longitudinal Chest X-Rays and Reports to Pre-fill Radiology Reports
Qingqing Zhu, Tejas Sudharshan Mathai, Pritam Mukherjee, Yifan Peng 0002, Ronald M. Summers, Zhiyong Lu
MICCAI (5)4
2023 Attend Who is Weak: Pruning-assisted Medical Image Localization under Sophisticated and Implicit Imbalances
abstract
Deep neural networks (DNNs) have rapidly become a de facto choice for medical image understanding tasks. However, DNNs are notoriously fragile to the class imbalance in image classification. We further point out that such imbalance fragility can be amplified when it comes to more sophisticated tasks such as pathology localization, as imbalances in such problems can have highly complex and often implicit forms of presence. For example, different pathology can have different sizes or colors (w.r.t.the background), different underlying demographic distributions, and in general different difficulty levels to recognize, even in a meticulously curated balanced distribution of training data. In this paper, we propose to use pruning to automatically and adaptively identify hard-to-learn (HTL) training samples, and improve pathology localization by attending them explicitly, during training in supervised, semi-supervised, and weakly-supervised settings. Our main inspiration is drawn from the recent finding that deep classification models have difficult-to-memorize samples and those may be effectively exposed through network pruning [15] - and we extend such observation beyond classification for the first time. We also present an interesting demographic analysis which illustrates HTLs ability to capture complex demographic imbalances. Our extensive experiments on the Skin Lesion Localization task in multiple training settings by paying additional attention to HTLs show significant improvement of localization performance by ~2-3%.
Ajay Jaiswal, Tianlong Chen 0001, Justin F. Rousseau, Yifan Peng 0002, Ying Ding 0001, Zhangyang Wang
WACV4
2023 EvidenceMap: a three-level knowledge representation for medical evidence computation and comprehension
abstract
OBJECTIVE: To develop a computable representation for medical evidence and to contribute a gold standard dataset of annotated randomized controlled trial (RCT) abstracts, along with a natural language processing (NLP) pipeline for transforming free-text RCT evidence in PubMed into the structured representation. MATERIALS AND METHODS: Our representation, EvidenceMap, consists of 3 levels of abstraction: Medical Evidence Entity, Proposition and Map, to represent the hierarchical structure of medical evidence composition. Randomly selected RCT abstracts were annotated following EvidenceMap based on the consensus of 2 independent annotators to train an NLP pipeline. Via a user study, we measured how the EvidenceMap improved evidence comprehension and analyzed its representative capacity by comparing the evidence annotation with EvidenceMap representation and without following any specific guidelines. RESULTS: Two corpora including 229 disease-agnostic and 80 COVID-19 RCT abstracts were annotated, yielding 12 725 entities and 1602 propositions. EvidenceMap saves users 51.9% of the time compared to reading raw-text abstracts. Most evidence elements identified during the freeform annotation were successfully represented by EvidenceMap, and users gave the enrollment, study design, and study Results sections mean 5-scale Likert ratings of 4.85, 4.70, and 4.20, respectively. The end-to-end evaluations of the pipeline show that the evidence proposition formulation achieves F1 scores of 0.84 and 0.86 in the adjusted random index score. CONCLUSIONS: EvidenceMap extends the participant, intervention, comparator, and outcome framework into 3 levels of abstraction for transforming free-text evidence from the clinical literature into a computable structure. It can be used as an interoperable format for better evidence retrieval and synthesis and an interpretable representation to efficiently comprehend RCT findings.
Tian Kang, Yingcheng Sun, Jae Hyun Kim, Casey N. Ta, Adler J. Perotte, Kayla Schiffer, Mutong Wu, Nour Fahmy, Yifan Peng 0002, Chunhua Weng
J. Am. Medical Informatics Assoc.10
2023 An NLP approach to identify SDoH-related circumstance and suicide crisis from death investigation narratives
abstract
OBJECTIVES: Suicide presents a major public health challenge worldwide, affecting people across the lifespan. While previous studies revealed strong associations between Social Determinants of Health (SDoH) and suicide deaths, existing evidence is limited by the reliance on structured data. To resolve this, we aim to adapt a suicide-specific SDoH ontology (Suicide-SDoHO) and use natural language processing (NLP) to effectively identify individual-level SDoH-related social risks from death investigation narratives. MATERIALS AND METHODS: We used the latest National Violent Death Report System (NVDRS), which contains 267 804 victim suicide data from 2003 to 2019. After adapting the Suicide-SDoHO, we developed a transformer-based model to identify SDoH-related circumstances and crises in death investigation narratives. We applied our model retrospectively to annotate narratives whose crisis variables were not coded in NVDRS. The crisis rates were calculated as the percentage of the group's total suicide population with the crisis present. RESULTS: The Suicide-SDoHO contains 57 fine-grained circumstances in a hierarchical structure. Our classifier achieves AUCs of 0.966 and 0.942 for classifying circumstances and crises, respectively. Through the crisis trend analysis, we observed that not everyone is equally affected by SDoH-related social risks. For the economic stability crisis, our result showed a significant increase in crisis rate in 2007-2009, parallel with the Great Recession. CONCLUSIONS: This is the first study curating a Suicide-SDoHO using death investigation narratives. We showcased that our model can effectively classify SDoH-related social risks through NLP approaches. We hope our study will facilitate the understanding of suicide crises and inform effective prevention strategies.
Song Wang 0026, Yifang Dang, Zhaoyi Sun, Ying Ding 0001, Jyotishman Pathak, Cui Tao, Yunyu Xiao, Yifan Peng 0002
J. Am. Medical Informatics Assoc.8
2023 Representing and utilizing clinical textual data for real world studies: An OHDSI approach
Vipina Kuttichi Keloth, Juan M. Banda, Michael J. Gurley, Paul M. Heider, Georgina Kennedy, Timothy A. Miller, Karthik Natarajan, Olga V. Patterson, Yifan Peng 0002, Kalpana Raja, Ruth M. Reeves, Masoud Rouhizadeh, Jianlin Shi, Yanshan Wang, Wei-Qi Wei, Andrew E. Williams, Rui Zhang 0028, Rimma Belenkaya, Christian G. Reich, Clair Blacketer, Patrick B. Ryan, George Hripcsak, Noémie Elhadad, Hua Xu 0001
J. Biomed. Informatics11
2023 A scoping review on multimodal deep learning in biomedical images and texts
Zhaoyi Sun, Mingquan Lin, Qingqing Zhu, Qianqian Xie, Fei Wang 0001, Zhiyong Lu, Yifan Peng 0002
J. Biomed. Informatics7
2023 Radiomics-Guided Global-Local Transformer for Weakly Supervised Pathology Localization in Chest X-Rays
abstract
Before the recent success of deep learning methods for automated medical image analysis, practitioners used handcrafted radiomic features to quantitatively describe local patches of medical images. However, extracting discriminative radiomic features relies on accurate pathology localization, which is difficult to acquire in real-world settings. Despite advances in disease classification and localization from chest X-rays, many approaches fail to incorporate clinically-informed domainspecific radiomic features. For these reasons, we propose a Radiomics-Guided Transformer (RGT) that fuses global image information with local radiomics-guided auxiliary information to provide accurate cardiopulmonary pathology localization and classification without any bounding box annotations. RGT consists of an image Transformer branch, a radiomics Transformer branch, and fusion layers that aggregate image and radiomics information. Using the learned self-attention of its image branch, RGT extracts a bounding box for which to compute radiomic features, which are further processed by the radiomics branch; learned image and radiomic features are then fused and mutually interact via cross-attention layers. Thus, RGT utilizes a novel end-to-end feedback loop that can bootstrap accurate pathology localization only using image-level disease labels. Experiments on the NIH ChestXRay dataset demonstrate that RGT outperforms prior works in weakly supervised disease localization (by an average margin of 3.6% over various intersection-over-union thresholds) and classification (by 1.1% in average area under the receiver operating characteristic curve). We publicly release our codes and pre-trained models at https://github.com/VITAGroup/chext.
Yan Han 0001, Gregory Holste, Ying Ding 0001, Ahmed H. Tewfik, Yifan Peng 0002, Zhangyang Wang
IEEE Trans. Medical Imaging5
2022 Prompt-based Learning for Assertion Classification in Clinical Notes
Song Wang 0026, Liyan Tang, Akash Majety, Justin F. Rousseau, George Shih, Ying Ding 0001, Yifan Peng 0002
AMIA7
2022 RoS-KD: A Robust Stochastic Knowledge Distillation Approach for Noisy Medical Imaging
abstract
AI-powered Medical Imaging has recently achieved enormous attention due to its ability to provide fast-paced healthcare diagnoses. However, it usually suffers from a lack of high-quality datasets due to high annotation cost, interobserver variability, human annotator error, and errors in computer-generated labels. Deep learning models trained on noisy labelled datasets are sensitive to the noise type and lead to less generalization on the unseen samples. To address this challenge, we propose a Robust Stochastic Knowledge Distillation (RoS-KD) framework which mimics the notion of learning a topic from multiple sources to ensure deterrence in learning noisy information. More specifically, RoS-KD learns a smooth, well-informed, and robust student manifold by distilling knowledge from multiple teachers trained on overlapping subsets of training data. Our extensive experiments on popular medical imaging classification tasks (cardiopulmonary disease and lesion classification) using real-world datasets, show the performance benefit of RoS-KD, its ability to distill knowledge from many popular large networks (ResNet-50, DenseNet-121, MobileNetV2) in a comparatively small network, and its robustness to adversarial attacks (PGD, FSGM). More specifically, RoS-KD achieves >2% and > 4% improvement on F1-score for lesion classification and cardiopulmonary disease classification tasks, respectively, when the underlying student is ResNet-18 against recent competitive knowledge distillation baseline. Additionally, on cardiopulmonary disease classification task, RoS-KD outperforms most of the SOTA baselines by ~1% gain in AUC score.
Ajay Jaiswal, Kumar Ashutosh, Justin F. Rousseau, Yifan Peng 0002, Zhangyang Wang, Ying Ding 0001
ICDM4
2022 Knowledge-Augmented Contrastive Learning for Abnormality Classification and Localization in Chest X-rays with Radiomics using a Feedback Loop
abstract
Accurate classification and localization of abnormalities in chest X-rays play an important role in clinical diagnosis and treatment planning. Building a highly accurate predictive model for these tasks usually requires a large number of manually annotated labels and pixel regions (bounding boxes) of abnormalities. However, it is expensive to acquire such annotations, especially the bounding boxes. Recently, contrastive learning has shown strong promise in leveraging unlabeled natural images to produce highly generalizable and discriminative features. However, extending its power to the medical image domain is under-explored and highly non-trivial, since medical images are much less amendable to data augmentations. In contrast, their prior knowledge, as well as radiomic features, is often crucial. To bridge this gap, we propose an end-to-end semi-supervised knowledge-augmented contrastive learning framework, that simultaneously performs disease classification and localization tasks. The key knob of our framework is a unique positive sampling approach tailored for the medical images, by seamlessly integrating radiomic features as a knowledge augmentation. Specifically, we first apply an image encoder to classify the chest X-rays and to generate the image features. We next leverage Grad-CAM to highlight the crucial (abnormal) regions for chest X-rays (even when unannotated), from which we extract radiomic features. The radiomic features are then passed through another dedicated encoder to act as the positive sample for the image features generated from the same chest X-ray. In this way, our framework constitutes a feedback loop for image and radiomic features to mutually reinforce each other. Their contrasting yields knowledge-augmented representations that are both robust and interpretable. Extensive experiments on the NIH Chest X-ray dataset demonstrate that our approach outperforms existing baselines in both classification and localization tasks.
Yan Han 0001, Chongyan Chen, Ahmed H. Tewfik, Benjamin S. Glicksberg, Ying Ding 0001, Yifan Peng 0002, Zhangyang Wang
WACV6
2022 Design and validation of a FHIR-based EHR-driven phenotyping toolbox
abstract
OBJECTIVES: To develop and validate a standards-based phenotyping tool to author electronic health record (EHR)-based phenotype definitions and demonstrate execution of the definitions against heterogeneous clinical research data platforms. MATERIALS AND METHODS: We developed an open-source, standards-compliant phenotyping tool known as the PhEMA Workbench that enables a phenotype representation using the Fast Healthcare Interoperability Resources (FHIR) and Clinical Quality Language (CQL) standards. We then demonstrated how this tool can be used to conduct EHR-based phenotyping, including phenotype authoring, execution, and validation. We validated the performance of the tool by executing a thrombotic event phenotype definition at 3 sites, Mayo Clinic (MC), Northwestern Medicine (NM), and Weill Cornell Medicine (WCM), and used manual review to determine precision and recall. RESULTS: An initial version of the PhEMA Workbench has been released, which supports phenotype authoring, execution, and publishing to a shared phenotype definition repository. The resulting thrombotic event phenotype definition consisted of 11 CQL statements, and 24 value sets containing a total of 834 codes. Technical validation showed satisfactory performance (both NM and MC had 100% precision and recall and WCM had a precision of 95% and a recall of 84%). CONCLUSIONS: We demonstrate that the PhEMA Workbench can facilitate EHR-driven phenotype definition, execution, and phenotype sharing in heterogeneous clinical research data environments. A phenotype definition that integrates with existing standards-compliant systems, and the use of a formal representation facilitates automation and can decrease potential for human error.
Pascal S. Brandt, Jennifer A. Pacheco, Prakash Adekkanattu, Evan Sholle, Sajjad Abedian, Daniel J. Stone, David Knaack, Jie Xu 0012, Yifan Peng 0002, Natalie C. Benda, Fei Wang 0001, Yuan Luo 0001, Guoqian Jiang, Jyotishman Pathak, Luke V. Rasmussen
J. Am. Medical Informatics Assoc.10
2022 Trustworthy assertion classification through prompting
Song Wang 0026, Liyan Tang, Akash Majety, Justin F. Rousseau, George Shih, Ying Ding 0001, Yifan Peng 0002
J. Biomed. Informatics7
2022 Global-Local attention network with multi-task uncertainty loss for abnormal lymph node detection in MR images
Shuai Wang 0003, Yingying Zhu 0003, Sungwon Lee 0003, Daniel C. Elton, Thomas C. Shen, Youbao Tang, Yifan Peng 0002, Zhiyong Lu, Ronald M. Summers
Medical Image Anal.7
2022 Robust convolutional neural networks against adversarial attacks on medical images
abstract
Convolutional neural networks (CNNs) have been widely applied to medical images. However, medical images are vulnerable to adversarial attacks by perturbations that are undetectable to human experts. This poses significant security risks and challenges to CNN-based applications in clinic practice. In this work, we quantify the scale of adversarial perturbation imperceptible to clinical practitioners and investigate the cause of the vulnerability in CNNs. Specifically, we discover that noise (i.e., irrelevant or corrupted discriminative information) in medical images might be a key contributor to performance deterioration of CNNs against adversarial perturbations, as noisy features are learned unconsciously by CNNs in feature representations and magnified by adversarial perturbations. In response, we propose a novel defense method by embedding sparsity denoising operators in CNNs for improved robustness. Tested with various state-of-the-art attacking methods on two distinct medical image modalities, we demonstrate that the proposed method can successfully defend against those unnoticeable adversarial attacks by retaining as much as over 90% of its original performance. We believe our findings are critical for improving and deploying CNN-based medical applications in real-world scenarios.
Xiaoshuang Shi, Yifan Peng 0002, Qingyu Chen 0001, Tiarnan D. Keenan, Alisa T. Thavikulwat, Sungwon Lee 0003, Yuxing Tang, Emily Y. Chew, Ronald M. Summers, Zhiyong Lu
Pattern Recognit.2
2021 Multi-task deep learning-based survival analysis on the prognosis of late AMD using the longitudinal data in AREDS
Gregory C. Ghahramani, Matthew Brendel, Mingquan Lin, Qingyu Chen 0001, Tiarnan D. Keenan, Kun Chen 0002, Emily Y. Chew, Zhiyong Lu, Yifan Peng 0002, Fei Wang 0001
AMIA9
2021 Using Radiomics as Prior Knowledge for Thorax Disease Classification and Localization in Chest X-rays
Yan Han 0001, Chongyan Chen, Liyan Tang, Mingquan Lin, Ajay Jaiswal, Song Wang 0026, Ahmed H. Tewfik, George Shih, Ying Ding 0001, Yifan Peng 0002
AMIA10
2021 SCALP - Supervised Contrastive Learning for Cardiopulmonary Disease Classification and Localization in Chest X-rays using Patient Metadata
abstract
Computer-aided diagnosis plays a salient role in more accessible and accurate cardiopulmonary diseases classification and localization on chest radiography. Millions of people get affected and die due to these diseases without an accurate and timely diagnosis. Recently proposed contrastive learning heavily relies on data augmentation, especially positive data augmentation. However, generating clinically-accurate data augmentations for medical images is extremely difficult because the common data augmentation methods in computer vision, such as sharp, blur, and crop operations, can severely alter the clinical settings of medical images. In this paper, we proposed a novel and simple data augmentation method based on patient metadata and supervised knowledge to create clinically accurate positive and negative augmentations for chest X-rays. We introduce an end-to-end framework, SCALP, which extends the self-supervised contrastive approach to a supervised setting. Specifically, SCALP pulls together chest X-rays from the same patient (positive keys) and pushes apart chest X-rays from different patients (negative keys). In addition, it uses ResNet-50 along with the triplet-attention mechanism to identify cardiopulmonary diseases, and Grad-CAM++ to highlight the abnormal regions. Our extensive experiments demonstrate that SCALP outperforms existing baselines with significant margins in both classification and localization tasks. Specifically, the average classification AUCs improve from 82.8% (SOTA using DenseNet-121) to 83.9% (SCALP using ResNet-50), while the localization results improve on average by 3.7% over different IoU thresholds.
Ajay Jaiswal, Cyprian Zander, Yan Han 0001, Justin F. Rousseau, Yifan Peng 0002, Ying Ding 0001
ICDM6
2021 Leveraging Deep Representations of Radiology Reports in Survival Analysis for Predicting Heart Failure Patient Mortality
abstract
Utilizing clinical texts in survival analysis is difficult because they are largely unstructured. Current automatic extraction models fail to capture textual information comprehensively since their labels are limited in scope. Furthermore, they typically require a large amount of data and high-quality expert annotations for training. In this work, we present a novel method of using BERT-based hidden layer representations of clinical texts as covariates for proportional hazards models to predict patient survival outcomes. We show that hidden layers yield notably more accurate predictions than predefined features, outperforming the previous baseline model by 5.7% on average across C-index and time-dependent AUC. We make our work publicly available at https://github.com/bionlplab/heart_failure_mortality.
Hyun Gi Lee, Evan Sholle, Ashley Beecy, Subhi J. Al'Aref, Yifan Peng 0002
NAACL-HLT5
2021 Multimodal, multitask, multiattention (M3) deep learning detection of reticular pseudodrusen: Toward automated and accessible classification of age-related macular degeneration
abstract
OBJECTIVE: Reticular pseudodrusen (RPD), a key feature of age-related macular degeneration (AMD), are poorly detected by human experts on standard color fundus photography (CFP) and typically require advanced imaging modalities such as fundus autofluorescence (FAF). The objective was to develop and evaluate the performance of a novel multimodal, multitask, multiattention (M3) deep learning framework on RPD detection. MATERIALS AND METHODS: A deep learning framework (M3) was developed to detect RPD presence accurately using CFP alone, FAF alone, or both, employing >8000 CFP-FAF image pairs obtained prospectively (Age-Related Eye Disease Study 2). The M3 framework includes multimodal (detection from single or multiple image modalities), multitask (training different tasks simultaneously to improve generalizability), and multiattention (improving ensembled feature representation) operation. Performance on RPD detection was compared with state-of-the-art deep learning models and 13 ophthalmologists; performance on detection of 2 other AMD features (geographic atrophy and pigmentary abnormalities) was also evaluated. RESULTS: For RPD detection, M3 achieved an area under the receiver-operating characteristic curve (AUROC) of 0.832, 0.931, and 0.933 for CFP alone, FAF alone, and both, respectively. M3 performance on CFP was very substantially superior to human retinal specialists (median F1 score = 0.644 vs 0.350). External validation (the Rotterdam Study) demonstrated high accuracy on CFP alone (AUROC, 0.965). The M3 framework also accurately detected geographic atrophy and pigmentary abnormalities (AUROC, 0.909 and 0.912, respectively), demonstrating its generalizability. CONCLUSIONS: This study demonstrates the successful development, robust evaluation, and external validation of a novel deep learning framework that enables accessible, accurate, and automated AMD diagnosis and prognosis.
Qingyu Chen 0001, Tiarnan D. Keenan, Alexis Allot, Yifan Peng 0002, Elvira Agrón, Amitha Domalpally, Caroline C. W. Klaver, Daniel T. Luttikhuizen, Marcus H. Colyer, Catherine Cukras, Henry E. Wiley, M. Teresa Magone, Chantal Cousineau-Krieger, Wai T. Wong, Yingying Zhu 0003, Emily Y. Chew, Zhiyong Lu
J. Am. Medical Informatics Assoc.4
2021 COVID-19-CT-CXR: A Freely Accessible and Weakly Labeled Chest X-Ray and CT Image Collection on COVID-19 From Biomedical Literature
abstract
The latest threat to global health is the COVID-19 outbreak. Although there exist large datasets of chest X-rays (CXR) and computed tomography (CT) scans, few COVID-19 image collections are currently available due to patient privacy. At the same time, there is a rapid growth of COVID-19-relevant articles in the biomedical literature, including those that report findings on radiographs. Here, we present COVID-19-CT-CXR, a public database of COVID-19 CXR and CT images, which are automatically extracted from COVID-19-relevant articles from the PubMed Central Open Access (PMC-OA) Subset. We extracted figures, associated captions, and relevant figure descriptions in the article and separated compound figures into subfigures. Because a large portion of figures in COVID-19 articles are not CXR or CT, we designed a deep-learning model to distinguish them from other figure types and to classify them accordingly. The final database includes 1,327 CT and 263 CXR images (as of May 9, 2020) with their relevant text. To demonstrate the utility of COVID-19-CT-CXR, we conducted four case studies. (1) We show that COVID-19-CT-CXR, when used as additional training data, is able to contribute to improved deep-learning (DL) performance for the classification of COVID-19 and non-COVID-19 CT. (2) We collected CT images of influenza, another common infectious respiratory illness that may present similarly to COVID-19, and fine-tuned a baseline deep neural network to distinguish a diagnosis of COVID-19, influenza, or normal or other types of diseases on CT. (3) We fine-tuned an unsupervised one-class classifier from non-COVID-19 CXR and performed anomaly detection to detect COVID-19 CXR. (4) From text-mined captions and figure descriptions, we compared 15 clinical symptoms and 20 clinical findings of COVID-19 versus those of influenza to demonstrate the disease differences in the scientific publications. Our database is unique, as the figures are retrieved along with relevant text with fine-grained descriptions, and it can be extended easily in the future. We believe that our work is complementary to existing resources and hope that it will contribute to medical image analysis of the COVID-19 pandemic. The dataset, code, and DL models are publicly available at https://github.com/ncbi-nlp/COVID-19-CT-CXR.
Yifan Peng 0002, Yuxing Tang, Sungwon Lee 0003, Yingying Zhu 0003, Ronald M. Summers, Zhiyong Lu
IEEE Trans. Big Data1
2021 Generalized Zero-Shot Chest X-Ray Diagnosis Through Trait-Guided Multi-View Semantic Embedding With Self-Training
abstract
Zero-shot learning (ZSL) is one of the most promising avenues of annotation-efficient machine learning. In the era of deep learning, ZSL techniques have achieved unprecedented success. However, the developments of ZSL methods have taken place mostly for natural images. ZSL for medical images has remained largely unexplored. We design a novel strategy for generalized zero-shot diagnosis of chest radiographs. In doing so, we leverage the potential of multi-view semantic embedding, a useful yet less-explored direction for ZSL. Our design also incorporates a self-training phase to tackle the problem of noisy labels alongside improving the performance for classes not seen during training. Through rigorous experiments, we show that our model trained on one dataset can produce consistent performance across test datasets from different sources including those with very different quality. Comparisons with a number of state-of-the-art techniques show the superiority of the proposed method for generalized zero-shot chest x-ray diagnosis.
Angshuman Paul, Thomas C. Shen, Sungwon Lee 0003, Niranjan Balachandar, Yifan Peng 0002, Zhiyong Lu, Ronald M. Summers
IEEE Trans. Medical Imaging5
2020 Detection of reticular pseudodrusen using deep learning
Qingyu Chen 0001, Tiarnan D. Keenan, Yifan Peng 0002, Elvira Agrón, Christopher Hwang, Alisa T. Thavikulwat, Debora Lee, Wai T. Wong, Emily Y. Chew, Zhiyong Lu
AMIA3
2019 A deep learning-based survival model for prediction of progression in late Age-related Macular Degeneration (AMD) from color fundus photographs
Yifan Peng 0002, Tiarnan D. Keenan, Qingyu Chen 0001, Elvira Agrón, Wai T. Wong, Emily Y. Chew, Zhiyong Lu
AMIA1
2019 Holistic and Comprehensive Annotation of Clinically Significant Findings on Diverse CT Images: Learning From Radiology Reports and Label Ontology
abstract
In radiologists' routine work, one major task is to read a medical image, e.g., a CT scan, find significant lesions, and describe them in the radiology report. In this paper, we study the lesion description or annotation problem. Given a lesion image, our aim is to predict a comprehensive set of relevant labels, such as the lesion's body part, type, and attributes, which may assist downstream fine-grained diagnosis. To address this task, we first design a deep learning module to extract relevant semantic labels from the radiology reports associated with the lesion images. With the images and text-mined labels, we propose a lesion annotation network (LesaNet) based on a multilabel convolutional neural network (CNN) to learn all labels holistically. Hierarchical relations and mutually exclusive relations between the labels are leveraged to improve the label prediction accuracy. The relations are utilized in a label expansion strategy and a reliable hard example mining algorithm. We also attach a simple score propagation layer on LesaNet to enhance recall and explore implicit relation between labels. Multilabel metric learning is combined with classification to enable interpretable prediction. We evaluated LesaNet on the public DeepLesion dataset, which contains over 32K diverse lesion images. Experiments show that LesaNet can precisely annotate the lesions using an ontology of 171 fine-grained labels with an average AUC of 0.9344.
Ke Yan 0006, Yifan Peng 0002, Veit Sandfort, Mohammadhadi Bagheri, Zhiyong Lu, Ronald M. Summers
CVPR2
2019 MULAN: Multitask Universal Lesion Analysis Network for Joint Lesion Detection, Tagging, and Segmentation
Ke Yan 0006, Youbao Tang, Yifan Peng 0002, Veit Sandfort, Mohammadhadi Bagheri, Zhiyong Lu, Ronald M. Summers
MICCAI (6)3
2019 ML-Net: multi-label classification of biomedical texts with deep neural networks
abstract
OBJECTIVE: In multi-label text classification, each textual document is assigned 1 or more labels. As an important task that has broad applications in biomedicine, a number of different computational methods have been proposed. Many of these methods, however, have only modest accuracy or efficiency and limited success in practical use. We propose ML-Net, a novel end-to-end deep learning framework, for multi-label classification of biomedical texts. MATERIALS AND METHODS: ML-Net combines a label prediction network with an automated label count prediction mechanism to provide an optimal set of labels. This is accomplished by leveraging both the predicted confidence score of each label and the deep contextual information (modeled by ELMo) in the target document. We evaluate ML-Net on 3 independent corpora in 2 text genres: biomedical literature and clinical notes. For evaluation, we use example-based measures, such as precision, recall, and the F measure. We also compare ML-Net with several competitive machine learning and deep learning baseline models. RESULTS: Our benchmarking results show that ML-Net compares favorably to state-of-the-art methods in multi-label classification of biomedical text. ML-Net is also shown to be robust when evaluated on different text genres in biomedicine. CONCLUSION: ML-Net is able to accuractely represent biomedical document context and dynamically estimate the label count in a more systematic and accurate manner. Unlike traditional machine learning methods, ML-Net does not require human effort for feature engineering and is a highly efficient and scalable approach to tasks with a large set of labels, so there is no need to build individual classifiers for each separate label.
Jingcheng Du, Qingyu Chen 0001, Yifan Peng 0002, Yang Xiang 0003, Cui Tao, Zhiyong Lu
J. Am. Medical Informatics Assoc.3
2018 TieNet: Text-Image Embedding Network for Common Thorax Disease Classification and Reporting in Chest X-Rays
abstract
Chest X-rays are one of the most common radiological examinations in daily clinical routines. Reporting thorax diseases using chest X-rays is often an entry-level task for radiologist trainees. Yet, reading a chest X-ray image remains a challenging job for learning-oriented machine intelligence, due to (1) shortage of large-scale machine-learnable medical image datasets, and (2) lack of techniques that can mimic the high-level reasoning of human radiologists that requires years of knowledge accumulation and professional training. In this paper, we show the clinical free-text radiological reportscan be utilized as a priori knowledge for tackling these two key problems. We propose a novel Text-Image Embedding network (TieNet) for extracting the distinctive image and text representations. Multi-level attention models are integrated into an end-to-end trainable CNN-RNN architecture for highlighting the meaningful text words and image regions. We first apply TieNet to classify the chest X-rays by using both image features and text embeddings extracted from associated reports. The proposed auto-annotation framework achieves high accuracy (over 0.9 on average in AUCs) in assigning disease labels for our hand-label evaluation dataset. Furthermore, we transform the TieNet into a chest X-ray reporting system. It simulates the reporting process and can output disease classification and a preliminary report together. The classification results are significantly improved (6% increase on average in AUCs) compared to the state-of-the-art baseline on an unseen and hand-labeled dataset (OpenI).
Xiaosong Wang 0001, Yifan Peng 0002, Le Lu 0001, Zhiyong Lu, Ronald M. Summers
CVPR2
2017 Text Mining Radiology Reports for Deep Learning Radiology Images
Yifan Peng 0002, Xiaosong Wang 0001, Le Lu 0001, Mohammadhadi Bagheri, Ronald M. Summers, Zhiyong Lu
AMIA1
2017 ChestX-Ray8: Hospital-Scale Chest X-Ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax Diseases
abstract
The chest X-ray is one of the most commonly accessible radiological examinations for screening and diagnosis of many lung diseases. A tremendous number of X-ray imaging studies accompanied by radiological reports are accumulated and stored in many modern hospitals Picture Archiving and Communication Systems (PACS). On the other side, it is still an open question how this type of hospital-size knowledge database containing invaluable imaging informatics (i.e., loosely labeled) can be used to facilitate the data-hungry deep learning paradigms in building truly large-scale high precision computer-aided diagnosis (CAD) systems. In this paper, we present a new chest X-ray database, namely ChestX-ray8, which comprises 108,948 frontal-view X-ray images of 32,717 unique patients with the text-mined eight disease image labels (where each image can have multi-labels), from the associated radiological reports using natural language processing. Importantly, we demonstrate that these commonly occurring thoracic diseases can be detected and even spatially-located via a unified weakly-supervised multi-label image classification and disease localization framework, which is validated using our proposed dataset. Although the initial quantitative results are promising as reported, deep convolutional neural network based reading chest X-rays (i.e., recognizing and locating the common disease patterns trained with only image-level labels) remains a strenuous task for fully-automated high precision CAD systems.
Xiaosong Wang 0001, Yifan Peng 0002, Le Lu 0001, Zhiyong Lu, Mohammadhadi Bagheri, Ronald M. Summers
CVPR2
2015 miRTex: A Text Mining System for miRNA-Gene Relation Extraction
abstract
MicroRNAs (miRNAs) regulate a wide range of cellular and developmental processes through gene expression suppression or mRNA degradation. Experimentally validated miRNA gene targets are often reported in the literature. In this paper, we describe miRTex, a text mining system that extracts miRNA-target relations, as well as miRNA-gene and gene-miRNA regulation relations. The system achieves good precision and recall when evaluated on a literature corpus of 150 abstracts with F-scores close to 0.90 on the three different types of relations. We conducted full-scale text mining using miRTex to process all the Medline abstracts and all the full-length articles in the PubMed Central Open Access Subset. The results for all the Medline abstracts are stored in a database for interactive query and file download via the website at http://proteininformationresource.org/mirtex. Using miRTex, we identified genes potentially regulated by miRNAs in Triple Negative Breast Cancer, as well as miRNA-gene relations that, in conjunction with kinase-substrate relations, regulate the response to abiotic stress in Arabidopsis thaliana. These two use cases demonstrate the usefulness of miRTex text mining in the analysis of miRNA-regulated biological processes.
Gang Li 0021, Karen E. Ross, Cecilia N. Arighi, Yifan Peng 0002, Cathy H. Wu, K. Vijay-Shanker
PLoS Comput. Biol.4
2014 A Generalizable NLP Framework for Fast Development of Pattern-based Biomedical Relation Extraction Systems
abstract
BACKGROUND: Text mining is increasingly used in the biomedical domain because of its ability to automatically gather information from large amount of scientific articles. One important task in biomedical text mining is relation extraction, which aims to identify designated relations among biological entities reported in literature. A relation extraction system achieving high performance is expensive to develop because of the substantial time and effort required for its design and implementation. Here, we report a novel framework to facilitate the development of a pattern-based biomedical relation extraction system. It has several unique design features: (1) leveraging syntactic variations possible in a language and automatically generating extraction patterns in a systematic manner, (2) applying sentence simplification to improve the coverage of extraction patterns, and (3) identifying referential relations between a syntactic argument of a predicate and the actual target expected in the relation extraction task. RESULTS: A relation extraction system derived using the proposed framework achieved overall F-scores of 72.66% for the Simple events and 55.57% for the Binding events on the BioNLP-ST 2011 GE test set, comparing favorably with the top performing systems that participated in the BioNLP-ST 2011 GE task. We obtained similar results on the BioNLP-ST 2013 GE test set (80.07% and 60.58%, respectively). We conducted additional experiments on the training and development sets to provide a more detailed analysis of the system and its individual modules. This analysis indicates that without increasing the number of patterns, simplification and referential relation linking play a key role in the effective extraction of biomedical relations. CONCLUSIONS: In this paper, we present a novel framework for fast development of relation extraction systems. The framework requires only a list of triggers as input, and does not need information from an annotated corpus. Thus, we reduce the involvement of domain experts, who would otherwise have to provide manual annotations and help with the design of hand crafted patterns. We demonstrate how our framework is used to develop a system which achieves state-of-the-art performance on a public benchmark corpus.
Yifan Peng 0002, Manabu Torii, Cathy H. Wu, K. Vijay-Shanker
BMC Bioinform.1
2012 iSimp: A sentence simplification system for biomedicail text
abstract
Text mining applications using natural language processing are often confronted with long and complicated sentences. This is observed particularly in the abstracts of scientific articles where authors summarize, in few sentences, the various facts described throughout the manuscript. Being rich in novel and important information, the abstract has been the primary target of biomedicai text mining applications. In this work, we aim to simplify complex sentences in abstracts of biomedicai text so that they can be readily processed by text mining applications. We focus on syntactic constructs that are frequently encountered in the biomedicai literature, such as coordinations, relative clauses, and appositions, with emphasis on their boundary detection. Our approach yielded good detection performance (average F-measure between 86.5% and 92.7%), and aided in improving biomedicai text mining applications, RLIMS-P and RankPref.
Yifan Peng 0002, Catalina O. Tudor, Manabu Torii, Cathy H. Wu, K. Vijay-Shanker
BIBM1