Qingyu Chen 0001

dblp:28/5691-1 · DBLP profile ↗
← Back
33ranked-venue papers
9as first author
26since 2021 · last 2026
0000-0002-6036-1516ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 25 · 9 first-author · 19 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 MED-COREASONER: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning
abstract
While reasoning-enhanced large language models perform strongly on English medical tasks, a persistent multilingual gap remains, with substantially weaker reasoning in local languages, limiting equitable global medical deployment. To bridge this gap, we introduce Med-CoReasoner, a language-informed co-reasoning framework that elicits parallel English and local-language reasoning, abstracts them into structured concepts, and integrates local clinical knowledge into an English logical scaffold via concept-level alignment and retrieval. This design combines the structural robustness of English reasoning with the practice-grounded expertise encoded in local languages. To evaluate multilingual medical reasoning beyond multiple-choice settings, we construct MultiMed-X, a benchmark covering seven languages with expert-annotated long-form question answering and natural language inference tasks, comprising 350 instances per language. Experiments across three benchmarks show that Med-CoReasoner improves multilingual reasoning performance by an average of 5%, with particularly substantial gains in low-resource languages. Moreover, model distillation and expert evaluation analysis further confirm that Med-CoReasoner produces clinically sound and culturally grounded reasoning traces.
Sherry T. Tong, Jiwoong Sohn, Ding Xia, Piyalitt Ittichaiwong, Kanyakorn Veerakanjana, Hyunjae Kim, Qingyu Chen 0001, Edison Marrese-Taylor, Kazuma Kobayashi, Akiko Aizawa, Irene Li
ACL (1)10
2026 LMOD\(\boldsymbol{+}\): A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology
abstract
The rising prevalence of vision-threatening eye diseases poses a major global health and economic burden, yet timely diagnosis remains limited by workforce shortages, diagnostic delays, and restricted access to specialized care. Artificial intelligence (AI) offers potential solutions. In particular, recent progress in foundation models and large language models-especially multimodal large language models (MLLMs)-has shown promise in medical image interpretation and automated clinical documentation. However, advancing MLLMs for ophthalmology is hindered by the lack of unified, comprehensive benchmark datasets for development and evaluation. Most existing benchmarks were designed for earlier models, which focused on narrow tasks or specific disease conditions. These benchmarks typically provide outputs in the form of disease labels rather than free-text responses. As a result, they are less suitable for assessing emerging generative models. In this work, we present LMOD+, a large-scale multimodal ophthalmology benchmark dataset comprising 32,633 instances with multi-granular annotations across 12 common ophthalmic conditions and 5 imaging modalities. The dataset integrates imaging, anatomical structures, demographics, and free-text annotations. It supports primary ophthalmic applications such as anatomical structure recognition, disease screening, disease staging, and demographic prediction for potential performance bias evaluation. Alongside the dataset, we introduce a systematic and unified data curation pipeline that repurposes existing or new datasets for MLLM development. LMOD+ extends our preliminary LMOD benchmark-the first multimodal ophthalmology benchmark for MLLMs-with three major enhancements. First, we expanded the dataset by nearly 50% (from 21,933 to 32,633 instances). The color fundus photography (CFP) modality, the most accessible imaging modality in ophthalmology, was significantly enlarged to cover a broader range of pathological conditions. Second, we broadened task coverage to include (a) 12 binary disease diagnosis tasks for prevalent conditions such as diabetic retinopathy, age-related macular degeneration, and retinal vein occlusion; (b) multi-class ophthalmic disease diagnosis; (c) disease severity classification, including a diabetic retinopathy staging task, which uses two internationally adopted grading standards: the international clinical diabetic retinopathy classification and the Scottish diabetic retinopathy grading scheme classification; and (d) demographic prediction (age and sex) to assess potential model bias. Third, we systematically evaluated 24 state-of-the-art MLLMs, including recent models from the InternVL, Qwen, and DeepSeek families. Our evaluations highlight both the promise and limitations of current MLLMs in ophthalmology. For example, Qwen-7B and InternVL achieved accuracies of 58.26% and 57.83% in disease screening under a zero-shot setting with a single model-a considerably more challenging paradigm than traditional fine-tuning, where separate models are trained for each specific task. InternVL also demonstrated potential in anatomical recognition. Nonetheless, overall performance remained suboptimal and often close to random baselines for challenging tasks such as disease staging, underscoring the substantial gap between general-domain MLLMs and the specialized requirements of ophthalmology. We publicly release the dataset, curation pipeline, and leaderboard to encourage community-wide development and evaluation of MLLMs, with the goal of advancing ophthalmic applications and ultimately reducing the global burden of vision-threatening diseases through AI. The dataset website, benchmark leaderboard, and download link are available at https://kfzyqin.github.io/lmod_plus.
Zhenyue Qin, Yang Liu 0249, Jinyu Ding, Anran Li 0001, Dylan Campbell, Xuansheng Wu, Ke Zou, Tiarnan D. Keenan, Emily Y. Chew, Zhiyong Lu, Ninghao Liu 0001, Xiuzhen Zhang 0001, Qingyu Chen 0001
ACM Trans. Comput. Heal.16
2026 Information extraction from clinical notes: are we ready to switch to large language models?
abstract
OBJECTIVES: To assess the performance, generalizability, and computational efficiency of instruction-tuned Large Language Model Meta AI (LLaMA)-2 and LLaMA-3 models compared to bidirectional encoder representations from transformers (BERT) for clinical information extraction (IE) tasks, specifically named entity recognition (NER) and relation extraction (RE). MATERIALS AND METHODS: We developed a comprehensive annotated corpus of 1588 clinical notes from 4 data sources-UT Physicians (UTP) (1342 notes), Transcribed Medical Transcription Sample Reports and Examples (MTSamples) (146), Medical Information Mart for Intensive Care (MIMIC)-III (50), and Informatics for Integrating Biology and the Bedside (i2b2) (50), capturing 4 clinical entities (problems, tests, medications, other treatments) and 16 modifiers (eg, negation, certainty). Large Language Model Meta AI-2 and LLaMA-3 were instruction-tuned for clinical NER and RE, and their performance was benchmarked against BERT. RESULTS: Large Language Model Meta AI models consistently outperformed BERT across datasets. In data-rich settings (eg, UTP), LLaMA achieved marginal gains (approximately 1% improvement for NER and 1.5%-3.7% for RE). Under limited data conditions (eg, MTSamples, MIMIC-III) and on the unseen i2b2 dataset, LLaMA-3-70B improved F1 scores by over 7% for NER and 4% for RE. However, performance gains came with increased computational costs, with LLaMA models requiring more memory and Graphics Processing Unit (GPU) hours and running up to 28 times slower than BERT. DISCUSSION: While LLaMA models offer enhanced performance, their higher computational demands and slower throughput highlight the need to balance performance with practical resource constraints. Application-specific considerations are essential when choosing between LLMs and BERT for clinical IE. CONCLUSION: Instruction-tuned LLaMA models show promise for clinical NER and RE tasks. However, the tradeoff between improved performance and increased computational cost must be carefully evaluated. We release our Kiwi package (https://kiwi.clinicalnlp.org/) to facilitate the application of both LLaMA and BERT models in clinical IE applications.
Xu Zuo, Yujia Zhou 0003, Xueqing Peng, Jimin Huang, Vipina Kuttichi Keloth, Vincent J. Zhang, Ruey-Ling Weng, Cathy Shyr, Qingyu Chen 0001, Xiaoqian Jiang, Kirk Roberts, Hua Xu 0001
J. Am. Medical Informatics Assoc.10
2025 GraphCheck: Breaking Long-Term Text Barriers with Extracted Knowledge Graph-Powered Fact-Checking
abstract
, a fact-checking framework that uses extracted knowledge graphs to enhance text representation. Graph Neural Networks further process these graphs as a soft prompt, enabling LLMs to incorporate structured knowledge more effectively. Enhanced with graph-based reasoning, GraphCheck captures multihop reasoning chains that are often overlooked by existing methods, enabling precise and efficient fact-checking in a single inference call. Experimental results on seven benchmarks spanning both general and medical domains demonstrate up to a 7.1% overall improvement over baseline models. Notably, GraphCheck outperforms existing specialized fact-checkers and achieves comparable performance with state-of-the-art LLMs, such as DeepSeek-V3 and OpenAI-o1, with significantly fewer parameters.
Yingjian Chen, Yinhong Liu, Jinxiang Xie, Rui Yang 0016, Yanran Fu, Peng Yuan Zhou, Qingyu Chen 0001, James Caverlee, Irene Li
ACL (1)9
2025 MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
abstract
Weihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing, Junjue Wang, Fan Gao, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen, Douglas Teodoro, Nan Liu, Randy Goebel, Lei Ma, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Weihao Xuan, Rui Yang 0016, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing 0001, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li 0079, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen 0001, Douglas Teodoro, Nan Liu 0003, Randy Goebel, Lei Ma 0003, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li
EMNLP23
2025 Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewards
abstract
Jaehoon Yun, Jiwoong Sohn, Jungwoo Park, Hyunjae Kim, Xiangru Tang, Daniel Shao, Yong Hoe Koo, Ko Minhyeok, Qingyu Chen, Mark Gerstein, Michael Moor, Jaewoo Kang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jaehoon Yun, Jiwoong Sohn, Jungwoo Park, Hyunjae Kim, Xiangru Tang, Daniel Shao, Yonghoe Koo, Minhyeok Ko, Qingyu Chen 0001, Mark Gerstein, Michael Moor, Jaewoo Kang
EMNLP9
2025 Word-Sequence Entropy: Towards uncertainty estimation in free-form medical question answering applications and beyond
Zhiyuan Wang 0007, Jinhao Duan, Chenxi Yuan, Qingyu Chen 0001, Tianlong Chen 0001, Yue Zhang 0025, Ren Wang 0008, Xiaoshuang Shi, Kaidi Xu
Eng. Appl. Artif. Intell.4
2024 MedCalc-Bench: Evaluating Large Language Models for Medical Calculations
abstract
Current benchmarks for evaluating large language models (LLMs) in medicine are primarily focused on question-answering involving domain knowledge and descriptive reasoning. While such qualitative capabilities are vital to medical diagnosis, in real-world scenarios, doctors frequently use clinical calculators that follow quantitative equations and rule-based reasoning paradigms for evidence-based decision support. To this end, we propose MedCalc-Bench, a first-of-its-kind dataset focused on evaluating the medical calculation capability of LLMs. MedCalc-Bench contains an evaluation set of over 1000 manually reviewed instances from 55 different medical calculation tasks. Each instance in MedCalc-Bench consists of a patient note, a question requesting to compute a specific medical value, a ground truth answer, and a step-by-step explanation showing how the answer is obtained. While our evaluation results show the potential of LLMs in this area, none of them are effective enough for clinical settings. Common issues include extracting the incorrect entities, not using the correct equation or rules for a calculation task, or incorrectly performing the arithmetic for the computation. We hope our study highlights the quantitative knowledge and reasoning gaps in LLMs within medical settings, encouraging future improvements of LLMs for various clinical calculation tasks. MedCalc-Bench is publicly available at: https://github.com/ncbi-nlp/MedCalc-Bench.
Nikhil Khandekar, Qiao Jin 0001, Guangzhi Xiong, Soren Dunn, Serina S. Applebaum, Zain Anwar, Maame Sarfo-Gyamfi, Conrad W. Safranek, Abid A Anwar, Aidan Gilson, Maxwell B. Singer, Amisha D. Dave, Andrew Taylor, Aidong Zhang 0001, Qingyu Chen 0001, Zhiyong Lu
NeurIPS16
2024 Opportunities and challenges for ChatGPT and large language models in biomedicine and health
abstract
ChatGPT has drawn considerable attention from both the general public and domain experts with its remarkable text generation capabilities. This has subsequently led to the emergence of diverse applications in the field of biomedicine and health. In this work, we examine the diverse applications of large language models (LLMs), such as ChatGPT, in biomedicine and health. Specifically we explore the areas of biomedical information retrieval, question answering, medical text summarization, information extraction, and medical education, and investigate whether LLMs possess the transformative power to revolutionize these tasks or whether the distinct complexities of biomedical domain presents unique challenges. Following an extensive literature survey, we find that significant advances have been made in the field of text generation tasks, surpassing the previous state-of-the-art methods. For other applications, the advances have been modest. Overall, LLMs have not yet revolutionized biomedicine, but recent rapid progress indicates that such methods hold great potential to provide valuable means for accelerating discovery and improving health. We also find that the use of LLMs, like ChatGPT, in the fields of biomedicine and health entails various risks and challenges, including fabricated information in its generated responses, as well as legal and privacy concerns associated with sensitive patient data. We believe this survey can provide a comprehensive and timely overview to biomedical researchers and healthcare practitioners on the opportunities and challenges associated with using ChatGPT and other LLMs for transforming biomedicine and health.
Shubo Tian, Qiao Jin 0001, Lana Yeganova, Po-Ting Lai, Qingqing Zhu, Xiuying Chen, Yifan Yang 0006, Qingyu Chen 0001, Won Kim 0003, Donald C. Comeau, Rezarta Islamaj Dogan, Aadit Kapoor, Xin Gao 0001, Zhiyong Lu
Briefings Bioinform.8
2024 GeneGPT: augmenting large language models with domain tools for improved access to biomedical information
abstract
MOTIVATION: While large language models (LLMs) have been successfully applied to various tasks, they still face challenges with hallucinations. Augmenting LLMs with domain-specific tools such as database utilities can facilitate easier and more precise access to specialized knowledge. In this article, we present GeneGPT, a novel method for teaching LLMs to use the Web APIs of the National Center for Biotechnology Information (NCBI) for answering genomics questions. Specifically, we prompt Codex to solve the GeneTuring tests with NCBI Web APIs by in-context learning and an augmented decoding algorithm that can detect and execute API calls. RESULTS: Experimental results show that GeneGPT achieves state-of-the-art performance on eight tasks in the GeneTuring benchmark with an average score of 0.83, largely surpassing retrieval-augmented LLMs such as the new Bing (0.44), biomedical LLMs such as BioMedLM (0.08) and BioGPT (0.04), as well as GPT-3 (0.16) and ChatGPT (0.12). Our further analyses suggest that: First, API demonstrations have good cross-task generalizability and are more useful than documentations for in-context learning; second, GeneGPT can generalize to longer chains of API calls and answer multi-hop questions in GeneHop, a novel dataset introduced in this work; finally, different types of errors are enriched in different tasks, providing valuable insights for future improvements. AVAILABILITY AND IMPLEMENTATION: The GeneGPT code and data are publicly available at https://github.com/ncbi/GeneGPT.
Qiao Jin 0001, Yifan Yang 0006, Qingyu Chen 0001, Zhiyong Lu
Bioinform.3
2024 Advancing entity recognition in biomedicine via instruction tuning of large language models
abstract
MOTIVATION: Large Language Models (LLMs) have the potential to revolutionize the field of Natural Language Processing, excelling not only in text generation and reasoning tasks but also in their ability for zero/few-shot learning, swiftly adapting to new tasks with minimal fine-tuning. LLMs have also demonstrated great promise in biomedical and healthcare applications. However, when it comes to Named Entity Recognition (NER), particularly within the biomedical domain, LLMs fall short of the effectiveness exhibited by fine-tuned domain-specific models. One key reason is that NER is typically conceptualized as a sequence labeling task, whereas LLMs are optimized for text generation and reasoning tasks. RESULTS: We developed an instruction-based learning paradigm that transforms biomedical NER from a sequence labeling task into a generation task. This paradigm is end-to-end and streamlines the training and evaluation process by automatically repurposing pre-existing biomedical NER datasets. We further developed BioNER-LLaMA using the proposed paradigm with LLaMA-7B as the foundational LLM. We conducted extensive testing on BioNER-LLaMA across three widely recognized biomedical NER datasets, consisting of entities related to diseases, chemicals, and genes. The results revealed that BioNER-LLaMA consistently achieved higher F1-scores ranging from 5% to 30% compared to the few-shot learning capabilities of GPT-4 on datasets with different biomedical entities. We show that a general-domain LLM can match the performance of rigorously fine-tuned PubMedBERT models and PMC-LLaMA, biomedical-specific language model. Our findings underscore the potential of our proposed paradigm in developing general-domain LLMs that can rival SOTA performances in multi-task, multi-domain scenarios in biomedical and health applications. AVAILABILITY AND IMPLEMENTATION: Datasets and other resources are available at https://github.com/BIDS-Xu-Lab/BioNER-LLaMA.
Vipina Kuttichi Keloth, Qianqian Xie, Xueqing Peng, Yan Wang 0015, Andrew Zheng, Melih Selek, Kalpana Raja, Chih-Hsuan Wei, Qiao Jin 0001, Zhiyong Lu, Qingyu Chen 0001, Hua Xu 0001
Bioinform.12
2024 PubMed Computed Authors in 2024: an open resource of disambiguated author names in biomedical literature
abstract
SUMMARY: Over 55% of author names in PubMed are ambiguous: the same name is shared by different individual researchers. This poses significant challenges on precise literature retrieval for author name queries, a common behavior in biomedical literature search. In response, we present a comprehensive dataset of disambiguated authors. Specifically, we complement the automatic PubMed Computed Authors algorithm with the latest ORCID data for improved accuracy. As a result, the enhanced algorithm achieves high performance in author name disambiguation, and subsequently our dataset contains more than 21 million disambiguated authors for over 35 million PubMed articles and is incrementally updated on a weekly basis. More importantly, we make the dataset publicly available for the community such that it can be utilized in a wide variety of potential applications beyond assisting PubMed's author name queries. Finally, we propose a set of guidelines for best practices of authors pertaining to use of their names. AVAILABILITY AND IMPLEMENTATION: The PubMed Computed Authors dataset is publicly available for bulk download at: https://ftp.ncbi.nlm.nih.gov/pub/lu/ComputedAuthors/. Additionally, it is available for query through web API at: https://www.ncbi.nlm.nih.gov/research/bionlp/APIs/authors/.
Shubo Tian, Qingyu Chen 0001, Donald C. Comeau, W. John Wilbur, Zhiyong Lu
Bioinform.2
2024 Improving large language models for clinical named entity recognition via prompt engineering
abstract
IMPORTANCE: The study highlights the potential of large language models, specifically GPT-3.5 and GPT-4, in processing complex clinical data and extracting meaningful information with minimal training data. By developing and refining prompt-based strategies, we can significantly enhance the models' performance, making them viable tools for clinical NER tasks and possibly reducing the reliance on extensive annotated datasets. OBJECTIVES: This study quantifies the capabilities of GPT-3.5 and GPT-4 for clinical named entity recognition (NER) tasks and proposes task-specific prompts to improve their performance. MATERIALS AND METHODS: We evaluated these models on 2 clinical NER tasks: (1) to extract medical problems, treatments, and tests from clinical notes in the MTSamples corpus, following the 2010 i2b2 concept extraction shared task, and (2) to identify nervous system disorder-related adverse events from safety reports in the vaccine adverse event reporting system (VAERS). To improve the GPT models' performance, we developed a clinical task-specific prompt framework that includes (1) baseline prompts with task description and format specification, (2) annotation guideline-based prompts, (3) error analysis-based instructions, and (4) annotated samples for few-shot learning. We assessed each prompt's effectiveness and compared the models to BioClinicalBERT. RESULTS: Using baseline prompts, GPT-3.5 and GPT-4 achieved relaxed F1 scores of 0.634, 0.804 for MTSamples and 0.301, 0.593 for VAERS. Additional prompt components consistently improved model performance. When all 4 components were used, GPT-3.5 and GPT-4 achieved relaxed F1 socres of 0.794, 0.861 for MTSamples and 0.676, 0.736 for VAERS, demonstrating the effectiveness of our prompt framework. Although these results trail BioClinicalBERT (F1 of 0.901 for the MTSamples dataset and 0.802 for the VAERS), it is very promising considering few training samples are needed. DISCUSSION: The study's findings suggest a promising direction in leveraging LLMs for clinical NER tasks. However, while the performance of GPT models improved with task-specific prompts, there's a need for further development and refinement. LLMs like GPT-4 show potential in achieving close performance to state-of-the-art models like BioClinicalBERT, but they still require careful prompt engineering and understanding of task-specific knowledge. The study also underscores the importance of evaluation schemas that accurately reflect the capabilities and performance of LLMs in clinical settings. CONCLUSION: While direct application of GPT models to clinical NER tasks falls short of optimal performance, our task-specific prompt framework, incorporating medical knowledge and training samples, significantly enhances GPT models' feasibility for potential clinical applications.
Qingyu Chen 0001, Jingcheng Du, Xueqing Peng, Vipina Kuttichi Keloth, Xu Zuo, Yujia Zhou 0003, Zehan Li, Xiaoqian Jiang, Zhiyong Lu, Kirk Roberts, Hua Xu 0001
J. Am. Medical Informatics Assoc.2
2024 Augmenting biomedical named entity recognition with general-domain resources
abstract
OBJECTIVE: Training a neural network-based biomedical named entity recognition (BioNER) model usually requires extensive and costly human annotations. While several studies have employed multi-task learning with multiple BioNER datasets to reduce human effort, this approach does not consistently yield performance improvements and may introduce label ambiguity in different biomedical corpora. We aim to tackle those challenges through transfer learning from easily accessible resources with fewer concept overlaps with biomedical datasets. METHODS: We proposed GERBERA, a simple-yet-effective method that utilized general-domain NER datasets for training. We performed multi-task learning to train a pre-trained biomedical language model with both the target BioNER dataset and the general-domain dataset. Subsequently, we fine-tuned the models specifically for the BioNER dataset. RESULTS: We systematically evaluated GERBERA on five datasets of eight entity types, collectively consisting of 81,410 instances. Despite using fewer biomedical resources, our models demonstrated superior performance compared to baseline models trained with additional BioNER datasets. Specifically, our models consistently outperformed the baseline models in six out of eight entity types, achieving an average improvement of 0.9% over the best baseline performance across eight entities. Our method was especially effective in amplifying performance on BioNER datasets characterized by limited data, with a 4.7% improvement in F1 scores on the JNLPBA-RNA dataset. CONCLUSION: This study introduces a new training method that leverages cost-effective general-domain NER datasets to augment BioNER models. This approach significantly improves BioNER model performance, making it a valuable asset for scenarios with scarce or costly biomedical datasets. We make data, codes, and models publicly available via https://github.com/qingyu-qc/bioner_gerbera.
Hyunjae Kim, Chih-Hsuan Wei, Jaewoo Kang, Zhiyong Lu, Hua Xu 0001, Qingyu Chen 0001
J. Biomed. Informatics9
2023 MedCPT: Contrastive Pre-trained Transformers with large-scale PubMed search logs for zero-shot biomedical information retrieval
abstract
MOTIVATION: Information retrieval (IR) is essential in biomedical knowledge acquisition and clinical decision support. While recent progress has shown that language model encoders perform better semantic retrieval, training such models requires abundant query-article annotations that are difficult to obtain in biomedicine. As a result, most biomedical IR systems only conduct lexical matching. In response, we introduce MedCPT, a first-of-its-kind Contrastively Pre-trained Transformer model for zero-shot semantic IR in biomedicine. RESULTS: To train MedCPT, we collected an unprecedented scale of 255 million user click logs from PubMed. With such data, we use contrastive learning to train a pair of closely integrated retriever and re-ranker. Experimental results show that MedCPT sets new state-of-the-art performance on six biomedical IR tasks, outperforming various baselines including much larger models, such as GPT-3-sized cpt-text-XL. In addition, MedCPT also generates better biomedical article and sentence representations for semantic evaluations. As such, MedCPT can be readily applied to various real-world biomedical IR tasks. AVAILABILITY AND IMPLEMENTATION: The MedCPT code and model are available at https://github.com/ncbi/MedCPT.
Qiao Jin 0001, Won Kim 0003, Qingyu Chen 0001, Donald C. Comeau, Lana Yeganova, W. John Wilbur, Zhiyong Lu
Bioinform.3
2023 AIONER: all-in-one scheme-based biomedical named entity recognition using deep learning
abstract
MOTIVATION: Biomedical named entity recognition (BioNER) seeks to automatically recognize biomedical entities in natural language text, serving as a necessary foundation for downstream text mining tasks and applications such as information extraction and question answering. Manually labeling training data for the BioNER task is costly, however, due to the significant domain expertise required for accurate annotation. The resulting data scarcity causes current BioNER approaches to be prone to overfitting, to suffer from limited generalizability, and to address a single entity type at a time (e.g. gene or disease). RESULTS: We therefore propose a novel all-in-one (AIO) scheme that uses external data from existing annotated resources to enhance the accuracy and stability of BioNER models. We further present AIONER, a general-purpose BioNER tool based on cutting-edge deep learning and our AIO schema. We evaluate AIONER on 14 BioNER benchmark tasks and show that AIONER is effective, robust, and compares favorably to other state-of-the-art approaches such as multi-task learning. We further demonstrate the practical utility of AIONER in three independent tasks to recognize entity types not previously seen in training data, as well as the advantages of AIONER over existing methods for processing biomedical text at a large scale (e.g. the entire PubMed data). AVAILABILITY AND IMPLEMENTATION: The source code, trained models and data for AIONER are freely available at https://github.com/ncbi/AIONER.
Ling Luo 0001, Chih-Hsuan Wei, Po-Ting Lai, Robert Leaman, Qingyu Chen 0001, Zhiyong Lu
Bioinform.5
2023 BioREx: Improving biomedical relation extraction by leveraging heterogeneous datasets
Po-Ting Lai, Chih-Hsuan Wei, Ling Luo 0001, Qingyu Chen 0001, Zhiyong Lu
J. Biomed. Informatics4
2022 Automated and Accessible Diagnosis of Age-related Macular Degeneration: a Comparative Analysis of the impact of machine learning models in clinical diagnostic Workflows
Qingyu Chen 0001, Tiarnan D. Keenan, Alexis Allot, Sanjeeb Bhandari, Geoff Broadhead, Chantal Cousineau-Krieger, Ellen Davis, William G. Gensheimer, David Grasic, Seema Gupta, Eleni Konstantinou, Tania Lamba, Michele Maiberger, Arnold Oshinsky, Brittany E. Powell, Boonkit Purt, Soo Shin, Hillary Steifel, Alisa T. Thavikulwat, Keith Wroblewski, Sirisha Koirala, Tom Murickan, Michael F. Chiang, Michelle R. Hribar, Emily Y. Chew, Zhiyong Lu
AMIA1
2022 Deep learning automated diagnosis and quantitative classification of cataract type and severity: quantifying the effectiveness and usability of deep learning-assisted disease diagnosis models with 14 ophthalmologists and multi-center validations
Qingyu Chen 0001, Tiarnan D. Keenan, Elvira Agrón, Amr S. Elsawy, Emily Y. Chew, Zhiyong Lu
AMIA1
2022 Robust convolutional neural networks against adversarial attacks on medical images
abstract
Convolutional neural networks (CNNs) have been widely applied to medical images. However, medical images are vulnerable to adversarial attacks by perturbations that are undetectable to human experts. This poses significant security risks and challenges to CNN-based applications in clinic practice. In this work, we quantify the scale of adversarial perturbation imperceptible to clinical practitioners and investigate the cause of the vulnerability in CNNs. Specifically, we discover that noise (i.e., irrelevant or corrupted discriminative information) in medical images might be a key contributor to performance deterioration of CNNs against adversarial perturbations, as noisy features are learned unconsciously by CNNs in feature representations and magnified by adversarial perturbations. In response, we propose a novel defense method by embedding sparsity denoising operators in CNNs for improved robustness. Tested with various state-of-the-art attacking methods on two distinct medical image modalities, we demonstrate that the proposed method can successfully defend against those unnoticeable adversarial attacks by retaining as much as over 90% of its original performance. We believe our findings are critical for improving and deploying CNN-based medical applications in real-world scenarios.
Xiaoshuang Shi, Yifan Peng 0002, Qingyu Chen 0001, Tiarnan D. Keenan, Alisa T. Thavikulwat, Sungwon Lee 0003, Yuxing Tang, Emily Y. Chew, Ronald M. Summers, Zhiyong Lu
Pattern Recognit.3
2022 LitMC-BERT: Transformer-Based Multi-Label Classification of Biomedical Literature With An Application on COVID-19 Literature Curation
abstract
The rapid growth of biomedical literature poses a significant challenge for curation and interpretation. This has become more evident during the COVID-19 pandemic. LitCovid, a literature database of COVID-19 related papers in PubMed, has accumulated over 200,000 articles with millions of accesses. Approximately 10,000 new articles are added to LitCovid every month. A main curation task in LitCovid is topic annotation where an article is assigned with up to eight topics, e.g., Treatment and Diagnosis. The annotated topics have been widely used both in LitCovid (e.g., accounting for ∼18% of total uses) and downstream studies such as network generation. However, it has been a primary curation bottleneck due to the nature of the task and the rapid literature growth. This study proposes LITMC-BERT, a transformer-based multi-label classification method in biomedical literature. It uses a shared transformer backbone for all the labels while also captures label-specific features and the correlations between label pairs. We compare LITMC-BERT with three baseline models on two datasets. Its micro-F1 and instance-based F1 are 5% and 4% higher than the current best results, respectively, and only requires ∼18% of the inference time than the Binary BERT baseline. The related datasets and models are available via https://github.com/ncbi/ml-transformer.
Qingyu Chen 0001, Jingcheng Du, Alexis Allot, Zhiyong Lu
IEEE ACM Trans. Comput. Biol. Bioinform.1
2021 AM2BERT: attention guided and regularized transformer-based multi-label classification model for COVID-19 literature curation
Qingyu Chen 0001, Jingcheng Du, Alexis Allot, Zhiyong Lu
AMIA1
2021 Deep learning detection of reticular pseudodrusen using multi-modal, multi-task, and multi-attention mechanisms: towards automated and accessible classification of age-related macular degeneration
Qingyu Chen 0001, Tiarnan D. Keenan, Emily Y. Chew, Zhiyong Lu
AMIA1
2021 Multi-task deep learning-based survival analysis on the prognosis of late AMD using the longitudinal data in AREDS
Gregory C. Ghahramani, Matthew Brendel, Mingquan Lin, Qingyu Chen 0001, Tiarnan D. Keenan, Kun Chen 0002, Emily Y. Chew, Zhiyong Lu, Yifan Peng 0002, Fei Wang 0001
AMIA4
2021 Long Covid: A Comprehensive Collection of Articles Regarding Long-Haul Symptoms in COVID-19 Survivors
Robert Leaman, Qingyu Chen 0001, Alexis Allot, Zhiyong Lu
AMIA2
2021 Multimodal, multitask, multiattention (M3) deep learning detection of reticular pseudodrusen: Toward automated and accessible classification of age-related macular degeneration
abstract
OBJECTIVE: Reticular pseudodrusen (RPD), a key feature of age-related macular degeneration (AMD), are poorly detected by human experts on standard color fundus photography (CFP) and typically require advanced imaging modalities such as fundus autofluorescence (FAF). The objective was to develop and evaluate the performance of a novel multimodal, multitask, multiattention (M3) deep learning framework on RPD detection. MATERIALS AND METHODS: A deep learning framework (M3) was developed to detect RPD presence accurately using CFP alone, FAF alone, or both, employing >8000 CFP-FAF image pairs obtained prospectively (Age-Related Eye Disease Study 2). The M3 framework includes multimodal (detection from single or multiple image modalities), multitask (training different tasks simultaneously to improve generalizability), and multiattention (improving ensembled feature representation) operation. Performance on RPD detection was compared with state-of-the-art deep learning models and 13 ophthalmologists; performance on detection of 2 other AMD features (geographic atrophy and pigmentary abnormalities) was also evaluated. RESULTS: For RPD detection, M3 achieved an area under the receiver-operating characteristic curve (AUROC) of 0.832, 0.931, and 0.933 for CFP alone, FAF alone, and both, respectively. M3 performance on CFP was very substantially superior to human retinal specialists (median F1 score = 0.644 vs 0.350). External validation (the Rotterdam Study) demonstrated high accuracy on CFP alone (AUROC, 0.965). The M3 framework also accurately detected geographic atrophy and pigmentary abnormalities (AUROC, 0.909 and 0.912, respectively), demonstrating its generalizability. CONCLUSIONS: This study demonstrates the successful development, robust evaluation, and external validation of a novel deep learning framework that enables accessible, accurate, and automated AMD diagnosis and prognosis.
Qingyu Chen 0001, Tiarnan D. Keenan, Alexis Allot, Yifan Peng 0002, Elvira Agrón, Amitha Domalpally, Caroline C. W. Klaver, Daniel T. Luttikhuizen, Marcus H. Colyer, Catherine Cukras, Henry E. Wiley, M. Teresa Magone, Chantal Cousineau-Krieger, Wai T. Wong, Yingying Zhu 0003, Emily Y. Chew, Zhiyong Lu
J. Am. Medical Informatics Assoc.1
2020 Detection of reticular pseudodrusen using deep learning
Qingyu Chen 0001, Tiarnan D. Keenan, Yifan Peng 0002, Elvira Agrón, Christopher Hwang, Alisa T. Thavikulwat, Debora Lee, Wai T. Wong, Emily Y. Chew, Zhiyong Lu
AMIA1
2020 Privacy concerns of the Australian My Health Record: Implications for other large-scale opt-out personal health records
Patrick Pang 0001, Dana McKay, Shanton Chang, Qingyu Chen 0001, Xiuzhen Zhang 0001, Lishan Cui
Inf. Process. Manag.4
2020 Better synonyms for enriching biomedical search
abstract
OBJECTIVE: In a biomedical literature search, the link between a query and a document is often not established, because they use different terms to refer to the same concept. Distributional word embeddings are frequently used for detecting related words by computing the cosine similarity between them. However, previous research has not established either the best embedding methods for detecting synonyms among related word pairs or how effective such methods may be. MATERIALS AND METHODS: In this study, we first create the BioSearchSyn set, a manually annotated set of synonyms, to assess and compare 3 widely used word-embedding methods (word2vec, fastText, and GloVe) in their ability to detect synonyms among related pairs of words. We demonstrate the shortcomings of the cosine similarity score between word embeddings for this task: the same scores have very different meanings for the different methods. To address the problem, we propose utilizing pool adjacent violators (PAV), an isotonic regression algorithm, to transform a cosine similarity into a probability of 2 words being synonyms. RESULTS: Experimental results using the BioSearchSyn set as a gold standard reveal which embedding methods have the best performance in identifying synonym pairs. The BioSearchSyn set also allows converting cosine similarity scores into probabilities, which provides a uniform interpretation of the synonymy score over different methods. CONCLUSIONS: We introduced the BioSearchSyn corpus of 1000 term pairs, which allowed us to identify the best embedding method for detecting synonymy for biomedical search. Using the proposed method, we created PubTermVariants2.0: a large, automatically extracted set of synonym pairs that have augmented PubMed searches since the spring of 2019.
Lana Yeganova, Sun Kim, Qingyu Chen 0001, Grigory Balasanov, W. John Wilbur, Zhiyong Lu
J. Am. Medical Informatics Assoc.3
2020 BioConceptVec: Creating and evaluating literature-based biomedical concept embeddings on a large scale
abstract
A massive number of biological entities, such as genes and mutations, are mentioned in the biomedical literature. The capturing of the semantic relatedness of biological entities is vital to many biological applications, such as protein-protein interaction prediction and literature-based discovery. Concept embeddings-which involve the learning of vector representations of concepts using machine learning models-have been employed to capture the semantics of concepts. To develop concept embeddings, named-entity recognition (NER) tools are first used to identify and normalize concepts from the literature, and then different machine learning models are used to train the embeddings. Despite multiple attempts, existing biomedical concept embeddings generally suffer from suboptimal NER tools, small-scale evaluation, and limited availability. In response, we employed high-performance machine learning-based NER tools for concept recognition and trained our concept embeddings, BioConceptVec, via four different machine learning models on ~30 million PubMed abstracts. BioConceptVec covers over 400,000 biomedical concepts mentioned in the literature and is of the largest among the publicly available biomedical concept embeddings to date. To evaluate the validity and utility of BioConceptVec, we respectively performed two intrinsic evaluations (identifying related concepts based on drug-gene and gene-gene interactions) and two extrinsic evaluations (protein-protein interaction prediction and drug-drug interaction extraction), collectively using over 25 million instances from nine independent datasets (17 million instances from six intrinsic evaluation tasks and 8 million instances from three extrinsic evaluation tasks), which is, by far, the most comprehensive to our best knowledge. The intrinsic evaluation results demonstrate that BioConceptVec consistently has, by a large margin, better performance than existing concept embeddings in identifying similar and related concepts. More importantly, the extrinsic evaluation results demonstrate that using BioConceptVec with advanced deep learning models can significantly improve performance in downstream bioinformatics studies and biomedical text-mining applications. Our BioConceptVec embeddings and benchmarking datasets are publicly available at https://github.com/ncbi-nlp/BioConceptVec.
Qingyu Chen 0001, Kyubum Lee, Shankai Yan, Sun Kim, Chih-Hsuan Wei, Zhiyong Lu
PLoS Comput. Biol.1
2019 A deep learning-based survival model for prediction of progression in late Age-related Macular Degeneration (AMD) from color fundus photographs
Yifan Peng 0002, Tiarnan D. Keenan, Qingyu Chen 0001, Elvira Agrón, Wai T. Wong, Emily Y. Chew, Zhiyong Lu
AMIA3
2019 ML-Net: multi-label classification of biomedical texts with deep neural networks
abstract
OBJECTIVE: In multi-label text classification, each textual document is assigned 1 or more labels. As an important task that has broad applications in biomedicine, a number of different computational methods have been proposed. Many of these methods, however, have only modest accuracy or efficiency and limited success in practical use. We propose ML-Net, a novel end-to-end deep learning framework, for multi-label classification of biomedical texts. MATERIALS AND METHODS: ML-Net combines a label prediction network with an automated label count prediction mechanism to provide an optimal set of labels. This is accomplished by leveraging both the predicted confidence score of each label and the deep contextual information (modeled by ELMo) in the target document. We evaluate ML-Net on 3 independent corpora in 2 text genres: biomedical literature and clinical notes. For evaluation, we use example-based measures, such as precision, recall, and the F measure. We also compare ML-Net with several competitive machine learning and deep learning baseline models. RESULTS: Our benchmarking results show that ML-Net compares favorably to state-of-the-art methods in multi-label classification of biomedical text. ML-Net is also shown to be robust when evaluated on different text genres in biomedicine. CONCLUSION: ML-Net is able to accuractely represent biomedical document context and dynamically estimate the label count in a more systematic and accurate manner. Unlike traditional machine learning methods, ML-Net does not require human effort for feature engineering and is a highly efficient and scalable approach to tasks with a large set of labels, so there is no need to build individual classifiers for each separate label.
Jingcheng Du, Qingyu Chen 0001, Yifan Peng 0002, Yang Xiang 0003, Cui Tao, Zhiyong Lu
J. Am. Medical Informatics Assoc.2
2016 Evaluation of CD-HIT for constructing non-redundant databases
abstract
CD-HIT is one of the most popular tools for reducing sequence redundancy, and is considered to be the state-of-art method. It tries to minimise redundancy by reducing an input database into several representative sequences, under a user-defined threshold of sequence identity. We present a comprehensive assessment of the redundancy in the outputs of CD-HIT, exploring the impact of different identity thresholds and new evaluation data on the redundancy. We demonstrate that the relationship between threshold and redundancies is surprising weak. Applications of CD-HIT that set low identity threshold values also may suffer from substantial degradation in both efficiency and accuracy.
Qingyu Chen 0001, Yu Wan 0001, Yang Lei 0003, Justin Zobel, Karin Verspoor
BIBM1