EDBT 2026 Demo / reviewers in the wild / expert
Shubo Tian
dblp:298/6926
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0001-6415-1439ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing biomedical relation extraction with directionalityabstractSUMMARY: Biological relation networks contain rich information for understanding the biological mechanisms behind the relationship of entities such as genes, proteins, diseases, and chemicals. The vast growth of biomedical literature poses significant challenges in updating the network knowledge. The recent Biomedical Relation Extraction Dataset (BioRED) provides valuable manual annotations, facilitating the development of machine learning and pre-trained language model approaches for automatically identifying novel document-level (inter-sentence context) relationships. Nonetheless, its annotations lack directionality (subject/object) for the entity roles, which is essential for studying complex biological networks. Herein, we annotate the entity roles of the relationships in the BioRED corpus and subsequently propose a novel multi-task language model with soft-prompt learning to jointly identify the relationship, novel findings, and entity roles. Our results include an enriched BioRED corpus with 10 864 directionality annotations. Moreover, our proposed method outperforms existing large language models, such as the state-of-the-art GPT-4 and Llama-3, on two benchmarking tasks. AVAILABILITY AND IMPLEMENTATION: Our source code and dataset are available at https://github.com/ncbi-nlp/BioREDirect. Po-Ting Lai, Chih-Hsuan Wei, Shubo Tian, Robert Leaman, Zhiyong Lu |
Bioinform. | 3 |
| 2024 | Response to Letter to Editor 'Timely need for navigating the potential and downsides of LLMs in healthcare and biomedicine'abstractWe thank the author of [1] for commending our article, ‘Opportunities and Challenges for ChatGPT and Large Language Models in Biomedicine and Health’ [2]. We appreciate the suggested addition of more LLMs and the discussion on emerging challenges and potential applications. However, most of the pre-trained LLMs listed in Table 1 of [1] fall outside the scope of our review, which is focused on generative AI. Only the last two models—Medical-mT5 and BioMistral—are relevant, but they were released after our article was published [3, 4]. In our publication, we have included extensive discussions on the limitations and challenges of LLMs pertinent to the biomedical and health domains [2]. The ‘new challenges’ identified in Table 2 of [1] appear more akin to potential applications rather than challenges. As such, what was outlined there is beyond the scope of our review. They are rather subjects of current and future exploration by the research community. This research is supported by the NIH Intramural Research Program, National Library of Medicine. Shubo Tian, Qiao Jin 0001, Zhiyong Lu |
Briefings Bioinform. | 1 |
| 2024 | Opportunities and challenges for ChatGPT and large language models in biomedicine and healthabstractChatGPT has drawn considerable attention from both the general public and domain experts with its remarkable text generation capabilities. This has subsequently led to the emergence of diverse applications in the field of biomedicine and health. In this work, we examine the diverse applications of large language models (LLMs), such as ChatGPT, in biomedicine and health. Specifically we explore the areas of biomedical information retrieval, question answering, medical text summarization, information extraction, and medical education, and investigate whether LLMs possess the transformative power to revolutionize these tasks or whether the distinct complexities of biomedical domain presents unique challenges. Following an extensive literature survey, we find that significant advances have been made in the field of text generation tasks, surpassing the previous state-of-the-art methods. For other applications, the advances have been modest. Overall, LLMs have not yet revolutionized biomedicine, but recent rapid progress indicates that such methods hold great potential to provide valuable means for accelerating discovery and improving health. We also find that the use of LLMs, like ChatGPT, in the fields of biomedicine and health entails various risks and challenges, including fabricated information in its generated responses, as well as legal and privacy concerns associated with sensitive patient data. We believe this survey can provide a comprehensive and timely overview to biomedical researchers and healthcare practitioners on the opportunities and challenges associated with using ChatGPT and other LLMs for transforming biomedicine and health. Shubo Tian, Qiao Jin 0001, Lana Yeganova, Po-Ting Lai, Qingqing Zhu, Xiuying Chen, Yifan Yang 0006, Qingyu Chen 0001, Won Kim 0003, Donald C. Comeau, Rezarta Islamaj Dogan, Aadit Kapoor, Xin Gao 0001, Zhiyong Lu |
Briefings Bioinform. | 1 |
| 2024 | PubMed Computed Authors in 2024: an open resource of disambiguated author names in biomedical literatureabstractSUMMARY: Over 55% of author names in PubMed are ambiguous: the same name is shared by different individual researchers. This poses significant challenges on precise literature retrieval for author name queries, a common behavior in biomedical literature search. In response, we present a comprehensive dataset of disambiguated authors. Specifically, we complement the automatic PubMed Computed Authors algorithm with the latest ORCID data for improved accuracy. As a result, the enhanced algorithm achieves high performance in author name disambiguation, and subsequently our dataset contains more than 21 million disambiguated authors for over 35 million PubMed articles and is incrementally updated on a weekly basis. More importantly, we make the dataset publicly available for the community such that it can be utilized in a wide variety of potential applications beyond assisting PubMed's author name queries. Finally, we propose a set of guidelines for best practices of authors pertaining to use of their names. AVAILABILITY AND IMPLEMENTATION: The PubMed Computed Authors dataset is publicly available for bulk download at: https://ftp.ncbi.nlm.nih.gov/pub/lu/ComputedAuthors/. Additionally, it is available for query through web API at: https://www.ncbi.nlm.nih.gov/research/bionlp/APIs/authors/. Shubo Tian, Qingyu Chen 0001, Donald C. Comeau, W. John Wilbur, Zhiyong Lu |
Bioinform. | 1 |
| 2024 | Large language models in biomedicine and health: current research landscape and future directionsabstractLarge language models in biomedicine and health: current research landscape and future directionsLarge language models (LLMs) are a specialized type of generative artificial intelligence (AI) focused on generating natural language text.These models are developed through extensive training on massive amounts of text data and use deep learning algorithms to generate new text that closely resembles human-generated text.Generative AI methods, including LLMs, are rapidly transforming various domains, including biomedicine and healthcare.[1][2][3][4][5][6] They have already demonstrated remarkable potential as a means to process and analyze large amounts of text, interpret natural language, and generate new content in these domains.For example, Nori et al reported that GPT-4 is able to correctly answer the majority of questions from medical practice licensing exams, comfortably obtaining a passing grade.7 Similarly, Stribling et al found that this model exceeded the average performance of students in the graduate medical sciences on the majority of examinations, including strong performance on short answer and essay questions.8 Even though passing the exam is not the same as applying the knowledge in a real-world setting, these results demonstrate that LLMs can generate appropriate multiple-choice and narrative responses to questions framed in natural language.ChatGPT, first released in November 2022, has garnered phenomenal attention from both the scientific community and a broader society.A keyword search of "large language models" OR "ChatGPT" in PubMed returned over 4500 articles that discuss the technology and its implications for various topics, including medical informatics, by the end of June 2024.In addition, LLM-based technologies have already been deployed in several healthcare systems and are offered as integrated products for use in the clinic within vendor electronic health record systems (for thoughts on initial evaluations of an early product, see Garcia et al 9 and Tai-Seale et al 10 ).This rapid adoption of LLMs like ChatGPT brings an unprecedented opportunity to use this novel AI technology to transform healthcare and medicine.Despite their potential benefits, LLMs can sometimes produce invalid and unsubstantiated responses, a phenomenon known as the "hallucination and confabulation issue" in the literature, or biased responses, due to the biases inherent in their training data.[11][12][13][14][15][16][17] With this great potential also comes the need for trustworthy and responsible development and use of technology.As we continue to explore the capabilities of ChatGPT and other LLMs, it is critical to address related ethical, legal, and social issues to ensure that the technology is used in ways that are safe, fair, trustworthy, and beneficial for all.In the context of biomedicine and healthcare, it is particularly important to engage stakeholders, such as AI researchers, developers of data-driven clinical decision support, care providers, and system implementers from both academic medical centers and industry, to ensure responsible use of LLMs for good.To accelerate research and development in this area, we issued a call for submissions in Summer 2023, specifically focusing on the intersection of biomedicine/health and LLMs, and invited contributions on all related aspects.We invited submissions that report on innovative informatics methods development and evaluation, as well as studies that demonstrate the effectiveness/limitations of LLMs methodologies in healthcare.We particularly encouraged submissions that address the challenges and opportunities of this intersection and offer new insights into how these fields can work together to advance healthcare.This editorial provides an overview of the papers accepted in this Focus Issue.We highlight major themes and unique aspects of the research papers in medical LLMs, discuss ongoing challenges, and recommend future research directions.Box 1 lists the relevant large language model terms and abbreviations used in this editorial. Overall statistics of the Focus IssueThis JAMIA Focus Issue on LLMs in biomedicine and health has drawn enthusiasm from many researchers across different research disciplines.In total, we received over 150 submissions from authors in 25 countries and regions across 6 continents worldwide.The rigorous JAMIA peer review process was applied to all submissions, 41 of which were ultimately accepted for publication in the Focus Issue (Table 1).The majority of the accepted papers were authored by authors in North America, followed by those in Asia and Europe (Figure 1A).The Focus Issue highlights the nature of multi-disciplinary collaboration in medical informatics research across the broad JAMIA community.The number of authors per paper varies from 1 to 23, with an average of 7.3.Many papers feature authors with diverse expertise from different departments and organizations.The authors' expertise spans a wide range of fields, including computer science, data science, informatics, statistics, medicine, nursing, clinical services, public health policies, and more.Several papers also demonstrate scientific collaborations across different sectors, including academia, government labs, research institutes, hospitals, and industry.Additionally, a few papers showcase international collaborations among authors. Zhiyong Lu, Yifan Peng 0002, Trevor Cohen, Marzyeh Ghassemi, Chunhua Weng, Shubo Tian |
J. Am. Medical Informatics Assoc. | 6 |
| 2023 | Parsing Clinical Trial Eligibility Criteria for Cohort Query by a Multi-Input Multi-Output Sequence Labeling ModelabstractTo enable electronic screening of eligible patients for clinical trials, free-text clinical trial eligibility criteria should be translated to a computable format. Natural language processing (NLP) techniques have the potential to automate this process. In this study, we explored a supervised multi-input multi-output (MIMO) sequence labelling model to parse eligibility criteria into combinations of fact and condition tuples. Our experiments on a small manually annotated training dataset showed that that the performance of the MIMO framework with a BERT-based encoder using all the input sequences achieved an overall lenient-level AUROC of 0.61. Although the performance is suboptimal, representing eligibility criteria into logical and semantically clear tuples can potentially make subsequent translation of these tuples into database queries more reliable. Shubo Tian, Pengfei Yin, Hansi Zhang, Arslan Erdengasileng, Jiang Bian 0001, Zhe He 0001 |
BIBM | 1 |
| 2022 | Using Twitter Data Analysis to Understand the Perceptions, Awareness, and Barriers to the Wide Use of Pre-Exposure Prophylaxis in the United StatesabstractUser-generated social media posts such as tweets can provide insights about the public's perception, cognitive, and behavioral responses to health-related issues. Pre-Exposure Prophylaxis (PrEP) is one of the most effective ways to reduce the risk of HIV infection. However, its utilization is low in the US, especially among populations disproportionately affected by HIV such as the age group of under 24 years old. It is therefore important to understand the barriers to the wider use of PrEP in the US using social media posts. In this study, we collected tweets from Twitter about PrEP in the past 4 years to identify such barriers by first identifying tweets about personal discussions, and then performing textual analysis using word analysis, UMLS semantic type analysis, and topic modeling. We found that the public often discussed advocacy, risks/benefits, access, pricing, insurance coverage, legislation, stigma, health education, and prevention of HIV. This result is consistent with the literature and can help identify strategies for promoting the use of PrEP, especially among young adults. Arslan Erdengasileng, Shubo Tian, Sara S. Green, Sylvie Naar, Zhe He 0001 |
BIBM | 2 |
| 2022 | A Machine-Learning Based Approach for Predicting Older Adults' Adherence to Technology-Based Cognitive Training
Zhe He 0001, Shubo Tian, Ankita Singh, Shayok Chakraborty, Mia Liza A. Lustria, Neil Charness, Nelson A. Roque, Erin Harrell, Walter R. Boot |
Inf. Process. Manag. | 2 |