VLDB 2026 Research / reviewers in the wild / expert
Ling Luo 0001
dblp:00/1811-1
· DBLP profile ↗
43ranked-venue papers
8as first author
32since 2021 · last 2026
0000-0002-5141-0259ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 37 · 8 first-author · 26 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reader comes first: A demand-oriented readability controllable summarization
Qinyu Han, Yuanyuan Sun 0002, Wenfei Liu, Ling Luo 0001, Hongfei Lin |
Expert Syst. Appl. | 5 |
| 2026 | Memory-KGC: Memory-augmented structural learning for Knowledge Graph Completion
Jiru Li, Yuanyuan Sun 0002, Bo Xu 0009, Dinghao Pan, Ling Luo 0001, Hongfei Lin |
Inf. Process. Manag. | 5 |
| 2026 | CNER-Omni: A unified dynamic modality learning framework for Chinese named entity recognition across text and speech
Jinzhong Ning, Wenxuan Mu, Yi-Jia Zhang 0001, Ling Luo 0001, Yuanyuan Sun 0002, Mingyu Lu, Hongfei Lin |
Neural Networks | 5 |
| 2025 | DALE: Semantically Disentangled LoRA Expert Mixture for Depression Detection in Psychiatric DialogueabstractMajor depressive disorder (MDD) is a significant global health burden; timely and precise diagnosis is essential to reduce relapse and mortality. However, existing approaches typically frame depression diagnosis as a monolithic classification problem, neglecting the multi-dimensional and hierarchical reasoning that underpins clinical interviews. We introduce DALE, a modular framework that mirrors psychiatrists'reasoning by explicitly disentangling four dimensions of patient narratives-psychological symptoms, somatic symptoms, protective factors and stressors. Leveraging GPT-4, we augment the dataset with dialogue-level annotations that map patient interview onto a list of attributes. On this auxiliary data, we train four domain-specialised LoRA adapters atop a frozen LLM; each adapter conducts a brief diagnostic dialogue and produces a concise report of its domain. A lightweight classifier then integrates these reports to generate a final summary and predict depression and suicide risk. Experiments on the D4 psychiatric dialogue benchmark show that DALE showing strong performance, while requiring far fewer trainable parameters, and yields interpretable, attribute-level evidence for its predictions. Dailin Li, Qinyu Han, Tengxiao Lv, Jian Wang 0021, Hongfei Lin, Ling Luo 0001, Yuanyuan Sun 0002 |
BIBM | 7 |
| 2025 | CAKADE: Improving ADE Detection on Social Media with LLMs via Counterfactual Augmentation and Knowledge-Enhanced Instruction TuningabstractAutomatic detection of Adverse Drug Events (ADEs) from social media has become increasingly important for post-market drug safety surveillance and pharmacovigilance. Although existing social media-based ADE detection methods have effectively addressed some challenges such as data sparsity and class imbalance, they still suffer from spurious correlations, where models tend to learn co-occurrence patterns between drugs and symptoms, leading them to incorrectly identify drug inefficacy or therapeutic intent as ADEs. To address this issue, we propose a structured medical knowledge-guided counterfactual generation method that leverages authoritative medical databases and large language models to construct clinically plausible counterfactual samples, thereby mitigating spurious correlations at the data level. Furthermore, we propose a knowledge-enhanced instruction-tuning strategy that injects drug indications as causal cues into model inputs to enhance its causal reasoning capabilities. Experimental results on two benchmark datasets demonstrate that our method consistently outperforms state-of-the-art models, effectively alleviating spurious correlations and exhibiting superior capabilities in drug-symptom relationship identification. Weiru Fu, Yunzhi Qiu, Ling Luo 0001, Jian Wang 0021, Hongfei Lin |
BIBM | 3 |
| 2025 | FocusMed: A Large Language Model-Based Framework for Enhancing Medical Question Summarization with Focus IdentificationabstractWith the rapid development of online medical platforms, consumer health questions (CHQs) are inefficient in diagnosis due to redundant information and frequent non-professional terms. The medical question summary (MQS) task aims to transform CHQs into streamlined doctors' frequently asked questions (FAQs), but existing methods still face challenges such as poor identification of question focus and low summary faithfulness. This paper explores the potential of large language models (LLMs) in the MQS task and finds that direct fine-tuning is prone to focus identification bias and generates summaries with low faithfulness. To this end, we propose an optimization framework based on core focus guidance. First, a prompt template is designed to drive the LLMs to extract the core focus from the CHQs that is faithful to the original text. Then, a fine-tuning dataset is constructed in combination with the original CHQ-FAQ pairs to improve the ability to identify the focus of the question. Finally, a multi-dimensional quality evaluation and selection mechanism is proposed to comprehensively improve the quality of the summary from multiple dimensions. We conduct comprehensive experiments on two widely-adopted MQS datasets using three established evaluation metrics. The proposed framework achieves state-of-the-art performance across all measures, demonstrating a significant boost in the model's ability to identify the critical focus of questions and a notable improvement in the faithfulness of the summary. The source codes are freely available at https://github.com/DUT-LiuChao/FocusMed. Ling Luo 0001, Tengxiao Lv, Huan Zhuang, Lejing Yu, Jian Wang 0021, Hongfei Lin |
BIBM | 2 |
| 2025 | A Unified Biomedical Named Entity Recognition Framework With Large Language ModelsabstractAccurate recognition of biomedical named entities is critical for medical information extraction and knowledge discovery. However, existing methods often struggle with nested entities, entity boundary ambiguity, and cross-lingual generalization. In this paper, we propose a unified Biomedical Named Entity Recognition (BioNER) framework based on Large Language Models (LLMs). We first reformulate BioNER as a text generation task and design a symbolic tagging strategy to jointly handle both flat and nested entities with explicit boundary annotation. To enhance multilingual and multi-task generalization, we perform bilingual joint fine-tuning across multiple Chinese and English datasets. Additionally, we introduce a contrastive learning-based entity selector that filters incorrect or spurious predictions by leveraging boundary-sensitive positive and negative samples. Experimental results on four benchmark datasets and two unseen corpora show that our method achieves state-of-the-art performance and robust zero-shot generalization across languages. The source codes are freely available at https://github.com/dreamer-tx/LLMNER. Tengxiao Lv, Ling Luo 0001, Huiyi Lv, Yuanyuan Sun 0002, Jian Wang 0021, Hongfei Lin |
BIBM | 2 |
| 2025 | MedTrust-RAG: Evidence Verification and Trust Alignment for Biomedical Question AnsweringabstractBiomedical question answering (QA) requires precise interpretation of complex medical knowledge. Large language models (LLMs) and retrieval-augmented generation (RAG) leverage external medical literature but often produce hallucinations due to noisy retrieval and insufficient verification. We propose MedTrust-Guided Iterative RAG, a framework that improves factual consistency and reduces hallucinations in medical QA. It introduces three innovations. First, citation-aware reasoning grounds generation in retrieved documents and uses Negative Knowledge Assertions when evidence is missing. Second, an iterative retrieval-verification process refines queries through Medical Gap Analysis. Third, the MedTrust-Align Module (MTAM) applies Direct Preference Optimization to align generation with verified evidence and suppress hallucination-prone patterns. Yingpeng Ning, Yuanyuan Sun 0002, Ling Luo 0001, Hongfei Lin |
BIBM | 3 |
| 2025 | Can Herpes Zoster Vaccine Reduce Alzheimer's Disease Risk? A KG and LLMS Synergistic Integration ApproachabstractCurrent research reveals that the herpes zoster vaccine(HZV) can effectively prevent and mitigate Alzheimer's disease(AD). However, the underlying mechanisms linking the HZV and AD remain incompletely understood. We employ a literature-based discovery (LBD) paradigm to investigate these mechanisms. Traditional knowledge graph-based approaches demonstrate strong performance in implicit knowledge discovery tasks. However, existing methods still face critical challenges: (1) The diverse representations of biomedical entities often degrade knowledge graph quality, such as introducing path redundancy; (2) Current path-ranking mechanisms lack rigorous biological plausibility assessment, resulting in a high number of false-positive paths being retained. To address the aforementioned challenges, this paper proposes a Synergistic Integration framework of Knowledge graphs and Large language models (SIKL). First, the PubTator3.0 tool and manual rule-based methods are employed to extract entities and relational triples from abstract content. A knowledge graph is constructed with medical subject headings as nodes, integrating multiple attributes and relations, effectively mitigating path redundancy caused by diverse entity expressions. Next, a twostage LLM-driven path filtering module is designed. In the first stage, a retrieval-augmented large language model assesses whether candidate paths meet causality conditions, performing preliminary filtering to eliminate false positives. The second stage leverages a chain-of-thought large language model to conduct step-by-step reasoning on remaining paths, further evaluating their biological plausibility. Finally, a comprehensive path ranking strategy combines literature support counts and LLM-generated scores to output Top-p high-confidence hypothetical paths. Experimental results reveals that HZV may delay AD progression through pathways such as neuroinflammatory regulation, with partial mechanisms supported by literature. Yunzhi Qiu, Weiru Fu, Ling Luo 0001, Hongfei Lin |
BIBM | 4 |
| 2025 | LCDL: Classification of ICD codes based on disease label co-occurrence dependency and LongFormer with medical knowledge
Hongfei Lin, Yi-Jia Zhang 0001, Di Zhao 0003, Ling Luo 0001 |
Artif. Intell. Medicine | 6 |
| 2025 | ADENER: A syntax-augmented grid-tagging model for Adverse Drug Event extraction in social media
Weiru Fu, Ling Luo 0001, Hongfei Lin |
J. Biomed. Informatics | 3 |
| 2024 | From Retrieval to Generation: A Simple and Unified Generative Model for End-to-End Task-Oriented DialogueabstractRetrieving appropriate records from the external knowledge base to generate informative responses is the core capability of end-to-end task-oriented dialogue systems (EToDs). Most of the existing methods additionally train the retrieval model or use the memory network to retrieve the knowledge base, which decouples the knowledge retrieval task from the response generation task, making it difficult to jointly optimize and failing to capture the internal relationship between the two tasks. In this paper, we propose a simple and unified generative model for task-oriented dialogue systems, which recasts the EToDs task as a single sequence generation task and uses maximum likelihood training to train the two tasks in a unified manner. To prevent the generation of non-existent records, we design the prefix trie to constrain the model generation, which ensures consistency between the generated records and the existing records in the knowledge base. Experimental results on three public benchmark datasets demonstrate that our method achieves robust performance on generating system responses and outperforms the baseline systems. To facilitate future research in this area, the code is available at https://github.com/dzy1011/Uni-ToD. Zeyuan Ding, Ling Luo 0001, Yuanyuan Sun 0002, Hongfei Lin |
AAAI | 3 |
| 2024 | Biomedical Event Extraction as Semantic SegmentationabstractIn the biomedical field, information is widely distributed across numerous pieces of literature. Extracting events between entities from biomedical texts has garnered significant attention in recent years. However, previous research primarily focus on extracting flat biomedical events, with less attention given to nested biomedical events. Moreover, existing methods for extracting nested events often overlook the long-distance dependencies and global information between trigger words and arguments within events, and they lack sufficient interaction with event type information. To address these issues, we propose a semantic segmentation-based method for extracting nested biomedical events. We introduce U-Net to capture global information and interdependencies between event entities. Additionally, we map event types to natural language text and combine them with sentences for encoding to enhance interaction. We also employ two auxiliary tasks to improve the identification of trigger words and arguments. Finally, events are extracted by identifying the four vertices of the segmented region. Experimental results on two benchmark datasets show that our method excels in recognizing nested biomedical events and outperforms current state-of-the-art methods. Liangyu Gao, Jinzhong Ning, Lei Wang 0085, Yin Zhang 0009, Ling Luo 0001, Bo Xu 0009, Jian Wang 0021, Zhehuan Zhao, Yuanyuan Sun 0002, Hongfei Lin |
BIBM | 8 |
| 2024 | Document-level Biomedical Relation Extraction Based on Relation-guided Entity-level GraphsabstractThe task of document-level biomedical relation extraction involves identifying relational facts between entities across sentences, given specific entities. However, most current methods overlook the associations between entity pairs and generate fixed entity representations merely through mentions, leading to irrelevant mentions interfering with the determination of relational facts. Additionally, these methods fail to consider the global information and dependencies between relational entities. To address these issues, we propose a document-level relation extraction model based on relation-guided entity-level graphs. Our model aggregates all mentions of the same entity through a relation-guided attention mechanism to obtain flexible entity representations. Furthermore, by using U-Net to generate entity-level feature graphs, it facilitates global interactions and dependency capture between entity pairs. Experimental results on two benchmark datasets demonstrate the advantages of our approach in document-level biomedical relation extraction. Liangyu Gao, Haixin Tan, Lei Wang 0085, Yin Zhang 0009, Ling Luo 0001, Bo Xu 0009, Jian Wang 0021, Zhehuan Zhao, Yuanyuan Sun 0002, Hongfei Lin |
BIBM | 7 |
| 2024 | Document Embeddings Enhance Biomedical Retrieval-Augmented GenerationabstractLarge language models (LLMs) perform well in many NLP tasks but frequently generate inaccurate information in the biomedical domain, due to hallucination issues. Retrieval-Augmented Generation (RAG) has been introduced to address this issue by integrating external knowledge, enhancing the factual accuracy of outputs. However, naive RAG encounters challenges in effectively utilizing retrieved content, particularly in specialized domains like biomedicine. LLMs often struggle to integrate retrieved content as irrelevant information can interfere with the model’s judgment. Even if relevant documents are retrieved, the model may be unable to accurately comprehend and utilize the domain-specific features due to its inherent knowledge limitations. To overcome these limitations, we propose Document Embeddings Enhanced Biomedical RAG (DEEB-RAG), a framework that incorporates document embeddings along with the original retrieved text. DEEB-RAG uses MedCPT to generate document embeddings and these embeddings are then aligned with the LLM’s semantic space using a two-stage training process on a simple projector. Experimental results on biomedical QA datasets show that DEEB-RAG improves accuracy, with an average performance increase of 2.3% over naive RAG. This demonstrates DEEB-RAG’s ability to mitigate the challenges of utilizing complex biomedical information, thereby enhancing the reliability and effectiveness of LLMs in biomedical domain. Yongle Kong, Ling Luo 0001, Zeyuan Ding, Lei Wang 0085, Yin Zhang 0009, Bo Xu 0009, Jian Wang 0021, Yuanyuan Sun 0002, Zhehuan Zhao, Hongfei Lin |
BIBM | 3 |
| 2024 | Biomedical Document-level Relation Extraction with Coreference and Anaphor GraphsabstractBiomedical document-level relation extraction is a crucial technology for mining the biomedical relationships necessary for clinical diagnosis, treatment, and medical discovery. Although existing intrasentential relation extraction methods have achieved significant results, the complexity and scattered nature of information in biomedical literature require relation extraction techniques to effectively handle cross-sentence information. For example, existing methods have not been able to explicitly model the phenomena of coreference and anaphor in documents, thus affecting the model’s understanding of complex semantics within the document. To address this issue, we propose a new document-level relation extraction model with coreference and anaphor graphs. By abstracting the document into an undirected graph that includes coreference and anaphor information, the framework effectively models the interactions between entities and leverages graph convolutional network in conjunction with pretrained language model to dynamically understand graph structures. Additionally, the shift from fine-grained entity-pair level to coarse-grained document-level training and inference significantly enhances the model’s efficiency while maintaining high extraction performance. Extensive experiments demonstrate that our model achieves a 5.3% increase in F1-score over baseline models on the BioRED dataset with higher efficiency, confirming its effectiveness in handling relation extraction tasks in complex biomedical literature. Jiru Li, Yuanyuan Sun 0002, Ling Luo 0001, Lei Wang 0085, Yin Zhang 0009, Bo Xu 0009, Jian Wang 0021, Zhehuan Zhao, Hongfei Lin |
BIBM | 4 |
| 2024 | An Improved Method for Phenotype Concept Recognition Using Rich HPO InformationabstractAutomatically identifying human phenotype ontology (HPO) concepts from text is important for disease analysis. Existing ontology-driven methods for phenotype concept recognition mainly rely on concept names and synonym information from the ontology, without fully exploiting the rich ontology information. In this paper, we present an improved phenotype concept recognition method by incorporating rich HPO information. We first design prompts with HPO information and use a cutting-edge large language model GPT-4 to generate synonym augmentation for expanding distant supervised training data. We then propose an ontology vector-enhanced phenotype concept classification model to efficiently integrate the taxonomic hierarchical structure of HPO. Additionally, we employ noisy data augmentation to improve the model’s recognition ability in noisy texts and implement a negation detection function. Experimental results on three standard corpora and two typo corpora show our method compares favorably to previous methods and achieves a significant improvement in noisy texts. The source code and data are freely available at https://github.com/DUTIR-BioNLP/PhenoTagger-Updates. Jiewei Qi, Ling Luo 0001, Jian Wang 0021, Huiwei Zhou, Hongfei Lin |
BIBM | 2 |
| 2024 | Efficient Knowledge Graph Embedding Framework to Alleviate Data Sparsity for Polypharmacy Side Effects PredictionabstractPolypharmacy is the combined use of multiple drugs for the treatment of diseases, which also often comes with a higher risk of side effects. In the medical industry, acquiring rich and comprehensive information about the side effects of multiple drug therapy becomes a crucial task. However, data collection for many side effects is often sparse, so the features of these data cannot be adequately learned, resulting in poor performance in side effects prediction. In this paper, we propose a framework based on knowledge graph embedding (KGE) models which improves KGE by using LTE operations and subsampling methods (called LTESampleKGE). LTESampleKGE consists of two main modules i.e., Entity embedding enhancement module and KGE subsampling module. The former applies linear transformation to entity representation instead of GCN structure to enhance entity embedding, while the latter utilizes subsampling methods for KGE negative sampling (NS) loss to pay more attention to sparse data. Thus, LTESampleKGE can effectively alleviate the problem of data sparsity in the polypharmacy side effects prediction task. Experimental evaluations indicate that our method demonstrates superior performance compared with baseline models. For example, LTESampleKGE outperforms MSTE by 1.20% in PR-AUC score on TWOSIDES dataset and by 0.46% in AP@n score on Drugbank dataset. Senbo Tu, Lei Wang 0085, Yin Zhang 0009, Ling Luo 0001, Bo Xu 0009, Jian Wang 0021, Zhehuan Zhao, Hongfei Lin |
BIBM | 6 |
| 2024 | EDNER: Edge Detection for Named Entity Recognition
Liangyu Gao, Ling Luo 0001, Wenfei Liu, Hongfei Lin, Jian Wang 0021 |
NLPCC (2) | 3 |
| 2024 | Learning to explain is a good biomedical few-shot learnerabstractMOTIVATION: Significant progress has been achieved in biomedical text mining using deep learning methods, which rely heavily on large amounts of high-quality data annotated by human experts. However, the reality is that obtaining high-quality annotated data is extremely challenging due to data scarcity (e.g. rare or new diseases), data privacy and security concerns, and the high cost of data annotation. Additionally, nearly all researches focus on predicting labels without providing corresponding explanations. Therefore, in this paper, we investigate a more realistic scenario, biomedical few-shot learning, and explore the impact of interpretability on biomedical few-shot learning. RESULTS: We present LetEx-Learning to explain-a novel multi-task generative approach that leverages reasoning explanations from large language models (LLMs) to enhance the inductive reasoning ability of few-shot learning. Our approach includes (1) collecting high-quality explanations by devising a suite of complete workflow based on LLMs through CoT prompting and self-training strategies, (2) converting various biomedical NLP tasks into a text-to-text generation task in a unified manner, where collected explanations serve as additional supervision between text-label pairs by multi-task training. Experiments are conducted on three few-shot settings across six biomedical benchmark datasets. The results show that learning to explain improves the performances of diverse biomedical NLP tasks in low-resource scenario, outperforming strong baseline models significantly by up to 6.41%. Notably, the proposed method makes the 220M LetEx perform superior reasoning explanation ability against LLMs. AVAILABILITY AND IMPLEMENTATION: Our source code and data are available at https://github.com/cpmss521/LetEx. Jian Wang 0021, Ling Luo 0001, Hongfei Lin |
Bioinform. | 3 |
| 2024 | Taiyi: a bilingual fine-tuned large language model for diverse biomedical tasksabstractOBJECTIVE: Most existing fine-tuned biomedical large language models (LLMs) focus on enhancing performance in monolingual biomedical question answering and conversation tasks. To investigate the effectiveness of the fine-tuned LLMs on diverse biomedical natural language processing (NLP) tasks in different languages, we present Taiyi, a bilingual fine-tuned LLM for diverse biomedical NLP tasks. MATERIALS AND METHODS: We first curated a comprehensive collection of 140 existing biomedical text mining datasets (102 English and 38 Chinese datasets) across over 10 task types. Subsequently, these corpora were converted to the instruction data used to fine-tune the general LLM. During the supervised fine-tuning phase, a 2-stage strategy is proposed to optimize the model performance across various tasks. RESULTS: Experimental results on 13 test sets, which include named entity recognition, relation extraction, text classification, and question answering tasks, demonstrate that Taiyi achieves superior performance compared to general LLMs. The case study involving additional biomedical NLP tasks further shows Taiyi's considerable potential for bilingual biomedical multitasking. CONCLUSION: Leveraging rich high-quality biomedical corpora and developing effective fine-tuning strategies can significantly improve the performance of LLMs within the biomedical domain. Taiyi shows the bilingual multitasking capability through supervised fine-tuning. However, those tasks such as information extraction that are not generation tasks in nature remain challenging for LLM-based generative approaches, and they still underperform the conventional discriminative approaches using smaller language models. Ling Luo 0001, Jinzhong Ning, Yingwen Zhao, Zeyuan Ding, Weiru Fu, Qinyu Han, Guangtao Xu, Yunzhi Qiu, Dinghao Pan, Jiru Li, Wenduo Feng, Senbo Tu, Jian Wang 0021, Yuanyuan Sun 0002, Hongfei Lin |
J. Am. Medical Informatics Assoc. | 1 |
| 2023 | Joint Biomedical Entity and Relation Extraction Based on Triple Region VerticesabstractAutomatic extraction of biomedical entities and their relations plays a significant role in biomedical curation tasks. Currently, the table-filling methods have received lots of attention in the general domain. However, the presence of complex lengthy sentences and overlapping relations in biomedical texts makes automatic extraction a challenging task. To address this challenge, we propose a joint extraction table-filling method based on the vertices of the triple region. We extract triples by using multi-label classification to mark the boundaries of the triples, fully utilizing the boundary information of the entities. To incorporate the information of the distance between entity pairs, distance embedding is introduced and dilated convolutions are utilized to capture multi-scale contextual information. We evaluated our model on the CHEMPROT and DDIExtraction2013 datasets. The experimental results demonstrate that our model achieves the state-of-the-art performance on both datasets. Jinzhong Ning, Ling Luo 0001, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
BIBM | 5 |
| 2023 | Joint Biomedical Entity and Relation Extraction with Unified Interaction MapsabstractAutomatic extraction of entities and their relations from unstructured literature to form structured triples is essential for biomedical knowledge construction. Although most existing joint methods have effectively addressed some challenging problems in the biomedical corpora, i.e., the prevalent overlapping issue, they still suffer from a lack of consideration for the intrinsic correlations between entities and relations, as well as low computational efficiency. In this paper, we present a joint entity and relation extraction model with unified interaction maps. Specifically, we concatenate all relations in the natural language form with the input text to integrate the semantic information of relations through a deep Transformer-based encoder. In addition, we apply unified interaction maps to capture the correlations, which can naturally handle the overlapping issue. Extensive experiments on the CHEMPROT and DDIExtraction2013 datasets demonstrate the effectiveness of our model, achieving the state-of-the-art performance with higher efficiency. Haixin Tan, Zeyuan Ding, Ling Luo 0001, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
BIBM | 5 |
| 2023 | Term-BLAST-like alignment tool for concept recognition in noisy clinical textsabstractMOTIVATION: Methods for concept recognition (CR) in clinical texts have largely been tested on abstracts or articles from the medical literature. However, texts from electronic health records (EHRs) frequently contain spelling errors, abbreviations, and other nonstandard ways of representing clinical concepts. RESULTS: Here, we present a method inspired by the BLAST algorithm for biosequence alignment that screens texts for potential matches on the basis of matching k-mer counts and scores candidates based on conformance to typical patterns of spelling errors derived from 2.9 million clinical notes. Our method, the Term-BLAST-like alignment tool (TBLAT) leverages a gold standard corpus for typographical errors to implement a sequence alignment-inspired method for efficient entity linkage. We present a comprehensive experimental comparison of TBLAT with five widely used tools. Experimental results show an increase of 10% in recall on scientific publications and 20% increase in recall on EHR records (when compared against the next best method), hence supporting a significant enhancement of the entity linking task. The method can be used stand-alone or as a complement to existing approaches. AVAILABILITY AND IMPLEMENTATION: Fenominal is a Java library that implements TBLAT for named CR of Human Phenotype Ontology terms and is available at https://github.com/monarch-initiative/fenominal under the GNU General Public License v3.0. Tudor Groza, Honghan Wu, Marcel E. Dinger, Daniel Danis, Coleman Hilton, Anita Bagley, Jon R. Davids, Ling Luo 0001, Zhiyong Lu, Peter N. Robinson |
Bioinform. | 8 |
| 2023 | AIONER: all-in-one scheme-based biomedical named entity recognition using deep learningabstractMOTIVATION: Biomedical named entity recognition (BioNER) seeks to automatically recognize biomedical entities in natural language text, serving as a necessary foundation for downstream text mining tasks and applications such as information extraction and question answering. Manually labeling training data for the BioNER task is costly, however, due to the significant domain expertise required for accurate annotation. The resulting data scarcity causes current BioNER approaches to be prone to overfitting, to suffer from limited generalizability, and to address a single entity type at a time (e.g. gene or disease). RESULTS: We therefore propose a novel all-in-one (AIO) scheme that uses external data from existing annotated resources to enhance the accuracy and stability of BioNER models. We further present AIONER, a general-purpose BioNER tool based on cutting-edge deep learning and our AIO schema. We evaluate AIONER on 14 BioNER benchmark tasks and show that AIONER is effective, robust, and compares favorably to other state-of-the-art approaches such as multi-task learning. We further demonstrate the practical utility of AIONER in three independent tasks to recognize entity types not previously seen in training data, as well as the advantages of AIONER over existing methods for processing biomedical text at a large scale (e.g. the entire PubMed data). AVAILABILITY AND IMPLEMENTATION: The source code, trained models and data for AIONER are freely available at https://github.com/ncbi/AIONER. Ling Luo 0001, Chih-Hsuan Wei, Po-Ting Lai, Robert Leaman, Qingyu Chen 0001, Zhiyong Lu |
Bioinform. | 1 |
| 2023 | GNorm2: an improved gene name recognition and normalization systemabstractMOTIVATION: Gene name normalization is an important yet highly complex task in biomedical text mining research, as gene names can be highly ambiguous and may refer to different genes in different species or share similar names with other bioconcepts. This poses a challenge for accurately identifying and linking gene mentions to their corresponding entries in databases such as NCBI Gene or UniProt. While there has been a body of literature on the gene normalization task, few have addressed all of these challenges or make their solutions publicly available to the scientific community. RESULTS: Building on the success of GNormPlus, we have created GNorm2: a more advanced tool with optimized functions and improved performance. GNorm2 integrates a range of advanced deep learning-based methods, resulting in the highest levels of accuracy and efficiency for gene recognition and normalization to date. Our tool is freely available for download. AVAILABILITY AND IMPLEMENTATION: https://github.com/ncbi/GNorm2. Chih-Hsuan Wei, Ling Luo 0001, Rezarta Islamaj Dogan, Po-Ting Lai, Zhiyong Lu |
Bioinform. | 2 |
| 2023 | BioREx: Improving biomedical relation extraction by leveraging heterogeneous datasets
Po-Ting Lai, Chih-Hsuan Wei, Ling Luo 0001, Qingyu Chen 0001, Zhiyong Lu |
J. Biomed. Informatics | 3 |
| 2022 | BioRED: A Comprehensive Biomedical Relation Extraction Dataset
Ling Luo 0001, Po-Ting Lai, Chih-Hsuan Wei, Cecilia N. Arighi, Zhiyong Lu |
AMIA | 1 |
| 2022 | PhenoGene: Disease-gene prioritization using graph embedding on patient phenotypic profiles
Shankai Yan, Ling Luo 0001, Daniel Veltri, Andrew J. Oler, Rajarshi Ghosh, Chih-Hsuan Wei, Morgan Similuk, Zhiyong Lu |
AMIA | 2 |
| 2022 | BioRED: a rich biomedical relation extraction datasetabstractAutomated relation extraction (RE) from biomedical literature is critical for many downstream text mining applications in both research and real-world settings. However, most existing benchmarking datasets for biomedical RE only focus on relations of a single type (e.g. protein-protein interactions) at the sentence level, greatly limiting the development of RE systems in biomedicine. In this work, we first review commonly used named entity recognition (NER) and RE datasets. Then, we present a first-of-its-kind biomedical relation extraction dataset (BioRED) with multiple entity types (e.g. gene/protein, disease, chemical) and relation pairs (e.g. gene-disease; chemical-chemical) at the document level, on a set of 600 PubMed abstracts. Furthermore, we label each relation as describing either a novel finding or previously known background knowledge, enabling automated algorithms to differentiate between novel and background information. We assess the utility of BioRED by benchmarking several existing state-of-the-art methods, including Bidirectional Encoder Representations from Transformers (BERT)-based models, on the NER and RE tasks. Our results show that while existing approaches can reach high performance on the NER task (F-score of 89.3%), there is much room for improvement for the RE task, especially when extracting novel relations (F-score of 47.7%). Our experiments also demonstrate that such a rich dataset can successfully facilitate the development of more accurate, efficient and robust RE systems for biomedicine. Availability: The BioRED dataset and annotation guidelines are freely available at https://ftp.ncbi.nlm.nih.gov/pub/lu/BioRED/. Ling Luo 0001, Po-Ting Lai, Chih-Hsuan Wei, Cecilia N. Arighi, Zhiyong Lu |
Briefings Bioinform. | 1 |
| 2022 | PhenoRerank: A re-ranking model for phenotypic concept recognition pre-trained on human phenotype ontologyabstractThe study aims at developing a neural network model to improve the performance of Human Phenotype Ontology (HPO) concept recognition tools. We used the terms, definitions, and comments about the phenotypic concepts in the HPO database to train our model. The document to be analyzed is first split into sentences and annotated with a base method to generate candidate concepts. The sentences, along with the candidate concepts, are then fed into the pre-trained model for re-ranking. Our model comprises the pre-trained BlueBERT and a feature selection module, followed by a contrastive loss. We re-ranked the results generated by three robust HPO annotation tools and compared the performance against most of the existing approaches. The experimental results show that our model can improve the performance of the existing methods. Significantly, it boosted 3.0% and 5.6% in F1 score on the two evaluated datasets compared with the base methods. It removed more than 80% of the false positives predicted by the base methods, resulting in up to 18% improvement in precision. Our model utilizes the descriptive data in the ontology and the contextual information in the sentences for re-ranking. The results indicate that the additional information and the re-ranking model can significantly enhance the precision of HPO concept recognition compared with the base method. Shankai Yan, Ling Luo 0001, Po-Ting Lai, Daniel Veltri, Andrew J. Oler, Sandhya Xirasagar, Rajarshi Ghosh, Morgan Similuk, Peter N. Robinson, Zhiyong Lu |
J. Biomed. Informatics | 2 |
| 2021 | PhenoTagger: a hybrid method for phenotype concept recognition using human phenotype ontologyabstractMOTIVATION: Automatic phenotype concept recognition from unstructured text remains a challenging task in biomedical text mining research. Previous works that address the task typically use dictionary-based matching methods, which can achieve high precision but suffer from lower recall. Recently, machine learning-based methods have been proposed to identify biomedical concepts, which can recognize more unseen concept synonyms by automatic feature learning. However, most methods require large corpora of manually annotated data for model training, which is difficult to obtain due to the high cost of human annotation. RESULTS: In this article, we propose PhenoTagger, a hybrid method that combines both dictionary and machine learning-based methods to recognize Human Phenotype Ontology (HPO) concepts in unstructured biomedical text. We first use all concepts and synonyms in HPO to construct a dictionary, which is then used to automatically build a distantly supervised training dataset for machine learning. Next, a cutting-edge deep learning model is trained to classify each candidate phrase (n-gram from input sentence) into a corresponding concept label. Finally, the dictionary and machine learning-based prediction results are combined for improved performance. Our method is validated with two HPO corpora, and the results show that PhenoTagger compares favorably to previous methods. In addition, to demonstrate the generalizability of our method, we retrained PhenoTagger using the disease ontology MEDIC for disease concept recognition to investigate the effect of training on different ontologies. Experimental results on the NCBI disease corpus show that PhenoTagger without requiring manually annotated training data achieves competitive performance as compared with state-of-the-art supervised methods. AVAILABILITYAND IMPLEMENTATION: The source code, API information and data for PhenoTagger are freely available at https://github.com/ncbi-nlp/PhenoTagger. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ling Luo 0001, Shankai Yan, Po-Ting Lai, Daniel Veltri, Andrew J. Oler, Sandhya Xirasagar, Rajarshi Ghosh, Morgan Similuk, Peter N. Robinson, Zhiyong Lu |
Bioinform. | 1 |
| 2020 | Exploiting sequence labeling framework to extract document-level relations from biomedical textsabstractBACKGROUND: Both intra- and inter-sentential semantic relations in biomedical texts provide valuable information for biomedical research. However, most existing methods either focus on extracting intra-sentential relations and ignore inter-sentential ones or fail to extract inter-sentential relations accurately and regard the instances containing entity relations as being independent, which neglects the interactions between relations. We propose a novel sequence labeling-based biomedical relation extraction method named Bio-Seq. In the method, sequence labeling framework is extended by multiple specified feature extractors so as to facilitate the feature extractions at different levels, especially at the inter-sentential level. Besides, the sequence labeling framework enables Bio-Seq to take advantage of the interactions between relations, and thus, further improves the precision of document-level relation extraction. RESULTS: Our proposed method obtained an F1-score of 63.5% on BioCreative V chemical disease relation corpus, and an F1-score of 54.4% on inter-sentential relations, which was 10.5% better than the document-level classification baseline. Also, our method achieved an F1-score of 85.1% on n2c2-ADE sub-dataset. CONCLUSION: Sequence labeling method can be successfully used to extract document-level relations, especially for boosting the performance on inter-sentential relation extraction. Our work can facilitate the research on document-level biomedical text mining. Zhiheng Li 0004, Yang Xiang 0003, Ling Luo 0001, Yuanyuan Sun 0002, Hongfei Lin |
BMC Bioinform. | 4 |
| 2020 | Exploiting adversarial transfer learning for adverse drug reaction detection from texts
Zhiheng Li 0004, Ling Luo 0001, Yang Xiang 0003, Hongfei Lin |
J. Biomed. Informatics | 3 |
| 2020 | A neural network-based joint learning approach for biomedical entity and relation extraction from biomedical literature
Ling Luo 0001, Mingyu Cao, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin |
J. Biomed. Informatics | 1 |
| 2018 | A multi-task learning based approach to biomedical entity relation extraction
Ling Luo 0001, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021, Liang Yang 0003, Kan Xu, Yi-Jia Zhang 0001 |
BIBM | 3 |
| 2018 | HMNPPID: A Database of Protein-protein Interactions Associated with Human Malignant Neoplasms
Zhehuan Zhao, Ling Luo 0001, Zhiheng Li 0004, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021, Yi-Jia Zhang 0001 |
BIBM | 4 |
| 2018 | Protein-Protein Interaction Article Classification: A Knowledge-enriched Self-Attention Convolutional Neural Network Approach
Ling Luo 0001, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021, Liang Yang 0003, Kan Xu, Yi-Jia Zhang 0001 |
BIBM | 1 |
| 2018 | An attention-based BiLSTM-CRF approach to document-level chemical named entity recognitionabstractMotivation: In biomedical research, chemical is an important class of entities, and chemical named entity recognition (NER) is an important task in the field of biomedical information extraction. However, most popular chemical NER methods are based on traditional machine learning and their performances are heavily dependent on the feature engineering. Moreover, these methods are sentence-level ones which have the tagging inconsistency problem. Results: In this paper, we propose a neural network approach, i.e. attention-based bidirectional Long Short-Term Memory with a conditional random field layer (Att-BiLSTM-CRF), to document-level chemical NER. The approach leverages document-level global information obtained by attention mechanism to enforce tagging consistency across multiple instances of the same token in a document. It achieves better performances with little feature engineering than other state-of-the-art methods on the BioCreative IV chemical compound and drug name recognition (CHEMDNER) corpus and the BioCreative V chemical-disease relation (CDR) task corpus (the F-scores of 91.14 and 92.57%, respectively). Availability and implementation: Data and code are available at https://github.com/lingluodlut/Att-ChemdNER. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Ling Luo 0001, Yin Zhang 0009, Lei Wang 0085, Hongfei Lin, Jian Wang 0021 |
Bioinform. | 1 |
| 2017 | An attention-based effective neural model for drug-drug interactions extractionabstractBACKGROUND: Drug-drug interactions (DDIs) often bring unexpected side effects. The clinical recognition of DDIs is a crucial issue for both patient safety and healthcare cost control. However, although text-mining-based systems explore various methods to classify DDIs, the classification performance with regard to DDIs in long and complex sentences is still unsatisfactory. METHODS: In this study, we propose an effective model that classifies DDIs from the literature by combining an attention mechanism and a recurrent neural network with long short-term memory (LSTM) units. In our approach, first, a candidate-drug-oriented input attention acting on word-embedding vectors automatically learns which words are more influential for a given drug pair. Next, the inputs merging the position- and POS-embedding vectors are passed to a bidirectional LSTM layer whose outputs at the last time step represent the high-level semantic information of the whole sentence. Finally, a softmax layer performs DDI classification. RESULTS: Experimental results from the DDIExtraction 2013 corpus show that our system performs the best with respect to detection and classification (84.0% and 77.3%, respectively) compared with other state-of-the-art methods. In particular, for the Medline-2013 dataset with long and complex sentences, our F-score far exceeds those of top-ranking systems by 12.6%. CONCLUSIONS: Our approach effectively improves the performance of DDI classification tasks. Experimental analysis demonstrates that our model performs better with respect to recognizing not only close-range but also long-range patterns among words, especially for long, complex and compound sentences. Wei Zheng 0003, Hongfei Lin, Ling Luo 0001, Zhehuan Zhao, Zhengguang Li, Yi-Jia Zhang 0001, Jian Wang 0021 |
BMC Bioinform. | 3 |
| 2016 | ML-CNN: A novel deep learning based disease named entity recognition architectureabstractIn this paper, we present a deep learning based disease named entity recognition architecture. First, the word-level embedding, character-level embedding and lexicon feature embedding are concatenated as input. Then multiple convolutional layers are stacked over the input to extract useful features automatically. Finally, multiple label strategy, which is firstly introduced, is applied to the output layer to capture the correlation information between neighboring labels. Experimental results on both NCBI and CDR corpora show that ML-CNN can achieve the state-of-the-art performance. Zhehuan Zhao, Ling Luo 0001, Yin Zhang 0009, Lei Wang 0085, Hongfei Lin, Jian Wang 0021 |
BIBM | 3 |
| 2016 | Drug drug interaction extraction from biomedical literature using syntax convolutional neural networkabstractMOTIVATION: Detecting drug-drug interaction (DDI) has become a vital part of public health safety. Therefore, using text mining techniques to extract DDIs from biomedical literature has received great attentions. However, this research is still at an early stage and its performance has much room to improve. RESULTS: In this article, we present a syntax convolutional neural network (SCNN) based DDI extraction method. In this method, a novel word embedding, syntax word embedding, is proposed to employ the syntactic information of a sentence. Then the position and part of speech features are introduced to extend the embedding of each word. Later, auto-encoder is introduced to encode the traditional bag-of-words feature (sparse 0-1 vector) as the dense real value vector. Finally, a combination of embedding-based convolutional features and traditional features are fed to the softmax classifier to extract DDIs from biomedical literature. Experimental results on the DDIExtraction 2013 corpus show that SCNN obtains a better performance (an F-score of 0.686) than other state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: The source code is available for academic use at http://202.118.75.18:8080/DDI/SCNN-DDI.zip CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Zhehuan Zhao, Ling Luo 0001, Hongfei Lin, Jian Wang 0021 |
Bioinform. | 3 |
| 2015 | Deep neural network based protein-protein interaction extraction from biomedical literatureabstractThis paper presents a deep neural network-based protein-protein interactions (PPIs) information extraction approach which can learn complex and abstract features automatically from unlabeled data by unsupervised representation learning methods. This approach first employs the training algorithm of auto-encoders to initialize the parameters of a deep multilayer neural network. Then the gradient descent method using back-propagation is applied to train this deep multilayer neural network model. Experimental results on five public PPI corpora show that our method can achieve better performance than can a multilayer neural network. In addition, the performance comparison with APG also verifies the effectiveness of our method. Zhehuan Zhao, Ling Luo 0001, Hongfei Lin, Jian Wang 0021 |
BIBM | 3 |