VLDB 2026 Research / reviewers in the wild / expert
Lei Wang 0085
dblp:w/LeiWang85
· DBLP profile ↗
37ranked-venue papers
0as first author
16since 2021 · last 2024
0000-0002-8420-6860ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 37 · 16 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Biomedical Event Extraction as Semantic SegmentationabstractIn the biomedical field, information is widely distributed across numerous pieces of literature. Extracting events between entities from biomedical texts has garnered significant attention in recent years. However, previous research primarily focus on extracting flat biomedical events, with less attention given to nested biomedical events. Moreover, existing methods for extracting nested events often overlook the long-distance dependencies and global information between trigger words and arguments within events, and they lack sufficient interaction with event type information. To address these issues, we propose a semantic segmentation-based method for extracting nested biomedical events. We introduce U-Net to capture global information and interdependencies between event entities. Additionally, we map event types to natural language text and combine them with sentences for encoding to enhance interaction. We also employ two auxiliary tasks to improve the identification of trigger words and arguments. Finally, events are extracted by identifying the four vertices of the segmented region. Experimental results on two benchmark datasets show that our method excels in recognizing nested biomedical events and outperforms current state-of-the-art methods. Liangyu Gao, Jinzhong Ning, Lei Wang 0085, Yin Zhang 0009, Ling Luo 0001, Bo Xu 0009, Jian Wang 0021, Zhehuan Zhao, Yuanyuan Sun 0002, Hongfei Lin |
BIBM | 5 |
| 2024 | Document-level Biomedical Relation Extraction Based on Relation-guided Entity-level GraphsabstractThe task of document-level biomedical relation extraction involves identifying relational facts between entities across sentences, given specific entities. However, most current methods overlook the associations between entity pairs and generate fixed entity representations merely through mentions, leading to irrelevant mentions interfering with the determination of relational facts. Additionally, these methods fail to consider the global information and dependencies between relational entities. To address these issues, we propose a document-level relation extraction model based on relation-guided entity-level graphs. Our model aggregates all mentions of the same entity through a relation-guided attention mechanism to obtain flexible entity representations. Furthermore, by using U-Net to generate entity-level feature graphs, it facilitates global interactions and dependency capture between entity pairs. Experimental results on two benchmark datasets demonstrate the advantages of our approach in document-level biomedical relation extraction. Liangyu Gao, Haixin Tan, Lei Wang 0085, Yin Zhang 0009, Ling Luo 0001, Bo Xu 0009, Jian Wang 0021, Zhehuan Zhao, Yuanyuan Sun 0002, Hongfei Lin |
BIBM | 4 |
| 2024 | Document Embeddings Enhance Biomedical Retrieval-Augmented GenerationabstractLarge language models (LLMs) perform well in many NLP tasks but frequently generate inaccurate information in the biomedical domain, due to hallucination issues. Retrieval-Augmented Generation (RAG) has been introduced to address this issue by integrating external knowledge, enhancing the factual accuracy of outputs. However, naive RAG encounters challenges in effectively utilizing retrieved content, particularly in specialized domains like biomedicine. LLMs often struggle to integrate retrieved content as irrelevant information can interfere with the model’s judgment. Even if relevant documents are retrieved, the model may be unable to accurately comprehend and utilize the domain-specific features due to its inherent knowledge limitations. To overcome these limitations, we propose Document Embeddings Enhanced Biomedical RAG (DEEB-RAG), a framework that incorporates document embeddings along with the original retrieved text. DEEB-RAG uses MedCPT to generate document embeddings and these embeddings are then aligned with the LLM’s semantic space using a two-stage training process on a simple projector. Experimental results on biomedical QA datasets show that DEEB-RAG improves accuracy, with an average performance increase of 2.3% over naive RAG. This demonstrates DEEB-RAG’s ability to mitigate the challenges of utilizing complex biomedical information, thereby enhancing the reliability and effectiveness of LLMs in biomedical domain. Yongle Kong, Ling Luo 0001, Zeyuan Ding, Lei Wang 0085, Yin Zhang 0009, Bo Xu 0009, Jian Wang 0021, Yuanyuan Sun 0002, Zhehuan Zhao, Hongfei Lin |
BIBM | 5 |
| 2024 | Biomedical Document-level Relation Extraction with Coreference and Anaphor GraphsabstractBiomedical document-level relation extraction is a crucial technology for mining the biomedical relationships necessary for clinical diagnosis, treatment, and medical discovery. Although existing intrasentential relation extraction methods have achieved significant results, the complexity and scattered nature of information in biomedical literature require relation extraction techniques to effectively handle cross-sentence information. For example, existing methods have not been able to explicitly model the phenomena of coreference and anaphor in documents, thus affecting the model’s understanding of complex semantics within the document. To address this issue, we propose a new document-level relation extraction model with coreference and anaphor graphs. By abstracting the document into an undirected graph that includes coreference and anaphor information, the framework effectively models the interactions between entities and leverages graph convolutional network in conjunction with pretrained language model to dynamically understand graph structures. Additionally, the shift from fine-grained entity-pair level to coarse-grained document-level training and inference significantly enhances the model’s efficiency while maintaining high extraction performance. Extensive experiments demonstrate that our model achieves a 5.3% increase in F1-score over baseline models on the BioRED dataset with higher efficiency, confirming its effectiveness in handling relation extraction tasks in complex biomedical literature. Jiru Li, Yuanyuan Sun 0002, Ling Luo 0001, Lei Wang 0085, Yin Zhang 0009, Bo Xu 0009, Jian Wang 0021, Zhehuan Zhao, Hongfei Lin |
BIBM | 5 |
| 2024 | Efficient Knowledge Graph Embedding Framework to Alleviate Data Sparsity for Polypharmacy Side Effects PredictionabstractPolypharmacy is the combined use of multiple drugs for the treatment of diseases, which also often comes with a higher risk of side effects. In the medical industry, acquiring rich and comprehensive information about the side effects of multiple drug therapy becomes a crucial task. However, data collection for many side effects is often sparse, so the features of these data cannot be adequately learned, resulting in poor performance in side effects prediction. In this paper, we propose a framework based on knowledge graph embedding (KGE) models which improves KGE by using LTE operations and subsampling methods (called LTESampleKGE). LTESampleKGE consists of two main modules i.e., Entity embedding enhancement module and KGE subsampling module. The former applies linear transformation to entity representation instead of GCN structure to enhance entity embedding, while the latter utilizes subsampling methods for KGE negative sampling (NS) loss to pay more attention to sparse data. Thus, LTESampleKGE can effectively alleviate the problem of data sparsity in the polypharmacy side effects prediction task. Experimental evaluations indicate that our method demonstrates superior performance compared with baseline models. For example, LTESampleKGE outperforms MSTE by 1.20% in PR-AUC score on TWOSIDES dataset and by 0.46% in AP@n score on Drugbank dataset. Senbo Tu, Lei Wang 0085, Yin Zhang 0009, Ling Luo 0001, Bo Xu 0009, Jian Wang 0021, Zhehuan Zhao, Hongfei Lin |
BIBM | 3 |
| 2024 | Predicting Protein Functions Based on Heterogeneous Graph Attention TechniqueabstractIn bioinformatics, protein function prediction stands as a fundamental area of research and plays a crucial role in addressing various biological challenges, such as the identification of potential targets for drug discovery and the elucidation of disease mechanisms. However, known functional annotation databases usually provide positive experimental annotations that proteins carry out a given function, and rarely record negative experimental annotations that proteins do not carry out a given function. Therefore, existing computational methods based on deep learning models focus on these positive annotations for prediction and ignore these scarce but informative negative annotations, leading to an underestimation of precision. To address this issue, we introduce a deep learning method that utilizes a heterogeneous graph attention technique. The method first constructs a heterogeneous graph that covers the protein-protein interaction network, ontology structure, and positive and negative annotation information. Then, it learns embedding representations of proteins and ontology terms by using the heterogeneous graph attention technique. Finally, it leverages these learned representations to reconstruct the positive protein-term associations and score unobserved functional annotations. It can enhance the predictive performance by incorporating these known limited negative annotations into the constructed heterogeneous graph. Experimental results on three species (i.e., Human, Mouse, and Arabidopsis) demonstrate that our method can achieve better performance in predicting new protein annotations than state-of-the-art methods. Yingwen Zhao, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Joint Biomedical Entity and Relation Extraction Based on Triple Region VerticesabstractAutomatic extraction of biomedical entities and their relations plays a significant role in biomedical curation tasks. Currently, the table-filling methods have received lots of attention in the general domain. However, the presence of complex lengthy sentences and overlapping relations in biomedical texts makes automatic extraction a challenging task. To address this challenge, we propose a joint extraction table-filling method based on the vertices of the triple region. We extract triples by using multi-label classification to mark the boundaries of the triples, fully utilizing the boundary information of the entities. To incorporate the information of the distance between entity pairs, distance embedding is introduced and dilated convolutions are utilized to capture multi-scale contextual information. We evaluated our model on the CHEMPROT and DDIExtraction2013 datasets. The experimental results demonstrate that our model achieves the state-of-the-art performance on both datasets. Jinzhong Ning, Ling Luo 0001, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
BIBM | 6 |
| 2023 | Joint Biomedical Entity and Relation Extraction with Unified Interaction MapsabstractAutomatic extraction of entities and their relations from unstructured literature to form structured triples is essential for biomedical knowledge construction. Although most existing joint methods have effectively addressed some challenging problems in the biomedical corpora, i.e., the prevalent overlapping issue, they still suffer from a lack of consideration for the intrinsic correlations between entities and relations, as well as low computational efficiency. In this paper, we present a joint entity and relation extraction model with unified interaction maps. Specifically, we concatenate all relations in the natural language form with the input text to integrate the semantic information of relations through a deep Transformer-based encoder. In addition, we apply unified interaction maps to capture the correlations, which can naturally handle the overlapping issue. Extensive experiments on the CHEMPROT and DDIExtraction2013 datasets demonstrate the effectiveness of our model, achieving the state-of-the-art performance with higher efficiency. Haixin Tan, Zeyuan Ding, Ling Luo 0001, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
BIBM | 6 |
| 2023 | Improving Protein Function Prediction by Adaptively Fusing Information From Protein Sequences and Biomedical LiteratureabstractProteins are the main undertakers of life activities, and accurately predicting their biological functions can help human better understand life mechanism and promote the development of themselves. With the rapid development of high-throughput technologies, an abundance of proteins are discovered. However, the gap between proteins and function annotations is still huge. To accelerate the process of protein function prediction, some computational methods taking advantage of multiple data have been proposed. Among these methods, the deep-learning-based methods are currently the most popular for their capability of learning information automatically from raw data. However, due to the diversity and scale difference between data, it is challenging for existing deep learning methods to capture related information from different data effectively. In this paper, we introduce a deep learning method that can adaptively learn information from protein sequences and biomedical literature, namely DeepAF. DeepAF first extracts the two kinds of information by using different extractors, which are built based on pre-trained language models and can capture rudimentary biological knowledge. Then, to integrate those information, it performs an adaptive fusion layer based on a Cross-attention mechanism that considers the knowledge of mutual interactions between two information. Finally, based on the mixed information, DeepAF utilizes logistic regression to obtain prediction scores. The experimental results on the datasets of two species (i.e., Human and Yeast) show that DeepAF outperforms other state-of-the-art approaches. Yingwen Zhao, Yongkai Hong, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | BioNER-CFEM: Biomedical Named Entity Recognition Based on Character Feature Enhancement with Multimodal MethodabstractBiomedical named entity recognition (Bio-NER) is an essential task for biomedical information extraction. In this paper, we regard word-level features and character-level features as two different modalities from a novel perspective and propose a biomedical named entity recognition model based on character feature enhancement with multimodal method (called BioNER-CFEM). BioNER-CFEM can not only capture interactions between modalities, but also learn interactions within modalities. In addition, our proposed cross-attention based sparse selection mechanism can effectively alleviate the noise in the interaction process of the two ‘modalities’. Experimental results show the effectiveness of BioNER-CFEM for the Bio-NER task: it achieves performance boost over SOTA models with competitive efficiency on all six Bio-NER datasets, i.e., $+0.89, +0.64, +0.40$, $+1.40, +5.57, +2.81$ on NCBI-Disease, BC5CDR-Disease, BC5CDR-Chem, BC2GM, JNLPBA, BC4CHEMD, respectively. Jinzhong Ning, Jiru Li, Yuanyuan Sun 0002, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
BIBM | 5 |
| 2022 | Adaptive Multi-view Graph Convolutional Network for Gene Ontology Annotations of ProteinsabstractGene Ontology (GO) containing a set of standard concepts (or terms) is launched to unify the functional descriptions of proteins. Developing computational models based on GO to automatically annotate protein functions has been a longstanding active research area. In this paper, we propose a novel method to adaptively fuse functional and topological information between GO Terms. Our method is composed of a pre-trained language model for encoding protein sequences and an adaptive multi-view graph convolutional network (Multi-view GCN) for representing GO terms. Particularly, the Multi-view GCN considers multiple views from functional information, topological structures, and their combinations, and extracts multiple corresponding representations of GO terms. Then, an attention mechanism is applied to adaptively learn the importance weights of these representations. Finally, the predicted scores are calculated by using a dot product between protein sequence features and GO term representations. Experimental results on the datasets of two species (i.e., Human and Yeast) show that our method outperforms other state-of-the-art methods. The code of our proposed method is available at: https://github.com/Candyperfect/Master. Yingwen Zhao, Yongkai Hong, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
BIBM | 4 |
| 2022 | MRC4BioER: Joint extraction of biomedical entities and relations in the machine reading comprehension framework
Cong Sun 0004, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
J. Biomed. Informatics | 3 |
| 2021 | SGAT: a Self-supervised Graph Attention Network for Biomedical Relation ExtractionabstractThe goal of relation extraction task is to classify texts containing entity pairs into predefined relation types. Biomedical relation extraction can extract high-quality information from massive medical texts, which plays an important role in biomedical research. In this paper, we propose a self-supervised graph attention network to extract biomedical relations from the complex and noisy biomedical texts. The model incorporates self-supervision within the standard graph attention mechanism. Specifically, the model applies the graph attention mechanism to reduce the influence of noisy words and introduces dependency-based parse trees to construct a self-supervised task. With the supervision of dependency-based parse trees, the graph attention network can not only improve its capacity of learning syntactic information but also alleviate its lack of interpretability. Additionally, we use Gumbel Tree-GRU to obtain sentence information for relation classification. Our model achieves state-of-the-art performance on the DDIExtraction 2013 and ChemProt datasets, respectively, which suggests that our proposed model can effectively improve the performance of biomedical relation extraction. Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jinzhong Ning |
BIBM | 3 |
| 2021 | Deep learning with language models improves named entity recognition for PharmaCoNERabstractBACKGROUND: The recognition of pharmacological substances, compounds and proteins is essential for biomedical relation extraction, knowledge graph construction, drug discovery, as well as medical question answering. Although considerable efforts have been made to recognize biomedical entities in English texts, to date, only few limited attempts were made to recognize them from biomedical texts in other languages. PharmaCoNER is a named entity recognition challenge to recognize pharmacological entities from Spanish texts. Because there are currently abundant resources in the field of natural language processing, how to leverage these resources to the PharmaCoNER challenge is a meaningful study. METHODS: Inspired by the success of deep learning with language models, we compare and explore various representative BERT models to promote the development of the PharmaCoNER task. RESULTS: The experimental results show that deep learning with language models can effectively improve model performance on the PharmaCoNER dataset. Our method achieves state-of-the-art performance on the PharmaCoNER dataset, with a max F1-score of 92.01%. CONCLUSION: For the BERT models on the PharmaCoNER dataset, biomedical domain knowledge has a greater impact on model performance than the native language (i.e., Spanish). The BERT models can obtain competitive performance by using WordPiece to alleviate the out of vocabulary limitation. The performance on the BERT model can be further improved by constructing a specific vocabulary based on domain knowledge. Moreover, the character case also has a certain impact on model performance. Cong Sun 0004, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
BMC Bioinform. | 3 |
| 2021 | Biomedical named entity recognition using BERT in the machine reading comprehension framework
Cong Sun 0004, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
J. Biomed. Informatics | 3 |
| 2021 | Lexicon Knowledge Boosted Interaction Graph Network for Adverse Drug Reaction Recognition From Social MediaabstractThe World Health Organization underlines the significance of adverse drug reaction (ADR) reports for patients' safety. Actually, many potential ADRs tend to be under-reported in post-market ADR surveillance. Recognizing ADRs from social media is indispensably important and could complement post-market ADR surveillance for more effective pharmacovigilance studies. However, previous approaches pose two challenges: 1) ADRs show high expression variability in social media, and thus, many potential ADRs are out-of-lexicon ones, which are difficult to be recognized, and 2) most phrasal ADRs are non-standard mentions and their boundaries are difficult to identify accurately. To tackle these challenges, we design three interaction graphs and propose a neural network approach, i.e., Interaction Graph Network (IGN). Specifically, to recognize more out-of-lexicon ADRs, besides the mentions in ADR lexicon, noun phrases in the input sentence are regarded as candidate phrases and their features are taken into considerations. Moreover, in an attempt to accurately identify ADR boundaries, three word-phrase interaction graphs are designed to represent lexicon knowledge and are encoded using graph attention networks (GATs) to directly integrate various boundary and contextual information of candidate phrases into ADR recognition. Experimental results on two benchmark datasets show that IGN can recognize ADR accurately and consistently outperforms other state-of-the-art approaches. Zhiheng Li 0004, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Extracting biomedical relations via a multi-head attention based graph convolutional networkabstractAutomatic extraction of biomedical relations is important for many tasks, such as drug discovery, protein prediction and knowledge graph construction. However, due to the complex and noisy expressions in biomedical texts, existing traditional neural networks, such as recurrent neural networks and convolutional neural networks, fail to capture syntactic information effectively. In this paper, we introduce a multi-head attention mechanism into graph convolutional networks to extract biomedical relations. In our method, the graph convolutional network is exploited to encode the dependency structure of an input sentence and the multi-head attention mechanism is utilized to alleviate the influence of noisy words. We evaluated our method on the ChemProt corpus and the protein-protein interaction corpus which includes five separate sub-datasets and it achieves F-scores of 67.37% and 84.8% on ChemProt corpus and PPI corpora, respectively. The experimental results suggest that Our model can not only alleviate the influence of noisy words, but also obtain more semantic and syntactic information from dependency graph than previous proposed models. Erniu Wang, Fan Wang 0005, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
BIBM | 4 |
| 2020 | Star-BiLSTM-LAN for Document-level Mutation-Disease Relation Extraction from Biomedical LiteratureabstractRelations between mutations and diseases hiding in biomedical literature are valuable for the analysis and interpretation of many complex diseases, which can help explore more effective treatment options for corresponding diseases. Most current document-level mutation-disease relation extraction methods are based on classification approaches and suffer the lack of ability to extract inter-sentential relations. To solve this problem, we regard extracting document-level mutation-disease relations as a sequence tagging task and propose a neural network-based method called Star-BiLSTM-LAN. By combining the star transformer and the Bi-directional Long Short-Term Memory network, this method achieves a strong ability to capture semantic and syntactic information at the document level from different aspects, and it can discover internal representations that prove useful for the task of interest. Star-BiLSTM-LAN is evaluated on EMU BCa and PCa datasets, and achieves the state-of-the-art F-scores of 89.20% and 90.43%, which are 4.70% and 2.43% higher than the baseline, respectively. Also, the proposed method achieves an F-score of 94.40% on BRONCO dataset. Yuan Xu 0025, Yawen Song, Zhiheng Li 0004, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
BIBM | 5 |
| 2020 | Chemical-protein interaction extraction via Gaussian probability distribution and external biomedical knowledgeabstractMOTIVATION: The biomedical literature contains a wealth of chemical-protein interactions (CPIs). Automatically extracting CPIs described in biomedical literature is essential for drug discovery, precision medicine, as well as basic biomedical research. Most existing methods focus only on the sentence sequence to identify these CPIs. However, the local structure of sentences and external biomedical knowledge also contain valuable information. Effective use of such information may improve the performance of CPI extraction. RESULTS: In this article, we propose a novel neural network-based approach to improve CPI extraction. Specifically, the approach first employs BERT to generate high-quality contextual representations of the title sequence, instance sequence and knowledge sequence. Then, the Gaussian probability distribution is introduced to capture the local structure of the instance. Meanwhile, the attention mechanism is applied to fuse the title information and biomedical knowledge, respectively. Finally, the related representations are concatenated and fed into the softmax function to extract CPIs. We evaluate our proposed model on the CHEMPROT corpus. Our proposed model is superior in performance as compared with other state-of-the-art models. The experimental results show that the Gaussian probability distribution and external knowledge are complementary to each other. Integrating them can effectively improve the CPI extraction performance. Furthermore, the Gaussian probability distribution can effectively improve the extraction performance of sentences with overlapping relations in biomedical relation extraction tasks. AVAILABILITY AND IMPLEMENTATION: Data and code are available at https://github.com/CongSun-dlut/CPI_extraction. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Cong Sun 0004, Leilei Su, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
Bioinform. | 4 |
| 2020 | A neural network-based joint learning approach for biomedical entity and relation extraction from biomedical literature
Ling Luo 0001, Mingyu Cao, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin |
J. Biomed. Informatics | 4 |
| 2020 | Attention guided capsule networks for chemical-protein interaction extraction
Cong Sun 0004, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
J. Biomed. Informatics | 3 |
| 2018 | A multi-task learning based approach to biomedical entity relation extraction
Ling Luo 0001, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021, Liang Yang 0003, Kan Xu, Yi-Jia Zhang 0001 |
BIBM | 4 |
| 2018 | HMNPPID: A Database of Protein-protein Interactions Associated with Human Malignant Neoplasms
Zhehuan Zhao, Ling Luo 0001, Zhiheng Li 0004, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021, Yi-Jia Zhang 0001 |
BIBM | 6 |
| 2018 | PC-SENE: A node embedding based method for protein complex detection
Shengtian Sang, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Bo Xu 0009, Yi-Jia Zhang 0001, Liang Yang 0003, Kan Xu, Jian Wang 0021 |
BIBM | 4 |
| 2018 | Protein-Protein Interaction Article Classification: A Knowledge-enriched Self-Attention Convolutional Neural Network Approach
Ling Luo 0001, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021, Liang Yang 0003, Kan Xu, Yi-Jia Zhang 0001 |
BIBM | 3 |
| 2018 | A Knowledge Graph based Bidirectional Recurrent Neural Network Method for Literature-based Discovery
Shengtian Sang, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021, Liang Yang 0003, Kan Xu, Yi-Jia Zhang 0001 |
BIBM | 4 |
| 2018 | Hierarchical Recurrent Convolutional Neural Network for Chemical-protein Relation Extraction from Biomedical Literature
Cong Sun 0004, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021, Liang Yang 0003, Kan Xu, Yi-Jia Zhang 0001 |
BIBM | 3 |
| 2018 | An attention-based BiLSTM-CRF approach to document-level chemical named entity recognitionabstractMotivation: In biomedical research, chemical is an important class of entities, and chemical named entity recognition (NER) is an important task in the field of biomedical information extraction. However, most popular chemical NER methods are based on traditional machine learning and their performances are heavily dependent on the feature engineering. Moreover, these methods are sentence-level ones which have the tagging inconsistency problem. Results: In this paper, we propose a neural network approach, i.e. attention-based bidirectional Long Short-Term Memory with a conditional random field layer (Att-BiLSTM-CRF), to document-level chemical NER. The approach leverages document-level global information obtained by attention mechanism to enforce tagging consistency across multiple instances of the same token in a document. It achieves better performances with little feature engineering than other state-of-the-art methods on the BioCreative IV chemical compound and drug name recognition (CHEMDNER) corpus and the BioCreative V chemical-disease relation (CDR) task corpus (the F-scores of 91.14 and 92.57%, respectively). Availability and implementation: Data and code are available at https://github.com/lingluodlut/Att-ChemdNER. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Ling Luo 0001, Yin Zhang 0009, Lei Wang 0085, Hongfei Lin, Jian Wang 0021 |
Bioinform. | 5 |
| 2018 | Identifying protein complexes based on node embeddings obtained from protein-protein interaction networksabstractBACKGROUND: Protein complexes are one of the keys to deciphering the behavior of a cell system. During the past decade, most computational approaches used to identify protein complexes have been based on discovering densely connected subgraphs in protein-protein interaction (PPI) networks. However, many true complexes are not dense subgraphs and these approaches show limited performances for detecting protein complexes from PPI networks. RESULTS: To solve these problems, in this paper we propose a supervised learning method based on network node embeddings which utilizes the informative properties of known complexes to guide the search process for new protein complexes. First, node embeddings are obtained from human protein interaction network. Then the protein interactions are weighted through the similarities between node embeddings. After that, the supervised learning method is used to detect protein complexes. Then the random forest model is used to filter the candidate complexes in order to obtain the final predicted complexes. Experimental results on real human and yeast protein interaction networks show that our method effectively improves the performance for protein complex detection. CONCLUSIONS: We provided a new method for identifying protein complexes from human and yeast protein interaction networks, which has great potential to benefit the field of protein complex detection. Shengtian Sang, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021, Bo Xu 0008 |
BMC Bioinform. | 5 |
| 2018 | SemaTyP: a knowledge graph based literature mining method for drug discoveryabstractBACKGROUND: Drug discovery is the process through which potential new medicines are identified. High-throughput screening and computer-aided drug discovery/design are the two main drug discovery methods for now, which have successfully discovered a series of drugs. However, development of new drugs is still an extremely time-consuming and expensive process. Biomedical literature contains important clues for the identification of potential treatments. It could support experts in biomedicine on their way towards new discoveries. METHODS: Here, we propose a biomedical knowledge graph-based drug discovery method called SemaTyP, which discovers candidate drugs for diseases by mining published biomedical literature. We first construct a biomedical knowledge graph with the relations extracted from biomedical abstracts, then a logistic regression model is trained by learning the semantic types of paths of known drug therapies' existing in the biomedical knowledge graph, finally the learned model is used to discover drug therapies for new diseases. RESULTS: The experimental results show that our method could not only effectively discover new drug therapies for new diseases, but also could provide the potential mechanism of action of the candidate drugs. CONCLUSIONS: In this paper we propose a novel knowledge graph based literature mining method for drug discovery. It could be a supplementary method for current drug discovery methods. Shengtian Sang, Lei Wang 0085, Hongfei Lin, Jian Wang 0021 |
BMC Bioinform. | 3 |
| 2017 | A hybrid protein-protein interaction triple extraction method for biomedical literatureabstractProtein-protein interaction extraction research can be widely applied to the field of life science research. However, most of the machine learning based methods focus on binary PPI relation extraction, which loses rich relationship type information that is critical to the PPIs study. The rule based open information extraction methods can extract the PPI triple (i.e. “protein1, interaction word, protein2”), but suffers from low recall rate problem. In this paper, we propose a hybrid protein-protein interaction triple extraction method. In this method, firstly, machine learning techniques are used to recognize protein entities and extract relational protein pairs. Then, the syntactic patterns and a dictionary are employed to find out corresponding interaction words that represent the relationships between two proteins. This method obtains an F-score of 40.18% on the AImed corpus, which is much higher than the result achieved by the rule based Stanford open information extraction method. Zhehuan Zhao, Cong Sun 0004, Lei Wang 0085, Hongfei Lin |
BIBM | 4 |
| 2016 | Epistasis detection using a permutation-based Gradient Boosting MachineabstractDetecting single nucleotide polymorphism (SNP) epistasis contributes to understand disease susceptibility and discover disease pathogenesis underlying complex disease. In this paper, we propose an approach called permutation-based Gradient Boosting Machine (pGBM) to detect pure epistasis by estimating the power of a GBM classifier which is influenced by permuting SNP pairs. pGBM is based on two permutation strategies and gradient boosting machine model. To extend pGBM to detect pure epistasis well on unbalanced dataset, average AUC difference value is chosen as the metric that quantifies the SNP interactions intensity. The experiment results demonstrate that our method has a high success rate with both balanced/unbalanced simulation and real dataset. In addition, pGBM shows great potential to detect pure SNP epistasis to uncover more complex disease pathogenesis. Kai Che, Maozu Guo 0001, Lei Wang 0085, Yin Zhang 0009 |
BIBM | 5 |
| 2016 | CIDExtractor: A chemical-induced disease relation extraction system for biomedical literatureabstractAdverse drug reactions between chemicals and diseases make chemical-disease relations (CDR) become a research focus. In this paper, we present a chemical-induced disease (CID) relation extraction system, CIDExtractor, to extract CID relations from biomedical literature. CIDExtractor first employs a sentence-level classifier to extract the CID relations located in the same sentence. To construct the classifier, a sentence-level training set is manually annotated and then Co-Training algorithm is used to exploit the unlabeled data with the feature kernel and graph kernel as two independent views. Then CIDExtractor uses a document-level classifier to extract the CID relations spanning multiple sentences. The classifier utilizes the document level information (features) of the chemical and disease pair. Finally, some post-processing rules are applied to the union set of two classifiers and generate the final outputs. Experimental results on the test set of BioCreative V CDR CID subtask show that CIDExtractor can achieve better performance (an F-score of 67.72%) than the state-of-the-art methods. The online CIDExtractor demonstration system is available at http://202.118.75.18:8888/cdr-dut-ir/cid.html. Zhiheng Li 0004, Hongfei Lin, Jian Wang 0021, Yingyi Gui, Yin Zhang 0009, Lei Wang 0085 |
BIBM | 7 |
| 2016 | Constructing an integrated gene similarity network for the identification of disease genesabstractDiscovering novel genes that are involved in human diseases is a challenging task. In recent years, several computational approaches have been proposed to prioritize candidate disease genes. Most of these methods are mainly based on protein-protein interaction (PPI) networks. However, since these PPI networks contain false positives and only cover less half of known human genes, their reliability and coverage are both very low. Therefore, it is highly necessary to fuse multiple genomic data to construct a reliable gene similarity network and then infer disease genes on the whole genomic scale. Here, we proposed a novel method, named RWRB, to infer causal genes of interested disease. First, we construct five individual gene (protein) similarity networks based on multiple genomic data of human genes. Then, an integrated gene similarity network (IGSN) is reconstructed based on similarity network fusion (SNF) method. Finally, we employ the random walk with restart algorithm on the phenotype-gene bilayer network, which combines phenotype similarity network, IGSN as well as the phenotype-gene association network, to prioritize candidate disease genes. We investigate the effectiveness of RWRB through leave-one-out cross-validation methods in inferring phenotype-gene relationships. Results show that RWRB is more accurate than state-of-the-art methods on most evaluation metrics. Further analysis shows that the success of RWRB is benefited from IGSN which has a wider coverage and higher reliability comparing with current PPI networks. Zhen Tian 0004, Maozu Guo 0001, Chunyu Wang 0002, Linlin Xing, Lei Wang 0085, Yin Zhang 0009 |
BIBM | 5 |
| 2016 | Reconstructing gene regulatory network based on candidate auto selection methodabstractThe reconstruction of gene regulatory network (GRN) is a great challenge in systems biology and bioinformatics, and methods based on Bayesian network (BN) draw most of attention because of its inherent probability characteristics. As NP-hard problems, most of the BN methods often adopt the heuristic search, but they are time-consuming for biological networks with a large number of nodes. To solve this problem, this paper presents a Candidate Auto Selection algorithm (CAS) based on mutual information and breakpoint detection to limit the search space in order to accelerate the learning process. The proposed algorithm automatically restricts the neighbors of each node to a small set of candidates before structure learning. Then based on CAS algorithm, we propose a globally optimal greedy search method (CAS+G), which focuses on finding the high-scoring network structure, and a local learning method (CAS+L), which focuses on faster learning the structure with small loss of quality. Results show that the proposed CAS algorithm can effectively identify the neighbor nodes of each node. In the experiments, the CAS+G method outperforms the state-of-the-art method on simulation data for inferring GRNs, and the CAS+L method is significantly faster than the state-of-the-art method with little loss of accuracy. Hence, the CAS based algorithms are more suitable for GRN inference. Linlin Xing, Maozu Guo 0001, Chunyu Wang 0002, Lei Wang 0085, Yin Zhang 0009 |
BIBM | 5 |
| 2016 | ML-CNN: A novel deep learning based disease named entity recognition architectureabstractIn this paper, we present a deep learning based disease named entity recognition architecture. First, the word-level embedding, character-level embedding and lexicon feature embedding are concatenated as input. Then multiple convolutional layers are stacked over the input to extract useful features automatically. Finally, multiple label strategy, which is firstly introduced, is applied to the output layer to capture the correlation information between neighboring labels. Experimental results on both NCBI and CDR corpora show that ML-CNN can achieve the state-of-the-art performance. Zhehuan Zhao, Ling Luo 0001, Yin Zhang 0009, Lei Wang 0085, Hongfei Lin, Jian Wang 0021 |
BIBM | 5 |
| 2016 | Disease-specific protein complex detection in the human protein interaction network with a supervised learning methodabstractHigh-throughput experimental techniques have produced a large amount of human protein-protein interactions, making it possible to construct a large-scale human PPI network and detect human protein complexes from the network with computational approaches. However, most of current complex detection methods are based on graph theory which can't utilize the information of the known complexes. In this paper, we present a supervised learning method to detect protein complexes in a human PPI network. In this method, biological characteristics and properties of the network are taken into consideration to construct a rich feature set to train a regression model for protein complex detection. In addition, the specific disease related PPIs are extracted from biomedical literatures and then integrated into the original PPI network for detecting the disease-specific protein complexes more effectively. Experimental results show that the performance of our method is superior to other existing state-of-the-art methods. Furthermore, through the analysis of the breast cancer specific complexes detected with our method, more biological insights for breast cancer (e.g., some candidate susceptible genes of breast cancer) are provided. Yingyi Gui, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021 |
BIBM | 5 |