EDBT 2026 Demo / reviewers in the wild / expert
Zhehuan Zhao
dblp:123/7904
· DBLP profile ↗
32ranked-venue papers
8as first author
17since 2021 · last 2025
0000-0002-5001-1352ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ITM-BERT: A Novel Index-Based Triple Match Bert Framework for Biomedical Relation ExtractionabstractAutomatically extracting relationships between entity pairs from vast biomedical texts presents a challenging research direction with substantial practical significance. For example, the correct identification of the interaction between drugs and drugs could help doctors guide patients to take the correct medication, thereby avoiding the occurrence of medical errors. However, the same statements containing different types of entity pair relationships are easily misclassified into the same category by classifiers because they share consistent semantic information. In order to solve the problems, a new Indexbased Triple Match BERT(ITM-BERT) method is proposed in this paper. Triple loss training strategy is introduced and triple data generation decision is made. If an instance has an identical opposite instance, no index is used to form the triple data. If not, the index method is used to match the global semantic information similarity, and the triple data is constructed according to the rules. After the construction is completed, the triple loss method is used for training, with the purpose of increasing the distance between different class instances, making the distance between the same class instances closer, and finally classifying the entity relationship category correctly. In this paper, our approach is validated on widely used benchmark datasets such as DDI2013, BioInfer, and AIMed. The results show that our approach is advanced compared with other models. Bo Xu 0009, Zhikui Chen, Zhehuan Zhao, Jianhua Luo, Linlin Tian |
BIBM | 4 |
| 2025 | Multi-View Community-Contrastive Graph Attention Network for Fmri-Based Alzheimer's Disease ClassificationabstractFunctional brain network analysis based on fMRI is a vital tool for understanding neural mechanisms and diagnosing neurological disorders. However, existing approaches often overlook the joint modeling of topological structures and attribute information in brain connectivity graphs, limiting their performance in disease classification. To address this issue, we propose a novel Graph Neural Network framework combining multi-view modeling and supervised contrastive learning for fMRI-based Alzheimer's disease classification. Specifically, our method employs a Graph Attention Network (GAT) to encode structural relationships and a Multi-Layer Perceptron (MLP) to capture node attribute features. Furthermore, unsupervised clustering is utilized to extract community-level representations, capturing mesoscale brain network organization. To enhance feature robustness and discriminability, we introduce a dual-view supervised contrastive learning strategy. Extensive experiments on a cohort of 480 subjects from the Alzheimer's Disease Neuroimaging Initiative (ADNI) demonstrate that our proposed model consistently outperforms state-of-the-art methods across key metrics, including accuracy, recall, F1-score, and AUC, highlighting its robustness and effectiveness for Alzheimer's disease diagnosis. Bo Xu 0008, Baijiang Xu, Zihan Yuan, Jinshi Yu, Zhehuan Zhao, Lin Lin 0008 |
BIBM | 7 |
| 2025 | NaviPath: A Novel Knowledge Graph-Based RAG Framework for Medical QAabstractLarge Language Models (LLMs) have shown strong abilities in language understanding and reasoning, drawing increasing attention in medical question answering (QA). While retrieval-augmented generation (RAG) methods improve factual accuracy, existing approaches still struggle to retrieve and organize relevant knowledge effectively. To overcome this, we propose NaviPath, a knowledge graph-based RAG framework for medical QA. It enhances LLM responses through a structured prompt built in three steps: (1) extended entity retrieval, (2) multiperspective reasoning path construction, and (3) natural language transformation for better comprehension. Experiments on two medical QA benchmarks show that NaviPath achieves state-of-the-art performance in diagnostic accuracy and factual consistency. The implementation is available at https://github.com/zyr319/Navipath.git. Zhehuan Zhao, Yuran Zhang, Bo Xu 0008, Ludan Zhang, Yu Liu 0035, Shimin Shan, Jian Wang 0021, Hongfei Lin |
BIBM | 1 |
| 2025 | Multi-event Temporal Relation Extraction by Ranking
Zhehuan Zhao, Bo Xu 0009 |
NLPCC (4) | 1 |
| 2024 | Biomedical Event Extraction as Semantic SegmentationabstractIn the biomedical field, information is widely distributed across numerous pieces of literature. Extracting events between entities from biomedical texts has garnered significant attention in recent years. However, previous research primarily focus on extracting flat biomedical events, with less attention given to nested biomedical events. Moreover, existing methods for extracting nested events often overlook the long-distance dependencies and global information between trigger words and arguments within events, and they lack sufficient interaction with event type information. To address these issues, we propose a semantic segmentation-based method for extracting nested biomedical events. We introduce U-Net to capture global information and interdependencies between event entities. Additionally, we map event types to natural language text and combine them with sentences for encoding to enhance interaction. We also employ two auxiliary tasks to improve the identification of trigger words and arguments. Finally, events are extracted by identifying the four vertices of the segmented region. Experimental results on two benchmark datasets show that our method excels in recognizing nested biomedical events and outperforms current state-of-the-art methods. Liangyu Gao, Jinzhong Ning, Lei Wang 0085, Yin Zhang 0009, Ling Luo 0001, Bo Xu 0009, Jian Wang 0021, Zhehuan Zhao, Yuanyuan Sun 0002, Hongfei Lin |
BIBM | 12 |
| 2024 | Document-level Biomedical Relation Extraction Based on Relation-guided Entity-level GraphsabstractThe task of document-level biomedical relation extraction involves identifying relational facts between entities across sentences, given specific entities. However, most current methods overlook the associations between entity pairs and generate fixed entity representations merely through mentions, leading to irrelevant mentions interfering with the determination of relational facts. Additionally, these methods fail to consider the global information and dependencies between relational entities. To address these issues, we propose a document-level relation extraction model based on relation-guided entity-level graphs. Our model aggregates all mentions of the same entity through a relation-guided attention mechanism to obtain flexible entity representations. Furthermore, by using U-Net to generate entity-level feature graphs, it facilitates global interactions and dependency capture between entity pairs. Experimental results on two benchmark datasets demonstrate the advantages of our approach in document-level biomedical relation extraction. Liangyu Gao, Haixin Tan, Lei Wang 0085, Yin Zhang 0009, Ling Luo 0001, Bo Xu 0009, Jian Wang 0021, Zhehuan Zhao, Yuanyuan Sun 0002, Hongfei Lin |
BIBM | 11 |
| 2024 | Document Embeddings Enhance Biomedical Retrieval-Augmented GenerationabstractLarge language models (LLMs) perform well in many NLP tasks but frequently generate inaccurate information in the biomedical domain, due to hallucination issues. Retrieval-Augmented Generation (RAG) has been introduced to address this issue by integrating external knowledge, enhancing the factual accuracy of outputs. However, naive RAG encounters challenges in effectively utilizing retrieved content, particularly in specialized domains like biomedicine. LLMs often struggle to integrate retrieved content as irrelevant information can interfere with the model’s judgment. Even if relevant documents are retrieved, the model may be unable to accurately comprehend and utilize the domain-specific features due to its inherent knowledge limitations. To overcome these limitations, we propose Document Embeddings Enhanced Biomedical RAG (DEEB-RAG), a framework that incorporates document embeddings along with the original retrieved text. DEEB-RAG uses MedCPT to generate document embeddings and these embeddings are then aligned with the LLM’s semantic space using a two-stage training process on a simple projector. Experimental results on biomedical QA datasets show that DEEB-RAG improves accuracy, with an average performance increase of 2.3% over naive RAG. This demonstrates DEEB-RAG’s ability to mitigate the challenges of utilizing complex biomedical information, thereby enhancing the reliability and effectiveness of LLMs in biomedical domain. Yongle Kong, Ling Luo 0001, Zeyuan Ding, Lei Wang 0085, Yin Zhang 0009, Bo Xu 0009, Jian Wang 0021, Yuanyuan Sun 0002, Zhehuan Zhao, Hongfei Lin |
BIBM | 11 |
| 2024 | Biomedical Document-level Relation Extraction with Coreference and Anaphor GraphsabstractBiomedical document-level relation extraction is a crucial technology for mining the biomedical relationships necessary for clinical diagnosis, treatment, and medical discovery. Although existing intrasentential relation extraction methods have achieved significant results, the complexity and scattered nature of information in biomedical literature require relation extraction techniques to effectively handle cross-sentence information. For example, existing methods have not been able to explicitly model the phenomena of coreference and anaphor in documents, thus affecting the model’s understanding of complex semantics within the document. To address this issue, we propose a new document-level relation extraction model with coreference and anaphor graphs. By abstracting the document into an undirected graph that includes coreference and anaphor information, the framework effectively models the interactions between entities and leverages graph convolutional network in conjunction with pretrained language model to dynamically understand graph structures. Additionally, the shift from fine-grained entity-pair level to coarse-grained document-level training and inference significantly enhances the model’s efficiency while maintaining high extraction performance. Extensive experiments demonstrate that our model achieves a 5.3% increase in F1-score over baseline models on the BioRED dataset with higher efficiency, confirming its effectiveness in handling relation extraction tasks in complex biomedical literature. Jiru Li, Yuanyuan Sun 0002, Ling Luo 0001, Lei Wang 0085, Yin Zhang 0009, Bo Xu 0009, Jian Wang 0021, Zhehuan Zhao, Hongfei Lin |
BIBM | 11 |
| 2024 | Efficient Knowledge Graph Embedding Framework to Alleviate Data Sparsity for Polypharmacy Side Effects PredictionabstractPolypharmacy is the combined use of multiple drugs for the treatment of diseases, which also often comes with a higher risk of side effects. In the medical industry, acquiring rich and comprehensive information about the side effects of multiple drug therapy becomes a crucial task. However, data collection for many side effects is often sparse, so the features of these data cannot be adequately learned, resulting in poor performance in side effects prediction. In this paper, we propose a framework based on knowledge graph embedding (KGE) models which improves KGE by using LTE operations and subsampling methods (called LTESampleKGE). LTESampleKGE consists of two main modules i.e., Entity embedding enhancement module and KGE subsampling module. The former applies linear transformation to entity representation instead of GCN structure to enhance entity embedding, while the latter utilizes subsampling methods for KGE negative sampling (NS) loss to pay more attention to sparse data. Thus, LTESampleKGE can effectively alleviate the problem of data sparsity in the polypharmacy side effects prediction task. Experimental evaluations indicate that our method demonstrates superior performance compared with baseline models. For example, LTESampleKGE outperforms MSTE by 1.20% in PR-AUC score on TWOSIDES dataset and by 0.46% in AP@n score on Drugbank dataset. Senbo Tu, Lei Wang 0085, Yin Zhang 0009, Ling Luo 0001, Bo Xu 0009, Jian Wang 0021, Zhehuan Zhao, Hongfei Lin |
BIBM | 10 |
| 2024 | Beyond Linguistic Cues: Fine-grained Conversational Emotion Recognition via Belief-Desire ModellingabstractEmotion recognition in conversation (ERC) is essential for dialogue systems to identify the emotions expressed by speakers. Although previous studies have made significant progress, accurate recognition and interpretation of similar fine-grained emotion properly accounting for individual variability remains a challenge. One particular under-explored area is the role of individual beliefs and desires in modelling emotion. Inspired by the Belief-Desire Theory of Emotion, we propose a novel method for conversational emotion recognition that incorporates both belief and desire to accurately identify emotions. We extract emotion-eliciting events from utterances and construct graphs that represent beliefs and desires in conversations. By applying message passing between nodes, our graph effectively models the utterance context, speaker’s global state, and the interaction between emotional beliefs, desires, and utterances. We evaluate our model’s performance by conducting extensive experiments on four popular ERC datasets and comparing it with multiple state-of-the-art models. The experimental results demonstrate the superiority of our proposed model and validate the effectiveness of each module in the model. Bo Xu 0009, Longjiao Li, Wei Luo 0001, Mehdi Naseriparsa, Zhehuan Zhao, Hongfei Lin, Feng Xia 0001 |
LREC/COLING | 5 |
| 2024 | ESCP: Enhancing Emotion Recognition in Conversation with Speech and Contextual PrefixesabstractEmotion Recognition in Conversation (ERC) aims to analyze the speaker’s emotional state in a conversation. Fully mining the information in multimodal and historical utterances plays a crucial role in the performance of the model. However, recent works in ERC focus on historical utterances modeling and generally concatenate the multimodal features directly, which neglects mining deep multimodal information and brings redundancy at the same time. To address the shortcomings of existing models, we propose a novel model, termed Enhancing Emotion Recognition in Conversation with Speech and Contextual Prefixes (ESCP). ESCP employs a directed acyclic graph (DAG) to model historical utterances in a conversation and incorporates a contextual prefix containing the sentiment and semantics of historical utterances. By adding speech and contextual prefixes, the inter- and intra-modal emotion information is efficiently modeled using the prior knowledge of the large-scale pre-trained model. Experiments conducted on several public benchmarks demonstrate that the proposed approach achieves state-of-the-art (SOTA) performances. These results affirm the effectiveness of the novel ESCP model and underscore the significance of incorporating speech and contextual prefixes to guide the pre-trained model. Xiujuan Xu, Xiaoxiao Shi, Zhehuan Zhao, Yu Liu 0035 |
LREC/COLING | 3 |
| 2024 | MLGAT: Multi-Scale Line Graph Attention Network for Emotion Recognition in ConversationabstractEmotion Recognition in Conversation (ERC) plays an important role in intelligent human-computer interaction. Highly accurate emotion recognition helps to improve the ability of machines to serve humans. Recent works exhibit poor generalization ability and low recognition accuracy, and better performances can only be presented on specific datasets. To be able to adapt to complex real-world application scenarios, we propose a novel model, termed Multi-scale Line Graph ATtention Network for Emotion Recognition in Conversation (MLGAT). MLGAT mines the emotional information dependent on the target utterance by focusing on local context and global context at different scales. Experiments show that our model achieves the second-highest performance among all current methods on both IEMOCAP and MELD datasets. Wa-F1 scores are 71.49% and 75.08%, respectively, which are only slightly different from their respective SOTAs (state of the art). Moreover, the model can show high accuracy in a few classes without additional measures. Xiujuan Xu, Xiaoxiao Shi, Zhehuan Zhao, Yu Liu 0035 |
ECAI | 3 |
| 2024 | Semantics Driven Multi-View Knowledge Graph Embedding for Cross-Lingual Entity AlignmentabstractCross-lingual entity alignment (EA) is a critical step in the integration of multilingual knowledge, which aims to match entities with the same meaning in different knowledge graphs (KGs). Recently, based on GCN models and pre-trained language models (PLMs), EA has achieved breakthrough performance by utilizing graph structures and auxiliary semantic information. However, existing EA methods rely heavily on artificially exploring and designing the interaction of graph structures and auxiliary semantic information, which limits their applicability in real-world situations. In this work, we proposed a simple but effective Semantics Driven Multi-view Knowledge Graph Embedding for cross-lingual entity alignment (SDMKGE). Our proposed SDMKGE utilizes two Siamese Networks based on PLMs to encode the semantics of entities and structures separately, which effectively reduces the difficulty of feature aggregation. We use three well-known datasets to evaluate our SDMKGE. Experimental results demonstrate that our framework outperforms the state-of-the-art EA methods. Xin Zhang 0136, Yu Liu 0035, Zhehuan Zhao |
ICASSP | 3 |
| 2023 | Biomedical Named Entity Recognition Through Deep Reinforcement LearningabstractBiomedical Named Entity Recognition (BioNER) is a crucial task in extracting entities from biomedical literature. It plays a key role as the initial step in various biomedical information extraction tasks. Pre-training models have gained popularity in BioNER. However, these models require a large number of parameters, which poses a barrier for many researchers due to the hardware resource requirements. To address this challenge, we propose REIN-NER, a novel reinforcement learning-based approach for BioNER. REIN-NER employs a two-stage training mechanism: Firstly, the basic named recognition model (BaseNER) is pre-trained to capture the mapping knowledge from input words to the corresponding labels. Secondly, after initializing the agent model (AgentNER) with the BaseNER, the reinforcement training is carried out to further obtain the dependency knowledge between the output labels. Both BaseNER and AgentNER are built with lightweight Bi-LSTM networks, which significantly reduce parameter sizes compared to pre-training models. Remarkably, REIN-NER achieves superior performance with F-scores of 84.04% and 93.34% on the CDR-disease and CDR-chemical corpora, respectively, outperforming BERT-based models. Our work presents a pioneering exploration of reinforcement learning in biomedical named entity recognition and demonstrates its effectiveness through state-of-the-art results. Zhehuan Zhao, Bo Xu 0009, Yuying Zou, Hongfei Lin |
BIBM | 1 |
| 2023 | Dual Relation-Aware Entity Alignment for Knowledge GraphabstractEntity alignment (EA) is a vital step for knowledge fusion, which aims to discover entities with the same meaning from the different knowledge graphs. Several researchers attempted to obtain enhanced entity embeddings by using the strategies that relation-aware entity embeddings for EA. However, they ignored the effect of aggregating different range neighborhoods of entities on relation embeddings. In this study, we propose a novel Dual Relation-aware Entity Alignment framework named DRAEA for EA. Specifically, we utilize entity embeddings based on 1-hop neighborhoods and 2-hop neighbor-hoods to represent different relation embeddings, respectively. Then we aggregate different relation embeddings back to entity embeddings by using the self-attention mechanism. Last, an effective entity-assisted global alignment strategy is designed to accomplish the EA tasks. Experimental results on three real-world datasets show that DRAEA outperforms the state-of-the-art methods. Xin Zhang 0136, Yu Liu 0035, Zhehuan Zhao |
IJCNN | 3 |
| 2022 | MET-Meme: A Multimodal Meme Dataset Rich in MetaphorsabstractMemes have become the popular means of communication for Internet users worldwide. Understanding the Internet meme is one of the most tricky challenges in natural language processing (NLP) tasks due to its convenient non-standard writing and network vocabulary. Recently, many linguists suggested that memes contain rich metaphorical information. However, the existing researches ignore this key feature. Therefore, to incorporate informative metaphors into the meme analysis, we introduce a novel multimodal meme dataset called MET-Meme, which is rich in metaphorical features. It contains 10045 text-image pairs, with manual annotations of the metaphor occurrence, sentiment categories, intentions, and offensiveness degree. Moreover, we propose a range of strong baselines to demonstrate the importance of combining metaphorical features for meme sentiment analysis and semantic understanding tasks, respectively. MET-Meme, and its code are released publicly for research in \urlhttps://github.com/liaolianfoka/MET-Meme-A-Multi-modal-Meme-Dataset-Rich-in-Metaphors. Bo Xu 0009, Junzhe Zheng, Mehdi Naseriparsa, Zhehuan Zhao, Hongfei Lin, Feng Xia 0001 |
SIGIR | 5 |
| 2021 | TL-BERT: A Novel Biomedical Relation Extraction ApproachabstractAutomatically extracting entity-pair interactions from biomedical literature plays an important role in promoting the development of the biomedical field. For instance, the interactions between drugs can guide patients to take drugs correctly and avoid clinical adverse drug reactions; The interactions between proteins can help researchers design therapeutic drugs and discover disease mechanisms. However, it is found that relation instances of different classes generated from the same sentence are easily classified into the same class since their context information is almost the same. To address this issue, Triplet Loss based BERT (TL-BERT) approach is proposed in this paper, where the Triplet Loss training strategy is first introduced into the biomedical relation extraction field. Triplet Loss training strategy will increase the distances between these instances generated from the same sentence but belonging to different classes, and decrease the distances between these instances generated from different sentences but belonging to the same class. As a result, our approach can classify these instances generated from the same sentence more correctly. TL-BERT was evaluated on AIMed, BioInfer, and DDI Extraction-2013 corpus, the experimental results demonstrate that the Triplet Loss training strategy can improve the performance on both Protein-Protein interactions extraction tasks and Drug-Drug interactions detection tasks. Zhehuan Zhao, Yuying Zou, Bo Xu 0009, Jian Wang 0021, Hongfei Lin, Shimin Shan, Yu Liu 0035 |
BIBM | 1 |
| 2020 | A Context-Based Network For Referring Image SegmentationabstractReferring image segmentation is an important task aiming at segmenting out the object referred by a natural language expression. Current works usually employ the methods of concatenating the visual and linguistic features. They underestimate the importance of language-to-vision and object-to-object relationships when the natural language expression has multiple entities. Therefore, we propose a new network named Context-Based Network(CBN) to improve the accuracy of locating the correct referent. The CBN is composed of two modules: Intra Relation Selection(Intra-RS) and Inter Relation Selection(Inter-RS). The Intra-RS can capture object-to-object relationships in an embedding visual and linguistic feature space and the Inter-RS uses the multi-scale linguistic features as a guide to match the most similar region from the image feature maps. Besides, we apply spatial pyramid pooling to get global information to solve the limited receptive field problem. Experimental results on four public datasets showed that CBN achieved comparable performance to the other state-of-art methods. Yu Liu 0035, Kaiping Xu, Zhehuan Zhao, Sipei Liu |
ICIP | 4 |
| 2020 | DRGCN: Deep Relation GCN for Group Activity Recognition
Yiqiang Feng, Shimin Shan, Yu Liu 0035, Zhehuan Zhao, Kaiping Xu |
ICONIP (4) | 4 |
| 2019 | Disease Gene Prediction Based on Heterogeneous Probabilistic Hypergraph RankingabstractIn order to save time and cost, many disease gene prediction methods have been proposed in recent years. However, the traditional network model uses a binary relationship to represent the relationship between different proteins or gene molecules and phenotypes, which leads to the loss of information. Recently, hypergraph shows that it can overcome this loss of information to some extent and preserve the multivariate relationship, so we transformed the disease gene prediction problem into the problem of ranking the multivariate-relationship object. In this paper, we propose a method of Heterogeneous Probabilistic Hypergraph Ranking (HPHR) to predict disease genes. Firstly, fix a graph centroid for each hyperedge and according to different associations, and add other nodes related to the graph centroid to hyperedges with a certain probability. Then transform the problem of predicting disease genes into the problem of ranking heterogeneous objects, and the candidate genes are sorted by hypergraph ranking. The method is then applied to the integrated disease gene network. Compared with other prediction methods achieved better results, which was verified by this experiment. Feng Ding 0004, Xiangjie Kong 0001, Zhehuan Zhao, Feng Xia 0001, Anfu Liu, Chenxu Bai, Bo Xu 0008, Shengtian Sang, Hongfei Lin, Jian Wang 0021 |
BIBM | 3 |
| 2019 | Distantly Supervised Relation Extraction through a Trade-off MechanismabstractDistantly supervised relation extraction can label large amounts of unstructured text without human annotations for training. However, distant supervision inevitably accompanies with the wrong labeling problem, which can deteriorate the performance of relation extraction. What's more, the entity-pair information, which can enrich instance information, is still underutilized. In the light of these issues, we propose TMNN, a novel Neural Network framework with a Trade-off Mechanism, which combines the feature of text and entity pair on the sentence level to predict relations. Our proposed trade-off mechanism is a probability generation module to dynamically adjust the weights of text and corresponding entity pair for each sentence. Experimental results on a widely used dataset show that the proposed method reduces the noisy labels and achieves substantial improvement over the state-of-the-art methods. Yu Liu 0035, Kai Wang 0057, Zhehuan Zhao, Quan Z. Sheng |
IJCNN | 4 |
| 2019 | Neural network-based approaches for biomedical relation classification: A reviewabstractThe explosive growth of biomedical literature has created a rich source of knowledge, such as that on protein-protein interactions (PPIs) and drug-drug interactions (DDIs), locked in unstructured free text. Biomedical relation classification aims to automatically detect and classify biomedical relations, which has great benefits for various biomedical research and applications. In the past decade, significant progress has been made in biomedical relation classification. With the advance of neural network methodology, neural network-based approaches have been applied in biomedical relation classification and achieved state-of-the-art performance for some public datasets and shared tasks. In this review, we describe the recent advancement of neural network-based approaches for classifying biomedical relations. We summarize the available corpora and introduce evaluation metrics. We present the general framework for neural network-based approaches in biomedical relation extraction and pretrained word embedding resources. We discuss neural network-based approaches, including convolutional neural networks (CNNs) and recurrent neural networks (RNNs). We conclude by describing the remaining challenges and outlining future directions. Yi-Jia Zhang 0001, Hongfei Lin, Jian Wang 0021, Yuanyuan Sun 0002, Bo Xu 0009, Zhehuan Zhao |
J. Biomed. Informatics | 7 |
| 2018 | HMNPPID: A Database of Protein-protein Interactions Associated with Human Malignant Neoplasms
Zhehuan Zhao, Ling Luo 0001, Zhiheng Li 0004, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021, Yi-Jia Zhang 0001 |
BIBM | 3 |
| 2018 | Full-attention Based Drug Drug Interaction Extraction Exploiting User-generated Content
Bo Xu 0009, Xiufeng Shi, Zhehuan Zhao, Wei Zheng 0003, Hongfei Lin, Jian Wang 0021, Feng Xia 0001 |
BIBM | 3 |
| 2018 | Protein complexes identification based on go attributed network embeddingabstractBACKGROUND: Identifying protein complexes from protein-protein interaction (PPI) network is one of the most important tasks in proteomics. Existing computational methods try to incorporate a variety of biological evidences to enhance the quality of predicted complexes. However, it is still a challenge to integrate different types of biological information into the complexes discovery process under a unified framework. Recently, attributed network embedding methods have be proved to be remarkably effective in generating vector representations for nodes in the network. In the transformed vector space, both the topological proximity and node attributed affinity between different nodes are preserved. Therefore, such attributed network embedding methods provide us a unified framework to integrate various biological evidences into the protein complexes identification process. RESULTS: In this article, we propose a new method called GANE to predict protein complexes based on Gene Ontology (GO) attributed network embedding. Firstly, it learns the vector representation for each protein from a GO attributed PPI network. Based on the pair-wise vector representation similarity, a weighted adjacency matrix is constructed. Secondly, it uses the clique mining method to generate candidate cores. Consequently, seed cores are obtained by ranking candidate cores based on their densities on the weighted adjacency matrix and removing redundant cores. For each seed core, its attachments are the proteins with correlation score that is larger than a given threshold. The combination of a seed core and its attachment proteins is reported as a predicted protein complex by the GANE algorithm. For performance evaluation, we compared GANE with six protein complex identification methods on five yeast PPI networks. Experimental results showes that GANE performs better than the competing algorithms in terms of different evaluation metrics. CONCLUSIONS: GANE provides a framework that integrate many valuable and different biological information into the task of protein complex identification. The protein vector representation learned from our attributed PPI network can also be used in other tasks, such as PPI prediction and disease gene prediction. Bo Xu 0008, Wei Zheng 0003, Yi-Jia Zhang 0001, Zhehuan Zhao, Zengyou He |
BMC Bioinform. | 6 |
| 2017 | A hybrid protein-protein interaction triple extraction method for biomedical literatureabstractProtein-protein interaction extraction research can be widely applied to the field of life science research. However, most of the machine learning based methods focus on binary PPI relation extraction, which loses rich relationship type information that is critical to the PPIs study. The rule based open information extraction methods can extract the PPI triple (i.e. “protein1, interaction word, protein2”), but suffers from low recall rate problem. In this paper, we propose a hybrid protein-protein interaction triple extraction method. In this method, firstly, machine learning techniques are used to recognize protein entities and extract relational protein pairs. Then, the syntactic patterns and a dictionary are employed to find out corresponding interaction words that represent the relationships between two proteins. This method obtains an F-score of 40.18% on the AImed corpus, which is much higher than the result achieved by the rule based Stanford open information extraction method. Zhehuan Zhao, Cong Sun 0004, Lei Wang 0085, Hongfei Lin |
BIBM | 1 |
| 2017 | An attention-based effective neural model for drug-drug interactions extractionabstractBACKGROUND: Drug-drug interactions (DDIs) often bring unexpected side effects. The clinical recognition of DDIs is a crucial issue for both patient safety and healthcare cost control. However, although text-mining-based systems explore various methods to classify DDIs, the classification performance with regard to DDIs in long and complex sentences is still unsatisfactory. METHODS: In this study, we propose an effective model that classifies DDIs from the literature by combining an attention mechanism and a recurrent neural network with long short-term memory (LSTM) units. In our approach, first, a candidate-drug-oriented input attention acting on word-embedding vectors automatically learns which words are more influential for a given drug pair. Next, the inputs merging the position- and POS-embedding vectors are passed to a bidirectional LSTM layer whose outputs at the last time step represent the high-level semantic information of the whole sentence. Finally, a softmax layer performs DDI classification. RESULTS: Experimental results from the DDIExtraction 2013 corpus show that our system performs the best with respect to detection and classification (84.0% and 77.3%, respectively) compared with other state-of-the-art methods. In particular, for the Medline-2013 dataset with long and complex sentences, our F-score far exceeds those of top-ranking systems by 12.6%. CONCLUSIONS: Our approach effectively improves the performance of DDI classification tasks. Experimental analysis demonstrates that our model performs better with respect to recognizing not only close-range but also long-range patterns among words, especially for long, complex and compound sentences. Wei Zheng 0003, Hongfei Lin, Ling Luo 0001, Zhehuan Zhao, Zhengguang Li, Yi-Jia Zhang 0001, Jian Wang 0021 |
BMC Bioinform. | 4 |
| 2016 | ML-CNN: A novel deep learning based disease named entity recognition architectureabstractIn this paper, we present a deep learning based disease named entity recognition architecture. First, the word-level embedding, character-level embedding and lexicon feature embedding are concatenated as input. Then multiple convolutional layers are stacked over the input to extract useful features automatically. Finally, multiple label strategy, which is firstly introduced, is applied to the output layer to capture the correlation information between neighboring labels. Experimental results on both NCBI and CDR corpora show that ML-CNN can achieve the state-of-the-art performance. Zhehuan Zhao, Ling Luo 0001, Yin Zhang 0009, Lei Wang 0085, Hongfei Lin, Jian Wang 0021 |
BIBM | 1 |
| 2016 | Drug drug interaction extraction from biomedical literature using syntax convolutional neural networkabstractMOTIVATION: Detecting drug-drug interaction (DDI) has become a vital part of public health safety. Therefore, using text mining techniques to extract DDIs from biomedical literature has received great attentions. However, this research is still at an early stage and its performance has much room to improve. RESULTS: In this article, we present a syntax convolutional neural network (SCNN) based DDI extraction method. In this method, a novel word embedding, syntax word embedding, is proposed to employ the syntactic information of a sentence. Then the position and part of speech features are introduced to extend the embedding of each word. Later, auto-encoder is introduced to encode the traditional bag-of-words feature (sparse 0-1 vector) as the dense real value vector. Finally, a combination of embedding-based convolutional features and traditional features are fed to the softmax classifier to extract DDIs from biomedical literature. Experimental results on the DDIExtraction 2013 corpus show that SCNN obtains a better performance (an F-score of 0.686) than other state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: The source code is available for academic use at http://202.118.75.18:8080/DDI/SCNN-DDI.zip CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Zhehuan Zhao, Ling Luo 0001, Hongfei Lin, Jian Wang 0021 |
Bioinform. | 1 |
| 2016 | A graph kernel based on context vectors for extracting drug-drug interactions
Wei Zheng 0003, Hongfei Lin, Zhehuan Zhao, Bo Xu 0009, Yi-Jia Zhang 0001, Jian Wang 0021 |
J. Biomed. Informatics | 3 |
| 2015 | Deep neural network based protein-protein interaction extraction from biomedical literatureabstractThis paper presents a deep neural network-based protein-protein interactions (PPIs) information extraction approach which can learn complex and abstract features automatically from unlabeled data by unsupervised representation learning methods. This approach first employs the training algorithm of auto-encoders to initialize the parameters of a deep multilayer neural network. Then the gradient descent method using back-propagation is applied to train this deep multilayer neural network model. Experimental results on five public PPI corpora show that our method can achieve better performance than can a multilayer neural network. In addition, the performance comparison with APG also verifies the effectiveness of our method. Zhehuan Zhao, Ling Luo 0001, Hongfei Lin, Jian Wang 0021 |
BIBM | 1 |
| 2012 | PPIExtractor: A protein-protein interaction Extractor for biomédical literatureabstractKnowledge about protein-protein interactions (PPIs) unveils the molecular mechanisms of biological processes. In this paper, we present a PPI extraction system, termed PPIExtractor, which automatically extracts PPIs from biomedical text and visualizes them. Given a Medline record dataset, PPIExtractor first applies Feature Coupling Generalization (FCG) to tag protein names, next uses the extended semantic similarity-based method to normalize them, then combines feature-based, convolution tree and graph kernels to extract PPIs, and finally visualizes the PPI network. Experimental evaluations show that PPIExtractor can achieve state-of-the-art performance on a DIP subset with respect to comparable evaluations. PPIExtractor is freely available for academic purposes at: http://202.118.75.18:8080/PPIExtractor/. Zhehuan Zhao, Yuncui Hu, Hongfei Lin |
BIBM | 2 |