Huiwei Zhou

dblp:02/7677 · DBLP profile ↗
← Back
20ranked-venue papers
14as first author
10since 2021 · last 2026
0000-0002-7766-4177ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 11 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Modeling temporal self and interactive evolution for biomedical hypothesis generation
Hongyun Zeng, Huiwei Zhou, Weihong Yao
J. Biomed. Informatics2
2024 An Improved Method for Phenotype Concept Recognition Using Rich HPO Information
abstract
Automatically identifying human phenotype ontology (HPO) concepts from text is important for disease analysis. Existing ontology-driven methods for phenotype concept recognition mainly rely on concept names and synonym information from the ontology, without fully exploiting the rich ontology information. In this paper, we present an improved phenotype concept recognition method by incorporating rich HPO information. We first design prompts with HPO information and use a cutting-edge large language model GPT-4 to generate synonym augmentation for expanding distant supervised training data. We then propose an ontology vector-enhanced phenotype concept classification model to efficiently integrate the taxonomic hierarchical structure of HPO. Additionally, we employ noisy data augmentation to improve the model’s recognition ability in noisy texts and implement a negation detection function. Experimental results on three standard corpora and two typo corpora show our method compares favorably to previous methods and achieves a significant improvement in noisy texts. The source code and data are freely available at https://github.com/DUTIR-BioNLP/PhenoTagger-Updates.
Jiewei Qi, Ling Luo 0001, Jian Wang 0021, Huiwei Zhou, Hongfei Lin
BIBM5
2024 Sequential and Repetitive Pattern Learning for Temporal Knowledge Graph Reasoning
abstract
Temporal Knowledge Graph (TKG) reasoning has received a growing interest recently, especially in forecasting the future facts based on the historical KG sequences. Existing studies typically utilize a recurrent neural network to learn the evolutional representations of entities for temporal reasoning. However, these methods are hard to capture the complex temporal evolutional patterns such as sequential and repetitive patterns accurately. To tackle this challenge, we propose a novel Sequential and Repetitive Pattern Learning (SRPL) method, which comprehensively captures both the sequential and repetitive patterns. Specifically, a Dependency-aware Sequential Pattern Learning (DSPL) component expresses the temporal dependencies of each historical timestamp as embeddings for accurately capturing the sequential patterns across temporally adjacent facts. A Time-interval guided Repetitive Pattern Learning (TRPL) component models the irregular time intervals between historical repetitive facts for capturing the repetitive patterns. Extensive experiments on four representative benchmarks demonstrate that our proposed method outperforms state-of-the-art methods in all metrics by an obvious margin, especially on GDELT dataset, where performance improvement of MRR reaches up to 18.84%.
Xuefei Li 0005, Huiwei Zhou, Weihong Yao, Wenchu Li, Yingyu Lin
LREC/COLING2
2024 Temporal attention networks for biomedical hypothesis generation
Huiwei Zhou, Haibin Jiang, Lanlan Wang, Weihong Yao, Yingyu Lin
J. Biomed. Informatics1
2024 Contrasting Multi-Source Temporal Knowledge Graphs for Biomedical Hypothesis Generation
abstract
Hypothesis Generation (HG) aims to expedite biomedical researches by generating novel hypotheses from existing scientific literature. Most existing studies focused on modeling static snapshots of the corpus, neglecting the temporal evolution of scientific terms. Despite recent efforts to learn term evolution from Knowledge Bases (KBs) for HG, the temporal information from multi-source KBs is still overlooked, which contains important, up-to-date knowledge. In this paper, an innovative Temporal Contrastive Learning (TCL) framework is introduced to uncover latent associations between entities by jointly modeling their co-evolution across multi-source temporal KBs. Specifically, we first construct a temporal relation graph based on PubMed papers and a biomedical relation database (such as Comparative Toxicogenomics Database (CTD)). Then the constructed temporal relation graph and a temporal concept graph (such as Medical Subject Headings (MeSH)) are used to train two GCN-based recurrent networks for learning the entity temporal evolutional embeddings, respectively. Finally, a cross-view temporal prediction task is designed for learning knowledge enriched temporal embeddings by contrasting the temporal embeddings learned from the two Temporal Knowledge Graphs (TKGs). Findings from experiments conducted on three real-world biomedical term relationship datasets demonstrate that the proposed approach is clearly superior to approaches based on single TKG, achieving the state-of-the-art performance.
Huiwei Zhou, Wenchu Li, Weihong Yao, Yingyu Lin
IEEE ACM Trans. Comput. Biol. Bioinform.1
2024 Generating Biomedical Hypothesis With Spatiotemporal Transformers
abstract
Generating biomedical hypotheses is a difficult task as it requires uncovering the implicit associations between massive scientific terms from a large body of published literature. A recent line of Hypothesis Generation (HG) approaches - temporal graph-based approaches - have shown great success in modeling temporal evolution of term-pair relationships. However, these approaches model the temporal evolution of each term or term-pair with Recurrent Neural Network (RNN) independently, which neglects the rich covariation among all terms or term-pairs while ignoring direct dependencies between any two timesteps in a temporal sequence. To address this problem, we propose a Spatiotemporal Transformer-based Hypothesis Generation (STHG) method to interleave spatial covariation and temporal progression in a unified framework for constructing direct connections between any two term-pairs while modeling the temporal relevance between any two timesteps. Experiments on three biomedical relationship datasets show that STHG outperforms the state-of-the-art methods.
Huiwei Zhou, Lanlan Wang, Weihong Yao, Wenchu Li, Hongyun Zeng
IEEE J. Biomed. Health Informatics1
2024 Intricate Spatiotemporal Dependency Learning for Temporal Knowledge Graph Reasoning
abstract
Knowledge Graph (KG) reasoning has been an interesting topic in recent decades. Most current researches focus on predicting the missing facts for incomplete KG. Nevertheless, Temporal KG (TKG) reasoning, which is to forecast future facts, still faces with a dilemma due to the complex interactions between entities over time. This article proposes a novel intricate Spatiotemporal Dependency learning Network (STDN) based on Graph Convolutional Network (GCN) to capture the underlying correlations of an entity at different timestamps. Specifically, we first learn an adaptive adjacency matrix to depict the direct dependencies from the temporally adjacent facts of an entity, obtaining its previous context embedding. Then, a Spatiotemporal feature Encoding GCN (STE-GCN) is proposed to capture the latent spatiotemporal dependencies of the entity, getting the spatiotemporal embedding. Finally, a time gate unit is used to integrate the previous context embedding and the spatiotemporal embedding at the current timestamp to update the entity evolutional embedding for predicting future facts. STDN could generate the more expressive embeddings for capturing the intricate spatiotemporal dependencies in TKG. Extensive experiments on WIKI, ICEWS14, and ICEWS18 datasets prove our STDN has the advantage over state-of-the-art baselines for the temporal reasoning task.
Xuefei Li 0005, Huiwei Zhou, Weihong Yao, Wenchu Li, Baojie Liu, Yingyu Lin
ACM Trans. Knowl. Discov. Data2
2022 Learning temporal difference embeddings for biomedical hypothesis generation
abstract
MOTIVATION: Hypothesis generation (HG) refers to the discovery of meaningful implicit connections between disjoint scientific terms, which is of great significance for drug discovery, prediction of drug side effects and precision treatment. More recently, a few initial studies attempt to model the dynamic meaning of the terms or term pairs for HG. However, most existing methods still fail to accurately capture and utilize the dynamic evolution of scientific term relations. RESULTS: This article proposes a novel temporal difference embedding (TDE) learning framework to model the temporal difference information evolution of term-pair relations for predicting future interactions. Specifically, the HG problem is formulated as a future connectivity prediction task on a temporal sequence of a dynamic attributed graph. Our approach models both the local neighbor changes of the term-pairs and the changes of the global graph structure over time, learning local and global TDE of node-pairs, respectively. Future term-pair relations can be inferred in a recurrent network based on the local and global TDE. Experiments on three real-world biomedical term relationship datasets show the effectiveness and superiority of the proposed approach. AVAILABILITY AND IMPLEMENTATION: The data and source codes related to TDE are publicly available at https://github.com/Huiweizhou/TDE. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Huiwei Zhou, Haibin Jiang, Weihong Yao, Xun Du
Bioinform.1
2022 A new feature selection method based on feature distinguishing ability and network influence
Yanpeng Qi, Benzhe Su, Xiaohui Lin 0002, Huiwei Zhou
J. Biomed. Informatics4
2021 Improving the recall of biomedical named entity recognition with label re-correction and knowledge distillation
abstract
BACKGROUND: Biomedical named entity recognition is one of the most essential tasks in biomedical information extraction. Previous studies suffer from inadequate annotated datasets, especially the limited knowledge contained in them. METHODS: To remedy the above issue, we propose a novel Biomedical Named Entity Recognition (BioNER) framework with label re-correction and knowledge distillation strategies, which could not only create large and high-quality datasets but also obtain a high-performance recognition model. Our framework is inspired by two points: (1) named entity recognition should be considered from the perspective of both coverage and accuracy; (2) trustable annotations should be yielded by iterative correction. Firstly, for coverage, we annotate chemical and disease entities in a large-scale unlabeled dataset by PubTator to generate a weakly labeled dataset. For accuracy, we then filter it by utilizing multiple knowledge bases to generate another weakly labeled dataset. Next, the two datasets are revised by a label re-correction strategy to construct two high-quality datasets, which are used to train two recognition models, respectively. Finally, we compress the knowledge in the two models into a single recognition model with knowledge distillation. RESULTS: Experiments on the BioCreative V chemical-disease relation corpus and NCBI Disease corpus show that knowledge from large-scale datasets significantly improves the performance of BioNER, especially the recall of it, leading to new state-of-the-art results. CONCLUSIONS: We propose a framework with label re-correction and knowledge distillation strategies. Comparison results show that the two perspectives of knowledge in the two re-corrected datasets respectively are complementary and both effective for BioNER.
Huiwei Zhou, Zhe Liu 0020, Chengkun Lang, Yibin Xu, Yingyu Lin, Junjie Hou
BMC Bioinform.1
2020 Two-perspective Biomedical Named Entity Recognition with Weakly Labeled Data Correction
abstract
Biomedical Named Entity Recognition (BioNER) is one of the most essential tasks in biomedical information extraction. Previous studies suffer from inadequate annotation datasets, especially the limited knowledge inside. This paper proposes a two-perspective named entity recognition method with Weakly Labeled (WL) data correction. Firstly, from the perspective of coverage and accuracy, we utilize PubTator and multiple knowledge bases to construct two large-scale WL datasets, which are then revised by their corresponding label correction models respectively, obtaining two high-quality datasets. Finally, we compress the knowledge in the two datasets into a BioNER model with partial label integrating. Our approach achieves new state-of-the-art performances on three BioNER datasets.
Huiwei Zhou, Zhe Liu 0020, Chengkun Lang, Yibin Xu
BIBM1
2020 Global Context-enhanced Graph Convolutional Networks for Document-level Relation Extraction
abstract
Document-level Relation Extraction (RE) is particularly challenging due to complex semantic interactions among multiple entities in a document.Among exiting approaches, Graph Convolutional Networks (GCN) is one of the most effective approaches for document-level RE.However, traditional GCN simply takes word nodes and adjacency matrix to represent graphs, which is difficult to establish direct connections between distant entity pairs.In this paper, we propose Global Context-enhanced Graph Convolutional Networks (GCGCN), a novel model which is composed of entities as nodes and context of entity pairs as edges between nodes to capture rich global context information of entities in a document.Two hierarchical blocks, Context-aware Attention Guided Graph Convolution (CAGGC) for partially connected graphs and Multi-head Attention Guided Graph Convolution (MAGGC) for fully connected graphs, could take progressively more global context into account.Meantime, we leverage a large-scale distantly supervised dataset to pre-train a GCGCN model with curriculum learning, which is then fine-tuned on the human-annotated dataset for further improving document-level RE performance.The experimental results on DocRED show that our model could effectively capture rich global context information in the document, leading to a state-of-the-art result.
Huiwei Zhou, Yibin Xu, Weihong Yao, Zhe Liu 0020, Chengkun Lang, Haibin Jiang
COLING1
2020 Knowledge-enhanced biomedical named entity recognition and normalization: application to proteins and genes
abstract
BACKGROUND: Automated biomedical named entity recognition and normalization serves as the basis for many downstream applications in information management. However, this task is challenging due to name variations and entity ambiguity. A biomedical entity may have multiple variants and a variant could denote several different entity identifiers. RESULTS: To remedy the above issues, we present a novel knowledge-enhanced system for protein/gene named entity recognition (PNER) and normalization (PNEN). On one hand, a large amount of entity name knowledge extracted from biomedical knowledge bases is used to recognize more entity variants. On the other hand, structural knowledge of entities is extracted and encoded as identifier (ID) embeddings, which are then used for better entity normalization. Moreover, deep contextualized word representations generated by pre-trained language models are also incorporated into our knowledge-enhanced system for modeling multi-sense information of entities. Experimental results on the BioCreative VI Bio-ID corpus show that our proposed knowledge-enhanced system achieves 0.871 F1-score for PNER and 0.445 F1-score for PNEN, respectively, leading to a new state-of-the-art performance. CONCLUSIONS: We propose a knowledge-enhanced system that combines both entity knowledge and deep contextualized word representations. Comparison results show that entity knowledge is beneficial to the PNER and PNEN task and can be well combined with contextualized information in our system for further improvement.
Huiwei Zhou, Shixian Ning, Zhe Liu 0020, Chengkun Lang, Zhuang Liu 0001, Bizun Lei
BMC Bioinform.1
2019 Knowledge-guided convolutional networks for chemical-disease relation extraction
abstract
BACKGROUND: Automatic extraction of chemical-disease relations (CDR) from unstructured text is of essential importance for disease treatment and drug development. Meanwhile, biomedical experts have built many highly-structured knowledge bases (KBs), which contain prior knowledge about chemicals and diseases. Prior knowledge provides strong support for CDR extraction. How to make full use of it is worth studying. RESULTS: This paper proposes a novel model called "Knowledge-guided Convolutional Networks (KCN)" to leverage prior knowledge for CDR extraction. The proposed model first learns knowledge representations including entity embeddings and relation embeddings from KBs. Then, entity embeddings are used to control the propagation of context features towards a chemical-disease pair with gated convolutions. After that, relation embeddings are employed to further capture the weighted context features by a shared attention pooling. Finally, the weighted context features containing additional knowledge information are used for CDR extraction. Experiments on the BioCreative V CDR dataset show that the proposed KCN achieves 71.28% F1-score, which outperforms most of the state-of-the-art systems. CONCLUSIONS: This paper proposes a novel CDR extraction model KCN to make full use of prior knowledge. Experimental results demonstrate that KCN could effectively integrate prior knowledge and contexts for the performance improvement.
Huiwei Zhou, Chengkun Lang, Zhuang Liu 0001, Shixian Ning, Yingyu Lin
BMC Bioinform.1
2019 A new data analysis method based on feature linear combination
Xiaohui Lin 0002, Huiwei Zhou
J. Biomed. Informatics6
2019 Knowledge-aware attention network for protein-protein interaction extraction
Huiwei Zhou, Zhuang Liu 0001, Shixian Ning, Chengkun Lang, Yingyu Lin
J. Biomed. Informatics1
2019 Combining Context and Knowledge Representations for Chemical-Disease Relation Extraction
abstract
Automatically extracting the relationships between chemicals and diseases is significantly important to various areas of biomedical research and health care. Biomedical experts have built many large-scale knowledge bases (KBs) to advance the development of biomedical research. KBs contain huge amounts of structured information about entities and relationships, therefore plays a pivotal role in chemical-disease relation (CDR) extraction. However, previous researches pay less attention to the prior knowledge existing in KBs. This paper proposes a neural network-based attention model (NAM) for CDR extraction, which makes full use of context information in documents and prior knowledge in KBs. For a pair of entities in a document, an attention mechanism is employed to select important context words with respect to the relation representations learned from KBs. Experiments on the BioCreative V CDR dataset show that combining context and knowledge representations through the attention mechanism, could significantly improve the CDR extraction performance while achieve comparable results with state-of-the-art systems.
Huiwei Zhou, Shixian Ning, Zhuang Liu 0001, Chengkun Lang, Yingyu Lin, Degen Huang
IEEE ACM Trans. Comput. Biol. Bioinform.1
2018 Chemical-induced disease relation extraction with dependency information and prior knowledge
Huiwei Zhou, Shixian Ning, Zhuang Liu 0001, Chengkun Lang, Yingyu Lin
J. Biomed. Informatics1
2015 Learning Bilingual Sentiment Word Embeddings for Cross-language Sentiment Classification
abstract
HuiWei Zhou, Long Chen, Fulin Shi, Degen Huang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Huiwei Zhou, Long Chen 0019, Fulin Shi, Degen Huang
ACL (1)1
2014 Cross-Lingual Sentiment Classification Based on Denoising Autoencoder
Huiwei Zhou, Long Chen 0019, Degen Huang
NLPCC1