VLDB 2026 Research / reviewers in the wild / expert
Weihong Yao
dblp:68/6288
· DBLP profile ↗
10ranked-venue papers
0as first author
7since 2021 · last 2026
0009-0002-3145-3494ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Modeling temporal self and interactive evolution for biomedical hypothesis generation
Hongyun Zeng, Huiwei Zhou, Weihong Yao |
J. Biomed. Informatics | 3 |
| 2024 | Sequential and Repetitive Pattern Learning for Temporal Knowledge Graph ReasoningabstractTemporal Knowledge Graph (TKG) reasoning has received a growing interest recently, especially in forecasting the future facts based on the historical KG sequences. Existing studies typically utilize a recurrent neural network to learn the evolutional representations of entities for temporal reasoning. However, these methods are hard to capture the complex temporal evolutional patterns such as sequential and repetitive patterns accurately. To tackle this challenge, we propose a novel Sequential and Repetitive Pattern Learning (SRPL) method, which comprehensively captures both the sequential and repetitive patterns. Specifically, a Dependency-aware Sequential Pattern Learning (DSPL) component expresses the temporal dependencies of each historical timestamp as embeddings for accurately capturing the sequential patterns across temporally adjacent facts. A Time-interval guided Repetitive Pattern Learning (TRPL) component models the irregular time intervals between historical repetitive facts for capturing the repetitive patterns. Extensive experiments on four representative benchmarks demonstrate that our proposed method outperforms state-of-the-art methods in all metrics by an obvious margin, especially on GDELT dataset, where performance improvement of MRR reaches up to 18.84%. Xuefei Li 0005, Huiwei Zhou, Weihong Yao, Wenchu Li, Yingyu Lin |
LREC/COLING | 3 |
| 2024 | Temporal attention networks for biomedical hypothesis generation
Huiwei Zhou, Haibin Jiang, Lanlan Wang, Weihong Yao, Yingyu Lin |
J. Biomed. Informatics | 4 |
| 2024 | Contrasting Multi-Source Temporal Knowledge Graphs for Biomedical Hypothesis GenerationabstractHypothesis Generation (HG) aims to expedite biomedical researches by generating novel hypotheses from existing scientific literature. Most existing studies focused on modeling static snapshots of the corpus, neglecting the temporal evolution of scientific terms. Despite recent efforts to learn term evolution from Knowledge Bases (KBs) for HG, the temporal information from multi-source KBs is still overlooked, which contains important, up-to-date knowledge. In this paper, an innovative Temporal Contrastive Learning (TCL) framework is introduced to uncover latent associations between entities by jointly modeling their co-evolution across multi-source temporal KBs. Specifically, we first construct a temporal relation graph based on PubMed papers and a biomedical relation database (such as Comparative Toxicogenomics Database (CTD)). Then the constructed temporal relation graph and a temporal concept graph (such as Medical Subject Headings (MeSH)) are used to train two GCN-based recurrent networks for learning the entity temporal evolutional embeddings, respectively. Finally, a cross-view temporal prediction task is designed for learning knowledge enriched temporal embeddings by contrasting the temporal embeddings learned from the two Temporal Knowledge Graphs (TKGs). Findings from experiments conducted on three real-world biomedical term relationship datasets demonstrate that the proposed approach is clearly superior to approaches based on single TKG, achieving the state-of-the-art performance. Huiwei Zhou, Wenchu Li, Weihong Yao, Yingyu Lin |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2024 | Generating Biomedical Hypothesis With Spatiotemporal TransformersabstractGenerating biomedical hypotheses is a difficult task as it requires uncovering the implicit associations between massive scientific terms from a large body of published literature. A recent line of Hypothesis Generation (HG) approaches - temporal graph-based approaches - have shown great success in modeling temporal evolution of term-pair relationships. However, these approaches model the temporal evolution of each term or term-pair with Recurrent Neural Network (RNN) independently, which neglects the rich covariation among all terms or term-pairs while ignoring direct dependencies between any two timesteps in a temporal sequence. To address this problem, we propose a Spatiotemporal Transformer-based Hypothesis Generation (STHG) method to interleave spatial covariation and temporal progression in a unified framework for constructing direct connections between any two term-pairs while modeling the temporal relevance between any two timesteps. Experiments on three biomedical relationship datasets show that STHG outperforms the state-of-the-art methods. Huiwei Zhou, Lanlan Wang, Weihong Yao, Wenchu Li, Hongyun Zeng |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Intricate Spatiotemporal Dependency Learning for Temporal Knowledge Graph ReasoningabstractKnowledge Graph (KG) reasoning has been an interesting topic in recent decades. Most current researches focus on predicting the missing facts for incomplete KG. Nevertheless, Temporal KG (TKG) reasoning, which is to forecast future facts, still faces with a dilemma due to the complex interactions between entities over time. This article proposes a novel intricate Spatiotemporal Dependency learning Network (STDN) based on Graph Convolutional Network (GCN) to capture the underlying correlations of an entity at different timestamps. Specifically, we first learn an adaptive adjacency matrix to depict the direct dependencies from the temporally adjacent facts of an entity, obtaining its previous context embedding. Then, a Spatiotemporal feature Encoding GCN (STE-GCN) is proposed to capture the latent spatiotemporal dependencies of the entity, getting the spatiotemporal embedding. Finally, a time gate unit is used to integrate the previous context embedding and the spatiotemporal embedding at the current timestamp to update the entity evolutional embedding for predicting future facts. STDN could generate the more expressive embeddings for capturing the intricate spatiotemporal dependencies in TKG. Extensive experiments on WIKI, ICEWS14, and ICEWS18 datasets prove our STDN has the advantage over state-of-the-art baselines for the temporal reasoning task. Xuefei Li 0005, Huiwei Zhou, Weihong Yao, Wenchu Li, Baojie Liu, Yingyu Lin |
ACM Trans. Knowl. Discov. Data | 3 |
| 2022 | Learning temporal difference embeddings for biomedical hypothesis generationabstractMOTIVATION: Hypothesis generation (HG) refers to the discovery of meaningful implicit connections between disjoint scientific terms, which is of great significance for drug discovery, prediction of drug side effects and precision treatment. More recently, a few initial studies attempt to model the dynamic meaning of the terms or term pairs for HG. However, most existing methods still fail to accurately capture and utilize the dynamic evolution of scientific term relations. RESULTS: This article proposes a novel temporal difference embedding (TDE) learning framework to model the temporal difference information evolution of term-pair relations for predicting future interactions. Specifically, the HG problem is formulated as a future connectivity prediction task on a temporal sequence of a dynamic attributed graph. Our approach models both the local neighbor changes of the term-pairs and the changes of the global graph structure over time, learning local and global TDE of node-pairs, respectively. Future term-pair relations can be inferred in a recurrent network based on the local and global TDE. Experiments on three real-world biomedical term relationship datasets show the effectiveness and superiority of the proposed approach. AVAILABILITY AND IMPLEMENTATION: The data and source codes related to TDE are publicly available at https://github.com/Huiweizhou/TDE. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Huiwei Zhou, Haibin Jiang, Weihong Yao, Xun Du |
Bioinform. | 3 |
| 2020 | Global Context-enhanced Graph Convolutional Networks for Document-level Relation ExtractionabstractDocument-level Relation Extraction (RE) is particularly challenging due to complex semantic interactions among multiple entities in a document.Among exiting approaches, Graph Convolutional Networks (GCN) is one of the most effective approaches for document-level RE.However, traditional GCN simply takes word nodes and adjacency matrix to represent graphs, which is difficult to establish direct connections between distant entity pairs.In this paper, we propose Global Context-enhanced Graph Convolutional Networks (GCGCN), a novel model which is composed of entities as nodes and context of entity pairs as edges between nodes to capture rich global context information of entities in a document.Two hierarchical blocks, Context-aware Attention Guided Graph Convolution (CAGGC) for partially connected graphs and Multi-head Attention Guided Graph Convolution (MAGGC) for fully connected graphs, could take progressively more global context into account.Meantime, we leverage a large-scale distantly supervised dataset to pre-train a GCGCN model with curriculum learning, which is then fine-tuned on the human-annotated dataset for further improving document-level RE performance.The experimental results on DocRED show that our model could effectively capture rich global context information in the document, leading to a state-of-the-art result. Huiwei Zhou, Yibin Xu, Weihong Yao, Zhe Liu 0020, Chengkun Lang, Haibin Jiang |
COLING | 3 |
| 2019 | The Robust Classification Model Based on Combinatorial FeaturesabstractAnalyzing the disease data from the view of combinatorial features may better characterize the disease phenotype. In this study, a novel method is proposed to construct feature combinations and a classification model (CFC-CM) by mining key feature relationships. CFC-CM iteratively tests for differences in the feature relationship between different groups. To do this, it uses a modified $k$k-top-scoring pair (M-$k$k-TSP) algorithm and then selects the most discriminative feature pairs in the current feature set to infer the combinatorial features and build the classification model. Compared with support vector machines, random forests, least absolute shrinkage and selection operator, elastic net, and M-$k$k-TSP, the superior performance of CFC-CM on nine public gene expression datasets validates its potential for more precise identification of complex diseases. Subsequently, CFC-CM was applied to two metabolomics datasets, it obtained accuracy rates of $88.73\pm 2.06\%$88.73±2.06% and $79.11\pm 2.70\%$79.11±2.70% in distinguishing between hepatocellular carcinoma and hepatic cirrhosis groups and between acute kidney injury (AKI) and non-AKI samples, results superior to those of the other five methods. In summary, the better results of CFC-CM show that in contrast to molecules and combinations constituted by just two features, the combinations inferred by appropriate number of features could better identify the complex diseases. Xiaohui Lin 0002, Xin Huang 0015, Lina Zhou, Weihong Yao, Xingyuan Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2016 | The feature selection algorithm based on feature overlapping and group overlappingabstractIn systems biology, filtering the discriminative features from complex high-dimensional data is a crucial issue. This paper proposes a feature selection algorithm based on feature overlapping and group overlapping (FS-FOGO) to calculate the feature importance. FS-FOGO weighs feature from two aspects: overlapping degree based on the ratio of overlapping area on the effective range of each class and the overlapping degree based on the proportion of heterogeneous samples in every sample's nearest neighbors. To show the validation of FS-FOGO, it is compared with effective range based gene selection (ERGS), which calculates the feature weights based on overlapping area of the effective range, on six public biological data sets and one serum metabolomics data set about liver disease. Naive Bayes and Support Vector Machine are used as classifiers, respectively. The experiment results show that the top ranked features by FS-FOGO are more discriminative and get higher classification accuracy rates than those by ERGS in most cases. And in the metabolomics data, the top ranked metabolites by FS-FOGO could separate different liver diseases well. Xiaohui Lin 0002, Meng Fan, Lishuang Li, Weihong Yao |
BIBM | 6 |