Xiaowei Zhao 0004

dblp:02/8134-4 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0002-9868-5102ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 RaHDCL: Relation-aware hypergraph diffusion contrastive learning for drug-drug interaction event prediction
Xiaosa Zhao, Xiaowei Zhao 0004, Minghao Yin
Knowl. Based Syst.4
2025 CBKG-DTI: Multi-Level Knowledge Distillation and Biomedical Knowledge Graph for Drug-Target Interaction Prediction
abstract
The prediction of drug-target interactions (DTIs) has emerged as a vital step in drug discovery. Recently, biomedical knowledge graph enables the utilization of multi-omics resources for modelling complex biological systems and further improves overall performance of specific predictive task. However, due to the scale and generalization of biomedical knowledge graph, it is necessary to capture task-specific knowledge from biomedical knowledge graph for DTI prediction. Moreover, although biomedical knowledge graph has rich interactions between biological entities, there still needs to contain unignorable structural information of drugs or targets in the multi-modal fusion manner. To this end, we develop a novel DTI identification framework, CBKG-DTI, which aims to distill task-specific knowledge from the complex knowledge graph to the lightweight DTI prediction model. Specifically, CBKG-DTI first introduces a hierarchy-aware knowledge graph embedding as teacher model to capture semantic hierarchy information of biomedical knowledge graph. Then, to further improve model performance, CBKG-DTI integrates information from multiple aspects such as relational information and structural information by constructing a heterogeneous network and then employs a heterogeneous graph attention network framework as the lightweight student model. Moreover, we design a multi-level distillation mechanism to improve the representation and prediction ability of the lightweight student model via capturing the representation and logit distribution of the teacher model. Finally, we conduct the extensive comparison experiments and can reach the AUC of 0.9751 and the AUPR of 0.6310 under 5-fold cross validation. This not only demonstrates the superiority of CBKG-DTI in DTI prediction, but also, more importantly, validate the effectiveness of the framework capturing task-specific knowledge from biomedical knowledge graph.
Xiaosa Zhao, Qixian Wang, Ye Zhang 0014, Minghao Yin, Xiaowei Zhao 0004
IEEE J. Biomed. Health Informatics6
2023 Multi-view contrastive heterogeneous graph attention network for lncRNA-disease association prediction
abstract
MOTIVATION: Exploring the potential long noncoding RNA (lncRNA)-disease associations (LDAs) plays a critical role for understanding disease etiology and pathogenesis. Given the high cost of biological experiments, developing a computational method is a practical necessity to effectively accelerate experimental screening process of candidate LDAs. However, under the high sparsity of LDA dataset, many computational models hardly exploit enough knowledge to learn comprehensive patterns of node representations. Moreover, although the metapath-based GNN has been recently introduced into LDA prediction, it discards intermediate nodes along the meta-path and results in information loss. RESULTS: This paper presents a new multi-view contrastive heterogeneous graph attention network (GAT) for lncRNA-disease association prediction, MCHNLDA for brevity. Specifically, MCHNLDA firstly leverages rich biological data sources of lncRNA, gene and disease to construct two-view graphs, feature structural graph of feature schema view and lncRNA-gene-disease heterogeneous graph of network topology view. Then, we design a cross-contrastive learning task to collaboratively guide graph embeddings of the two views without relying on any labels. In this way, we can pull closer the nodes of similar features and network topology, and push other nodes away. Furthermore, we propose a heterogeneous contextual GAT, where long short-term memory network is incorporated into attention mechanism to effectively capture sequential structure information along the meta-path. Extensive experimental comparisons against several state-of-the-art methods show the effectiveness of proposed framework.The code and data of proposed framework is freely available at https://github.com/zhaoxs686/MCHNLDA.
Xiaosa Zhao, Jun Wu 0020, Xiaowei Zhao 0004, Minghao Yin
Briefings Bioinform.3
2022 Heterogeneous graph attention network based on meta-paths for lncRNA-disease association prediction
abstract
MOTIVATION: Discovering long noncoding RNA (lncRNA)-disease associations is a fundamental and critical part in understanding disease etiology and pathogenesis. However, only a few lncRNA-disease associations have been identified because of the time-consuming and expensive biological experiments. As a result, an efficient computational method is of great importance and urgently needed for identifying potential lncRNA-disease associations. With the ability of exploiting node features and relationships in network, graph-based learning models have been commonly utilized by these biomolecular association predictions. However, the capability of these methods in comprehensively fusing node features, heterogeneous topological structures and semantic information is distant from optimal or even satisfactory. Moreover, there are still limitations in modeling complex associations between lncRNAs and diseases. RESULTS: In this paper, we develop a novel heterogeneous graph attention network framework based on meta-paths for predicting lncRNA-disease associations, denoted as HGATLDA. At first, we conduct a heterogeneous network by incorporating lncRNA and disease feature structural graphs, and lncRNA-disease topological structural graph. Then, for the heterogeneous graph, we conduct multiple metapath-based subgraphs and then utilize graph attention network to learn node embeddings from neighbors of these homogeneous and heterogeneous subgraphs. Next, we implement attention mechanism to adaptively assign weights to multiple metapath-based subgraphs and get more semantic information. In addition, we combine neural inductive matrix completion to reconstruct lncRNA-disease associations, which is applied for capturing complicated associations between lncRNAs and diseases. Moreover, we incorporate cost-sensitive neural network into the loss function to tackle the commonly imbalance problem in lncRNA-disease association prediction. Finally, extensive experimental results demonstrate the effectiveness of our proposed framework.
Xiaosa Zhao, Xiaowei Zhao 0004, Minghao Yin
Briefings Bioinform.2
2022 SSKM_Succ: A Novel Succinylation Sites Prediction Method Incorporating K-Means Clustering With a New Semi-Supervised Learning Algorithm
abstract
Protein succinylation is a type of post-translational modification (PTM) that occurs on lysine sites and plays a key role in protein conformation regulation and cellular function control. When training in computational method, it is difficult to designate negative samples because of the uncertainty of non-succinylation lysine sites, and if not handled properly, it may affect the performance of computational models dramatically. Therefore, we propose a new semi-supervised learning method to identify reliable non-succinylation lysine sites as negative samples. This method, named SSKM_Succ, also employs K-means clustering to divide data into 5 clusters. Besides, information of proximal PTMs and three kinds of sequence features (grey pseudo amino acid composition, K-space and position-special amino acid propensity) are utilized to formulate protein. Then, we perform a two-step feature selection to remove redundant features and construct the optimization model for each cluster. Finally, support vector machine is applied to construct a prediction model for each cluster. Promising results are obtained by this method with an accuracy of 80.18 percent for succinylation sites on the independent testing dataset. Meanwhile, we compare the result with other existing tools, and it shows that our method is promising for predicting succinylation sites. Through analysis, we further verify that succinylated protein has potential effects on amino acid degradation and fatty acid metabolism, and speculate that protein succinylation may be closely related to neurodegenerative diseases. The code of SSKM_Succ is available on the web https://github.com/yangyq505/SSKM_Succ.git.
Zhiqiang Ma 0003, Xiaowei Zhao 0004, Minghao Yin
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 A Novel Method for Identification of Glutarylation Sites Combining Borderline-SMOTE With Tomek Links Technique in Imbalanced Data
abstract
Glutarylation is a type of post-translational modification that occurs on lysine residues. It plays an irreplaceable role in various cellular functions. Therefore, identification of glutarylation sites is significant for understanding the molecular mechanism of glutarylation. In this study, we proposed a method named DEXGB_Glu to identify lysine glutarylation sites using XGBoost as classifier which was optimized by differential evolution algorithm. Aiming at the imbalance between positive samples and negative samples, Borderline-SMOTE method was employed to synthesize positive samples, increasing their amount equal to negative samples. Then, Tomek links technique was applied to filter out noise data. Analysis of this method and its results showed that differential evolution algorithm obviously improved the performance and the combination of Borderline-SMOTE and Tomek links effectively solved the imbalance between positive samples and negative samples. Finally, the performance of this method was much better than other methods in prediction of glutarylation sites. The data and code are available on https://github.com/ningq669/DEXGB_Glu.
Xiaowei Zhao 0004, Zhiqiang Ma 0003
IEEE ACM Trans. Comput. Biol. Bioinform.2
2021 An Ensemble Deep Learning based Predictor for Simultaneously Identifying Protein Ubiquitylation and SUMOylation Sites
abstract
BACKGROUND: Several computational tools for predicting protein Ubiquitylation and SUMOylation sites have been proposed to study their regulatory roles in gene location, gene expression, and genome replication. However, existing methods generally rely on feature engineering, and ignore the natural similarity between the two types of protein translational modification. This study is the first all-in-one deep network to predict protein Ubiquitylation and SUMOylation sites from protein sequences as well as their crosstalk sites simultaneously. Our deep learning architecture integrates several meta classifiers that apply deep neural networks to protein sequence information and physico-chemical properties, which were trained on multi-label classification mode for simultaneously identifying protein Ubiquitylation and SUMOylation as well as their crosstalk sites. RESULTS: The promising AUCs of our method on Ubiquitylation, SUMOylation and crosstalk sites achieved 0.838, 0.888, and 0.862 respectively on tenfold cross-validation. The corresponding APs reached 0.683, 0.804 and 0.552, which also validated our effectiveness. CONCLUSIONS: The proposed architecture managed to classify ubiquitylated and SUMOylated lysine residues along with their crosstalk sites, and outperformed other well-known Ubiquitylation and SUMOylation site prediction tools.
Fei He 0003, Xiaowei Zhao 0004
BMC Bioinform.4
2021 MHRWR: Prediction of lncRNA-Disease Associations Based on Multiple Heterogeneous Networks
abstract
In the last few years, accumulating evidences had demonstrated that long non-coding RNAs (lncRNAs) participated in the regulation of target gene expression and played an important role in biological processes and human disease development. Thus, prediction of the associations between lncRNAs and disease had become a hot research in the fields of human sophisticated diseases. Most of these methods considered the information of two networks (lncRNA, disease) while neglected other networks. In this study, we designed a multi-layer network by integrating the similarity networks of lncRNAs, diseases and genes, and the known association networks of lncRNA-disease, lncRNAs-gene, and disease-gene, and then we developed a model called MHRWR for predicting the lncRNA-disease potential associations based on random walk with restart. The performance of MHRWR was evaluated by experimentally verified lncRNA-disease associations based on leave-one-out cross validation. MHRWR obtained a reliable AUC value of 0.91344, which significantly outperformed some previous methods. To further validate the reproducibility of performance, we used the model of MHRWR to verify related lncRNAs of colon cancer, colorectal cancer and lung adenocarcinoma in the case studies. The codes of MHRWR is available on: https://github.com/yangyq505/MHRWR.
Xiaowei Zhao 0004, Yiqin Yang, Minghao Yin
IEEE ACM Trans. Comput. Biol. Bioinform.1
2021 Exploring BCI Control in Smart Environments: Intention Recognition Via EEG Representation Enhancement Learning
abstract
The brain–computer interface (BCI) control technology that utilizes motor imagery to perform the desired action instead of manual operation will be widely used in smart environments. However, most of the research lacks robust feature representation of multi-channel EEG series, resulting in low intention recognition accuracy. This article proposes an EEG2Image based Denoised-ConvNets (called EID) to enhance feature representation of the intention recognition task. Specifically, we perform signal decomposition, slicing, and image mapping to decrease the noise from the irrelevant frequency bands. After that, we construct the Denoised-ConvNets structure to learn the colorspace and spatial variations of image objects without cropping new training images precisely. Toward further utilizing the color and spatial transformation layers, the colorspace and colored area of image objects have been enhanced and enlarged, respectively. In the multi-classification scenario, extensive experiments on publicly available EEG datasets confirm that the proposed method has better performance than state-of-the-art methods.
Lin Yue, Sen Wang 0001, Robert Boots, Guodong Long, Weitong Chen 0001, Xiaowei Zhao 0004
ACM Trans. Knowl. Discov. Data7
2020 Simplifying Reinforced Feature Selection via Restructured Choice Strategy of Single Agent
abstract
Feature selection aims to select a subset of features to optimize the performances of downstream predictive tasks. Recently, multi-agent reinforced feature selection (MARFS) has been introduced to automate feature selection, by creating agents for each feature to select or deselect corresponding features. Although MARFS enjoys the automation of the selection process, MARFS suffers from not just the data complexity in terms of contents and dimensionality, but also the exponentially-increasing computational costs with regard to the number of agents. The raised concern leads to a new research question: Can we simplify the selection process of agents under reinforcement learning context so as to improve the efficiency and costs of feature selection? To address the question, we develop a single-agent reinforced feature selection approach integrated with restructured choice strategy. Specifically, the restructured choice strategy includes: 1) we exploit only one single agent to handle the selection task of multiple features, instead of using multiple agents. 2) we develop a scanning method to empower the single agent to make multiple selection/deselection decisions in each round of scanning. 3) we exploit the relevance to predictive labels of features to prioritize the scanning orders of the agent for multiple features. 4) we propose a convolutional auto-encoder algorithm, integrated with the encoded index information of features, to improve state representation. 5) we design a reward scheme that take into account both prediction accuracy and feature redundancy to facilitate the exploration process. Finally, we present extensive experimental results to demonstrate the efficiency and effectiveness of the proposed method.
Xiaosa Zhao, Kunpeng Liu 0001, Wei Fan 0010, Lu Jiang 0007, Xiaowei Zhao 0004, Minghao Yin, Yanjie Fu
ICDM5
2019 Protein Ubiquitylation and Sumoylation Site Prediction Based on Ensemble and Transfer Learning
abstract
Ubiquitylation, a typical post-translational modification (PTM), plays an important role in signal transduction, apoptosis and cell proliferation. A ubiquitylation like PTM, sumoylation also may affect gene mapping, expression and genomic replication. Over the past two decades, machine learning has been widely employed in protein ubiquitylation and sumoylation site prediction tools. These existing tools require feature engineering, but failed to provide general interpretable features and probably underutilized the growing amount of data. This prompted us to propose a deep learning-based model that integrates multiple convolution and fully-connected layers of seven supervised learning sub-models to extract deep representations from protein sequences and physico-chemical properties (PCPs). Especially, we divided PCPs into 6 clusters and customized deep networks accordingly for handling the high correlations among one cluster. A stacking ensemble strategy was applied to combine these deep representations to make prediction. Furthermore, with the advantage of transfer learning, our deep learning model can work well on protein sumoylation site prediction as well after fine-tuning. On the high-quality annotated database Swiss-Prot, our model outperformed several well-known ubiquitylation and sumoylation site prediction tools. Our code is freely available at https://github.com/ruiwcoding/DeepUbiSumoPre.
Fei He 0003, Yanxin Gao, Duolin Wang, Dong Xu 0002, Xiaowei Zhao 0004
BIBM7
2019 Analysis and prediction of human acetylation using a cascade classifier based on support vector machine
abstract
BACKGROUND: Acetylation on lysine is a widespread post-translational modification which is reversible and plays a crucial role in some biological activities. To better understand the mechanism, it is necessary to identify acetylation sites in proteins accurately. Computational methods are popular because they are more convenient and faster than experimental methods. In this study, we proposed a new computational method to predict acetylation sites in human by combining sequence features and structural features including physicochemical property (PCP), position specific score matrix (PSSM), auto covariation (AC), residue composition (RC), secondary structure (SS) and accessible surface area (ASA), which can well characterize the information of acetylated lysine sites. Besides, a two-step feature selection was applied, which combined mRMR and IFS. It finally trained a cascade classifier based on SVM, which successfully solved the imbalance between positive samples and negative samples and covered all negative sample information. RESULTS: The performance of this method is measured with a specificity of 72.19% and a sensibility of 76.71% on independent dataset which shows that a cascade SVM classifier outperforms single SVM classifier. CONCLUSIONS: In addition to the analysis of experimental results, we also made a systematic and comprehensive analysis of the acetylation data.
Jinchao Ji, Zhiqiang Ma 0003, Xiaowei Zhao 0004
BMC Bioinform.5
2018 Detecting Succinylation sites from protein sequences using ensemble support vector machine
abstract
BACKGROUND: Lysine succinylation is a new kind of post-translational modification which plays a key role in protein conformation regulation and cellular function control. To understand the mechanism of succinylation profoundly, it is necessary to identify succinylation sites in proteins accurately. However, traditional methods, experimental approaches, are labor-intensive and time-consuming. Computational prediction methods have been proposed recent years, and they are popular because of their convenience and high speed. In this study, we developed a new method to predict succinylation sites in protein combining multiple features, including amino acid composition, binary encoding, physicochemical property and grey pseudo amino acid composition, with a feature selection scheme (information gain). And then, it was trained using SVM (Support Vector Machine) and an ensemble learning algorithm. RESULTS: The performance of this method was measured with an accuracy of 89.14% and a MCC (Matthew Correlation Coefficient) of 0.79 using 10-fold cross validation on training dataset and an accuracy of 84.5% and a MCC of 0.2 on independent dataset. CONCLUSIONS: The conclusions made from this study can help to understand more of the succinylation mechanism. These results suggest that our method was very promising for predicting succinylation sites. The source code and data of this paper are freely available at https://github.com/ningq669/PSuccE .
Xiaosa Zhao, Lingling Bao, Zhiqiang Ma 0003, Xiaowei Zhao 0004
BMC Bioinform.5
2017 A multimodal deep architecture for large-scale protein ubiquitylation site prediction
abstract
In eukaryotes, protein ubiquitylation is an important type of post-translation modification, in which the ubiquitin conjugates to a substrate protein. To have a better insight of the mechanisms underlying ubiquitylation, a key step is to identify protein ubiquitylation sites. Many existing computational methods are based on feature engineering, which may lead to biased and incomplete features. Deep learning provides multiple-layer networks and non-linear mapping operations to detect potential complex patterns in a data-driven way, especially for large-scale data. It provides a promising new method to predict ubiquitylation sites. In this paper, we proposed a multimodal deep architecture for protein ubiquitylation sites prediction. First, we designed different multiple layers to extract hidden informative patterns from three modalities, namely protein fragments, physico-chemical properties, and sequence profiles. Then, the deep representations corresponding to three modalities were merged to implement the classification. On the available largest scale protein ubiquitylation site database PLMD, the performance of our proposed method was measured with 66.7% sensitivity, 66.4% specificity, 66.43% accuracy, and 0.221 MCC value. A range of comparative experiments also showed that our proposed architecture outperformed several popular protein ubiquitylation site prediction tools. Our source code is freely available at https://github.com/jiagenlee/deepUbiquitylation.
Fei He 0003, Lingling Bao, Jiagen Li, Dong Xu 0002, Xiaowei Zhao 0004
BIBM6