Shaowu Zhang 0001

dblp:122/5958-1 · also Shao-Wu Zhang 0001, ShaoWu Zhang 0001 · DBLP profile ↗
← Back
45ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0003-1305-7447ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 38 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 7 · 5 since 2021
YearPublicationVenuePosition
2026 Dynamical Causality Under Latent Confounders for Biological Network Reconstruction
abstract
Causal interaction inference is prone to spurious causal interactions, due to the substantial confounders in a biological system. While many existing methods attempt to address misidentification challenges, there remains a notable lack of effective methods to infer causal interaction under latent/unobserved confounders. In this work, we propose a method to overcome such challenges to infer dynamical causality under invisible confounders (CIC) and further reconstruct the latent confounders from time-series data by developing an orthogonal decomposition theorem in a delay embedding space. This theoretical foundation ensures the causal detection for any high-dimensional system even with only two observed variables under many latent confounders, which is a long-standing problem in the field. In addition to the latent confounder problem, such a decomposition makes the coupled variables separable in the embedding space, thus also solving the non-separability problem of causal inference. Extensive validation of the CIC method is carried out using various real datasets, which all demonstrates its effectiveness to reconstruct real biological networks and unobserved confounders.
Jinling Yan, Shaowu Zhang 0001, Chihao Zhang 0002, Weitian Huang, Jifan Shi, Luonan Chen
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Few-shot drug synergy prediction via rapid cross-tier adaptation meta-optimization
abstract
Drug combination therapy offers key advantages over monotherapy in personalized oncology by reducing drug resistance and toxicity. However, predicting synergistic effects for rare cell lines remains challenging, as existing methods suffer from poor generalizability in data-scarce scenarios owing to their reliance on large training datasets and inability to effectively transfer knowledge across distinct cellular contexts. Here, we present MetaSynergy, a Rapid Cross-tier Adaptation Meta-Optimization (R-CAMO)-based framework for few-shot drug synergy prediction through cross-domain knowledge transfer and meta-optimized adaptation. We first designed a multimodal feature learning architecture integrating drug molecular graphs with cell line omics profiles, then implemented a stage-wise training strategy based on R-CAMO for few-shot drug synergy prediction: (i) cross-domain pretraining establishes meta-initialized representations by transferring knowledge from data-rich cell lines to scarce target domains, enhancing feature representation capability in data-scarce scenarios. (ii) Cross-tier meta-optimization enables rapid adaptation to data-scarce scenarios: the inner-tier refines task-specific parameters of the prediction network on the target domain, while the outer-tier meta-learns task-shared, generalizable parameters by minimizing the cross-cell line prediction loss. (iii) Fine-tuning further refines task-specific parameters, improving generalizability to novel drug combinations within the same cellular context. Experimental results demonstrate that MetaSynergy achieves excellent performance in few-shot, zero-shot and low-similarity tasks, surpassing most baseline methods and highlighting its robustness and generalizability. Ablation studies confirmed the pivotal role of R-CAMO strategy in data-scarce cell lines. Furthermore, MetaSynergy successfully identified novel synergistic drug combinations in several understudied malignancies, underscoring its potential in precision oncology.
Yue-Hua Feng, Ze-Lin Feng, Xiao-Ying Yan, Shaowu Zhang 0001, Jianyu Shi
Briefings Bioinform.4
2025 MTGCL: Multi-Task Graph Contrastive Learning for Identifying Cancer Driver Genes From Multi-Omics Data
abstract
Identification of cancer driver genes is crucial for understanding the molecular mechanisms of cancer. To address the limitations of graph convolutional networks-based cancer driver gene identification methods, including biased prediction results caused by convolutional layer structures that focus more on either the structural characteristics (e.g., degree) or biological features (e.g., mutation frequency) of nodes in the network, as well as sparse supervisory information, we propose a method called Multi-Task Graph Contrastive Learning (MTGCL) for the identification of cancer driver genes. MTGCL designs a new graph convolutional layer structure which can improve the performance of cancer driver gene identification by effectively integrating graph structure topology information and node features information, while a semi-supervised graph contrastive learning task is presented as a regularizer within a multi-task learning paradigm to enhance the performance of the main task of driver gene identification by utilizing a small portion of labeled nodes and a large amount of unlabeled nodes information. The experimental results on pan-cancer and some specific cancers demonstrate the effectiveness of MTGCL. In addition, we also find the features of different mutation types derived from somatic mutation data can effectively improve the performance of identifying driver genes for some specific cancer types.
Ming-Yu Xie, Shaowu Zhang 0001, Yan Li 0111
IEEE Trans. Comput. Biol. Bioinform.2
2024 MSCLK: Multi-scale fully separable convolution neural network with large kernels for early diagnosis of Alzheimer's disease
Run-Feng Tian, Jia-Ni Li, Shaowu Zhang 0001
Expert Syst. Appl.3
2024 Identifying cooperating cancer driver genes in individual patients through hypergraph random walk
Shaowu Zhang 0001, Ming-Yu Xie, Yan Li 0111
J. Biomed. Informatics2
2023 Diagnosis of Alzheimer's disease by joining dual attention CNN and MLP based on structural MRIs, clinical and genetic data
Yan-Rui Qiang, Shaowu Zhang 0001, Jia-Ni Li, Yan Li 0111, Qin-Yi Zhou
Artif. Intell. Medicine2
2023 A social theory-enhanced graph representation learning framework for multitask prediction of drug-drug interactions
abstract
Current machine learning-based methods have achieved inspiring predictions in the scenarios of mono-type and multi-type drug-drug interactions (DDIs), but they all ignore enhancive and depressive pharmacological changes triggered by DDIs. In addition, these pharmacological changes are asymmetric since the roles of two drugs in an interaction are different. More importantly, these pharmacological changes imply significant topological patterns among DDIs. To address the above issues, we first leverage Balance theory and Status theory in social networks to reveal the topological patterns among directed pharmacological DDIs, which are modeled as a signed and directed network. Then, we design a novel graph representation learning model named SGRL-DDI (social theory-enhanced graph representation learning for DDI) to realize the multitask prediction of DDIs. SGRL-DDI model can capture the task-joint information by integrating relation graph convolutional networks with Balance and Status patterns. Moreover, we utilize task-specific deep neural networks to perform two tasks, including the prediction of enhancive/depressive DDIs and the prediction of directed DDIs. Based on DDI entries collected from DrugBank, the superiority of our model is demonstrated by the comparison with other state-of-the-art methods. Furthermore, the ablation study verifies that Balance and Status patterns help characterize directed pharmacological DDIs, and that the joint of two tasks provides better DDI representations than individual tasks. Last, we demonstrate the practical effectiveness of our model by a version-dependent test, where 88.47 and 81.38% DDI out of newly added entries provided by the latest release of DrugBank are validated in two predicting tasks respectively.
Yue-Hua Feng, Shaowu Zhang 0001, Yi-Yang Feng, Qing-Qing Zhang, Ming-Hui Shi, Jianyu Shi
Briefings Bioinform.2
2023 PhenoDriver: interpretable framework for studying personalized phenotype-associated driver genes in breast cancer
abstract
Identifying personalized cancer driver genes and further revealing their oncogenic mechanisms is critical for understanding the mechanisms of cell transformation and aiding clinical diagnosis. Almost all existing methods primarily focus on identifying driver genes at the cohort or individual level but fail to further uncover their underlying oncogenic mechanisms. To fill this gap, we present an interpretable framework, PhenoDriver, to identify personalized cancer driver genes, elucidate their roles in cancer development and uncover the association between driver genes and clinical phenotypic alterations. By analyzing 988 breast cancer patients, we demonstrate the outstanding performance of PhenoDriver in identifying breast cancer driver genes at the cohort level compared to other state-of-the-art methods. Otherwise, our PhenoDriver can also effectively identify driver genes with both recurrent and rare mutations in individual patients. We further explore and reveal the oncogenic mechanisms of some known and unknown breast cancer driver genes (e.g. TP53, MAP3K1, HTT, etc.) identified by PhenoDriver, and construct their subnetworks for regulating clinical abnormal phenotypes. Notably, most of our findings are consistent with existing biological knowledge. Based on the personalized driver profiles, we discover two existing and one unreported breast cancer subtypes and uncover their molecular mechanisms. These results intensify our understanding for breast cancer mechanisms, guide therapeutic decisions and assist in the development of targeted anticancer therapies.
Yan Li 0111, Shaowu Zhang 0001, Ming-Yu Xie
Briefings Bioinform.2
2023 A novel heterophilic graph diffusion convolutional network for identifying cancer driver genes
abstract
Identifying cancer driver genes plays a curial role in the development of precision oncology and cancer therapeutics. Although a plethora of methods have been developed to tackle this problem, the complex cancer mechanisms and intricate interactions between genes still make the identification of cancer driver genes challenging. In this work, we propose a novel machine learning method of heterophilic graph diffusion convolutional networks (called HGDCs) to boost cancer-driver gene identification. Specifically, HGDC first introduces graph diffusion to generate an auxiliary network for capturing the structurally similar nodes in a biomolecular network. Then, HGDC designs an improved message aggregation and propagation scheme to adapt to the heterophilic setting of biomolecular networks, alleviating the problem of driver gene features being smoothed by its neighboring dissimilar genes. Finally, HGDC uses a layer-wise attention classifier to predict the probability of one gene being a cancer driver gene. In the comparison experiments with other existing state-of-the-art methods, our HGDC achieves outstanding performance in identifying cancer driver genes. The experimental results demonstrate that HGDC not only effectively identifies well-known driver genes on different networks but also novel candidate cancer genes. Moreover, HGDC can effectively prioritize cancer driver genes for individual patients. Particularly, HGDC can identify patient-specific additional driver genes, which work together with the well-known driver genes to cooperatively promote tumorigenesis.
Shaowu Zhang 0001, Ming-Yu Xie, Yan Li 0111
Briefings Bioinform.2
2023 Few-Shot Drug Synergy Prediction With a Prior-Guided Hypernetwork Architecture
abstract
Predicting drug synergy is critical to tailoring feasible drug combination treatment regimens for cancer patients. However, most of the existing computational methods only focus on data-rich cell lines, and hardly work on data-poor cell lines. To this end, here we proposed a novel few-shot drug synergy prediction method (called HyperSynergy) for data-poor cell lines by designing a prior-guided Hypernetwork architecture, in which the meta-generative network based on the task embedding of each cell line generates cell line dependent parameters for the drug synergy prediction network. In HyperSynergy model, we designed a deep Bayesian variational inference model to infer the prior distribution over the task embedding to quickly update the task embedding with a few labeled drug synergy samples, and presented a three-stage learning strategy to train HyperSynergy for quickly updating the prior distribution by a few labeled drug synergy samples of each data-poor cell line. Moreover, we proved theoretically that HyperSynergy aims to maximize the lower bound of log-likelihood of the marginal distribution over each data-poor cell line. The experimental results show that our HyperSynergy outperforms other state-of-the-art methods not only on data-poor cell lines with a few samples (e.g., 10, 5, 0), but also on data-rich cell lines.
Qing-Qing Zhang, Shaowu Zhang 0001, Yue-Hua Feng, Jianyu Shi
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 m 6 Aexpress-BHM: predicting m6A regulation of gene expression in multiple-groups context by a Bayesian hierarchical mixture model
abstract
As the most abundant RNA modification, N6-methyladenosine (m6A) plays an important role in various RNA activities including gene expression and translation. With the rapid application of MeRIP-seq technology, samples of multiple groups, such as the involved multiple viral/ bacterial infection or distinct cell differentiation stages, are extracted from same experimental unit. However, our current knowledge about how the dynamic m6A regulating gene expression and the role in certain biological processes (e.g. immune response in this complex context) is largely elusive due to lack of effective tools. To address this issue, we proposed a Bayesian hierarchical mixture model (called m6Aexpress-BHM) to predict m6A regulation of gene expression (m6A-reg-exp) in multiple groups of MeRIP-seq experiment with limited samples. Comprehensive evaluations of m6Aexpress-BHM on the simulated data demonstrate its high predicting precision and robustness. Applying m6Aexpress-BHM on three real-world datasets (i.e. Flaviviridae infection, infected time-points of bacteria and differentiation stages of dendritic cells), we predicted more m6A-reg-exp genes with positive regulatory mode that significantly participate in innate immune or adaptive immune pathways, revealing the underlying mechanism of the regulatory function of m6A during immune response. In addition, we also found that m6A may influence the expression of PD-1/PD-L1 via regulating its interacted genes. These results demonstrate the power of m6Aexpress-BHM, helping us understand the m6A regulatory function in immune system.
Shaowu Zhang 0001
Briefings Bioinform.2
2022 Prioritization of cancer driver gene with prize-collecting steiner tree by introducing an edge weighted strategy in the personalized gene interaction network
abstract
BACKGROUND: Cancer is a heterogeneous disease in which tumor genes cooperate as well as adapt and evolve to the changing conditions for individual patients. It is a meaningful task to discover the personalized cancer driver genes that can provide diagnosis and target drug for individual patients. However, most of existing methods mainly ranks potential personalized cancer driver genes by considering the patient-specific nodes information on the gene/protein interaction network. These methods ignore the personalized edge weight information in gene interaction network, leading to false positive results. RESULTS: In this work, we presented a novel algorithm (called PDGPCS) to predict the Personalized cancer Driver Genes based on the Prize-Collecting Steiner tree model by considering the personalized edge weight information. PDGPCS first constructs the personalized weighted gene interaction network by integrating the personalized gene expression data and prior known gene/protein interaction network knowledge. Then the gene mutation data and pathway data are integrated to quantify the impact of each mutant gene on every dysregulated pathway with the prize-collecting Steiner tree model. Finally, according to the mutant gene's aggregated impact score on all dysregulated pathways, the mutant genes are ranked for prioritizing the personalized cancer driver genes. Experimental results on four TCGA cancer datasets show that PDGPCS has better performance than other personalized driver gene prediction methods. In addition, we verified that the personalized edge weight of gene interaction network can improve the prediction performance. CONCLUSIONS: PDGPCS can more accurately identify the personalized driver genes and takes a step further toward personalized medicine and treatment. The source code of PDGPCS can be freely downloaded from https://github.com/NWPU-903PR/PDGPCS .
Shaowu Zhang 0001, Yan Li 0111, Weifeng Guo
BMC Bioinform.1
2022 Extracting ROI-Based Contourlet Subband Energy Feature From the sMRI Image for Alzheimer's Disease Classification
abstract
Structural magnetic resonance imaging (sMRI)-based Alzheimer's disease (AD) classification and its prodromal stage-mild cognitive impairment (MCI) classification have attracted many attentions and been widely investigated in recent years. Owing to the high dimensionality, representation of the sMRI image becomes a difficult issue in AD classification. Furthermore, regions of interest (ROI) reflected in the sMRI image are not characterized properly by spatial analysis techniques, which has been a main cause of weakening the discriminating ability of the extracted spatial feature. In this study, we propose a ROI-based contourlet subband energy (ROICSE) feature to represent the sMRI image in the frequency domain for AD classification. Specifically, a preprocessed sMRI image is first segmented into 90 ROIs by a constructed brain mask. Instead of extracting features from the 90 ROIs in the spatial domain, the contourlet transform is performed on each of these ROIs to obtain their energy subbands. And then for an ROI, a subband energy (SE) feature vector is constructed to capture its energy distribution and contour information. Afterwards, SE feature vectors of the 90 ROIs are concatenated to form a ROICSE feature of the sMRI image. Finally, support vector machine (SVM) classifier is used to classify 880 subjects from ADNI and OASIS databases. Experimental results show that the ROICSE approach outperforms six other state-of-the-art methods, demonstrating that energy and contour information of the ROI are important to capture differences between the sMRI images of AD and HC subjects. Meanwhile, brain regions related to AD can also be found using the ROICSE feature, indicating that the ROICSE feature can be a promising assistant imaging marker for the AD diagnosis via the sMRI image. Code and Sample IDs of this paper can be downloaded at https://github.com/NWPU-903PR/ROICSE.git.
Jinwang Feng, Shaowu Zhang 0001, Luonan Chen
IEEE ACM Trans. Comput. Biol. Bioinform.2
2022 Prediction of Transcription Factor Binding Sites With an Attention Augmented Convolutional Neural Network
abstract
Identification of transcription factor binding sites (TFBSs) is essential for revealing the rules of protein-DNA binding. Although some computational methods have been presented to predict TFBSs using epigenomic and sequence features, most of them ignore the common features among cross-cell types. It is still unclear to what extent the common features could help for this task. To this end, we proposed a new method (named Attention-augmented Convolutional Neural Network, or ACNN) to predict TFBSs. ACNN uses attention-augmented convolutional layers to capture global and local contexts in DNA sequences and employs the convolutional layers to capture features of histone modification markers. In addition, ACNN adopts the private and shared convolutional neural network (CNN) modules to learn specific and common features, respectively. To encourage the shared CNN module to learn the common features, adversarial training is applied in ACNN. The results on 253 ChIP-seq datasets show that ACNN outperforms other existing methods. The attention-augmented convolutional layers and adversarial training mechanism in ACNN can effectively improve the prediction performance. Moreover, in the case of limited labeled data, ACNN also performs better than a baseline method. We further visualize the convolution kernels as motifs to explain the interpretability of ACNN.
Fang Jing, Shaowu Zhang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2021 Resilience function uncovers the critical transitions in cancer initiation
abstract
Considerable evidence suggests that during the progression of cancer initiation, the state transition from wellness to disease is not necessarily smooth but manifests switch-like nonlinear behaviors, preventing the cancer prediction and early interventional therapy for patients. Understanding the mechanism of such wellness-to-disease transitions is a fundamental and challenging task. Despite the advances in flux theory of nonequilibrium dynamics and 'critical slowing down'-based system resilience theory, a system-level approach still lacks to fully describe this state transition. Here, we present a novel framework (called bioRFR) to quantify such wellness-to-disease transition during cancer initiation through uncovering the biological system's resilience function from gene expression data. We used bioRFR to reconstruct the biologically and dynamically significant resilience functions for cancer initiation processes (e.g. BRCA, LUSC and LUAD). The resilience functions display the similar resilience pattern with hysteresis feature but different numbers of tipping points, which implies that once the cell become cancerous, it is very difficult or even impossible to reverse to the normal state. More importantly, bioRFR can measure the severe degree of cancer patients and identify the personalized key genes that are associated with the individual system's state transition from normal to tumor in resilience perspective, indicating that bioRFR can contribute to personalized medicine and targeted cancer therapy.
Yan Li 0111, Shaowu Zhang 0001
Briefings Bioinform.2
2021 Identifying driver genes for individual patients through inductive matrix completion
abstract
MOTIVATION: The driver genes play a key role in the evolutionary process of cancer. Effectively identifying these driver genes is crucial to cancer diagnosis and treatment. However, due to the high heterogeneity of cancers, it remains challenging to identify the driver genes for individual patients. Although some computational methods have been proposed to tackle this problem, they seldom consider the fact that the genes functionally similar to the well-established driver genes may likely play similar roles in cancer process, which potentially promotes the driver gene identification. Thus, here we developed a novel approach of IMCDriver to promote the driver gene identification both for cohorts and individual patients. RESULTS: IMCDriver first considers the well-established driver genes as prior information, and adopts the using multi-omics data (e.g. somatic mutation, gene expression and protein-protein interaction) to compute the similarity between patients/genes. Then, IMCDriver prioritizes the personalized mutated genes according to their functional similarity to the well-established driver genes via Inductive Matrix Completion. Finally, IMCDriver identifies the highly rank-ordered genes as the personalized driver genes. The results on five cancer datasets from the Cancer Genome Consortium show that our IMCDriver outperforms other existing state-of-the-art methods both in the cohort and patient-specific driver gene identification. IMCDriver also reveals some novel driver genes that potentially drive cancer development. In addition, even for the driver genes rarely mutated among a population, IMCDriver can still identify them and prioritize them with high priorities. AVAILABILITY AND IMPLEMENTATION: Code available at https://github.com/NWPU-903PR/IMCDriver. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Shaowu Zhang 0001, Yan Li 0111
Bioinform.2
2021 Funm6AViewer: a web server and R package for functional analysis of context-specific m6A RNA methylation
abstract
MOTIVATION: N 6-methyladenosine (m6A) is the most abundant mammalian mRNA methylation with versatile functions. To date, although a number of bioinformatics tools have been developed for location discovery of m6A modification, functional understanding is still quite limited. As the focus of RNA epigenetics gradually shifts from site discovery to functional studies, there is an urgent need for user-friendly tools to identify and explore the functional relevance of context-specific m6A methylation to gain insights into the epitranscriptome layer of gene expression regulation. RESULTS: We introduced here Funm6AViewer, a novel platform to identify, prioritize and visualize the functional gene interaction networks mediated by dynamic m6A RNA methylation unveiled from a case control study. By taking the differential RNA methylation data and differential gene expression data, both of which can be inferred from the widely used MeRIP-seq data, as the inputs, Funm6AViewer enables a series of analysis, including: (i) examining the distribution of differential m6A sites, (ii) prioritizing the genes mediated by dynamic m6A methylation and (iii) characterizing functionally the gene regulatory networks mediated by condition-specific m6A RNA methylation. Funm6AViewer should effectively facilitate the understanding of the epitranscriptome circuitry mediated by this reversible RNA modification. AVAILABILITY AND IMPLEMENTATION: Funm6AViewer is available both as a convenient web server (https://www.xjtlu.edu.cn/biologicalsciences/funm6aviewer) with graphical interface and as an independent R package (https://github.com/NWPU-903PR/Funm6AViewer) for local usage. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Songyao Zhang, Shaowu Zhang 0001, Yujiao Tang, Xiaonan Fan 0001, Jia Meng 0001
Bioinform.2
2021 Alzheimer's disease classification using features extracted from nonsubsampled contourlet subband-based individual networks
Jinwang Feng, Shaowu Zhang 0001, Luonan Chen
Neurocomputing2
2021 Performance assessment of sample-specific network control methods for bulk and single-cell biological data analysis
abstract
In the past few years, a wealth of sample-specific network construction methods and structural network control methods has been proposed to identify sample-specific driver nodes for supporting the Sample-Specific network Control (SSC) analysis of biological networked systems. However, there is no comprehensive evaluation for these state-of-the-art methods. Here, we conducted a performance assessment for 16 SSC analysis workflows by using the combination of 4 sample-specific network reconstruction methods and 4 representative structural control methods. This study includes simulation evaluation of representative biological networks, personalized driver genes prioritization on multiple cancer bulk expression datasets with matched patient samples from TCGA, and cell marker genes and key time point identification related to cell differentiation on single-cell RNA-seq datasets. By widely comparing analysis of existing SSC analysis workflows, we provided the following recommendations and banchmarking workflows. (i) The performance of a network control method is strongly dependent on the up-stream sample-specific network method, and Cell-Specific Network construction (CSN) method and Single-Sample Network (SSN) method are the preferred sample-specific network construction methods. (ii) After constructing the sample-specific networks, the undirected network-based control methods are more effective than the directed network-based control methods. In addition, these data and evaluation pipeline are freely available on https://github.com/WilfongGuo/Benchmark_control.
Weifeng Guo, Xiangtian Yu, Jing J. Liang, Shaowu Zhang 0001, Tao Zeng 0003
PLoS Comput. Biol.5
2021 An Integrative Framework for Combining Sequence and Epigenomic Data to Predict Transcription Factor Binding Sites Using Deep Learning
abstract
Knowing the transcription factor binding sites (TFBSs) is essential for modeling the underlying binding mechanisms and follow-up cellular functions. Convolutional neural networks (CNNs) have outperformed methods in predicting TFBSs from the primary DNA sequence. In addition to DNA sequences, histone modifications and chromatin accessibility are also important factors influencing their activity. They have been explored to predict TFBSs recently. However, current methods rarely take into account histone modifications and chromatin accessibility using CNN in an integrative framework. To this end, we developed a general CNN model to integrate these data for predicting TFBSs. We systematically benchmarked a series of architecture variants by changing network structure in terms of width and depth, and explored the effects of sample length at flanking regions. We evaluated the performance of the three types of data and their combinations using 256 ChIP-seq experiments and also compared it with competing machine learning methods. We find that contributions from these three types of data are complementary to each other. Moreover, the integrative CNN framework is superior to traditional machine learning methods with significant improvements.
Fang Jing, Shaowu Zhang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2020 Identification of Alzheimer's disease based on wavelet transformation energy feature of the structural MRI image and NN classifier
Jinwang Feng, Shaowu Zhang 0001, Luonan Chen
Artif. Intell. Medicine2
2020 Network control principles for identifying personalized driver genes in cancer
abstract
To understand tumor heterogeneity in cancer, personalized driver genes (PDGs) need to be identified for unraveling the genotype-phenotype associations corresponding to particular patients. However, most of the existing driver-focus methods mainly pay attention on the cohort information rather than on individual information. Recent developing computational approaches based on network control principles are opening a new way to discover driver genes in cancer, particularly at an individual level. To provide comprehensive perspectives of network control methods on this timely topic, we first considered the cancer progression as a network control problem, in which the expected PDGs are altered genes by oncogene activation signals that can change the individual molecular network from one health state to the other disease state. Then, we reviewed the network reconstruction methods on single samples and introduced novel network control methods on single-sample networks to identify PDGs in cancer. Particularly, we gave a performance assessment of the network structure control-based PDGs identification methods on multiple cancer datasets from TCGA, for which the data and evaluation package also are publicly available. Finally, we discussed future directions for the application of network control methods to identify PDGs in cancer and diverse biological processes.
Weifeng Guo, Shaowu Zhang 0001, Tao Zeng 0003, Tatsuya Akutsu, Luonan Chen
Briefings Bioinform.2
2020 DPDDI: a deep predictor for drug-drug interactions
abstract
BACKGROUND: The treatment of complex diseases by taking multiple drugs becomes increasingly popular. However, drug-drug interactions (DDIs) may give rise to the risk of unanticipated adverse effects and even unknown toxicity. DDI detection in the wet lab is expensive and time-consuming. Thus, it is highly desired to develop the computational methods for predicting DDIs. Generally, most of the existing computational methods predict DDIs by extracting the chemical and biological features of drugs from diverse drug-related properties, however some drug properties are costly to obtain and not available in many cases. RESULTS: In this work, we presented a novel method (namely DPDDI) to predict DDIs by extracting the network structure features of drugs from DDI network with graph convolution network (GCN), and the deep neural network (DNN) model as a predictor. GCN learns the low-dimensional feature representations of drugs by capturing the topological relationship of drugs in DDI network. DNN predictor concatenates the latent feature vectors of any two drugs as the feature vector of the corresponding drug pairs to train a DNN for predicting the potential drug-drug interactions. Experiment results show that, the newly proposed DPDDI method outperforms four other state-of-the-art methods; the GCN-derived latent features include more DDI information than other features derived from chemical, biological or anatomical properties of drugs; and the concatenation feature aggregation operator is better than two other feature aggregation operators (i.e., inner product and summation). The results in case studies confirm that DPDDI achieves reasonable performance in predicting new DDIs. CONCLUSION: We proposed an effective and robust method DPDDI to predict the potential DDIs by utilizing the DDI network information without considering the drug properties (i.e., drug chemical and biological properties). The method should also be useful in other DDI-related scenarios, such as the detection of unexpected side effects, and the guidance of drug combination.
Yue-Hua Feng, Shaowu Zhang 0001, Jianyu Shi
BMC Bioinform.2
2020 Prediction of enhancer-promoter interactions using the cross-cell type information and domain adversarial neural network
abstract
BACKGROUND: Enhancer-promoter interactions (EPIs) play key roles in transcriptional regulation and disease progression. Although several computational methods have been developed to predict such interactions, their performances are not satisfactory when training and testing data from different cell lines. Currently, it is still unclear what extent a across cell line prediction can be made based on sequence-level information. RESULTS: In this work, we present a novel Sequence-based method (called SEPT) to predict the enhancer-promoter interactions in new cell line by using the cross-cell information and Transfer learning. SEPT first learns the features of enhancer and promoter from DNA sequences with convolutional neural network (CNN), then designing the gradient reversal layer of transfer learning to reduce the cell line specific features meanwhile retaining the features associated with EPIs. When the locations of enhancers and promoters are provided in new cell line, SEPT can successfully recognize EPIs in this new cell line based on labeled data of other cell lines. The experiment results show that SEPT can effectively learn the latent import EPIs-related features between cell lines and achieves the best prediction performance in terms of AUC (the area under the receiver operating curves). CONCLUSIONS: SEPT is an effective method for predicting the EPIs in new cell line. Domain adversarial architecture of transfer learning used in SEPT can learn the latent EPIs shared features among cell lines from all other existing labeled data. It can be expected that SEPT will be of interest to researchers concerned with biological interaction prediction.
Fang Jing, Shaowu Zhang 0001
BMC Bioinform.2
2020 smsMap: mapping single molecule sequencing reads by locating the alignment starting positions
abstract
BACKGROUND: Single Molecule Sequencing (SMS) technology can produce longer reads with higher sequencing error rate. Mapping these reads to a reference genome is often the most fundamental and computing-intensive step for downstream analysis. Most existing mapping tools generally adopt the traditional seed-and-extend strategy, and the candidate aligned regions for each query read are selected either by counting the number of matched seeds or chaining a group of seeds. However, for all the existing mapping tools, the coverage ratio of the alignment region to the query read is lower, and the read alignment quality and efficiency need to be improved. Here, we introduce smsMap, a novel mapping tool that is specifically designed to map the long reads of SMS to a reference genome. RESULTS: smsMap was evaluated with other existing seven SMS mapping tools (e.g., BLASR, minimap2, and BWA-MEM) on both simulated and real-life SMS datasets. The experimental results show that smsMap can efficiently achieve higher aligned read coverage ratio and has higher sensitivity that can align more sequences and bases to the reference genome. Additionally, smsMap is more robust to sequencing errors. CONCLUSIONS: smsMap is computationally efficient to align SMS reads, especially for the larger size of the reference genome (e.g., H. sapiens genome with over 3 billion base pairs). The source code of smsMap can be freely downloaded from https://github.com/NWPU-903PR/smsMap .
Ze-Gang Wei, Shaowu Zhang 0001
BMC Bioinform.2
2019 FunDMDeep-m6A: identification and prioritization of functional differential m6A methylation genes
abstract
MOTIVATION: As the most abundant mammalian mRNA methylation, N6-methyladenosine (m6A) exists in >25% of human mRNAs and is involved in regulating many different aspects of mRNA metabolism, stem cell differentiation and diseases like cancer. However, our current knowledge about dynamic changes of m6A levels and how the change of m6A levels for a specific gene can play a role in certain biological processes like stem cell differentiation and diseases like cancer is largely elusive. RESULTS: To address this, we propose in this paper FunDMDeep-m6A a novel pipeline for identifying context-specific (e.g. disease versus normal, differentiated cells versus stem cells or gene knockdown cells versus wild-type cells) m6A-mediated functional genes. FunDMDeep-m6A includes, at the first step, DMDeep-m6A a novel method based on a deep learning model and a statistical test for identifying differential m6A methylation (DmM) sites from MeRIP-Seq data at a single-base resolution. FunDMDeep-m6A then identifies and prioritizes functional DmM genes (FDmMGenes) by combing the DmM genes (DmMGenes) with differential expression analysis using a network-based method. This proposed network method includes a novel m6A-signaling bridge (MSB) score to quantify the functional significance of DmMGenes by assessing functional interaction of DmMGenes with their signaling pathways using a heat diffusion process in protein-protein interaction (PPI) networks. The test results on 4 context-specific MeRIP-Seq datasets showed that FunDMDeep-m6A can identify more context-specific and functionally significant FDmMGenes than m6A-Driver. The functional enrichment analysis of these genes revealed that m6A targets key genes of many important context-related biological processes including embryonic development, stem cell differentiation, transcription, translation, cell death, cell proliferation and cancer-related pathways. These results demonstrate the power of FunDMDeep-m6A for elucidating m6A regulatory functions and its roles in biological processes and diseases. AVAILABILITY AND IMPLEMENTATION: The R-package for DMDeep-m6A is freely available from https://github.com/NWPU-903PR/DMDeepm6A1.0. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Songyao Zhang, Shaowu Zhang 0001, Xiaonan Fan 0001, Jia Meng 0001, Yufei Huang 0001
Bioinform.2
2019 Prediction of lncRNA-disease associations by integrating diverse heterogeneous information sources with RWR algorithm and positive pointwise mutual information
abstract
BACKGROUND: Long non-coding RNAs play an important role in human complex diseases. Identification of lncRNA-disease associations will gain insight into disease-related lncRNAs and benefit disease diagnoses and treatment. However, using experiments to explore the lncRNA-disease associations is expensive and time consuming. RESULTS: In this study, we developed a novel method to identify potential lncRNA-disease associations by Integrating Diverse Heterogeneous Information sources with positive pointwise Mutual Information and Random Walk with restart algorithm (namely IDHI-MIRW). IDHI-MIRW first constructs multiple lncRNA similarity networks and disease similarity networks from diverse lncRNA-related and disease-related datasets, then implements the random walk with restart algorithm on these similarity networks for extracting the topological similarities which are fused with positive pointwise mutual information to build a large-scale lncRNA-disease heterogeneous network. Finally, IDHI-MIRW implemented random walk with restart algorithm on the lncRNA-disease heterogeneous network to infer potential lncRNA-disease associations. CONCLUSIONS: Compared with other state-of-the-art methods, IDHI-MIRW achieves the best prediction performance. In case studies of breast cancer, stomach cancer, and colorectal cancer, 36/45 (80%) novel lncRNA-disease associations predicted by IDHI-MIRW are supported by recent literatures. Furthermore, we found lncRNA LINC01816 is associated with the survival of colorectal cancer patients. IDHI-MIRW is freely available at https://github.com/NWPU-903PR/IDHI-MIRW .
Xiaonan Fan 0001, Shaowu Zhang 0001, Songyao Zhang, Kunju Zhu, Songjian Lu
BMC Bioinform.2
2019 LPI-BLS: Predicting lncRNA-protein interactions with a broad learning system-based stacked ensemble classifier
Xiaonan Fan 0001, Shaowu Zhang 0001
Neurocomputing2
2019 A novel network control model for identifying personalized driver genes in cancer
abstract
Although existing computational models have identified many common driver genes, it remains challenging to identify the personalized driver genes by using samples of an individual patient. Recently, the methods of exploiting the structure-based control principles of complex networks provide new clues for identifying minimum number of driver nodes to drive the state transition of large-scale complex networks from an initial state to the desired state. However, the structure-based network control methods cannot be directly applied to identify the personalized driver genes due to the unknown network dynamics of the personalized system. Here we proposed the personalized network control model (PNC) to identify the personalized driver genes by employing the structure-based network control principle on genetic data of individual patients. In PNC model, we firstly presented a paired single sample network construction method to construct the personalized state transition network for capturing the phenotype transitions between healthy and disease states. Then, we designed a novel structure-based network control method from the Feedback Vertex Sets-based control perspective to identify the personalized driver genes. The wide experimental results on 13 cancer datasets from The Cancer Genome Atlas firstly showed that PNC model outperforms current state-of-the-art methods, in terms of F-measures for identifying cancer driver genes enriched in the gold-standard cancer driver gene lists. Furthermore, these results showed that personalized driver genes can be explored by their network characteristics even when they are hidden factors in transcription and mutation profiles. Our PNC gives novel insights and useful tools into understanding the tumor heterogeneity in cancer. The PNC package and data resources used in this work can be freely downloaded from https://github.com/NWPU-903PR/PNC.
Weifeng Guo, Shaowu Zhang 0001, Tao Zeng 0003, Yan Li 0111, Jianxi Gao, Luonan Chen
PLoS Comput. Biol.2
2019 Global analysis of N6-methyladenosine functions and its disease association using deep learning and network-based methods
abstract
N6-methyladenosine (m6A) is the most abundant methylation, existing in >25% of human mRNAs. Exciting recent discoveries indicate the close involvement of m6A in regulating many different aspects of mRNA metabolism and diseases like cancer. However, our current knowledge about how m6A levels are controlled and whether and how regulation of m6A levels of a specific gene can play a role in cancer and other diseases is mostly elusive. We propose in this paper a computational scheme for predicting m6A-regulated genes and m6A-associated disease, which includes Deep-m6A, the first model for detecting condition-specific m6A sites from MeRIP-Seq data with a single base resolution using deep learning and Hot-m6A, a new network-based pipeline that prioritizes functional significant m6A genes and its associated diseases using the Protein-Protein Interaction (PPI) and gene-disease heterogeneous networks. We applied Deep-m6A and this pipeline to 75 MeRIP-seq human samples, which produced a compact set of 709 functionally significant m6A-regulated genes and nine functionally enriched subnetworks. The functional enrichment analysis of these genes and networks reveal that m6A targets key genes of many critical biological processes including transcription, cell organization and transport, and cell proliferation and cancer-related pathways such as Wnt pathway. The m6A-associated disease analysis prioritized five significantly associated diseases including leukemia and renal cell carcinoma. These results demonstrate the power of our proposed computational scheme and provide new leads for understanding m6A regulatory functions and its roles in diseases.
Songyao Zhang, Shaowu Zhang 0001, Xiaonan Fan 0001, Jia Meng 0001, Yidong Chen 0002, Shou-Jiang Gao, Yufei Huang 0001
PLoS Comput. Biol.2
2018 Combining Sequence and Epigenomic Data to Predict Transcription Factor Binding Sites Using Deep Learning
Fang Jing, Shaowu Zhang 0001
ISBRA2
2018 Discovering personalized driver mutation profiles of single samples in cancer by network control strategy
abstract
Motivation: It is a challenging task to discover personalized driver genes that provide crucial information on disease risk and drug sensitivity for individual patients. However, few methods have been proposed to identify the personalized-sample driver genes from the cancer omics data due to the lack of samples for each individual. To circumvent this problem, here we present a novel single-sample controller strategy (SCS) to identify personalized driver mutation profiles from network controllability perspective. Results: SCS integrates mutation data and expression data into a reference molecular network for each patient to obtain the driver mutation profiles in a personalized-sample manner. This is the first such a computational framework, to bridge the personalized driver mutation discovery problem and the structural network controllability problem. The key idea of SCS is to detect those mutated genes which can achieve the transition from the normal state to the disease state based on each individual omics data from network controllability perspective. We widely validate the driver mutation profiles of our SCS from three aspects: (i) the improved precision for the predicted driver genes in the population compared with other driver-focus methods; (ii) the effectiveness for discovering the personalized driver genes and (iii) the application to the risk assessment through the integration of the driver mutation signature and expression data, respectively, across the five distinct benchmarks from The Cancer Genome Atlas. In conclusion, our SCS makes efficient and robust personalized driver mutation profiles predictions, opening new avenues in personalized medicine and targeted cancer therapy. Availability and implementation: The MATLAB-package for our SCS is freely available from http://sysbio.sibcb.ac.cn/cb/chenlab/software.htm. Contact: [email protected] or [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Weifeng Guo, Shaowu Zhang 0001, Li-Li Liu, Tao Zeng 0003, Luonan Chen
Bioinform.2
2018 NPBSS: a new PacBio sequencing simulator for generating the continuous long reads with an empirical model
abstract
BACKGROUND: PacBio sequencing platform offers longer read lengths than the second-generation sequencing technologies. It has revolutionized de novo genome assembly and enabled the automated reconstruction of reference-quality genomes. Due to its extremely wide range of application areas, fast sequencing simulation systems with high fidelity are in great demand to facilitate the development and comparison of subsequent analysis tools. Although there are several available simulators (e.g., PBSIM, SimLoRD and FASTQSim) that target the specific generation of PacBio libraries, the error rate of simulated sequences is not well matched to the quality value of raw PacBio datasets, especially for PacBio's continuous long reads (CLR). RESULTS: By analyzing the characteristic features of CLR data from PacBio SMRT (single molecule real time) sequencing, we developed a new PacBio sequencing simulator (called NPBSS) for producing CLR reads. NPBSS simulator firstly samples the read sequences according to the read length logarithmic normal distribution, and choses different base quality values with different proportions. Then, NPBSS computes the overall error probability of each base in the read sequence with an empirical model, and calculates the deletion, substitution and insertion probabilities with the overall error probability to generate the PacBio CLR reads. Alignment results demonstrate that NPBSS fits the error rate of the PacBio CLR reads better than PBSIM and FASTQSim. In addition, the assembly results also show that simulated sequences of NPBSS are more like real PacBio CLR data. CONCLUSION: NPBSS simulator is convenient to use with efficient computation and flexible parameters setting. Its generating PacBio CLR reads are more like real PacBio datasets.
Ze-Gang Wei, Shaowu Zhang 0001
BMC Bioinform.2
2018 trumpet: transcriptome-guided quality assessment of m6A-seq data
abstract
Methylated RNA immunoprecipitation sequencing (MeRIP-seq or m 6 A-seq) has been extensively used for profiling transcriptome-wide distribution of RNA N6-Methyl-Adnosine methylation. However, due to the intrinsic properties of RNA molecules and the intricate procedures of this technique, m 6 A-seq data often suffer from various flaws. A convenient and comprehensive tool is needed to assess the quality of m 6 A-seq data to ensure that they are suitable for subsequent analysis. From a technical perspective, m 6 A-seq can be considered as a combination of ChIP-seq and RNA-seq; hence, by effectively combing the data quality assessment metrics of the two techniques, we developed the trumpet R package for evaluation of m 6 A-seq data quality. The trumpet package takes the aligned BAM files from m 6 A-seq data together with the transcriptome information as the inputs to generate a quality assessment report in the HTML format. The trumpet R package makes a valuable tool for assessing the data quality of m 6 A-seq, and it is also applicable to other fragmented RNA immunoprecipitation sequencing techniques, including m 1 A-seq, CeU-Seq, Ψ-seq, etc.
Shaowu Zhang 0001, Lin Zhang 0015, Jia Meng 0001
BMC Bioinform.2
2017 QNB: differential RNA methylation analysis for count-based small-sample sequencing data with a quad-negative binomial model
abstract
BACKGROUND: As a newly emerged research area, RNA epigenetics has drawn increasing attention recently for the participation of RNA methylation and other modifications in a number of crucial biological processes. Thanks to high throughput sequencing techniques, such as, MeRIP-Seq, transcriptome-wide RNA methylation profile is now available in the form of count-based data, with which it is often of interests to study the dynamics at epitranscriptomic layer. However, the sample size of RNA methylation experiment is usually very small due to its costs; and additionally, there usually exist a large number of genes whose methylation level cannot be accurately estimated due to their low expression level, making differential RNA methylation analysis a difficult task. RESULTS: We present QNB, a statistical approach for differential RNA methylation analysis with count-based small-sample sequencing data. Compared with previous approaches such as DRME model based on a statistical test covering the IP samples only with 2 negative binomial distributions, QNB is based on 4 independent negative binomial distributions with their variances and means linked by local regressions, and in the way, the input control samples are also properly taken care of. In addition, different from DRME approach, which relies only the input control sample only for estimating the background, QNB uses a more robust estimator for gene expression by combining information from both input and IP samples, which could largely improve the testing performance for very lowly expressed genes. CONCLUSION: A-Seq, Par-CLIP, RIP-Seq, etc.
Shaowu Zhang 0001, Yufei Huang 0001, Jia Meng 0001
BMC Bioinform.2
2016 A novel algorithm for calling mRNA m6A peaks by modeling biological variances in MeRIP-seq data
abstract
MOTIVATION: N(6)-methyl-adenosine (m(6)A) is the most prevalent mRNA methylation but precise prediction of its mRNA location is important for understanding its function. A recent sequencing technology, known as Methylated RNA Immunoprecipitation Sequencing technology (MeRIP-seq), has been developed for transcriptome-wide profiling of m(6)A. We previously developed a peak calling algorithm called exomePeak. However, exomePeak over-simplifies data characteristics and ignores the reads' variances among replicates or reads dependency across a site region. To further improve the performance, new model is needed to address these important issues of MeRIP-seq data. RESULTS: We propose a novel, graphical model-based peak calling method, MeTPeak, for transcriptome-wide detection of m(6)A sites from MeRIP-seq data. MeTPeak explicitly models read count of an m(6)A site and introduces a hierarchical layer of Beta variables to capture the variances and a Hidden Markov model to characterize the reads dependency across a site. In addition, we developed a constrained Newton's method and designed a log-barrier function to compute analytically intractable, positively constrained Beta parameters. We applied our algorithm to simulated and real biological datasets and demonstrated significant improvement in detection performance and robustness over exomePeak. Prediction results on publicly available MeRIP-seq datasets are also validated and shown to be able to recapitulate the known patterns of m(6)A, further validating the improved performance of MeTPeak. AVAILABILITY AND IMPLEMENTATION: The package 'MeTPeak' is implemented in R and C ++, and additional details are available at https://github.com/compgenomics/MeTPeak CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jia Meng 0001, Shaowu Zhang 0001, Yidong Chen 0002, Yufei Huang 0001
Bioinform.3
2016 Inference of Gene Regulatory Network Based on Local Bayesian Networks
abstract
The inference of gene regulatory networks (GRNs) from expression data can mine the direct regulations among genes and gain deep insights into biological processes at a network level. During past decades, numerous computational approaches have been introduced for inferring the GRNs. However, many of them still suffer from various problems, e.g., Bayesian network (BN) methods cannot handle large-scale networks due to their high computational complexity, while information theory-based methods cannot identify the directions of regulatory interactions and also suffer from false positive/negative problems. To overcome the limitations, in this work we present a novel algorithm, namely local Bayesian network (LBN), to infer GRNs from gene expression data by using the network decomposition strategy and false-positive edge elimination scheme. Specifically, LBN algorithm first uses conditional mutual information (CMI) to construct an initial network or GRN, which is decomposed into a number of local networks or GRNs. Then, BN method is employed to generate a series of local BNs by selecting the k-nearest neighbors of each gene as its candidate regulatory genes, which significantly reduces the exponential search space from all possible GRN structures. Integrating these local BNs forms a tentative network or GRN by performing CMI, which reduces redundant regulations in the GRN and thus alleviates the false positive problem. The final network or GRN can be obtained by iteratively performing CMI and local BN on the tentative network. In the iterative process, the false or redundant regulations are gradually removed. When tested on the benchmark GRN datasets from DREAM challenge as well as the SOS DNA repair network in E.coli, our results suggest that LBN outperforms other state-of-the-art methods (ARACNE, GENIE3 and NARROMI) significantly, with more accurate and robust performance. In particular, the decomposition strategy with local Bayesian networks not only effectively reduce the computational cost of BN due to much smaller sizes of local GRNs, but also identify the directions of the regulations.
Shaowu Zhang 0001, Weifeng Guo, Ze-Gang Wei, Luonan Chen
PLoS Comput. Biol.2
2016 m6A-Driver: Identifying Context-Specific mRNA m6A Methylation-Driven Gene Interaction Networks
abstract
As the most prevalent mammalian mRNA epigenetic modification, N6-methyladenosine (m6A) has been shown to possess important post-transcriptional regulatory functions. However, the regulatory mechanisms and functional circuits of m6A are still largely elusive. To help unveil the regulatory circuitry mediated by mRNA m6A methylation, we develop here m6A-Driver, an algorithm for predicting m6A-driven genes and associated networks, whose functional interactions are likely to be actively modulated by m6A methylation under a specific condition. Specifically, m6A-Driver integrates the PPI network and the predicted differential m6A methylation sites from methylated RNA immunoprecipitation sequencing (MeRIP-Seq) data using a Random Walk with Restart (RWR) algorithm and then builds a consensus m6A-driven network of m6A-driven genes. To evaluate the performance, we applied m6A-Driver to build the context-specific m6A-driven networks for 4 known m6A (de)methylases, i.e., FTO, METTL3, METTL14 and WTAP. Our results suggest that m6A-Driver can robustly and efficiently identify m6A-driven genes that are functionally more enriched and associated with higher degree of differential expression than differential m6A methylated genes. Pathway analysis of the constructed context-specific m6A-driven gene networks further revealed the regulatory circuitry underlying the dynamic interplays between the methyltransferases and demethylase at the epitranscriptomic layer of gene regulation.
Songyao Zhang, Shaowu Zhang 0001, Jia Meng 0001, Yufei Huang 0001
PLoS Comput. Biol.2
2015 Sketching the distribution of transcriptomic features on RNA transcripts with Travis coordinates
abstract
Biological features, such as, genes, transcription factor binding sites, SNPs, etc., are usually denoted with genome-based coordinates as the genomic features. While genome-based representation is usually very effective, it can be tedious to examine the distribution of RNA-related genomic features on RNA transcripts with existing tools due to the conversion and comparison between genome-based coordinates to RNA-based coordinates. We developed here an open source R package Travis for sketching the transcriptomic view of genomic features so as to facilitate the analysis of RNA-related but genome-based coordinates. Internally, Travis package extracts the coordinates relative to the landmarks of transcripts, with which the distribution of RNA-related genomic features can then be conveniently analyzed. We demonstrated the usage of Travis package in analyzing post-transcriptional RNA modifications (5-MethylCytosine and N6-MethylAdenosine) derived from high-throughput sequencing approaches (MeRIP-Seq and RNA BS-Seq). The Travis R package is now publicly available from GitHub: https://github.com/lzcyzm/Travis.
Lin Zhang 0015, Hui Liu 0024, Shaowu Zhang 0001, Yufei Huang 0001, Jia Meng 0001
BIBM6
2013 Unveiling the dynamics in RNA epigenetic regulations
abstract
Despite the prevalent studies of DNA/Chromatin related epigenetics, such as, histone modifications and DNA methylation, RNA epigenetics did not receive deserved attention due to the lack of high throughput approach for profiling epitranscriptome. Recently, a new affinity-based sequencing approach MeRIPseq was developed and applied to survey the global mRNA N6-methyladenosine (m6A) in mammalian cells. As a marriage of ChIPseq and RNAseq, MeRIPseq has the potential to study, for the first time, the transcriptome-wide distribution of different types of post-transcriptional RNA modifications. Yet, this technology introduced new computational challenges that have not been adequately addressed. We have previously developed a MATLAB-based package ‘exomePeak’ for detection of RNA methylation sites from MeRIPseq data. Here, we extend the features of exomePeak by including a novel computational framework that enables differential analysis to unveil the dynamics in RNA epigenetic regulations. The novel differential analysis monitors the percentage of modified RNA molecules among the total transcribed RNAs, which directly reflects the impact of RNA epigenetic regulations. In contrast, current available software packages developed for sequencing-based differential analysis such as DESeq or edgeR monitors the changes in the absolute amount of molecules, and, if applied to MeRIPseq data, might be dominated by transcriptional gene differential expression. The algorithm is implemented as an R-package ‘exomePeak’ and freely available. It takes directly the aligned BAM files as input, statistically supports biological replicates, corrects PCR artifacts, and outputs exome-based results in BED format, which is compatible with all major genome browsers for convenient visualization and manipulation. Examples are also provided to depict how exomePeak R-package is integrated with exiting tools for MeRIPseq based peak calling and differential analysis. Particularly, the rationales behind each processing step as well as the specific method used, the best practice, and possible alternative strategies are briefly discussed. The algorithm was applied to the human HepG2 cell MeRIPseq data sets and detects more than 16000 RNA m6A sites, many of which are differentially methylated under ultraviolet radiation. The challenges and potentials of MeRIPseq in epitranscriptome studies are discussed in the end.
Jia Meng 0001, Hui Liu 0024, Lin Zhang 0015, Shaowu Zhang 0001, Manjeet K. Rao, Yidong Chen 0002, Yufei Huang 0001
BIBM5
2013 Detecting seasonal marine microbial communities with symmetrical non-negative matrix factorization
abstract
With the development of high-throughput and low-cost sequencing technology, a large amount of marine microbial sequences is generated. So, it is possible to research more uncultivated marine microbes. The marine microbial diversity, the association patterns among marine microbial species and environment factors are hidden in these large amount sequences. Understanding these association patterns has a high potential for exploiting the marine resources. Yet, very few marine microbial association patterns are well characterized even with the weight of research effort presently devoted to this field. In this paper, with the 16S rRNA tag pyrosequencing data taken monthly over 6 years at a temperate marine coastal sits in West English Channel, we first introduced a neighbor-seeds based heuristic clustering method called as NbHCluster by incorporating an adaptive neighbor set expanding procedure and a greedy heuristic clustering procedure, to generate the operational taxonomic units (OTUs), and utilized the mutual information (MI) algorithm to construct the spring, summer, fall, and winter seasonal marine association networks of microbe and environmental factors. Then, we used the fuzzy clustering framework by defining a clique-node similarity matrix and adopting the symmetrical non-negative matrix factorization method, to detect the association community patterns and structures in the four seasonal marine networks. The results show that the four seasonal marine microbial association networks have characters of complex networks, and the marine microbial association patterns are related with the seasonal variability; the same environmental factor influence different species in the four seasons; and the correlative relationships are stronger between OTUs (taxa) than with environmental factors.
Shaowu Zhang 0001, Ze-Gang Wei, Wei Chen 0145
BIBM1
2010 Prediction of Protein-RNA interaction site using SVM-KNN algorithm with spatial information
abstract
Protein-RNA interactions are vitally important to a number of fundamental cellular processes, including regulation of gene expression such as RNA splicing, transport and translation, protein synthesis and assembly of ribosome. More detailed information on the Protein-RNA interaction is helpful for comprehending the function notation and molecular regulatory mechanism, meanwhile, knowing the knowledge of Protein-RNA recognition can also help the biological scientist and researcher understand the site-directed mutagenesis and drug design. In the present work, we proposed a computational approach, based on SVM-KNN algorithm, with evolutionary information of spatial neighbour residues for prediction of protein-RNA interaction sites. The overall success rate obtained by 5-fold cross-validation is 78.00%, which is comparable or better than other existing methods, indicating our method is very promising for identifying and predicting protein-RNA interaction sites.
Wei Chen 0145, Shaowu Zhang 0001, Yongmei Cheng, Quan Pan 0001
BIBM2
2010 PPLook: an automated data mining tool for protein-protein interaction
abstract
BACKGROUND: Extracting and visualizing of protein-protein interaction (PPI) from text literatures are a meaningful topic in protein science. It assists the identification of interactions among proteins. There is a lack of tools to extract PPI, visualize and classify the results. RESULTS: We developed a PPI search system, termed PPLook, which automatically extracts and visualizes protein-protein interaction (PPI) from text. Given a query protein name, PPLook can search a dataset for other proteins interacting with it by using a keywords dictionary pattern-matching algorithm, and display the topological parameters, such as the number of nodes, edges, and connected components. The visualization component of PPLook enables us to view the interaction relationship among the proteins in a three-dimensional space based on the OpenGL graphics interface technology. PPLook can also provide the functions of selecting protein semantic class, counting the number of semantic class proteins which interact with query protein, counting the literature number of articles appearing the interaction relationship about the query protein. Moreover, PPLook provides heterogeneous search and a user-friendly graphical interface. CONCLUSIONS: PPLook is an effective tool for biologists and biosystem developers who need to access PPI information from the literature. PPLook is freely available for non-commercial users at http://meta.usc.edu/softs/PPLook.
Shaowu Zhang 0001, Yao-Jun Li, Li C. Xia, Quan Pan 0001
BMC Bioinform.1
2008 Prediction of Protein Homo-oligomer Types with a Novel Approach of Glide Zoom Window Feature Extraction
Qi-Peng Li, Shaowu Zhang 0001, Quan Pan 0001
ICIC (1)2
2003 Classification of protein quaternary structure with support vector machine
abstract
MOTIVATION: Since the gap between sharply increasing known sequences and slow accumulation of known structures is becoming large, an automatic classification process based on the primary sequences and known three-dimensional structure becomes indispensable. The classification of protein quaternary structure based on the primary sequences can provide some useful information for the biologists. So a fully automatic and reliable classification system is needed. This work tries to look for the effective methods of extracting attribute and the algorithm for classifying the quaternary structure from the primary sequences. RESULTS: Both of the support vector machine (SVM) and the covariant discriminant algorithms have been first introduced to predict quaternary structure properties from the protein primary sequences. The amino acid composition and the auto-correlation functions based on the amino acid index profile of the primary sequence have been taken into account in the algorithms. We have analyzed 472 amino acid indices and selected the four amino acid indices as the examples, which have the best performance. Thus the five attribute parameter data sets (COMP, FASG, NISK, WOLS and KYTJ) were established from the protein primary sequences. The COMP attribute data set is composed of amino acid composition, and the FASG, NISK, WOLS and KYTJ attribute data sets are composed of the amino acid composition and the auto-correlation functions of the corresponding amino acid residue index. The overall accuracies of SVM are 78.5, 87.5, 83.2, 81.7 and 81.9%, respectively, for COMP, FASG, NISK, WOLS and KYTJ data sets in jackknife test, which are 19.6, 7.8, 15.5, 13.1 and 15.8%, respectively, higher than that of the covariant discriminant algorithm in the same test. The results show that SVM may be applied to discriminate between the primary sequences of homodimers and non-homodimers and the two protein sequence descriptors can reflect the quaternary structure information. Compared with previous Robert Garian's investigation, the performance of SVM is almost equal to that of the Decision tree models, and the methods of extracting feature vector from the primary sequences are superior to Robert's binning function method. AVAILABILITY: Programs are available on request from the authors.
Shaowu Zhang 0001, Quan Pan 0001, Hongcai Zhang, Yun-Long Zhang, Hai-Yu Wang
Bioinform.1