Tao Zeng 0003

dblp:63/3737-3 · DBLP profile ↗
← Back
24ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-0295-3994ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 23 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2025 Hi-C3: a statistical inference-based model for reconstructing higher-order cell-cell communication networks
abstract
Multicellular organisms are composed of diverse cell types that must coordinate their behaviors through communication. Cell-cell communication (CCC) is essential for growth, development, differentiation, and immune response. Recent computational methods have leveraged single-cell RNA sequencing (scRNA-seq) to infer CCC via ligand-receptor interactions (LRIs), with most approaches focusing on pairwise interactions. However, many biological processes are driven by the coordinated action of multiple cell types, underscoring the need to model higher-order cellular interactions beyond pairwise interactions. Inspired by principles of network diffusion and epidemic dynamics, we first model the receptor expression as: a Poisson-distributed random variable biologically regulated by the collective signaling of multiple ligand-producing cell types. Then, we propose Hi-C3, a unified statistical inference-based framework for inferring both conventional pairwise and new higher-order CCC patterns from scRNA-seq data, which is solved via an efficient likelihood-based (EM) algorithm. Particularly, Hi-C3 employed a modified PageRank algorithm to assess the importance of individual cells or cell types within the higher-order network, revealing key cellular communication hubs supported by independent spatial and biological evidence. When applied to diverse datasets from Arabidopsis thaliana and colorectal cancer, Hi-C3 achieved comparable performance to state-of-the-art methods in the inferring pairwise communication while uniquely uncovering complex higher-order cellular communication structures. Collectively, Hi-C3 offers a powerful statistical model and computational framework for uncovering complex and coordinated multicellular signaling structures disregarded by pairwise communication inference methods, and offers novel biology network insights into the logic of cellular organization and communication in both development and disease.
Yuyan Tong, Renhao Hong, Weihao Deng, Tao Zeng 0003, Rui Liu 0009
Briefings Bioinform.7
2023 Latent space search based multimodal optimization with personalized edge-network biomarker for multi-purpose early disease prediction
abstract
Considering that cancer is resulting from the comutation of several essential genes of individual patients, researchers have begun to focus on identifying personalized edge-network biomarkers (PEBs) using personalized edge-network analysis for clinical practice. However, most of existing methods ignored the optimization of PEBs when multimodal biomarkers exist in multi-purpose early disease prediction (MPEDP). To solve this problem, this study proposes a novel model (MMPDENB-RBM) that combines personalized dynamic edge-network biomarkers (PDENB) theory, multimodal optimization strategy and latent space search scheme to identify biomarkers with different configurations of PDENB modules (i.e. to effectively identify multimodal PDENBs). The application to the three largest cancer omics datasets from The Cancer Genome Atlas database (i.e. breast invasive carcinoma, lung squamous cell carcinoma and lung adenocarcinoma) showed that the MMPDENB-RBM model could more effectively predict critical cancer state compared with other advanced methods. And, our model had better convergence, diversity and multimodal property as well as effective optimization ability compared with the other state-of-art methods. Particularly, multimodal PDENBs identified were more enriched with different functional biomarkers simultaneously, such as tissue-specific synthetic lethality edge-biomarkers including cancer driver genes and disease marker genes. Importantly, as our aim, these multimodal biomarkers can perform diverse biological and biomedical significances for drug target screen, survival risk assessment and novel biomedical sight as the expected multi-purpose of personalized early disease prediction. In summary, the present study provides multimodal property of PDENBs, especially the therapeutic biomarkers with more biological significances, which can help with MPEDP of individual cancer patients.
Jing J. Liang, Zong-Wei Li, Ze-Ning Sun, Ying Bi 0001, Tao Zeng 0003, Weifeng Guo
Briefings Bioinform.6
2022 Vec2image: an explainable artificial intelligence model for the feature representation and classification of high-dimensional biological data by vector-to-image conversion
abstract
Feature representation and discriminative learning are proven models and technologies in artificial intelligence fields; however, major challenges for machine learning on large biological datasets are learning an effective model with mechanistical explanation on the model determination and prediction. To satisfy such demands, we developed Vec2image, an explainable convolutional neural network framework for characterizing the feature engineering, feature selection and classifier training that is mainly based on the collaboration of principal component coordinate conversion, deep residual neural networks and embedded k-nearest neighbor representation on pseudo images of high-dimensional biological data, where the pseudo images represent feature measurements and feature associations simultaneously. Vec2image has achieved better performance compared with other popular methods and illustrated its efficiency on feature selection in cell marker identification from tissue-specific single-cell datasets. In particular, in a case study on type 2 diabetes (T2D) by multiple human islet scRNA-seq datasets, Vec2image first displayed robust performance on T2D classification model building across different datasets, then a specific Vec2image model was trained to accurately recognize the cell state and efficiently rank feature genes relevant to T2D which uncovered potential T2D cellular pathogenesis; and next the cell activity changes, cell composition imbalances and cell-cell communication dysfunctions were associated to our finding T2D feature genes from both population-shared and individual-specific perspectives. Collectively, Vec2image is a new and efficient explainable artificial intelligence methodology that can be widely applied in human-readable classification and prediction on the basis of pseudo image representation of biological deep sequencing data.
Xiangtian Yu, Rui Liu 0009, Tao Zeng 0003
Briefings Bioinform.4
2022 Deep latent space fusion for adaptive representation of heterogeneous multi-omics data
abstract
The integration of multi-omics data makes it possible to understand complex biological organisms at the system level. Numerous integration approaches have been developed by assuming a common underlying data space. Due to the noise and heterogeneity of biological data, the performance of these approaches is greatly affected. In this work, we propose a novel deep neural network architecture, named Deep Latent Space Fusion (DLSF), which integrates the multi-omics data by learning consistent manifold in the sample latent space for disease subtypes identification. DLSF is built upon a cycle autoencoder with a shared self-expressive layer, which can naturally and adaptively merge nonlinear features at each omics level into one unified sample manifold and produce adaptive representation of heterogeneous samples at the multi-omics level. We have assessed DLSF on various biological and biomedical datasets to validate its effectiveness. DLSF can efficiently and accurately capture the intrinsic manifold of the sample structures or sample clusters compared with other state-of-the-art methods, and DLSF yielded more significant outcomes for biological significance, survival prognosis and clinical relevance in application of cancer study in The Cancer Genome Atlas. Notably, as a deep case study, we determined a new molecular subtype of kidney renal clear cell carcinoma that may benefit immunotherapy in the viewpoint of multi-omics, and we further found potential subtype-specific biomarkers from multiple omics data, which were validated by independent datasets. In addition, we applied DLSF to identify potential therapeutic agents of different molecular subtypes of chronic lymphocytic leukemia, demonstrating the scalability of DLSF in diverse omics data types and application scenarios.
Chengming Zhang 0003, Yabin Chen, Tao Zeng 0003, Chuanchao Zhang, Luonan Chen
Briefings Bioinform.3
2021 Performance assessment of sample-specific network control methods for bulk and single-cell biological data analysis
abstract
In the past few years, a wealth of sample-specific network construction methods and structural network control methods has been proposed to identify sample-specific driver nodes for supporting the Sample-Specific network Control (SSC) analysis of biological networked systems. However, there is no comprehensive evaluation for these state-of-the-art methods. Here, we conducted a performance assessment for 16 SSC analysis workflows by using the combination of 4 sample-specific network reconstruction methods and 4 representative structural control methods. This study includes simulation evaluation of representative biological networks, personalized driver genes prioritization on multiple cancer bulk expression datasets with matched patient samples from TCGA, and cell marker genes and key time point identification related to cell differentiation on single-cell RNA-seq datasets. By widely comparing analysis of existing SSC analysis workflows, we provided the following recommendations and banchmarking workflows. (i) The performance of a network control method is strongly dependent on the up-stream sample-specific network method, and Cell-Specific Network construction (CSN) method and Single-Sample Network (SSN) method are the preferred sample-specific network construction methods. (ii) After constructing the sample-specific networks, the undirected network-based control methods are more effective than the directed network-based control methods. In addition, these data and evaluation pipeline are freely available on https://github.com/WilfongGuo/Benchmark_control.
Weifeng Guo, Xiangtian Yu, Jing J. Liang, Shaowu Zhang 0001, Tao Zeng 0003
PLoS Comput. Biol.6
2020 Network control principles for identifying personalized driver genes in cancer
abstract
To understand tumor heterogeneity in cancer, personalized driver genes (PDGs) need to be identified for unraveling the genotype-phenotype associations corresponding to particular patients. However, most of the existing driver-focus methods mainly pay attention on the cohort information rather than on individual information. Recent developing computational approaches based on network control principles are opening a new way to discover driver genes in cancer, particularly at an individual level. To provide comprehensive perspectives of network control methods on this timely topic, we first considered the cancer progression as a network control problem, in which the expected PDGs are altered genes by oncogene activation signals that can change the individual molecular network from one health state to the other disease state. Then, we reviewed the network reconstruction methods on single samples and introduced novel network control methods on single-sample networks to identify PDGs in cancer. Particularly, we gave a performance assessment of the network structure control-based PDGs identification methods on multiple cancer datasets from TCGA, for which the data and evaluation package also are publicly available. Finally, we discussed future directions for the application of network control methods to identify PDGs in cancer and diverse biological processes.
Weifeng Guo, Shaowu Zhang 0001, Tao Zeng 0003, Tatsuya Akutsu, Luonan Chen
Briefings Bioinform.3
2020 Efficient Mining Multi-Mers in a Variety of Biological Sequences
abstract
Counting the occurrence frequency of each $k$k-mer in a biological sequence is a preliminary yet important step in many bioinformatics applications. However, most $k$k-mer counting algorithms rely on a given $k$k to produce single-length $k$k-mers, which is inefficient for sequence analysis for different $k$k. Moreover, existing $k$k-mer counters focus more on DNA and RNA sequences and less on protein ones. In practice, the analysis of $k$k-mers in protein sequences can provide substantial biological insights in structure, function, and evolution. To this end, an efficient algorithm, called MulMer (Multiple-Mer mining), is proposed to mine $k$k-mers of various lengths termed multi-mers via inverted-index technique, which is orders of magnitude faster than the conventional forward-index methods. Moreover, to the best of our knowledge, MulMer is the first able to mine multi-mers in a variety of sequences, including DNA, RNA, and protein sequences.
Jingsong Zhang, Jianmei Guo, Xiangtian Yu, Xiaoqing Yu, Weifeng Guo, Tao Zeng 0003, Luonan Chen
IEEE ACM Trans. Comput. Biol. Bioinform.7
2019 Dynamically characterizing individual clinical change by the steady state of disease-associated pathway
abstract
BACKGROUND: Along with the development of precision medicine, individual heterogeneity is attracting more and more attentions in clinical research and application. Although the biomolecular reaction seems to be some various when different individuals suffer a same disease (e.g. virus infection), the final pathogen outcomes of individuals always can be mainly described by two categories in clinics, i.e. symptomatic and asymptomatic. Thus, it is still a great challenge to characterize the individual specific intrinsic regulatory convergence during dynamic gene regulation and expression. Except for individual heterogeneity, the sampling time also increase the expression diversity, so that, the capture of similar steady biological state is a key to characterize individual dynamic biological processes. RESULTS: Assuming the similar biological functions (e.g. pathways) should be suitable to detect consistent functions rather than chaotic genes, we design and implement a new computational framework (ABP: Attractor analysis of Boolean network of Pathway). ABP aims to identify the dynamic phenotype associated pathways in a state-transition manner, using the network attractor to model and quantify the steady pathway states characterizing the final steady biological sate of individuals (e.g. normal or disease). By analyzing multiple temporal gene expression datasets of virus infections, ABP has shown its effectiveness on identifying key pathways associated with phenotype change; inferring the consensus functional cascade among key pathways; and grouping pathway activity states corresponding to disease states. CONCLUSIONS: Collectively, ABP can detect key pathways and infer their consensus functional cascade during dynamical process (e.g. virus infection), and can also categorize individuals with disease state well, which is helpful for disease classification and prediction.
Shaoyan Sun, Xiangtian Yu, Fengnan Sun, Tao Zeng 0003
BMC Bioinform.6
2019 A novel network control model for identifying personalized driver genes in cancer
abstract
Although existing computational models have identified many common driver genes, it remains challenging to identify the personalized driver genes by using samples of an individual patient. Recently, the methods of exploiting the structure-based control principles of complex networks provide new clues for identifying minimum number of driver nodes to drive the state transition of large-scale complex networks from an initial state to the desired state. However, the structure-based network control methods cannot be directly applied to identify the personalized driver genes due to the unknown network dynamics of the personalized system. Here we proposed the personalized network control model (PNC) to identify the personalized driver genes by employing the structure-based network control principle on genetic data of individual patients. In PNC model, we firstly presented a paired single sample network construction method to construct the personalized state transition network for capturing the phenotype transitions between healthy and disease states. Then, we designed a novel structure-based network control method from the Feedback Vertex Sets-based control perspective to identify the personalized driver genes. The wide experimental results on 13 cancer datasets from The Cancer Genome Atlas firstly showed that PNC model outperforms current state-of-the-art methods, in terms of F-measures for identifying cancer driver genes enriched in the gold-standard cancer driver gene lists. Furthermore, these results showed that personalized driver genes can be explored by their network characteristics even when they are hidden factors in transcription and mutation profiles. Our PNC gives novel insights and useful tools into understanding the tumor heterogeneity in cancer. The PNC package and data resources used in this work can be freely downloaded from https://github.com/NWPU-903PR/PNC.
Weifeng Guo, Shaowu Zhang 0001, Tao Zeng 0003, Yan Li 0111, Jianxi Gao, Luonan Chen
PLoS Comput. Biol.3
2018 Characterizing and Discriminating Individual Steady State of Disease-Associated Pathway
Shaoyan Sun, Xiangtian Yu, Fengnan Sun, Tao Zeng 0003
ICIC (1)6
2018 Discovering personalized driver mutation profiles of single samples in cancer by network control strategy
abstract
Motivation: It is a challenging task to discover personalized driver genes that provide crucial information on disease risk and drug sensitivity for individual patients. However, few methods have been proposed to identify the personalized-sample driver genes from the cancer omics data due to the lack of samples for each individual. To circumvent this problem, here we present a novel single-sample controller strategy (SCS) to identify personalized driver mutation profiles from network controllability perspective. Results: SCS integrates mutation data and expression data into a reference molecular network for each patient to obtain the driver mutation profiles in a personalized-sample manner. This is the first such a computational framework, to bridge the personalized driver mutation discovery problem and the structural network controllability problem. The key idea of SCS is to detect those mutated genes which can achieve the transition from the normal state to the disease state based on each individual omics data from network controllability perspective. We widely validate the driver mutation profiles of our SCS from three aspects: (i) the improved precision for the predicted driver genes in the population compared with other driver-focus methods; (ii) the effectiveness for discovering the personalized driver genes and (iii) the application to the risk assessment through the integration of the driver mutation signature and expression data, respectively, across the five distinct benchmarks from The Cancer Genome Atlas. In conclusion, our SCS makes efficient and robust personalized driver mutation profiles predictions, opening new avenues in personalized medicine and targeted cancer therapy. Availability and implementation: The MATLAB-package for our SCS is freely available from http://sysbio.sibcb.ac.cn/cb/chenlab/software.htm. Contact: [email protected] or [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Weifeng Guo, Shaowu Zhang 0001, Li-Li Liu, Tao Zeng 0003, Luonan Chen
Bioinform.8
2017 Mining K-mers of Various Lengths in Biological Sequences
Jingsong Zhang, Jianmei Guo, Xiaoqing Yu, Xiangtian Yu, Weifeng Guo, Tao Zeng 0003, Luonan Chen
ISBRA6
2017 Pattern fusion analysis by adaptive alignment of multiple heterogeneous omics data
abstract
MOTIVATION: Integrating different omics profiles is a challenging task, which provides a comprehensive way to understand complex diseases in a multi-view manner. One key for such an integration is to extract intrinsic patterns in concordance with data structures, so as to discover consistent information across various data types even with noise pollution. Thus, we proposed a novel framework called 'pattern fusion analysis' (PFA), which performs automated information alignment and bias correction, to fuse local sample-patterns (e.g. from each data type) into a global sample-pattern corresponding to phenotypes (e.g. across most data types). In particular, PFA can identify significant sample-patterns from different omics profiles by optimally adjusting the effects of each data type to the patterns, thereby alleviating the problems to process different platforms and different reliability levels of heterogeneous data. RESULTS: To validate the effectiveness of our method, we first tested PFA on various synthetic datasets, and found that PFA can not only capture the intrinsic sample clustering structures from the multi-omics data in contrast to the state-of-the-art methods, such as iClusterPlus, SNF and moCluster, but also provide an automatic weight-scheme to measure the corresponding contributions by data types or even samples. In addition, the computational results show that PFA can reveal shared and complementary sample-patterns across data types with distinct signal-to-noise ratios in Cancer Cell Line Encyclopedia (CCLE) datasets, and outperforms over other works at identifying clinically distinct cancer subtypes in The Cancer Genome Atlas (TCGA) datasets. AVAILABILITY AND IMPLEMENTATION: PFA has been implemented as a Matlab package, which is available at http://www.sysbio.ac.cn/cb/chenlab/images/PFApackage_0.1.rar . CONTACT: [email protected] , [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Chuanchao Zhang, Minrui Peng, Xiangtian Yu, Tao Zeng 0003, Juan Liu 0007, Luonan Chen
Bioinform.5
2017 Comparative network stratification analysis for identifying functional interpretable network biomarkers
abstract
BACKGROUND: A major challenge of bioinformatics in the era of precision medicine is to identify the molecular biomarkers for complex diseases. It is a general expectation that these biomarkers or signatures have not only strong discrimination ability, but also readable interpretations in a biological sense. Generally, the conventional expression-based or network-based methods mainly capture differential genes or differential networks as biomarkers, however, such biomarkers only focus on phenotypic discrimination and usually have less biological or functional interpretation. Meanwhile, the conventional function-based methods could consider the biomarkers corresponding to certain biological functions or pathways, but ignore the differential information of genes, i.e., disregard the active degree of particular genes involved in particular functions, thereby resulting in less discriminative ability on phenotypes. Hence, it is strongly demanded to develop elaborate computational methods to directly identify functional network biomarkers with both discriminative power on disease states and readable interpretation on biological functions. RESULTS: In this paper, we present a new computational framework based on an integer programming model, named as Comparative Network Stratification (CNS), to extract functional or interpretable network biomarkers, which are of strongly discriminative power on disease states and also readable interpretation on biological functions. In addition, CNS can not only recognize the pathogen biological functions disregarded by traditional Expression-based/Network-based methods, but also uncover the active network-structures underlying such dysregulated functions underestimated by traditional Function-based methods. To validate the effectiveness, we have compared CNS with five state-of-the-art methods, i.e. GSVA, Pathifier, stSVM, frSVM and AEP on four datasets of different complex diseases. The results show that CNS can enhance the discriminative power of network biomarkers, and further provide biologically interpretable information or disease pathogenic mechanism of these biomarkers. A case study on type 1 diabetes (T1D) demonstrates that CNS can identify many dysfunctional genes and networks previously disregarded by conventional approaches. CONCLUSION: Therefore, CNS is actually a powerful bioinformatics tool, which can identify functional or interpretable network biomarkers with both discriminative power on disease states and readable interpretation on biological functions. CNS was implemented as a Matlab package, which is available at http://www.sysbio.ac.cn/cb/chenlab/images/CNSpackage_0.1.rar .
Chuanchao Zhang, Juan Liu 0007, Tao Zeng 0003, Luonan Chen
BMC Bioinform.4
2017 Differential function analysis: identifying structure and activation variations in dysregulated pathways
Chuanchao Zhang, Juan Liu 0007, Tao Zeng 0003, Luonan Chen
Sci. China Inf. Sci.4
2016 Integration of multiple heterogeneous omics data
abstract
Integration of different genomic profiles is challenging to understand complex diseases in a multi-view manner. Computational method is needed to preserve useful information of data types as well as correct bias. Thus, we proposed a novel framework pattern fusion analysis (PFA), to fuse the local sample patterns into a global pattern of patients with respect to the underlying data, by adaptively aligning the information in each type of biological data. In particular, PFA can adjust the distinct data types and achieve more robust sample pattern within different profiles. To validate the effectiveness of PFA, we tested PFA on various synthetic datasets and found that PFA is able to effectively capture the intrinsic clustering structure than the state-of-the-art integrative methods, such as moCluster, iClusterPlus and SNF. Moreover, in a case study on kidney cancer, PFA not only identified the multi-way feature modules among the prior-known disease associated genes, methylations and miRNAs, but also outperformed in cancer subtypes identification and could get effective clinical prognosis prediction. Totally, PFA not only provides new insights on the more holistic & systems-level sample pattern, but also supplies a new way for selecting more informative types of biological data.
Chuanchao Zhang, Juan Liu 0007, Xiangtian Yu, Tao Zeng 0003, Luonan Chen
BIBM5
2016 Big-data-based edge biomarkers: study on dynamical drug sensitivity and resistance in individuals
abstract
Big-data-based edge biomarker is a new concept to characterize disease features based on biomedical big data in a dynamical and network manner, which also provides alternative strategies to indicate disease status in single samples. This article gives a comprehensive review on big-data-based edge biomarkers for complex diseases in an individual patient, which are defined as biomarkers based on network information and high-dimensional data. Specifically, we firstly introduce the sources and structures of biomedical big data accessible in public for edge biomarker and disease study. We show that biomedical big data are typically 'small-sample size in high-dimension space', i.e. small samples but with high dimensions on features (e.g. omics data) for each individual, in contrast to traditional big data in many other fields characterized as 'large-sample size in low-dimension space', i.e. big samples but with low dimensions on features. Then, we demonstrate the concept, model and algorithm for edge biomarkers and further big-data-based edge biomarkers. Dissimilar to conventional biomarkers, edge biomarkers, e.g. module biomarkers in module network rewiring-analysis, are able to predict the disease state by learning differential associations between molecules rather than differential expressions of molecules during disease progression or treatment in individual patients. In particular, in contrast to using the information of the common molecules or edges (i.e.molecule-pairs) across a population in traditional biomarkers including network and edge biomarkers, big-data-based edge biomarkers are specific for each individual and thus can accurately evaluate the disease state by considering the individual heterogeneity. Therefore, the measurement of big data in a high-dimensional space is required not only in the learning process but also in the diagnosing or predicting process of the tested individual. Finally, we provide a case study on analyzing the temporal expression data from a malaria vaccine trial by big-data-based edge biomarkers from module network rewiring-analysis. The illustrative results show that the identified module biomarkers can accurately distinguish vaccines with or without protection and outperformed previous reported gene signatures in terms of effectiveness and efficiency.
Tao Zeng 0003, Wanwei Zhang, Xiangtian Yu, Xiaoping Liu 0002, Meiyi Li, Luonan Chen
Briefings Bioinform.1
2015 Identification of phenotypic networks based on whole transcriptome by comparative network decomposition
abstract
Complex diseases are usually caused by the dysfunctions of the molecular system or molecular network rather than individual molecules. Generally, the conventional methods first obtain a disease-associated network based on expression data and then study its biological functions. However, such a network may be only a part of the system facilitating a biological function or may involve in multiple functions. In this paper, we present a computational framework based on an integer programming model, named as comparative network decomposition (CND), to jointly identify optimal structures of significant and moderate phenotypic functions/networks and their optimal combination by integrating gene expression, gene network and gene ontology together. Particularly, CND makes full use of dysfunctional information, e.g. both strong and weak changes on gene expressions and correlations, to extract various phenotypic networks, where one phenotypic network just corresponds to a specific biological function. A synthetic example clearly suggests that CND can identify multiple types of the disease-related phenotypic networks, rather than conventional approaches only exact significant phenotypic networks. As a proof-of-concept study to real data, CND is further used to identify the significant and moderate phenotypic networks for discriminating two different but associated diseases, e.g. subtypes of diabetes. In the comparison of type 1 and type 2 diabetes, the moderate and significant phenotypic networks can capture the disease-related biological functions and their corresponding networks. Therefore, CND is actually a powerful bioinformatics tool, which can investigate phenotype-associated genes and networks in a whole transcriptome and function-centered manner, and the comparative study of complex diseases with other works also demonstrates its effectiveness.
Chuanchao Zhang, Juan Liu 0007, Tao Zeng 0003, Luonan Chen
BIBM4
2015 Inferring Sequential Order of Somatic Mutations during Tumorgenesis based on Markov Chain Model
abstract
Tumors are developed and worsen with the accumulated mutations on DNA sequences during tumorigenesis. Identifying the temporal order of gene mutations in cancer initiation and development is a challenging topic. It not only provides a new insight into the study of tumorigenesis at the level of genome sequences but also is an effective tool for early diagnosis of tumors and preventive medicine. In this paper, we develop a novel method to accurately estimate the sequential order of gene mutations during tumorigenesis from genome sequencing data based on Markov chain model as TOMC (Temporal Order based on Markov Chain), and also provide a new criterion to further infer the order of samples or patients, which can characterize the severity or stage of the disease. We applied our method to the analysis of tumors based on several high-throughput datasets. Specifically, first, we revealed that tumor suppressor genes (TSG) tend to be mutated ahead of oncogenes, which are considered as important events for key functional loss and gain during tumorigenesis. Second, the comparisons of various methods demonstrated that our approach has clear advantages over the existing methods due to the consideration on the effect of mutation dependence among genes, such as co-mutation. Third and most important, our method is able to deduce the ordinal sequence of patients or samples to quantitatively characterize their severity of tumors. Therefore, our work provides a new way to quantitatively understand the development and progression of tumorigenesis based on high throughput sequencing data.
Hao Kang, Kwang-Hyun Cho, Xiaohua Douglas Zhang, Tao Zeng 0003, Luonan Chen
IEEE ACM Trans. Comput. Biol. Bioinform.4
2014 MMSE: A generalized coherence measure for identifying linear patterns
abstract
Biclustering is very useful in bioinformatics, information retrieval, electoral data analysis, dimension reduction, and so on. It is usually formulated as an optimization problem of searching maximal subsets of rows and columns satisfying some coherence criteria. The found submatrices are called as biclusters. There are several quantitative coherence measurements for linear patterns proposed. However, they are either lack the capability of properly evaluating all subtypes of linear patterns, or sensitive to the noise. In this paper, we propose a coherence measurement for the general linear patterns, the minimal mean squared error (MMSE). By using MMSE, the biclustering algorithms are expected to identify all types of linear patterns, including shifting (additive), scaling (multiplicative), and the general linear (the mixed form of shifting and scaling) ones, if they are presented in the data. Our comparative experimental results have highlighted that MMSE can actually help to identify significant general linear biclusters in artificial and real application data.
Shuhua Chen, Juan Liu 0007, Tao Zeng 0003
BIBM3
2014 Detecting tissue-specific early warning signals for complex diseases based on dynamical network biomarkers: study of type 2 diabetes by cross-tissue analysis
abstract
Identifying early warning signals of critical transitions during disease progression is a key to achieving early diagnosis of complex diseases. By exploiting rich information of high-throughput data, a novel model-free method has been developed to detect early warning signals of diseases. Its theoretical foundation is based on dynamical network biomarker (DNB), which is also called as the driver (or leading) network of the disease because components or molecules in DNB actually drive the whole system from one state (e.g. normal state) to another (e.g. disease state). In this article, we first reviewed the concept and main results of DNB theory, and then applied the new method to the analysis of type 2 diabetes mellitus (T2DM). Specifically, based on the temporal-spatial gene expression data of T2DM, we identified tissue-specific DNBs corresponding to the critical transitions occurring in liver, adipose and muscle during T2DM development and progression. Actually, we found that there are two different critical states during T2DM development characterized as responses to insulin resistance and serious inflammation, respectively. Interestingly, a new T2DM-associated function, i.e. steroid hormone biosynthesis, was discovered, and those related genes were significantly dysregulated in liver and adipose at the first critical transition during T2DM deterioration. Moreover, the dysfunction of genes related to responding hormone was also detected in muscle at the similar period. Based on the functional and network analysis on pathogenic molecular mechanism of T2DM, we showed that most of DNB genes, in particular the core ones, tended to be located at the upstream of biological pathways, which implied that DNB genes act as the causal factors rather than the consequence to drive the downstream molecules to change their transcriptional activities. This also validated our theoretical prediction of DNB as the driver network. As shown in this study, DNB can not only signal the emergence of the critical transitions for early diagnosis of diseases, but can also provide the causal network of the transitions for revealing molecular mechanisms of disease initiation and progression at a network level.
Meiyi Li, Tao Zeng 0003, Rui Liu 0009, Luonan Chen
Briefings Bioinform.2
2011 Prediction of Heme Binding Sites in Heme Proteins Using an Integrative Sequence Profile Coupling Evolutionary Information with Physicochemical Properties
abstract
Heme-protein interactions are essential for various biological processes such as electron transfer, catalysis, signal transduction and the control of gene expression. The knowledge of heme binding residues can provide crucial clues to understand the mechanism of heme- protein interactions and aid in functional annotation. In the present work, we propose a sequence-based approach for the accurate prediction of heme binding residues by a novel integrative sequence profile coupling position specific scoring matrices with heme specific physicochemical properties. Particularly, we design an intuitive feature selection scheme for informative physicochemical properties. As shown in the primary results, our integrative sequence profile approach for prediction of heme binding residues outperforms the conventional methods using amino acid and evolutionary information on the 5-fold cross validation and the independent test.
Yi Xiong 0002, Wen Zhang 0008, Tao Zeng 0003, Juan Liu 0007
BIBM3
2010 Discovering negative correlated gene sets from integrative gene expression data for cancer prognosis
abstract
Along with the emergence and development of translational biomedicine, more and more genetic information has been applied in clinical practice. In recent decade, the discovery of genetic biomarkers for cancer prognosis obtains increasing attentions and many methods have been developed. The "element" methods use one or two independent genes to judge the Boolean status of disease. The "set" methods use general genetic biomarkers to classify patients into different risks as a whole. And the advanced "sets" methods use a group of different gene sets as biomarkers. However, the existing methods always concern positive correlations among genes ignoring negative correlations. Whereas the negative regulation, negative feedback, and functional repression are actually the important clues in cancer expression profiles. Therefore, in this paper, we propose to mine negative correlated gene sets (NCGSs) from multiple datasets, and use them along with the pure positive correlated gene sets for prognosis classification. The exploring experimental results have shown the encouraging promotion of cancer prognosis accuracy with NCGSs.
Tao Zeng 0003, Juan Liu 0007
BIBM1
2010 Mixture classification model based on clinical markers for breast cancer prognosis
Tao Zeng 0003, Juan Liu 0007
Artif. Intell. Medicine1