EDBT 2026 Demo / reviewers in the wild / expert
Luonan Chen
dblp:83/4411
· DBLP profile ↗
109ranked-venue papers
9as first author
27since 2021 · last 2026
0000-0002-3960-0068ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 90 · 6 first-author · 19 since 2021Artificial intelligence and machine learning · 14 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamical Causality Under Latent Confounders for Biological Network ReconstructionabstractCausal interaction inference is prone to spurious causal interactions, due to the substantial confounders in a biological system. While many existing methods attempt to address misidentification challenges, there remains a notable lack of effective methods to infer causal interaction under latent/unobserved confounders. In this work, we propose a method to overcome such challenges to infer dynamical causality under invisible confounders (CIC) and further reconstruct the latent confounders from time-series data by developing an orthogonal decomposition theorem in a delay embedding space. This theoretical foundation ensures the causal detection for any high-dimensional system even with only two observed variables under many latent confounders, which is a long-standing problem in the field. In addition to the latent confounder problem, such a decomposition makes the coupled variables separable in the embedding space, thus also solving the non-separability problem of causal inference. Extensive validation of the CIC method is carried out using various real datasets, which all demonstrates its effectiveness to reconstruct real biological networks and unobserved confounders. Jinling Yan, Shaowu Zhang 0001, Chihao Zhang 0002, Weitian Huang, Jifan Shi, Luonan Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Brain-Inspired Chaotic Graph Backpropagation for Combinatorial OptimizationabstractGraph neural networks (GNNs) with unsupervised learning can provide high-quality approximate solutions to large-scale combinatorial optimization problems (COPs) with efficient time complexity, making them versatile for various applications. However, since this method maps the COP to the training process of a GNN, and the current mainstream backpropagation-based training algorithms are prone to falling into local minima, the optimization performance is still inferior to the current state-of-the-art (SOTA) COP methods. To address this issue, inspired by the possibility of learning through chaotic dynamics of the real brain, we introduce a chaotic training algorithm, i.e., chaotic graph backpropagation (CGBP), which introduces a local loss function in GNN that makes the training process not only chaotic but also highly efficient. Different from existing methods, we show that the global ergodicity and pseudorandomness with fractal structure of such chaotic dynamics enable CGBP to learn GNNs effectively and globally, thus solving the COP efficiently. We have applied CGBP to solve various COPs, such as the maximum independent set (MIS), maximum cut (MC), and graph coloring (GC). Results on several large-scale benchmark datasets showcase that CGBP can compete with or outperform SOTA methods. In addition, CGBP can be easily integrated into any existing learning method as an additional universal plug-in module to improve the searching ability and performance. Peng Tao 0009, Kazuyuki Aihara, Luonan Chen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Self-Assembling Graph PerceptronsabstractInspired by the workings of biological brains, humans have designed artificial neural networks (ANNs), sparking profound advancements across various fields. However, the biological brain possesses high plasticity, enabling it to develop simple, efficient, and powerful structures to cope with complex external environments. In contrast, the superior performance of ANNs often relies on meticulously crafted architectures, which can make them vulnerable when handling complex inputs. Moreover, overparameterization often characterizes the most advanced ANNs. This paper explores the path toward building streamlined and plastic ANNs. Firstly, we introduce the Graph Perceptron (GP), which extends the most fundamental ANN, the Multi-Layer Perceptron (MLP). Subsequently, we incorporate a self-assembly mechanism on top of GP called Self-Assembling Graph Perceptron (SAGP). During training, SAGP can autonomously adjust the network's number of neurons and synapses and their connectivity. SAGP achieves comparable or even superior performance with only about 5% of the size of an MLP. We also demonstrate the SAGP's advantages in enhancing model interpretability and feature selection. Bowen Deng 0002, Luonan Chen, Zibin Zheng, Chuan Chen 0001 |
NeurIPS | 4 |
| 2025 | MCGAE: unraveling tumor invasion through integrated multimodal spatial transcriptomicsabstractSpatially Resolved Transcriptomics (SRT) serves as a cornerstone in biomedical research, revealing the heterogeneity of tissue microenvironments. Integrating multimodal data including gene expression, spatial coordinates, and morphological information poses significant challenges for accurate spatial domain identification. Herein, we present the Multi-view Contrastive Graph Autoencoder (MCGAE), a cutting-edge deep computational framework specifically designed for the intricate analysis of spatial transcriptomics (ST) data. MCGAE advances the field by creating multi-view representations from gene expression and spatial adjacency matrices. Utilizing modular modeling, contrastive graph convolutional networks, and attention mechanisms, it generates modality-specific spatial representations and integrates them into a unified embedding. This integration process is further enriched by the inclusion of morphological image features, markedly enhancing the framework's capability to process multimodal data. Applied to both simulated and real SRT datasets, MCGAE demonstrates superior performance in spatial domain detection, data denoising, trajectory inference, and 3D feature extraction, outperforming existing methods. Specifically, in colorectal cancer liver metastases, MCGAE integrates histological and gene expression data to identify tumor invasion regions and characterize cellular molecular regulation. This breakthrough extends ST analysis and offers new tools for cancer and complex disease research. Chengming Zhang 0003, Zhaonan Liu, Kazuyuki Aihara, Chuanchao Zhang, Luonan Chen |
Briefings Bioinform. | 6 |
| 2025 | DNFE: Directed network flow entropy for detecting tipping points during biological processesabstractTypically, in dynamic biological processes, there is a critical state or tipping point that marks the transition from one stable state to another, surpassing which a considerable qualitative shift takes place. Identifying this tipping point and its driving network is essential to avert or delay disastrous outcomes. However, most traditional approaches built upon undirected networks still suffer from a lack of robustness and effectiveness when implemented based on high-dimensional small-sample data, especially for single-cell data. To address this challenge, we develop a directed network flow entropy (DNFE) method, which can transform measured omics data into a directed network. This method is applicable to both single-cell RNA-sequencing (scRNA-seq) and bulk data. Applying this algorithm to six real datasets, including three single-cell datasets, two bulk tumor datasets, and a blood dataset, the method is proved to be effective not only in identifying critical states, as well as their dynamic network biomarkers, but also in helping explore regulatory relationships between genes. Numerical simulation results demonstrate that the DNFE algorithm is robust across various noise levels and outperforms existing methods in detecting tipping points. Furthermore, the numerical simulations for 100-node and 1000-node gene regulatory networks illustrate the method's application for large-scale data. The DNFE method predicts active transcription factors, and further identified "dark genes", which are usually overlooked with traditional methods. Xueqing Peng, Peiluan Li, Luonan Chen |
PLoS Comput. Biol. | 4 |
| 2025 | Decoding the gut-brain axis: toward AI-driven integration of neuroimaging and gut microbiota in human healthabstractThe gut–brain axis (GBA) represents a complex, bidirectional communication network between the gut microbiome and the central nervous system, influencing both neurological health and the pathogenesis of various diseases. This review explores the integrative role of neuroimaging and machine learning (ML) in advancing our understanding of microbiota–brain interactions, emphasizing their combined potential to uncover novel biomarkers and therapeutic targets. Neuroimaging techniques, including functional magnetic resonance imaging (MRI), diffusion tensor imaging, and structural MRI, have revealed how gut microbiota imbalances, or dysbiosis, affect key brain networks and structural connectivity, contributing to cognitive dysfunction, emotional disturbances, and neurodegenerative conditions. ML methodologies, including deep learning (DL) and multimodal data fusion, are proving indispensable in extracting meaningful insights from high-dimensional neuroimaging and microbiome datasets. Supervised approaches, such as random forests and deep neural networks, have achieved high accuracy in predicting neurological outcomes based on microbial signatures, while unsupervised learning identifies distinct microbiota–brain connectivity patterns associated with disorders such as autism spectrum disorder and depression. Additionally, explainable AI (XAI) techniques are being increasingly applied to enhance the interpretability of ML-driven biomarker discovery, shedding light on the neuroprotective effects of butyrate-producing bacteria (e.g., Faecalibacterium , Roseburia ) and the potential for neuroinflammation linked to an overabundance of Proteobacteria . These findings point to the transformative potential of combining neuroimaging and ML in precision medicine, offering a new paradigm for the diagnosis and treatment of neurological disorders ranging from irritable bowel syndrome to Alzheimer’s disease. However, challenges related to data harmonization, generalizability across populations, and establishing causal relationships remain, necessitating further research to realize the full clinical potential of this approach. Dezhi Wu, Yueqiong Ni, Yurun Lu, Huating Li, Luonan Chen |
Vis. Comput. | 7 |
| 2024 | Survival Analysis of Histopathological Image Based on a Pretrained Hypergraph Model of Spatial Transcriptomics Data
Shangyan Cai, Weitian Huang, Weiting Yi, Hongmin Cai, Luonan Chen, Weifeng Su |
MICCAI (3) | 8 |
| 2024 | Multi-modal domain adaptation for revealing spatial functional landscape from spatially resolved transcriptomicsabstractSpatially resolved transcriptomics (SRT) has emerged as a powerful tool for investigating gene expression in spatial contexts, providing insights into the molecular mechanisms underlying organ development and disease pathology. However, the expression sparsity poses a computational challenge to integrate other modalities (e.g. histological images and spatial locations) that are simultaneously captured in SRT datasets for spatial clustering and variation analyses. In this study, to meet such a challenge, we propose multi-modal domain adaption for spatial transcriptomics (stMDA), a novel multi-modal unsupervised domain adaptation method, which integrates gene expression and other modalities to reveal the spatial functional landscape. Specifically, stMDA first learns the modality-specific representations from spatial multi-modal data using multiple neural network architectures and then aligns the spatial distributions across modal representations to integrate these multi-modal representations, thus facilitating the integration of global and spatially local information and improving the consistency of clustering assignments. Our results demonstrate that stMDA outperforms existing methods in identifying spatial domains across diverse platforms and species. Furthermore, stMDA excels in identifying spatially variable genes with high prognostic potential in cancer tissues. In conclusion, stMDA as a new tool of multi-modal data integration provides a powerful and flexible framework for analyzing SRT datasets, thereby advancing our understanding of intricate biological systems. Lequn Wang, Yaofeng Hu, Chuanchao Zhang, Qianqian Shi 0004, Luonan Chen |
Briefings Bioinform. | 6 |
| 2024 | Revealing cell-cell communication pathways with their spatially coupled gene programsabstractInference of cell-cell communication (CCC) provides valuable information in understanding the mechanisms of many important life processes. With the rise of spatial transcriptomics in recent years, many methods have emerged to predict CCCs using spatial information of cells. However, most existing methods only describe CCCs based on ligand-receptor interactions, but lack the exploration of their upstream/downstream pathways. In this paper, we proposed a new method to infer CCCs, called Intercellular Gene Association Network (IGAN). Specifically, it is for the first time that we can estimate the gene associations/network between two specific single spatially adjacent cells. By using the IGAN method, we can not only infer CCCs in an accurate manner, but also explore the upstream/downstream pathways of ligands/receptors from the network perspective, which are actually exhibited as a new panoramic cell-interaction-pathway graph, and thus provide extensive information for the regulatory mechanisms behind CCCs. In addition, IGAN can measure the CCC activity at single cell/spot resolution, and help to discover the CCC spatial heterogeneity. Interestingly, we found that CCC patterns from IGAN are highly consistent with the spatial microenvironment patterns for each cell type, which further indicated the accuracy of our method. Analyses on several public datasets validated the advantages of IGAN. Junchao Zhu, Luonan Chen |
Briefings Bioinform. | 3 |
| 2023 | Single-cell causal network inferred by cross-mapping entropyabstractGene regulatory networks (GRNs) reveal the complex molecular interactions that govern cell state. However, it is challenging for identifying causal relations among genes due to noisy data and molecular nonlinearity. Here, we propose a novel causal criterion, neighbor cross-mapping entropy (NME), for inferring GRNs from both steady data and time-series data. NME is designed to quantify 'continuous causality' or functional dependency from one variable to another based on their function continuity with varying neighbor sizes. NME shows superior performance on benchmark datasets, comparing with existing methods. By applying to scRNA-seq datasets, NME not only reliably inferred GRNs for cell types but also identified cell states. Based on the inferred GRNs and further their activity matrices, NME showed better performance in single-cell clustering and downstream analyses. In summary, based on continuous causality, NME provides a powerful tool in inferring causal regulations of GRNs between genes from scRNA-seq data, which is further exploited to identify novel cell types/states and predict cell type-specific network modules. Lin Li 0069, Peng Tao 0009, Luonan Chen |
Briefings Bioinform. | 6 |
| 2023 | Contrastively generative self-expression model for single-cell and spatial multimodal dataabstractAdvances in single-cell multi-omics technology provide an unprecedented opportunity to fully understand cellular heterogeneity. However, integrating omics data from multiple modalities is challenging due to the individual characteristics of each measurement. Here, to solve such a problem, we propose a contrastive and generative deep self-expression model, called single-cell multimodal self-expressive integration (scMSI), which integrates the heterogeneous multimodal data into a unified manifold space. Specifically, scMSI first learns each omics-specific latent representation and self-expression relationship to consider the characteristics of different omics data by deep self-expressive generative model. Then, scMSI combines these omics-specific self-expression relations through contrastive learning. In such a way, scMSI provides a paradigm to integrate multiple omics data even with weak relation, which effectively achieves the representation learning and data integration into a unified framework. We demonstrate that scMSI provides a cohesive solution for a variety of analysis tasks, such as integration analysis, data denoising, batch correction and spatial domain detection. We have applied scMSI on various single-cell and spatial multimodal datasets to validate its high effectiveness and robustness in diverse data types and application scenarios. Chengming Zhang 0003, Shijie Tang, Kazuyuki Aihara, Chuanchao Zhang, Luonan Chen |
Briefings Bioinform. | 6 |
| 2023 | Robust joint clustering of multi-omics single-cell data via multi-modal high-order neighborhood Laplacian matrix optimizationabstractMOTIVATION: Simultaneous profiling of multi-omics single-cell data represents exciting technological advancements for understanding cellular states and heterogeneity. Cellular indexing of transcriptomes and epitopes by sequencing allowed for parallel quantification of cell-surface protein expression and transcriptome profiling in the same cells; methylome and transcriptome sequencing from single cells allows for analysis of transcriptomic and epigenomic profiling in the same individual cells. However, effective integration method for mining the heterogeneity of cells over the noisy, sparse, and complex multi-modal data is in growing need. RESULTS: In this article, we propose a multi-modal high-order neighborhood Laplacian matrix optimization framework for integrating the multi-omics single-cell data: scHoML. Hierarchical clustering method was presented for analyzing the optimal embedding representation and identifying cell clusters in a robust manner. This novel method by integrating high-order and multi-modal Laplacian matrices would robustly represent the complex data structures and allow for systematic analysis at the multi-omics single-cell level, thus promoting further biological discoveries. AVAILABILITY AND IMPLEMENTATION: Matlab code is available at https://github.com/jianghruc/scHoML. Hao Jiang 0009, Senwen Zhan, Wai-Ki Ching, Luonan Chen |
Bioinform. | 4 |
| 2023 | Predicting time series by data-driven spatiotemporal information transformation
Peng Tao 0009, Xiaohu Hao, Luonan Chen |
Inf. Sci. | 4 |
| 2023 | Cell-type annotation with accurate unseen cell-type identification using multiple referencesabstractThe recent advances in single-cell RNA sequencing (scRNA-seq) techniques have stimulated efforts to identify and characterize the cellular composition of complex tissues. With the advent of various sequencing techniques, automated cell-type annotation using a well-annotated scRNA-seq reference becomes popular. But it relies on the diversity of cell types in the reference, which may not capture all the cell types present in the query data of interest. There are generally unseen cell types in the query data of interest because most data atlases are obtained for different purposes and techniques. Identifying previously unseen cell types is essential for improving annotation accuracy and uncovering novel biological discoveries. To address this challenge, we propose mtANN (multiple-reference-based scRNA-seq data annotation), a new method to automatically annotate query data while accurately identifying unseen cell types with the aid of multiple references. Key innovations of mtANN include the integration of deep learning and ensemble learning to improve prediction accuracy, and the introduction of a new metric that considers three complementary aspects to distinguish between unseen cell types and shared cell types. Additionally, we provide a data-driven method to adaptively select a threshold for identifying previously unseen cell types. We demonstrate the advantages of mtANN over state-of-the-art methods for unseen cell-type identification and cell-type annotation on two benchmark dataset collections, as well as its predictive power on a collection of COVID-19 datasets. The source code and tutorial are available at https://github.com/Zhangxf-ccnu/mtANN. Yi-Xuan Xiong, Meng-Guo Wang, Luonan Chen, Xiao-Fei Zhang |
PLoS Comput. Biol. | 3 |
| 2022 | Detecting the critical states during disease development based on temporal network flow entropyabstractComplex diseases progression can be generally divided into three states, which are normal state, predisease/critical state and disease state. The sudden deterioration of diseases can be viewed as a bifurcation or a critical transition. Therefore, hunting for the tipping point or critical state is of great importance to prevent the disease deterioration. However, it is still a challenging task to detect the critical states of complex diseases with high-dimensional data, especially based on an individual. In this study, we develop a new method based on network fluctuation of molecules, temporal network flow entropy (TNFE) or temporal differential network flow entropy, to detect the critical states of complex diseases on the basis of each individual. By applying this method to a simulated dataset and six real diseases, including respiratory viral infections and tumors with four time-course and two stage-course high-dimensional omics datasets, the critical states before deterioration were detected and their dynamic network biomarkers were identified successfully. The results on the simulated dataset indicate that the TNFE method is robust under different noise strengths, and is also superior to the existing methods on detecting the critical states. Moreover, the analysis on the real datasets demonstrated the effectiveness of TNFE for providing early-warning signals on various diseases. In addition, we also predicted disease deterioration risk and identified drug targets for cancers based on stage-wise data. Jinling Yan, Peiluan Li, Luonan Chen |
Briefings Bioinform. | 4 |
| 2022 | Transcriptome analysis method based on differential distribution evaluationabstractIdentifying differential genes over conditions provides insights into the mechanisms of biological processes and disease progression. Here we present an approach, the Kullback-Leibler divergence-based differential distribution (klDD), which provides a flexible framework for quantifying changes in higher-order statistical information of genes including mean and variance/covariation. The method can well detect subtle differences in gene expression distributions in contrast to mean or variance shifts of the existing methods. In addition to effectively identifying informational genes in terms of differential distribution, klDD can be directly applied to cancer subtyping, single-cell clustering and disease early-warning detection, which were all validated by various benchmark datasets. Yiwei Meng, Yanhong Huang, Xiao Chang, Xiaoping Liu 0002, Luonan Chen |
Briefings Bioinform. | 5 |
| 2022 | Deep latent space fusion for adaptive representation of heterogeneous multi-omics dataabstractThe integration of multi-omics data makes it possible to understand complex biological organisms at the system level. Numerous integration approaches have been developed by assuming a common underlying data space. Due to the noise and heterogeneity of biological data, the performance of these approaches is greatly affected. In this work, we propose a novel deep neural network architecture, named Deep Latent Space Fusion (DLSF), which integrates the multi-omics data by learning consistent manifold in the sample latent space for disease subtypes identification. DLSF is built upon a cycle autoencoder with a shared self-expressive layer, which can naturally and adaptively merge nonlinear features at each omics level into one unified sample manifold and produce adaptive representation of heterogeneous samples at the multi-omics level. We have assessed DLSF on various biological and biomedical datasets to validate its effectiveness. DLSF can efficiently and accurately capture the intrinsic manifold of the sample structures or sample clusters compared with other state-of-the-art methods, and DLSF yielded more significant outcomes for biological significance, survival prognosis and clinical relevance in application of cancer study in The Cancer Genome Atlas. Notably, as a deep case study, we determined a new molecular subtype of kidney renal clear cell carcinoma that may benefit immunotherapy in the viewpoint of multi-omics, and we further found potential subtype-specific biomarkers from multiple omics data, which were validated by independent datasets. In addition, we applied DLSF to identify potential therapeutic agents of different molecular subtypes of chronic lymphocytic leukemia, demonstrating the scalability of DLSF in diverse omics data types and application scenarios. Chengming Zhang 0003, Yabin Chen, Tao Zeng 0003, Chuanchao Zhang, Luonan Chen |
Briefings Bioinform. | 5 |
| 2022 | Single-cell RNA sequencing data analysis based on non-uniformε-neighborhood networkabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) technology provides the possibility to study cell heterogeneity and cell development on the resolution of individual cells. Arguably, three of the most important computational targets on scRNA-seq data analysis are data visualization, cell clustering and trajectory inference. Although a substantial number of algorithms have been developed, most of them do not treat the three targets in a systematic or consistent manner. RESULTS: In this article, we propose an efficient scRNA-seq analysis framework, which accomplishes the three targets consistently by non-uniform ε-neighborhood (NEN) network. First, a network is generated by our NEN method, which combines the advantages of both k-nearest neighbors (KNN) and ε-neighborhood (EN) to represent the manifold that data points reside in gene space. Then from such a network, we use its layout, its community and further its shortest path to achieve the purpose of scRNA-seq data visualization, clustering and trajectory inference. The results on both synthetic and real datasets indicate that our NEN method not only can visually provide the global topological structure of a dataset accurately compared with t-SNE (t-Distributed Stochastic Neighbor Embedding) and UMAP (Uniform Manifold Approximation and Projection), but also has superior performances on clustering and pseudotime ordering of cells over the existing approaches. AVAILABILITY AND IMPLEMENTATION: This analysis method has been made into a python package called ccnet and is freely available at https://github.com/Just-Jia/ccNet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Junbo Jia, Luonan Chen |
Bioinform. | 2 |
| 2022 | Identifying network biomarkers of cancer by sample-specific differential networkabstractAbundant datasets generated from various big science projects on diseases have presented great challenges and opportunities, which contributed to unfolding the complexity of diseases. The discovery of disease-associated molecular networks for each individual plays an important role in personalized therapy and precision treatment of cancer-based on the reference networks. However, there are no effective ways to distinguish the consistency of different reference networks. In this study, we developed a statistical method, i.e. a sample-specific differential network (SSDN), to construct and analyze such networks based on gene expression of a single sample against a reference dataset. We proved that the SSDN is structurally consistent even with different reference datasets if the reference dataset can follow certain conditions. The SSDN also can be used to identify patient-specific disease modules or network biomarkers as well as predict the potential driver genes of a tumor sample. Xiao Chang, Yanhong Huang, Shaoyan Sun, Luonan Chen, Xiaoping Liu 0002 |
BMC Bioinform. | 6 |
| 2022 | Predicting high-dimensional time series data with spatial, temporal and global information
Jining Wang, Chuan Chen 0001, Zibin Zheng, Luonan Chen |
Inf. Sci. | 4 |
| 2022 | Brain-inspired chaotic backpropagation for MLP
Peng Tao 0009, Luonan Chen |
Neural Networks | 3 |
| 2022 | Extracting ROI-Based Contourlet Subband Energy Feature From the sMRI Image for Alzheimer's Disease ClassificationabstractStructural magnetic resonance imaging (sMRI)-based Alzheimer's disease (AD) classification and its prodromal stage-mild cognitive impairment (MCI) classification have attracted many attentions and been widely investigated in recent years. Owing to the high dimensionality, representation of the sMRI image becomes a difficult issue in AD classification. Furthermore, regions of interest (ROI) reflected in the sMRI image are not characterized properly by spatial analysis techniques, which has been a main cause of weakening the discriminating ability of the extracted spatial feature. In this study, we propose a ROI-based contourlet subband energy (ROICSE) feature to represent the sMRI image in the frequency domain for AD classification. Specifically, a preprocessed sMRI image is first segmented into 90 ROIs by a constructed brain mask. Instead of extracting features from the 90 ROIs in the spatial domain, the contourlet transform is performed on each of these ROIs to obtain their energy subbands. And then for an ROI, a subband energy (SE) feature vector is constructed to capture its energy distribution and contour information. Afterwards, SE feature vectors of the 90 ROIs are concatenated to form a ROICSE feature of the sMRI image. Finally, support vector machine (SVM) classifier is used to classify 880 subjects from ADNI and OASIS databases. Experimental results show that the ROICSE approach outperforms six other state-of-the-art methods, demonstrating that energy and contour information of the ROI are important to capture differences between the sMRI images of AD and HC subjects. Meanwhile, brain regions related to AD can also be found using the ROICSE feature, indicating that the ROICSE feature can be a promising assistant imaging marker for the AD diagnosis via the sMRI image. Code and Sample IDs of this paper can be downloaded at https://github.com/NWPU-903PR/ROICSE.git. Jinwang Feng, Shaowu Zhang 0001, Luonan Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | Disease characterization using a partial correlation-based sample-specific networkabstractA single-sample network (SSN) is a biological molecular network constructed from single-sample data given a reference dataset and can provide insights into the mechanisms of individual diseases and aid in the development of personalized medicine. In this study, we proposed a computational method, a partial correlation-based single-sample network (P-SSN), which not only infers a network from each single-sample data given a reference dataset but also retains the direct interactions by excluding indirect interactions (https://github.com/hyhRise/P-SSN). By applying P-SSN to analyze tumor data from the Cancer Genome Atlas and single cell data, we validated the effectiveness of P-SSN in predicting driver mutation genes (DMGs), producing network distance, identifying subtypes and further classifying single cells. In particular, P-SSN is highly effective in predicting DMGs based on single-sample data. P-SSN is also efficient for subtyping complex diseases and for clustering single cells by introducing network distance between any two samples. Yanhong Huang, Xiao Chang, Luonan Chen, Xiaoping Liu 0002 |
Briefings Bioinform. | 4 |
| 2021 | Deep-joint-learning analysis model of single cell transcriptome and open chromatin accessibility dataabstractSimultaneous profiling transcriptomic and chromatin accessibility information in the same individual cells offers an unprecedented resolution to understand cell states. However, computationally effective methods for the integration of these inherent sparse and heterogeneous data are lacking. Here, we present a single-cell multimodal variational autoencoder model, which combines three types of joint-learning strategies with a probabilistic Gaussian Mixture Model to learn the joint latent features that accurately represent these multilayer profiles. Studies on both simulated datasets and real datasets demonstrate that it has more preferable capability (i) dissecting cellular heterogeneity in the joint-learning space, (ii) denoising and imputing data and (iii) constructing the association between multilayer omics data, which can be used for understanding transcriptional regulatory mechanisms. Chunman Zuo, Luonan Chen |
Briefings Bioinform. | 2 |
| 2021 | Revealing dynamic regulations and the related key proteins of myeloma-initiating cells by integrating experimental data into a systems biological modelabstractMOTIVATION: The growth and survival of myeloma cells are greatly affected by their surrounding microenvironment. To understand the molecular mechanism and the impact of stiffness on the fate of myeloma-initiating cells (MICs), we develop a systems biological model to reveal the dynamic regulations by integrating reverse-phase protein array data and the stiffness-associated pathway. RESULTS: We not only develop a stiffness-associated signaling pathway to describe the dynamic regulations of the MICs, but also clearly identify three critical proteins governing the MIC proliferation and death, including FAK, mTORC1 and NFκB, which are validated to be related with multiple myeloma by our immunohistochemistry experiment, computation and manually reviewed evidences. Moreover, we demonstrate that the systematic model performs better than widely used parameter estimation algorithms for the complicated signaling pathway. AVAILABILITY AND IMPLEMENTATION: We can not only use the systems biological model to infer the stiffness-associated genetic signaling pathway and locate the critical proteins, but also investigate the important pathways, proteins or genes for other type of the cancer. Thus, it holds universal scientific significance. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Le Zhang 0004, Guangdi Liu, Meijing Kong, Xiaobo Zhou 0001, Chuanwei Yang, Zhenzhou Yang, Luonan Chen |
Bioinform. | 10 |
| 2021 | Deep cross-omics cycle attention model for joint analysis of single-cell multi-omics dataabstractMOTIVATION: Joint profiling of single-cell transcriptomics and epigenomics data enables us to characterize cell states and transcriptomics regulatory programs related to cellular heterogeneity. However, the highly different features on sparsity, heterogeneity and dimensionality between multi-omics data have severely hindered its integrative analysis. RESULTS: We proposed deep cross-omics cycle attention (DCCA) model, a computational tool for joint analysis of single-cell multi-omics data, by combining variational autoencoders (VAEs) and attention-transfer. Specifically, we show that DCCA can leverage one omics data to fine-tune the network trained for another omics data, given a dataset of parallel multi-omics data within the same cell. Studies on both simulated and real datasets from various platforms, DCCA demonstrates its superior capability: (i) dissecting cellular heterogeneity; (ii) denoising and aggregating data and (iii) constructing the link between multi-omics data, which is used to infer new transcriptional regulatory relations. In our applications, DCCA was demonstrated to have a superior power to generate missing stages or omics in a biologically meaningful manner, which provides a new way to analyze and also understand complicated biological processes. AVAILABILITY AND IMPLEMENTATION: DCCA source code is available at https://github.com/cmzuo11/DCCA, and has been deposited in archived format at https://doi.org/10.5281/zenodo.4762065. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chunman Zuo, Luonan Chen |
Bioinform. | 3 |
| 2021 | Alzheimer's disease classification using features extracted from nonsubsampled contourlet subband-based individual networks
Jinwang Feng, Shaowu Zhang 0001, Luonan Chen |
Neurocomputing | 3 |
| 2020 | Network biomarker for quantifying regular state of a biological system, and dynamic network biomarker for quantifying critical state of a biological systemabstractWe defined two new types of biomarkers to quantify the states of biological systems based on network, in contrast to the traditional molecular biomarkers. Network biomarker is constructed to quantify regular state of a biological system, while dynamic network biomarker is to quantify the critical state or tipping point of a biological system. (1) Network biomarker (NB) is a subnetwork or network module, which is composed of a number of associations or regulations between molecules (or variables), rather than simply a number of molecules. Those associations (the second-order statistics) in the module are formed collectively as a biomarker, thus robustly and accurately quantifying the regular state of a biological system, completely different from the concentrations or densities of conventional molecular biomarkers (the first-order statistics). (2) Dynamic network biomarker (DNB) is a subnetwork or module, and is also composed of a number of associations or regulations between molecules but with three statistical conditions (in terms of variances and covariances), which are actually a number of strongly and collectively fluctuated molecules in the network. Theoretically, DNB is able to quantify the critical state or the tipping point of a biological system, thereby serving as a general early-warning signal to indicate an imminent state transition. A number of real datas are provided to validate the effectiveness of NB and DNB. Luonan Chen |
BIBM | 1 |
| 2020 | Identification of Alzheimer's disease based on wavelet transformation energy feature of the structural MRI image and NN classifier
Jinwang Feng, Shaowu Zhang 0001, Luonan Chen |
Artif. Intell. Medicine | 3 |
| 2020 | Network control principles for identifying personalized driver genes in cancerabstractTo understand tumor heterogeneity in cancer, personalized driver genes (PDGs) need to be identified for unraveling the genotype-phenotype associations corresponding to particular patients. However, most of the existing driver-focus methods mainly pay attention on the cohort information rather than on individual information. Recent developing computational approaches based on network control principles are opening a new way to discover driver genes in cancer, particularly at an individual level. To provide comprehensive perspectives of network control methods on this timely topic, we first considered the cancer progression as a network control problem, in which the expected PDGs are altered genes by oncogene activation signals that can change the individual molecular network from one health state to the other disease state. Then, we reviewed the network reconstruction methods on single samples and introduced novel network control methods on single-sample networks to identify PDGs in cancer. Particularly, we gave a performance assessment of the network structure control-based PDGs identification methods on multiple cancer datasets from TCGA, for which the data and evaluation package also are publicly available. Finally, we discussed future directions for the application of network control methods to identify PDGs in cancer and diverse biological processes. Weifeng Guo, Shaowu Zhang 0001, Tao Zeng 0003, Tatsuya Akutsu, Luonan Chen |
Briefings Bioinform. | 5 |
| 2020 | Quantifying Waddington's epigenetic landscape: a comparison of single-cell potency measuresabstractMOTIVATION: Estimating differentiation potency of single cells is a task of great biological and clinical significance, as it may allow identification of normal and cancer stem cell phenotypes. However, very few single-cell potency models have been proposed, and their robustness and reliability across independent studies have not yet been fully assessed. RESULTS: Using nine independent single-cell RNA-Seq experiments, we here compare four different single-cell potency models to each other, in their ability to discriminate cells that ought to differ in terms of differentiation potency. Two of the potency models approximate potency via network entropy measures that integrate the single-cell RNA-Seq profile of a cell with a protein interaction network. The comparison between the four models reveals that integration of RNA-Seq data with a protein interaction network dramatically improves the robustness and reliability of single-cell potency estimates. We demonstrate that underlying this robustness is a correlation relationship, according to which high differentiation potency is positively associated with overexpression of network hubs. We further show that overexpressed network hubs are strongly enriched for ribosomal mitochondrial proteins, suggesting that their mRNA levels may provide a universal marker of a cell's potency. Thus, this study provides novel systems-biological insight into cellular potency and may provide a foundation for improved models of differentiation potency with far-reaching implications for the discovery of novel stem cell or progenitor cell phenotypes. Jifan Shi, Andrew E. Teschendorff, Luonan Chen |
Briefings Bioinform. | 4 |
| 2020 | Corrigendum to: Single-sample landscape entropy reveals the imminent phase transition during disease progressionabstractBioinformatics (2009) doi:10.1093/bioinformatics/btz758 In the above article the funding source was given incorrectly as ‘Guangdong Natural Science Funds for Distinguished Young Scholar (No. 2019B151502062).’ This has now been corrected to ‘Guangdong Basic and Applied Basic Research Foundation (No. 2019B151502062).’ Rui Liu 0009, Pei Chen 0004, Luonan Chen |
Bioinform. | 3 |
| 2020 | Single-sample landscape entropy reveals the imminent phase transition during disease progressionabstractMOTIVATION: The time evolution or dynamic change of many biological systems during disease progression is not always smooth but occasionally abrupt, that is, there is a tipping point during such a process at which the system state shifts from the normal state to a disease state. It is challenging to predict such disease state with the measured omics data, in particular when only a single sample is available. RESULTS: In this study, we developed a novel approach, i.e. single-sample landscape entropy (SLE) method, to identify the tipping point during disease progression with only one sample data. Specifically, by evaluating the disorder of a network projected from a single-sample data, SLE effectively characterizes the criticality of this single sample network in terms of network entropy, thereby capturing not only the signals of the impending transition but also its leading network, i.e. dynamic network biomarkers. Using this method, we can characterize sample-specific state during disease progression and thus achieve the disease prediction of each individual by only one sample. Our method was validated by successfully identifying the tipping points just before the serious disease symptoms from four real datasets of individuals or subjects, including influenza virus infection, lung cancer metastasis, prostate cancer and acute lung injury. AVAILABILITY AND IMPLEMENTATION: https://github.com/rabbitpei/SLE. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rui Liu 0009, Pei Chen 0004, Luonan Chen |
Bioinform. | 3 |
| 2020 | A plausible accelerating function of intermediate states in cancer metastasisabstractEpithelial-to-mesenchymal transition (EMT) is a fundamental cellular process and plays an essential role in development, tissue regeneration, and cancer metastasis. Interestingly, EMT is not a binary process but instead proceeds with multiple partial intermediate states. However, the functions of these intermediate states are not fully understood. Here, we focus on a general question about how the number of partial EMT states affects cell transformation. First, by fitting a hidden Markov model of EMT with experimental data, we propose a statistical mechanism for EMT in which many unobservable microstates may exist within one of the observable macrostates. Furthermore, we find that increasing the number of intermediate states can accelerate the EMT process and that adding parallel paths or transition layers may accelerate the process even further. Last, a stabilized intermediate state traps cells in one partial EMT state. This work advances our understanding of the dynamics and functions of EMT plasticity during cancer metastasis. Hanah Goetz, Juan Ramon Melendez-Alvarez, Luonan Chen, Xiao-Jun Tian |
PLoS Comput. Biol. | 3 |
| 2020 | A Branch Point on Differentiation Trajectory is the Bifurcating Event Revealed by Dynamical Network Biomarker Analysis of Single-Cell DataabstractThe advance in single-cell profiling technologies and the development in computational algorithms provide the opportunity to reconstruct pseudo temporal trajectory with branch point of cellular development. On the other hand, theories such as dynamical network biomarkers (DNB) theory have been recently proposed to characterize the pre-transition state in biological systems. Few studies have validated whether the branch point identified in pseudo time is the critical point in dynamical system. In this study, the dynamical behavior of the branch point on the pseudo trajectory has been investigated. We study the pseudo temporal trajectories reconstructed by Wishbone and diffusion pseudotime analysis (DPT) algorithms, as well as the simulated trajectory. DNB theory is applied to justify the bifurcating event on the pseudo trajectories. Our results demonstrate that the branch point recovered by Wishbone and DPT algorithms is confirmed as a transition state in cell differentiation process by DNB theory. Furthermore, we show that an appropriate DNB group will amplify the comprehensive index of critical event as defined in DNB theory. Our study provides biological insights on pseudo trajectory with branch point in a dynamical view and also indicates that DNB theory may serve as a benchmark to check the validity of branch point. Xiangqi Bai, Xiawei Wang, Xiuqin Liu, Yuting Liu 0002, Luonan Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2020 | Quantifying Direct Dependencies in Biological Networks by Multiscale Association AnalysisabstractPartial correlation (PC) or conditional mutual information (CMI) is widely used in detecting direct dependencies between the observed variables in biological networks by eliminating indirect correlations/associations, but it fails whenever there are some strong correlations in a network. In this paper, we theoretically develop a multiscale association analysis to overcome this flaw. We propose a new measure, partial association (PA), based on the multiscale conditional mutual information. We show that linear PA and nonlinear PA have clear advantages over PC and CMI from both theoretical and computational aspects. Both simulated models and real omics datasets demonstrate that PA is superior to PC and CMI in terms of accuracy, and is a powerful tool to identify the direct associations or reconstruct molecular networks based on the observed data. Survival and functional analyses of the hub genes in the gene networks reconstructed from TCGA data for different cancers also validated the effectiveness of our method. Jifan Shi, Xiaoping Liu 0002, Luonan Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2020 | Efficient Mining Multi-Mers in a Variety of Biological SequencesabstractCounting the occurrence frequency of each $k$k-mer in a biological sequence is a preliminary yet important step in many bioinformatics applications. However, most $k$k-mer counting algorithms rely on a given $k$k to produce single-length $k$k-mers, which is inefficient for sequence analysis for different $k$k. Moreover, existing $k$k-mer counters focus more on DNA and RNA sequences and less on protein ones. In practice, the analysis of $k$k-mers in protein sequences can provide substantial biological insights in structure, function, and evolution. To this end, an efficient algorithm, called MulMer (Multiple-Mer mining), is proposed to mine $k$k-mers of various lengths termed multi-mers via inverted-index technique, which is orders of magnitude faster than the conventional forward-index methods. Moreover, to the best of our knowledge, MulMer is the first able to mine multi-mers in a variety of sequences, including DNA, RNA, and protein sequences. Jingsong Zhang, Jianmei Guo, Xiangtian Yu, Xiaoqing Yu, Weifeng Guo, Tao Zeng 0003, Luonan Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2019 | Systems biology intertwines with single cell and AIabstractA report of the 12th International Conference on Systems Biology (ISB2018), 18-21 August, Guiyang, China. Yong Wang 0001, Xiang-Sun Zhang, Luonan Chen |
BMC Bioinform. | 3 |
| 2019 | A novel network control model for identifying personalized driver genes in cancerabstractAlthough existing computational models have identified many common driver genes, it remains challenging to identify the personalized driver genes by using samples of an individual patient. Recently, the methods of exploiting the structure-based control principles of complex networks provide new clues for identifying minimum number of driver nodes to drive the state transition of large-scale complex networks from an initial state to the desired state. However, the structure-based network control methods cannot be directly applied to identify the personalized driver genes due to the unknown network dynamics of the personalized system. Here we proposed the personalized network control model (PNC) to identify the personalized driver genes by employing the structure-based network control principle on genetic data of individual patients. In PNC model, we firstly presented a paired single sample network construction method to construct the personalized state transition network for capturing the phenotype transitions between healthy and disease states. Then, we designed a novel structure-based network control method from the Feedback Vertex Sets-based control perspective to identify the personalized driver genes. The wide experimental results on 13 cancer datasets from The Cancer Genome Atlas firstly showed that PNC model outperforms current state-of-the-art methods, in terms of F-measures for identifying cancer driver genes enriched in the gold-standard cancer driver gene lists. Furthermore, these results showed that personalized driver genes can be explored by their network characteristics even when they are hidden factors in transcription and mutation profiles. Our PNC gives novel insights and useful tools into understanding the tumor heterogeneity in cancer. The PNC package and data resources used in this work can be freely downloaded from https://github.com/NWPU-903PR/PNC. Weifeng Guo, Shaowu Zhang 0001, Tao Zeng 0003, Yan Li 0111, Jianxi Gao, Luonan Chen |
PLoS Comput. Biol. | 6 |
| 2019 | Quantifying pluripotency landscape of cell differentiation from scRNA-seq data by continuous birth-death processabstractModeling cell differentiation from omics data is an essential problem in systems biology research. Although many algorithms have been established to analyze scRNA-seq data, approaches to infer the pseudo-time of cells or quantify their potency have not yet been satisfactorily solved. Here, we propose the Landscape of Differentiation Dynamics (LDD) method, which calculates cell potentials and constructs their differentiation landscape by a continuous birth-death process from scRNA-seq data. From the viewpoint of stochastic dynamics, we exploited the features of the differentiation process and quantified the differentiation landscape based on the source-sink diffusion process. In comparison with other scRNA-seq methods in seven benchmark datasets, we found that LDD could accurately and efficiently build the evolution tree of cells with pseudo-time, in particular quantifying their differentiation landscape in terms of potency. This study provides not only a computational tool to quantify cell potency or the Waddington potential landscape based on scRNA-seq data, but also novel insights to understand the cell differentiation process from a dynamic perspective. Jifan Shi, Luonan Chen, Kazuyuki Aihara |
PLoS Comput. Biol. | 3 |
| 2019 | Associating lncRNAs with small molecules via bilevel optimization reveals cancer-related lncRNAsabstractLong noncoding RNA (lncRNA) transcripts have emerging impacts in cancer studies, which suggests their potential as novel therapeutic agents. However, the molecular mechanism behind their treatment effects is still unclear. Here, we designed a computational model to Associate LncRNAs with Anti-Cancer Drugs (ALACD) based on a bilevel optimization model, which optimized the gene signature overlap in the upper level and imputed the missing lncRNA-gene association in the lower level. ALACD predicts genes coexpressed with lncRNAs mean while matching drug's gene signatures. This model allows us to borrow the target gene information of small molecules to understand the mechanisms of action of lncRNAs and their roles in cancer. The ALACD model was systematically applied to the 10 cancer types in The Cancer Genome Atlas (TCGA) that had matched lncRNA and mRNA expression data. Cancer type-specific lncRNAs and associated drugs were identified. These lncRNAs show significantly different expression levels in cancer patients. Follow-up functional and molecular pathway analysis suggest the gene signatures bridging drugs and lncRNAs are closely related to cancer development. Importantly, patient survival information and evidence from the literature suggest that the lncRNAs and drug-lncRNA associations identified by the ALACD model can provide an alternative choice for cancer targeting treatment and potential cancer pognostic biomarkers. The ALACD model is freely available at https://github.com/wangyc82/ALACD-v1. Yongcui Wang, Shilong Chen, Luonan Chen, Yong Wang 0001 |
PLoS Comput. Biol. | 3 |
| 2019 | A Semisupervised Classification Approach for Multidomain Networks With Domain SelectionabstractMultidomain network classification has attracted significant attention in data integration and machine learning, which can enhance network classification or prediction performance by integrating information from different sources. Despite the previous success, existing multidomain network learning methods usually assume that different views are available for the same set of instances, and thus, they seek a consistent classification result for all domains. However, in many real-world problems, each domain has its specific instance set, and one instance in one domain may correspond to multiple instances in another domain. Moreover, due to the rapid growth of data sources, different domains may not be relevant to each other, which asks for selecting domains relevant to the target/focused domain. A key challenge under this setting is how to achieve accurate prediction by integrating different data representations without losing data information. In this paper, we propose a semisupervised classification approach for a multidomain network based on label propagation, i.e., multidomain classification with domain selection (MCS), which can deal with the cross-domain information and different instance sets in domains. In particular, with sparse weight properties, the proposed MCS can automatically identify those domains relevant to our target domain by assigning them higher weights than the other irrelevant domains. This not only significantly improves a classification accuracy but also helps to obtain optimal network partition for the target domain. From the theoretical viewpoint, we equivalently decompose MCS into two simpler subproblems with analytical solutions, which can be efficiently solved by their computational procedures. Extensive experimental results on both synthetic and real-world data sets empirically demonstrate the advantages of the proposed approach in terms of both prediction performance and domain selection ability. Chuan Chen 0001, Jingxue Xin, Yong Wang 0001, Luonan Chen, Michael Kwok-Po Ng |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | Discovering personalized driver mutation profiles of single samples in cancer by network control strategyabstractMotivation: It is a challenging task to discover personalized driver genes that provide crucial information on disease risk and drug sensitivity for individual patients. However, few methods have been proposed to identify the personalized-sample driver genes from the cancer omics data due to the lack of samples for each individual. To circumvent this problem, here we present a novel single-sample controller strategy (SCS) to identify personalized driver mutation profiles from network controllability perspective. Results: SCS integrates mutation data and expression data into a reference molecular network for each patient to obtain the driver mutation profiles in a personalized-sample manner. This is the first such a computational framework, to bridge the personalized driver mutation discovery problem and the structural network controllability problem. The key idea of SCS is to detect those mutated genes which can achieve the transition from the normal state to the disease state based on each individual omics data from network controllability perspective. We widely validate the driver mutation profiles of our SCS from three aspects: (i) the improved precision for the predicted driver genes in the population compared with other driver-focus methods; (ii) the effectiveness for discovering the personalized driver genes and (iii) the application to the risk assessment through the integration of the driver mutation signature and expression data, respectively, across the five distinct benchmarks from The Cancer Genome Atlas. In conclusion, our SCS makes efficient and robust personalized driver mutation profiles predictions, opening new avenues in personalized medicine and targeted cancer therapy. Availability and implementation: The MATLAB-package for our SCS is freely available from http://sysbio.sibcb.ac.cn/cb/chenlab/software.htm. Contact: [email protected] or [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Weifeng Guo, Shaowu Zhang 0001, Li-Li Liu, Tao Zeng 0003, Luonan Chen |
Bioinform. | 9 |
| 2018 | Single cell clustering based on cell-pair differentiability correlation and variance analysisabstractMotivation: The rapid advancement of single cell technologies has shed new light on the complex mechanisms of cellular heterogeneity. Identification of intercellular transcriptomic heterogeneity is one of the most critical tasks in single-cell RNA-sequencing studies. Results: We propose a new cell similarity measure based on cell-pair differentiability correlation, which is derived from gene differential pattern among all cell pairs. Through plugging into the framework of hierarchical clustering with this new measure, we further develop a variance analysis based clustering algorithm 'Corr' that can determine cluster number automatically and identify cell types accurately. The robustness and superiority of the proposed algorithm are compared with representative algorithms: shared nearest neighbor (SNN)-Cliq and several other state-of-the-art clustering methods, on many benchmark or real single cell RNA-sequencing datasets in terms of both internal criteria (clustering number and accuracy) and external criteria (purity, adjusted rand index, F1-measure). Moreover, differentiability vector with our new measure provides a new means in identifying potential biomarkers from cancer related single cell datasets even with strong noise. Prognosis analyses from independent datasets of cancers confirmed the effectiveness of our 'Corr' method. Availability and implementation: The source code (Matlab) is available at http://sysbio.sibcb.ac.cn/cb/chenlab/soft/Corr--SourceCodes.zip. Supplementary information: Supplementary data are available at Bioinformatics online. Lydia L. Sohn, Luonan Chen |
Bioinform. | 4 |
| 2017 | Mining K-mers of Various Lengths in Biological Sequences
Jingsong Zhang, Jianmei Guo, Xiaoqing Yu, Xiangtian Yu, Weifeng Guo, Tao Zeng 0003, Luonan Chen |
ISBRA | 7 |
| 2017 | Pattern fusion analysis by adaptive alignment of multiple heterogeneous omics dataabstractMOTIVATION: Integrating different omics profiles is a challenging task, which provides a comprehensive way to understand complex diseases in a multi-view manner. One key for such an integration is to extract intrinsic patterns in concordance with data structures, so as to discover consistent information across various data types even with noise pollution. Thus, we proposed a novel framework called 'pattern fusion analysis' (PFA), which performs automated information alignment and bias correction, to fuse local sample-patterns (e.g. from each data type) into a global sample-pattern corresponding to phenotypes (e.g. across most data types). In particular, PFA can identify significant sample-patterns from different omics profiles by optimally adjusting the effects of each data type to the patterns, thereby alleviating the problems to process different platforms and different reliability levels of heterogeneous data. RESULTS: To validate the effectiveness of our method, we first tested PFA on various synthetic datasets, and found that PFA can not only capture the intrinsic sample clustering structures from the multi-omics data in contrast to the state-of-the-art methods, such as iClusterPlus, SNF and moCluster, but also provide an automatic weight-scheme to measure the corresponding contributions by data types or even samples. In addition, the computational results show that PFA can reveal shared and complementary sample-patterns across data types with distinct signal-to-noise ratios in Cancer Cell Line Encyclopedia (CCLE) datasets, and outperforms over other works at identifying clinically distinct cancer subtypes in The Cancer Genome Atlas (TCGA) datasets. AVAILABILITY AND IMPLEMENTATION: PFA has been implemented as a Matlab package, which is available at http://www.sysbio.ac.cn/cb/chenlab/images/PFApackage_0.1.rar . CONTACT: [email protected] , [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chuanchao Zhang, Minrui Peng, Xiangtian Yu, Tao Zeng 0003, Juan Liu 0007, Luonan Chen |
Bioinform. | 7 |
| 2017 | Comparative network stratification analysis for identifying functional interpretable network biomarkersabstractBACKGROUND: A major challenge of bioinformatics in the era of precision medicine is to identify the molecular biomarkers for complex diseases. It is a general expectation that these biomarkers or signatures have not only strong discrimination ability, but also readable interpretations in a biological sense. Generally, the conventional expression-based or network-based methods mainly capture differential genes or differential networks as biomarkers, however, such biomarkers only focus on phenotypic discrimination and usually have less biological or functional interpretation. Meanwhile, the conventional function-based methods could consider the biomarkers corresponding to certain biological functions or pathways, but ignore the differential information of genes, i.e., disregard the active degree of particular genes involved in particular functions, thereby resulting in less discriminative ability on phenotypes. Hence, it is strongly demanded to develop elaborate computational methods to directly identify functional network biomarkers with both discriminative power on disease states and readable interpretation on biological functions. RESULTS: In this paper, we present a new computational framework based on an integer programming model, named as Comparative Network Stratification (CNS), to extract functional or interpretable network biomarkers, which are of strongly discriminative power on disease states and also readable interpretation on biological functions. In addition, CNS can not only recognize the pathogen biological functions disregarded by traditional Expression-based/Network-based methods, but also uncover the active network-structures underlying such dysregulated functions underestimated by traditional Function-based methods. To validate the effectiveness, we have compared CNS with five state-of-the-art methods, i.e. GSVA, Pathifier, stSVM, frSVM and AEP on four datasets of different complex diseases. The results show that CNS can enhance the discriminative power of network biomarkers, and further provide biologically interpretable information or disease pathogenic mechanism of these biomarkers. A case study on type 1 diabetes (T1D) demonstrates that CNS can identify many dysfunctional genes and networks previously disregarded by conventional approaches. CONCLUSION: Therefore, CNS is actually a powerful bioinformatics tool, which can identify functional or interpretable network biomarkers with both discriminative power on disease states and readable interpretation on biological functions. CNS was implemented as a Matlab package, which is available at http://www.sysbio.ac.cn/cb/chenlab/images/CNSpackage_0.1.rar . Chuanchao Zhang, Juan Liu 0007, Tao Zeng 0003, Luonan Chen |
BMC Bioinform. | 5 |
| 2017 | Differential function analysis: identifying structure and activation variations in dysregulated pathways
Chuanchao Zhang, Juan Liu 0007, Tao Zeng 0003, Luonan Chen |
Sci. China Inf. Sci. | 5 |
| 2017 | Quantifying critical states of complex diseases using single-sample dynamic network biomarkersabstractDynamic network biomarkers (DNB) can identify the critical state or tipping point of a disease, thereby predicting rather than diagnosing the disease. However, it is difficult to apply the DNB theory to clinical practice because evaluating DNB at the critical state required the data of multiple samples on each individual, which are generally not available, and thus limit the applicability of DNB. In this study, we developed a novel method, i.e., single-sample DNB (sDNB), to detect early-warning signals or critical states of diseases in individual patients with only a single sample for each patient, thus opening a new way to predict diseases in a personalized way. In contrast to the information of differential expressions used in traditional biomarkers to "diagnose disease", sDNB is based on the information of differential associations, thereby having the ability to "predict disease" or "diagnose near-future disease". Applying this method to datasets for influenza virus infection and cancer metastasis led to accurate identification of the critical states or correct prediction of the immediate diseases based on individual samples. We successfully identified the critical states or tipping points just before the appearance of disease symptoms for influenza virus infection and the onset of distant metastasis for individual patients with cancer, thereby demonstrating the effectiveness and efficiency of our method for quantifying critical states at the single-sample level. Xiaoping Liu 0002, Xiao Chang, Rui Liu 0009, Xiangtian Yu, Luonan Chen, Kazuyuki Aihara |
PLoS Comput. Biol. | 5 |
| 2016 | Integration of multiple heterogeneous omics dataabstractIntegration of different genomic profiles is challenging to understand complex diseases in a multi-view manner. Computational method is needed to preserve useful information of data types as well as correct bias. Thus, we proposed a novel framework pattern fusion analysis (PFA), to fuse the local sample patterns into a global pattern of patients with respect to the underlying data, by adaptively aligning the information in each type of biological data. In particular, PFA can adjust the distinct data types and achieve more robust sample pattern within different profiles. To validate the effectiveness of PFA, we tested PFA on various synthetic datasets and found that PFA is able to effectively capture the intrinsic clustering structure than the state-of-the-art integrative methods, such as moCluster, iClusterPlus and SNF. Moreover, in a case study on kidney cancer, PFA not only identified the multi-way feature modules among the prior-known disease associated genes, methylations and miRNAs, but also outperformed in cancer subtypes identification and could get effective clinical prognosis prediction. Totally, PFA not only provides new insights on the more holistic & systems-level sample pattern, but also supplies a new way for selecting more informative types of biological data. Chuanchao Zhang, Juan Liu 0007, Xiangtian Yu, Tao Zeng 0003, Luonan Chen |
BIBM | 6 |
| 2016 | Big-data-based edge biomarkers: study on dynamical drug sensitivity and resistance in individualsabstractBig-data-based edge biomarker is a new concept to characterize disease features based on biomedical big data in a dynamical and network manner, which also provides alternative strategies to indicate disease status in single samples. This article gives a comprehensive review on big-data-based edge biomarkers for complex diseases in an individual patient, which are defined as biomarkers based on network information and high-dimensional data. Specifically, we firstly introduce the sources and structures of biomedical big data accessible in public for edge biomarker and disease study. We show that biomedical big data are typically 'small-sample size in high-dimension space', i.e. small samples but with high dimensions on features (e.g. omics data) for each individual, in contrast to traditional big data in many other fields characterized as 'large-sample size in low-dimension space', i.e. big samples but with low dimensions on features. Then, we demonstrate the concept, model and algorithm for edge biomarkers and further big-data-based edge biomarkers. Dissimilar to conventional biomarkers, edge biomarkers, e.g. module biomarkers in module network rewiring-analysis, are able to predict the disease state by learning differential associations between molecules rather than differential expressions of molecules during disease progression or treatment in individual patients. In particular, in contrast to using the information of the common molecules or edges (i.e.molecule-pairs) across a population in traditional biomarkers including network and edge biomarkers, big-data-based edge biomarkers are specific for each individual and thus can accurately evaluate the disease state by considering the individual heterogeneity. Therefore, the measurement of big data in a high-dimensional space is required not only in the learning process but also in the diagnosing or predicting process of the tested individual. Finally, we provide a case study on analyzing the temporal expression data from a malaria vaccine trial by big-data-based edge biomarkers from module network rewiring-analysis. The illustrative results show that the identified module biomarkers can accurately distinguish vaccines with or without protection and outperformed previous reported gene signatures in terms of effectiveness and efficiency. Tao Zeng 0003, Wanwei Zhang, Xiangtian Yu, Xiaoping Liu 0002, Meiyi Li, Luonan Chen |
Briefings Bioinform. | 6 |
| 2016 | Detecting critical state before phase transition of complex biological systems by hidden Markov modelabstractMOTIVATION: Identifying the critical state or pre-transition state just before the occurrence of a phase transition is a challenging task, because the state of the system may show little apparent change before this critical transition during the gradual parameter variations. Such dynamics of phase transition is generally composed of three stages, i.e. before-transition state, pre-transition state and after-transition state, which can be considered as three different Markov processes. RESULTS: By exploring the rich dynamical information provided by high-throughput data, we present a novel computational method, i.e. hidden Markov model (HMM) based approach, to detect the switching point of the two Markov processes from the before-transition state (a stationary Markov process) to the pre-transition state (a time-varying Markov process), thereby identifying the pre-transition state or early-warning signals of the phase transition. To validate the effectiveness, we apply this method to detect the signals of the imminent phase transitions of complex systems based on the simulated datasets, and further identify the pre-transition states as well as their critical modules for three real datasets, i.e. the acute lung injury triggered by phosgene inhalation, MCF-7 human breast cancer caused by heregulin and HCV-induced dysplasia and hepatocellular carcinoma. Both functional and pathway enrichment analyses validate the computational results. AVAILABILITY AND IMPLEMENTATION: The source code and some supporting files are available at https://github.com/rabbitpei/HMM_based-method CONTACTS: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Pei Chen 0004, Rui Liu 0009, Luonan Chen |
Bioinform. | 4 |
| 2016 | CMIP: a software package capable of reconstructing genome-wide regulatory networks using gene expression dataabstractBACKGROUND: A gene regulatory network (GRN) represents interactions of genes inside a cell or tissue, in which vertexes and edges stand for genes and their regulatory interactions respectively. Reconstruction of gene regulatory networks, in particular, genome-scale networks, is essential for comparative exploration of different species and mechanistic investigation of biological processes. Currently, most of network inference methods are computationally intensive, which are usually effective for small-scale tasks (e.g., networks with a few hundred genes), but are difficult to construct GRNs at genome-scale. RESULTS: Here, we present a software package for gene regulatory network reconstruction at a genomic level, in which gene interaction is measured by the conditional mutual information measurement using a parallel computing framework (so the package is named CMIP). The package is a greatly improved implementation of our previous PCA-CMI algorithm. In CMIP, we provide not only an automatic threshold determination method but also an effective parallel computing framework for network inference. Performance tests on benchmark datasets show that the accuracy of CMIP is comparable to most current network inference methods. Moreover, running tests on synthetic datasets demonstrate that CMIP can handle large datasets especially genome-wide datasets within an acceptable time period. In addition, successful application on a real genomic dataset confirms its practical applicability of the package. CONCLUSIONS: This new software package provides a powerful tool for genomic network reconstruction to biological community. The software can be accessed at http://www.picb.ac.cn/CMIP/ . Guangyong Zheng, Yaochen Xu, Zhi-Ping Liu, Luonan Chen, Xin-Guang Zhu |
BMC Bioinform. | 6 |
| 2016 | Inference of Gene Regulatory Network Based on Local Bayesian NetworksabstractThe inference of gene regulatory networks (GRNs) from expression data can mine the direct regulations among genes and gain deep insights into biological processes at a network level. During past decades, numerous computational approaches have been introduced for inferring the GRNs. However, many of them still suffer from various problems, e.g., Bayesian network (BN) methods cannot handle large-scale networks due to their high computational complexity, while information theory-based methods cannot identify the directions of regulatory interactions and also suffer from false positive/negative problems. To overcome the limitations, in this work we present a novel algorithm, namely local Bayesian network (LBN), to infer GRNs from gene expression data by using the network decomposition strategy and false-positive edge elimination scheme. Specifically, LBN algorithm first uses conditional mutual information (CMI) to construct an initial network or GRN, which is decomposed into a number of local networks or GRNs. Then, BN method is employed to generate a series of local BNs by selecting the k-nearest neighbors of each gene as its candidate regulatory genes, which significantly reduces the exponential search space from all possible GRN structures. Integrating these local BNs forms a tentative network or GRN by performing CMI, which reduces redundant regulations in the GRN and thus alleviates the false positive problem. The final network or GRN can be obtained by iteratively performing CMI and local BN on the tentative network. In the iterative process, the false or redundant regulations are gradually removed. When tested on the benchmark GRN datasets from DREAM challenge as well as the SOS DNA repair network in E.coli, our results suggest that LBN outperforms other state-of-the-art methods (ARACNE, GENIE3 and NARROMI) significantly, with more accurate and robust performance. In particular, the decomposition strategy with local Bayesian networks not only effectively reduce the computational cost of BN due to much smaller sizes of local GRNs, but also identify the directions of the regulations. Shaowu Zhang 0001, Weifeng Guo, Ze-Gang Wei, Luonan Chen |
PLoS Comput. Biol. | 5 |
| 2015 | Identification of phenotypic networks based on whole transcriptome by comparative network decompositionabstractComplex diseases are usually caused by the dysfunctions of the molecular system or molecular network rather than individual molecules. Generally, the conventional methods first obtain a disease-associated network based on expression data and then study its biological functions. However, such a network may be only a part of the system facilitating a biological function or may involve in multiple functions. In this paper, we present a computational framework based on an integer programming model, named as comparative network decomposition (CND), to jointly identify optimal structures of significant and moderate phenotypic functions/networks and their optimal combination by integrating gene expression, gene network and gene ontology together. Particularly, CND makes full use of dysfunctional information, e.g. both strong and weak changes on gene expressions and correlations, to extract various phenotypic networks, where one phenotypic network just corresponds to a specific biological function. A synthetic example clearly suggests that CND can identify multiple types of the disease-related phenotypic networks, rather than conventional approaches only exact significant phenotypic networks. As a proof-of-concept study to real data, CND is further used to identify the significant and moderate phenotypic networks for discriminating two different but associated diseases, e.g. subtypes of diabetes. In the comparison of type 1 and type 2 diabetes, the moderate and significant phenotypic networks can capture the disease-related biological functions and their corresponding networks. Therefore, CND is actually a powerful bioinformatics tool, which can investigate phenotype-associated genes and networks in a whole transcriptome and function-centered manner, and the comparative study of complex diseases with other works also demonstrates its effectiveness. Chuanchao Zhang, Juan Liu 0007, Tao Zeng 0003, Luonan Chen |
BIBM | 5 |
| 2015 | Identifying cancer-related microRNAs based on gene expression dataabstractMOTIVATION: MicroRNAs (miRNAs) are short non-coding RNAs that play important roles in post-transcriptional regulations as well as other important biological processes. Recently, accumulating evidences indicate that miRNAs are extensively involved in cancer. However, it is a big challenge to identify which miRNAs are related to which cancer considering the complex processes involved in tumors, where one miRNA may target hundreds or even thousands of genes and one gene may regulate multiple miRNAs. Despite integrative analysis of matched gene and miRNA expression data can help identify cancer-associated miRNAs, such kind of data is not commonly available. On the other hand, there are huge amount of gene expression data that are publicly accessible. It will significantly improve the efficiency of characterizing miRNA's function in cancer if we can identify cancer miRNAs directly from gene expression data. RESULTS: We present a novel computational framework to identify the cancer-related miRNAs based solely on gene expression profiles without requiring either miRNA expression data or the matched gene and miRNA expression data. The results on multiple cancer datasets show that our proposed method can effectively identify cancer-related miRNAs with higher precision compared with other popular approaches. Furthermore, some of our novel predictions are validated by both differentially expressed miRNAs and evidences from literature, implying the predictive power of our proposed method. In addition, we construct a cancer-miRNA-pathway network, which can help explain how miRNAs are involved in cancer. AVAILABILITY AND IMPLEMENTATION: The R code and data files for the proposed method are available at http://comp-sysbio.org/miR_Path/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: supplementary data are available at Bioinformatics online. Xing-Ming Zhao, Keqin Liu, Feng He 0004, Béatrice Duval, Jean-Michel Richer, De-Shuang Huang, Jin-Kao Hao, Luonan Chen |
Bioinform. | 10 |
| 2015 | Systematic computation with functional gene-sets among leukemic and hematopoietic stem cells reveals a favorable prognostic signature for acute myeloid leukemiaabstractBACKGROUND: Genes that regulate stem cell function are suspected to exert adverse effects on prognosis in malignancy. However, diverse cancer stem cell signatures are difficult for physicians to interpret and apply clinically. To connect the transcriptome and stem cell biology, with potential clinical applications, we propose a novel computational "gene-to-function, snapshot-to-dynamics, and biology-to-clinic" framework to uncover core functional gene-sets signatures. This framework incorporates three function-centric gene-set analysis strategies: a meta-analysis of both microarray and RNA-seq data, novel dynamic network mechanism (DNM) identification, and a personalized prognostic indicator analysis. This work uses complex disease acute myeloid leukemia (AML) as a research platform. RESULTS: We introduced an adjustable "soft threshold" to a functional gene-set algorithm and found that two different analysis methods identified distinct gene-set signatures from the same samples. We identified a 30-gene cluster that characterizes leukemic stem cell (LSC)-depleted cells and a 25-gene cluster that characterizes LSC-enriched cells in parallel; both mark favorable-prognosis in AML. Genes within each signature significantly share common biological processes and/or molecular functions (empirical p = 6e-5 and 0.03 respectively). The 25-gene signature reflects the abnormal development of stem cells in AML, such as AURKA over-expression. We subsequently determined that the clinical relevance of both signatures is independent of known clinical risk classifications in 214 patients with cytogenetically normal AML. We successfully validated the prognosis of both signatures in two independent cohorts of 91 and 242 patients respectively (log-rank p < 0.0015 and 0.05; empirical p < 0.015 and 0.08). CONCLUSION: The proposed algorithms and computational framework will harness systems biology research because they efficiently translate gene-sets (rather than single genes) into biological discoveries about AML and other complex diseases. Xinan Yang, Meiyi Li, Wanqi Zhu, Aurelie Desgardin, Kenan Onel, Jill de Jong, Luonan Chen, John M. Cunningham |
BMC Bioinform. | 9 |
| 2015 | Identifying module biomarker in type 2 diabetes mellitus by discriminative area of functional activityabstractBACKGROUND: Identifying diagnosis and prognosis biomarkers from expression profiling data is of great significance for achieving personalized medicine and designing therapeutic strategy in complex diseases. However, the reproducibility of identified biomarkers across tissues and experiments is still a challenge for this issue. RESULTS: We propose a strategy based on discriminative area of module activities to identify gene biomarkers which interconnect as a subnetwork or module by integrating gene expression data and protein-protein interactions. Then, we implement the procedure in T2DM as a case study and identify a module biomarker with 32 genes from mRNA expression data in skeletal muscle for T2DM. This module biomarker is enriched with known causal genes and related functions of T2DM. Further analysis shows that the module biomarker is of superior performance in classification, and has consistently high accuracies across tissues and experiments. CONCLUSION: The proposed approach can efficiently identify robust and functionally meaningful module biomarkers in T2DM, and could be employed in biomarker discovery of other complex diseases characterized by expression profiles. Lin Gao 0006, Zhi-Ping Liu, Luonan Chen |
BMC Bioinform. | 4 |
| 2015 | Inferring Sequential Order of Somatic Mutations during Tumorgenesis based on Markov Chain ModelabstractTumors are developed and worsen with the accumulated mutations on DNA sequences during tumorigenesis. Identifying the temporal order of gene mutations in cancer initiation and development is a challenging topic. It not only provides a new insight into the study of tumorigenesis at the level of genome sequences but also is an effective tool for early diagnosis of tumors and preventive medicine. In this paper, we develop a novel method to accurately estimate the sequential order of gene mutations during tumorigenesis from genome sequencing data based on Markov chain model as TOMC (Temporal Order based on Markov Chain), and also provide a new criterion to further infer the order of samples or patients, which can characterize the severity or stage of the disease. We applied our method to the analysis of tumors based on several high-throughput datasets. Specifically, first, we revealed that tumor suppressor genes (TSG) tend to be mutated ahead of oncogenes, which are considered as important events for key functional loss and gain during tumorigenesis. Second, the comparisons of various methods demonstrated that our approach has clear advantages over the existing methods due to the consideration on the effect of mutation dependence among genes, such as co-mutation. Third and most important, our method is able to deduce the ordinal sequence of patients or samples to quantitatively characterize their severity of tumors. Therefore, our work provides a new way to quantitatively understand the development and progression of tumorigenesis based on high throughput sequencing data. Hao Kang, Kwang-Hyun Cho, Xiaohua Douglas Zhang, Tao Zeng 0003, Luonan Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2015 | Guest Editorial for Special Section on ISB/TBC 2014abstractThe five papers in this special section were presented at the Eighth International Conference on Computational Systems Biology and Fourth Translational Bioinformatics Conference (ISB/TBC 2014), which was held in Qingdao, China, 24-27 October, 2014. Luonan Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2014 | DTMBIO 2014: International Workshop on Data and Text Mining in Biomedical InformaticsabstractHeld each year in conjunction with one of the largest data management conferences, CIKM, the Eighth ACM International Workshop on Data and Text Mining in Biomedical Informatics (DTMBIO 14) is organized to bring together researchers interested in development and application of cutting-edge biomedical and healthcare technology. The purpose of DTMBIO is to foster discussions regarding the state-of-the-art applications of data and text mining on biomedical research problems. DTMBIO 14 will help scientists navigate emerging trends and opportunities in the evolving area of informatics related techniques and problems in the context of biomedical research. Luonan Chen, Doheon Lee, Hua Xu 0001, Min Song 0001 |
CIKM | 1 |
| 2014 | Detecting tissue-specific early warning signals for complex diseases based on dynamical network biomarkers: study of type 2 diabetes by cross-tissue analysisabstractIdentifying early warning signals of critical transitions during disease progression is a key to achieving early diagnosis of complex diseases. By exploiting rich information of high-throughput data, a novel model-free method has been developed to detect early warning signals of diseases. Its theoretical foundation is based on dynamical network biomarker (DNB), which is also called as the driver (or leading) network of the disease because components or molecules in DNB actually drive the whole system from one state (e.g. normal state) to another (e.g. disease state). In this article, we first reviewed the concept and main results of DNB theory, and then applied the new method to the analysis of type 2 diabetes mellitus (T2DM). Specifically, based on the temporal-spatial gene expression data of T2DM, we identified tissue-specific DNBs corresponding to the critical transitions occurring in liver, adipose and muscle during T2DM development and progression. Actually, we found that there are two different critical states during T2DM development characterized as responses to insulin resistance and serious inflammation, respectively. Interestingly, a new T2DM-associated function, i.e. steroid hormone biosynthesis, was discovered, and those related genes were significantly dysregulated in liver and adipose at the first critical transition during T2DM deterioration. Moreover, the dysfunction of genes related to responding hormone was also detected in muscle at the similar period. Based on the functional and network analysis on pathogenic molecular mechanism of T2DM, we showed that most of DNB genes, in particular the core ones, tended to be located at the upstream of biological pathways, which implied that DNB genes act as the causal factors rather than the consequence to drive the downstream molecules to change their transcriptional activities. This also validated our theoretical prediction of DNB as the driver network. As shown in this study, DNB can not only signal the emergence of the critical transitions for early diagnosis of diseases, but can also provide the causal network of the transitions for revealing molecular mechanisms of disease initiation and progression at a network level. Meiyi Li, Tao Zeng 0003, Rui Liu 0009, Luonan Chen |
Briefings Bioinform. | 4 |
| 2014 | Identifying critical transitions of complex diseases based on a single sampleabstractMOTIVATION: Unlike traditional diagnosis of an existing disease state, detecting the pre-disease state just before the serious deterioration of a disease is a challenging task, because the state of the system may show little apparent change or symptoms before this critical transition during disease progression. By exploring the rich interaction information provided by high-throughput data, the dynamical network biomarker (DNB) can identify the pre-disease state, but this requires multiple samples to reach a correct diagnosis for one individual, thereby restricting its clinical application. RESULTS: In this article, we have developed a novel computational approach based on the DNB theory and differential distributions between the expressions of DNB and non-DNB molecules, which can detect the pre-disease state reliably even from a single sample taken from one individual, by compensating insufficient samples with existing datasets from population studies. Our approach has been validated by the successful identification of pre-disease samples from subjects or individuals before the emergence of disease symptoms for acute lung injury, influenza and breast cancer. Rui Liu 0009, Xiangtian Yu, Xiaoping Liu 0002, Dong Xu 0002, Kazuyuki Aihara, Luonan Chen |
Bioinform. | 6 |
| 2014 | Prediction and early diagnosis of complex diseases by edge-networkabstractMOTIVATION: In this article, we develop a novel edge-based network i.e. edge-network, to detect early signals of diseases by identifying the corresponding edge-biomarkers with their dynamical network biomarker score from dynamical network biomarkers. Specifically, we derive an edge-network based on the second-order statistics representation of gene expression profiles, which is able to accurately represent the stochastic dynamics of the original biological system (with Gaussian distribution assumption) by combining with the traditional node-network, which is based only on the first-order statistics representation of the noisy data. In other words, we show that the stochastic network of a biological system can be described by the integration of its node-network and its edge-network in an accurate manner. RESULTS: By applying edge-network analysis to gene expressions of healthy adults within live influenza experiment sampling at time points before the appearance of infection symptoms, we identified the edge-biomarkers (80 edges with 22 densely connected genes) discovered in edge-networks corresponding to symptomatic adults, which were used to predict the subsequent outcomes of influenza infection. In particular, we not only correctly predict the final infection outcome of each individual at an early time point before his/her clinic symptom but also reveal the key molecules during the disease progression. The prediction accuracy achieves ~90% under the leave-one-out cross-validation. Furthermore, we demonstrate the superiority of our method on disease classification and predication by comparing with the conventional node-biomarkers. Our edge-network analysis not only opens a new way to understand pathogenesis at a network level due to the new representation for a stochastic network, but also provides a powerful tool to make the early diagnosis of diseases. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiangtian Yu, Luonan Chen |
Bioinform. | 3 |
| 2014 | The Dynamics of DNA Methylation Covariation Patterns in CarcinogenesisabstractRecently it has been observed that cancer tissue is characterised by an increased variability in DNA methylation patterns. However, how the correlative patterns in genome-wide DNA methylation change during the carcinogenic progress has not yet been explored. Here we study genome-wide inter-CpG correlations in DNA methylation, in addition to single site variability, during cervical carcinogenesis. We demonstrate how the study of changes in DNA methylation covariation patterns across normal, intra-epithelial neoplasia and invasive cancer allows the identification of CpG sites that indicate the risk of neoplastic transformation in stages prior to neoplasia. Importantly, we show that the covariation in DNA methylation at these risk CpG loci is maximal immediately prior to the onset of cancer, supporting the view that high epigenetic diversity in normal cells increases the risk of cancer. Consistent with this, we observe that invasive cancers exhibit increased covariation in DNA methylation at the risk CpG sites relative to normal tissue, but lower levels relative to pre-cancerous lesions. We further show that the identified risk CpG sites undergo preferential DNA methylation changes in relation to human papilloma virus infection and age. Results are validated in independent data including prospectively collected samples prior to neoplastic transformation. Our data are consistent with a phase transition model of carcinogenesis, in which epigenetic diversity is maximal prior to the onset of cancer. The model and algorithm proposed here may allow, in future, network biomarkers predicting the risk of neoplastic transformation to be identified. Andrew E. Teschendorff, Xiaoping Liu 0002, Helena Carén, Steve M. Pollard, Stephan Beck 0002, Martin Widschwendter, Luonan Chen |
PLoS Comput. Biol. | 7 |
| 2013 | Identifying Critical Transitions of Biological Processes by Dynamical Network Biomarkers
Luonan Chen |
ISBRA | 1 |
| 2013 | NARROMI: a noise and redundancy reduction technique improves accuracy of gene regulatory network inferenceabstractMOTIVATION: Reconstruction of gene regulatory networks (GRNs) is of utmost interest to biologists and is vital for understanding the complex regulatory mechanisms within the cell. Despite various methods developed for reconstruction of GRNs from gene expression profiles, they are notorious for high false positive rate owing to the noise inherited in the data, especially for the dataset with a large number of genes but a small number of samples. RESULTS: In this work, we present a novel method, namely NARROMI, to improve the accuracy of GRN inference by combining ordinary differential equation-based recursive optimization (RO) and information theory-based mutual information (MI). In the proposed algorithm, the noisy regulations with low pairwise correlations are first removed by using MI, and the redundant regulations from indirect regulators are further excluded by RO to improve the accuracy of inferred GRNs. In particular, the RO step can help to determine regulatory directions without prior knowledge of regulators. The results on benchmark datasets from Dialogue for Reverse Engineering Assessments and Methods challenge and experimentally determined GRN of Escherichia coli show that NARROMI significantly outperforms other popular methods in terms of false positive rates and accuracy. AVAILABILITY: All the source data and code are available at: http://csb.shu.edu.cn/narromi.htm. Keqin Liu, Zhi-Ping Liu, Béatrice Duval, Jean-Michel Richer, Xing-Ming Zhao, Jin-Kao Hao, Luonan Chen |
Bioinform. | 8 |
| 2013 | NOA: a cytoscape plugin for network ontology analysisabstractSUMMARY: The Network Ontology Analysis (NOA) plugin for Cytoscape implements the NOA algorithm for network-based enrichment analysis, which extends Gene Ontology annotations to network links, or edges. The plugin facilitates the annotation and analysis of one or more networks in Cytoscape according to user-defined parameters. In addition to tables, the NOA plugin also presents results in the form of heatmaps and overview networks in Cytoscape, which can be exported for publication figures. AVAILABILITY: The NOA plugin is an open source, Java program for Cytoscape version 2.8 available via the Cytoscape App Store (http://apps.cytoscape.org/apps/noa) and plugin manager. A detailed user manual is available at http://nrnb.org/tools/noa. .ucsf.edu Chao Zhang 0032, Kristina Hanspers, Dong Xu 0002, Luonan Chen, Alexander R. Pico |
Bioinform. | 5 |
| 2013 | Research and applications: An integrated approach to identify causal network modules of complex diseases with application to colorectal cancerabstractBACKGROUND: Many methods have been developed to identify disease genes and further module biomarkers of complex diseases based on gene expression data. It is generally difficult to distinguish whether the variations in gene expression are causative or merely the effect of a disease. The limitation of relying on gene expression data alone highlights the need to develop new approaches that can explore various data to reflect the casual relationship between network modules and disease traits. METHODS: In this work, we developed a novel network-based approach to identify putative causal module biomarkers of complex diseases by integrating heterogeneous information, for example, epigenomic data, gene expression data, and protein-protein interaction network. We first formulated the identification of modules as a mathematical programming problem, which can be solved efficiently and effectively in an accurate manner. Then, we applied our approach to colorectal cancer (CRC) and identified several network modules that can serve as potential module biomarkers for characterizing CRC. Further validations using three additional gene expression datasets verified their candidate biomarker properties and the effectiveness of the method. Functional enrichment analysis also revealed that the identified modules are strongly related to hallmarks of cancer, and the enriched functions, such as inflammatory response, receptor and signaling pathways, are specific to CRC. RESULTS: Through constructing a transcription factor (TF)-module network, we found that aberrant DNA methylation of genes encoding TF considerably contributes to the activity change of some genes, which may function as causal genes of CRC, and that can also be exploited to develop efficient therapies or effective drugs. CONCLUSION: Our method can potentially be extended to the study of other complex diseases and the multiclassification problem. Zhenshu Wen, Zhi-Ping Liu, Zhengrong Liu, Luonan Chen |
J. Am. Medical Informatics Assoc. | 5 |
| 2012 | Inferring gene regulatory networks from gene expression data by path consistency algorithm based on conditional mutual informationabstractMOTIVATION: Reconstruction of gene regulatory networks (GRNs), which explicitly represent the causality of developmental or regulatory process, is of utmost interest and has become a challenging computational problem for understanding the complex regulatory mechanisms in cellular systems. However, all existing methods of inferring GRNs from gene expression profiles have their strengths and weaknesses. In particular, many properties of GRNs, such as topology sparseness and non-linear dependence, are generally in regulation mechanism but seldom are taken into account simultaneously in one computational method. RESULTS: In this work, we present a novel method for inferring GRNs from gene expression data considering the non-linear dependence and topological structure of GRNs by employing path consistency algorithm (PCA) based on conditional mutual information (CMI). In this algorithm, the conditional dependence between a pair of genes is represented by the CMI between them. With the general hypothesis of Gaussian distribution underlying gene expression data, CMI between a pair of genes is computed by a concise formula involving the covariance matrices of the related gene expression profiles. The method is validated on the benchmark GRNs from the DREAM challenge and the widely used SOS DNA repair network in Escherichia coli. The cross-validation results confirmed the effectiveness of our method (PCA-CMI), which outperforms significantly other previous methods. Besides its high accuracy, our method is able to distinguish direct (or causal) interactions from indirect associations. AVAILABILITY: All the source data and code are available at: http://csb.shu.edu.cn/subweb/grn.htm. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xing-Ming Zhao, Kun He 0007, Le Lu 0001, Yongwei Cao, Jingdong Liu, Jin-Kao Hao, Zhi-Ping Liu, Luonan Chen |
Bioinform. | 9 |
| 2012 | Identifying dysregulated pathways in cancers from pathway interaction networksabstractBACKGROUND: Cancers, a group of multifactorial complex diseases, are generally caused by mutation of multiple genes or dysregulation of pathways. Identifying biomarkers that can characterize cancers would help to understand and diagnose cancers. Traditional computational methods that detect genes differentially expressed between cancer and normal samples fail to work due to small sample size and independent assumption among genes. On the other hand, genes work in concert to perform their functions. Therefore, it is expected that dysregulated pathways will serve as better biomarkers compared with single genes. RESULTS: In this paper, we propose a novel approach to identify dysregulated pathways in cancer based on a pathway interaction network. Our contribution is three-fold. Firstly, we present a new method to construct pathway interaction network based on gene expression, protein-protein interactions and cellular pathways. Secondly, the identification of dysregulated pathways in cancer is treated as a feature selection problem, which is biologically reasonable and easy to interpret. Thirdly, the dysregulated pathways are identified as subnetworks from the pathway interaction networks, where the subnetworks characterize very well the functional dependency or crosstalk between pathways. The benchmarking results on several distinct cancer datasets demonstrate that our method can obtain more reliable and accurate results compared with existing state of the art methods. Further functional analysis and independent literature evidence also confirm that our identified potential pathogenic pathways are biologically reasonable, indicating the effectiveness of our method. CONCLUSIONS: Dysregulated pathways can serve as better biomarkers compared with single genes. In this work, by utilizing pathway interaction networks and gene expression data, we propose a novel approach that effectively identifies dysregulated pathways, which can not only be used as biomarkers to diagnose cancers but also serve as potential drug targets in the future. Keqin Liu, Zhi-Ping Liu, Jin-Kao Hao, Luonan Chen, Xing-Ming Zhao |
BMC Bioinform. | 4 |
| 2012 | Inferring a protein interaction map of Mycobacterium tuberculosis based on sequences and interologsabstractBACKGROUND: Mycobacterium tuberculosis is an infectious bacterium posing serious threats to human health. Due to the difficulty in performing molecular biology experiments to detect protein interactions, reconstruction of a protein interaction map of M. tuberculosis by computational methods will provide crucial information to understand the biological processes in the pathogenic microorganism, as well as provide the framework upon which new therapeutic approaches can be developed. RESULTS: In this paper, we constructed an integrated M. tuberculosis protein interaction network by machine learning and ortholog-based methods. Firstly, we built a support vector machine (SVM) method to infer the protein interactions of M. tuberculosis H37Rv by gene sequence information. We tested our predictors in Escherichia coli and mapped the genetic codon features underlying its protein interactions to M. tuberculosis. Moreover, the documented interactions of 14 other species were mapped to the interactome of M. tuberculosis by the interolog method. The ensemble protein interactions were validated by various functional relationships, i.e., gene coexpression, evolutionary relationship and functional similarity, extracted from heterogeneous data sources. The accuracy and validation demonstrate the effectiveness and efficiency of our framework. CONCLUSIONS: A protein interaction map of M. tuberculosis is inferred from genetic codons and interologs. The prediction accuracy and numerically experimental validation demonstrate the effectiveness and efficiency of our method. Furthermore, our methods can be straightforwardly extended to infer the protein interactions of other bacterial species. Zhi-Ping Liu, Yu-Qing Qiu, Ross K. K. Leung, Xiang-Sun Zhang, Stephen Kwok-Wing Tsui, Luonan Chen |
BMC Bioinform. | 7 |
| 2012 | Identifying disease genes and module biomarkers by differential interactionsabstractOBJECTIVE: A complex disease is generally caused by the mutation of multiple genes or by the dysfunction of multiple biological processes. Systematic identification of causal disease genes and module biomarkers can provide insights into the mechanisms underlying complex diseases, and help develop efficient therapies or effective drugs. MATERIALS AND METHODS: In this paper, we present a novel approach to predict disease genes and identify dysfunctional networks or modules, based on the analysis of differential interactions between disease and control samples, in contrast to the analysis of differential gene or protein expressions widely adopted in existing methods. RESULTS AND DISCUSSION: As an example, we applied our method to the study of three-stage microarray data for gastric cancer. We identified network modules or module biomarkers that include a set of genes related to gastric cancer, implying the predictive power of our method. The results on holdout validation data sets show that our identified module can serve as an effective module biomarker for accurately detecting or diagnosing gastric cancer, thereby validating the efficiency of our method. CONCLUSION: We proposed a new approach to detect module biomarkers for diseases, and the results on gastric cancer demonstrated that the differential interactions are useful to detect dysfunctional modules in the molecular interaction network, which in turn can be used as robust module biomarkers. Xiaoping Liu 0002, Zhi-Ping Liu, Xing-Ming Zhao, Luonan Chen |
J. Am. Medical Informatics Assoc. | 4 |
| 2012 | Guest Editorial: Bioinformatics and Computational Systems BiologyabstractThe articles in this special section include selected papers from the 2010 IEEE Conference on Bioinformatics and Biomedicine (BIBM). Luonan Chen, Michael Kwok-Po Ng |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2011 | Inferring Protein-Protein Interactions Based on Sequences and Interologs in Mycobacterium Tuberculosis
Zhi-Ping Liu, Yu-Qing Qiu, Ross K. K. Leung, Xiang-Sun Zhang, Stephen Kwok-Wing Tsui, Luonan Chen |
ICIC (3) | 7 |
| 2011 | Neural fate decisions mediated by trans-activation and cis-inhibition in Notch signalingabstractMOTIVATION: In the developing nervous system, the expression of proneural genes, i.e. Hes1, Neurogenin-2 (Ngn2) and Deltalike-1 (Dll1), oscillates in neural progenitors with a period of 2-3 h, but is persistent in post-mitotic neurons. Unlike the synchronization of segmentation clocks, oscillations in neural progenitors are asynchronous between cells. It is known that Notch signaling, in which Notch in a cell can be activated by Dll1 in neighboring cells (trans-activation) and can also be inhibited by Dll1 within the same cell (cis-inhibition), is important for neural fate decisions. There have been extensive studies of trans-activation, but the operating mechanisms and potential implications of cis-inhibition are less clear and need to be further investigated. RESULTS: In this article, we present a computational model for neural fate decisions based on intertwined dynamics with trans-activation and cis-inhibition involving the Hes1, Notch and Dll1 proteins. In agreement with experimental observations, the model predicts that both trans-activation and cis-inhibition play critical roles in regulating the choice between remaining as a progenitor and embarking on neural differentiation. In particular, trans-activation is essential for generation of oscillations in neural progenitors, and cis-inhibition is important for the asynchrony between adjacent cells, indicating that the asynchronous oscillations in neural progenitors depend on cooperation between trans-activation and cis-inhibition. In contrast, cis-inhibition plays more critical roles in embarking on neural differentiation by inactivating intercellular Notch signaling. The model presented here might be a good candidate for providing the first qualitative mechanism of neural fate decisions mediated by both trans-activation and cis-inhibition. Kaihui Liu, Luonan Chen, Kazuyuki Aihara |
Bioinform. | 3 |
| 2010 | Prediction of protein-RNA binding sites by a random forest method with combined featuresabstractMOTIVATION: Protein-RNA interactions play a key role in a number of biological processes, such as protein synthesis, mRNA processing, mRNA assembly, ribosome function and eukaryotic spliceosomes. As a result, a reliable identification of RNA binding site of a protein is important for functional annotation and site-directed mutagenesis. Accumulated data of experimental protein-RNA interactions reveal that a RNA binding residue with different neighbor amino acids often exhibits different preferences for its RNA partners, which in turn can be assessed by the interacting interdependence of the amino acid fragment and RNA nucleotide. RESULTS: In this work, we propose a novel classification method to identify the RNA binding sites in proteins by combining a new interacting feature (interaction propensity) with other sequence- and structure-based features. Specifically, the interaction propensity represents a binding specificity of a protein residue to the interacting RNA nucleotide by considering its two-side neighborhood in a protein residue triplet. The sequence as well as the structure-based features of the residues are combined together to discriminate the interaction propensity of amino acids with RNA. We predict RNA interacting residues in proteins by implementing a well-built random forest classifier. The experiments show that our method is able to detect the annotated protein-RNA interaction sites in a high accuracy. Our method achieves an accuracy of 84.5%, F-measure of 0.85 and AUC of 0.92 prediction of the RNA binding residues for a dataset containing 205 non-homologous RNA binding proteins, and also outperforms several existing RNA binding residue predictors, such as RNABindR, BindN, RNAProB and PPRint, and some alternative machine learning methods, such as support vector machine, naive Bayes and neural network in the comparison study. Furthermore, we provide some biological insights into the roles of sequences and structures in protein-RNA interactions by both evaluating the importance of features for their contributions in predictive accuracy and analyzing the binding patterns of interacting residues. AVAILABILITY: All the source data and code are available at http://www.aporc.org/doc/wiki/PRNA or http://www.sysbio.ac.cn/datatools.asp CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhi-Ping Liu, Ling-Yun Wu, Yong Wang 0001, Xiang-Sun Zhang, Luonan Chen |
Bioinform. | 5 |
| 2010 | Detecting disease associated modules and prioritizing active genes based on high throughput dataabstractBACKGROUND: The accumulation of high-throughput data greatly promotes computational investigation of gene function in the context of complex biological systems. However, a biological function is not simply controlled by an individual gene since genes function in a cooperative manner to achieve biological processes. In the study of human diseases, rather than to discover disease related genes, identifying disease associated pathways and modules becomes an essential problem in the field of systems biology. RESULTS: In this paper, we propose a novel method to detect disease related gene modules or dysfunctional pathways based on global characteristics of interactome coupled with gene expression data. Specifically, we exploit interacting relationships between genes to define a gene's active score function based on the kernel trick, which can represent nonlinear effects of gene cooperativity. Then, modules or pathways are inferred based on the active scores evaluated by the support vector regression in a global and integrative manner. The efficiency and robustness of the proposed method are comprehensively validated by using both simulated and real data with the comparison to existing methods. CONCLUSIONS: By applying the proposed method to two cancer related problems, i.e. breast cancer and prostate cancer, we successfully identified active modules or dysfunctional pathways related to these two types of cancers with literature confirmed evidences. We show that this network-based method is highly efficient and can be applied to a large-scale problem especially for human disease related modules or pathway extraction. Moreover, this method can also be used for prioritizing genes associated with a specific phenotype or disease. Yu-Qing Qiu, Xiang-Sun Zhang, Luonan Chen |
BMC Bioinform. | 4 |
| 2009 | Dynamically dysfunctional protein interactions in the development of Alzheimer's diseaseabstractAlzheimer's disease usually causes dementia in the old people and the symptom progression of the disease phenotype displays certain patterns. One possible reason is that the nerve cells in the brains of the patients degenerate at different stages. Here, we analyze the dynamics of disease progression based on its biomolecular network. We develop a novel computational method to integrate an ensemble protein network and the hippocampal gene expression data. Specifically, we construct the induced dynamical pathways which present particular characteristics at different disease stages from the control to disease samples. Based on the network-based method, we reveal that the active pathways tend to be more complicated during the development of disease. Also we find that the disease proteins performing important functions are always located in the cooperations of the identified pathways. These results also demonstrate that the network-based analysis can provide knowledge and evidences on the dynamics and pathological pathways of the complex Alzheimer's disease. Zhi-Ping Liu, Yong Wang 0001, Tieqiao Wen, Xiang-Sun Zhang, Weiming Xia, Luonan Chen |
SMC | 6 |
| 2009 | Modeling post-transcriptional regulation activity of small non-coding RNAs in Escherichia coliabstractBACKGROUND: Transcriptional regulation is a fundamental process in biological systems, where transcription factors (TFs) have been revealed to play crucial roles. In recent years, in addition to TFs, an increasing number of non-coding RNAs (ncRNAs) have been shown to mediate post-transcriptional processes and regulate many critical pathways in both prokaryotes and eukaryotes. On the other hand, with more and more high-throughput biological data becoming available, it is possible and imperative to quantitatively study gene regulation in a systematic and detailed manner. RESULTS: Most existing studies for inferring transcriptional regulatory interactions and the activity of TFs ignore the possible post-transcriptional effects of ncRNAs. In this work, we propose a novel framework to infer the activity of regulators including both TFs and ncRNAs by exploring the expression profiles of target genes and (post)transcriptional regulatory relationships. We model the integrated regulatory system by a set of biochemical reactions which lead to a log-bilinear problem. The inference process is achieved by an iterative algorithm, in which two linear programming models are efficiently solved. In contrast to available related studies, the effects of ncRNAs on transcription process are considered in this work, and thus more reasonable and accurate reconstruction can be expected. In addition, the approach is suitable for large-scale problems from the viewpoint of computation. Experiments on two synthesized data sets and a model system of Escherichia coli (E. coli) carbon source transition from glucose to acetate illustrate the effectiveness of our model and algorithm. CONCLUSION: Our results show that incorporating the post-transcriptional regulation of ncRNAs into system model can mine the hidden effects from the regulation activity of TFs in transcription processes and thus can uncover the biological mechanisms in gene regulation in a more accurate manner. The software for the algorithm in this paper is available upon request. Rui-Sheng Wang, Guangxu Jin, Xiang-Sun Zhang, Luonan Chen |
BMC Bioinform. | 4 |
| 2009 | Robustness of interval gene networks with multiple time-varying delays and noise
Jianwei Shen 0001, Baoguo Niu, Zengrong Liu, Luonan Chen |
Neurocomputing | 5 |
| 2009 | Disease-Aging Network Reveals Significant Roles of Aging Genes in Connecting Genetic DiseasesabstractOne of the challenging problems in biology and medicine is exploring the underlying mechanisms of genetic diseases. Recent studies suggest that the relationship between genetic diseases and the aging process is important in understanding the molecular mechanisms of complex diseases. Although some intricate associations have been investigated for a long time, the studies are still in their early stages. In this paper, we construct a human disease-aging network to study the relationship among aging genes and genetic disease genes. Specifically, we integrate human protein-protein interactions (PPIs), disease-gene associations, aging-gene associations, and physiological system-based genetic disease classification information in a single graph-theoretic framework and find that (1) human disease genes are much closer to aging genes than expected by chance; and (2) diseases can be categorized into two types according to their relationships with aging. Type I diseases have their genes significantly close to aging genes, while type II diseases do not. Furthermore, we examine the topological characters of the disease-aging network from a systems perspective. Theoretical results reveal that the genes of type I diseases are in a central position of a PPI network while type II are not; (3) more importantly, we define an asymmetric closeness based on the PPI network to describe relationships between diseases, and find that aging genes make a significant contribution to associations among diseases, especially among type I diseases. In conclusion, the network-based study provides not only evidence for the intricate relationship between the aging process and genetic diseases, but also biological implications for prying into the nature of human diseases. Yong Wang 0001, Luonan Chen, Xiang-Sun Zhang |
PLoS Comput. Biol. | 4 |
| 2009 | Evaluating Protein Similarity from Coarse StructuresabstractTo unscramble the relationship between protein function and protein structure, it is essential to assess the protein similarity from different aspects. Although many methods have been proposed for protein structure alignment or comparison, alternative similarity measures are still strongly demanded due to the requirement of fast screening and query in large-scale structure databases. In this paper, we first formulate a novel representation of a protein structure, i.e., Feature Sequence of Surface (FSS). Then, a new score scheme is developed to measure the similarity between two representations. To verify the proposed method, numerical experiments are conducted in four different protein data sets. We also classify SARS coronavirus to verify the effectiveness of the new method. Furthermore, preliminary results of fast classification of the whole CATH v2.5.1 database based on the new macrostructure similarity are given as a pilot study. We demonstrate that the proposed approach to measure the similarities between protein structures is simple to implement, computationally efficient, and surprisingly fast. In addition, the method itself provides a new and quantitative tool to view a protein structure. Yong Wang 0001, Ling-Yun Wu, Zhong-Wei Zhan, Xiang-Sun Zhang, Luonan Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2008 | Automatic Modeling of Signal Pathways from Protein-Protein Interaction Networks
Xing-Ming Zhao, Rui-Sheng Wang, Luonan Chen, Kazuyuki Aihara |
APBC | 3 |
| 2008 | Reconstruction of Regulator Activity in E. coli Post-Transcription ProcessesabstractTranscriptional regulation is a fundamental process in biological systems, where transcription factors (TFs) play crucial roles. Except for TFs, an increasing number of small non-coding RNAs (ncRNAs) have been shown to mediate post-transcriptional processes in both prokaryotes and eukaryotes. In this work, we propose a novel approach to infer the activities of regulators including TFs and ncRNAs by exploring target gene expression profiles and (post) transcriptional regulatory relationships. The inference process is efficiently achieved by an iteration algorithm, in which two linear programming models are iteratively solved. In contrast to the existing works, for the first time, the effects of ncRNAs on transcription process are considered and thus more reasonable inference can be expected. Experiments on a model system of E. coli carbon source transition from glucose to acetate illustrate the effectiveness of our method. Rui-Sheng Wang, Guangxu Jin, Xiang-Sun Zhang, Luonan Chen |
BIBM | 4 |
| 2008 | A Novel Classifier Based on Enhanced Lipschitz Embedding for Speech Emotion Recognition
Mingyu You, Guo-Zheng Li 0001, Luonan Chen, Jianhua Tao 0001 |
ICIC (1) | 3 |
| 2008 | Gene function prediction using labeled and unlabeled dataabstractBACKGROUND: In general, gene function prediction can be formalized as a classification problem based on machine learning technique. Usually, both labeled positive and negative samples are needed to train the classifier. For the problem of gene function prediction, however, the available information is only about positive samples. In other words, we know which genes have the function of interested, while it is generally unclear which genes do not have the function, i.e. the negative samples. If all the genes outside of the target functional family are seen as negative samples, the imbalanced problem will arise because there are only a relatively small number of genes annotated in each family. Furthermore, the classifier may be degraded by the false negatives in the heuristically generated negative samples. RESULTS: In this paper, we present a new technique, namely Annotating Genes with Positive Samples (AGPS), for defining negative samples in gene function prediction. With the defined negative samples, it is straightforward to predict the functions of unknown genes. In addition, the AGPS algorithm is able to integrate various kinds of data sources to predict gene functions in a reliable and accurate manner. With the one-class and two-class Support Vector Machines as the core learning algorithm, the AGPS algorithm shows good performances for function prediction on yeast genes. CONCLUSION: We proposed a new method for defining negative samples in gene function prediction. Experimental results on yeast genes show that AGPS yields good performances on both training and test sets. In addition, the overlapping between prediction results and GO annotations on unknown genes also demonstrates the effectiveness of the proposed method. Xing-Ming Zhao, Yong Wang 0001, Luonan Chen, Kazuyuki Aihara |
BMC Bioinform. | 3 |
| 2008 | Associative memory with a controlled chaotic neural network
Guoguang He, Luonan Chen, Kazuyuki Aihara |
Neurocomputing | 2 |
| 2008 | Clustering complex networks and biological networks by nonnegative matrix factorization with various similarity measures
Rui-Sheng Wang, Yong Wang 0001, Xiang-Sun Zhang, Luonan Chen |
Neurocomputing | 5 |
| 2008 | Modeling and Analyzing Biological Oscillations in Molecular NetworksabstractOne of the major challenges for postgenomic biology is to understand how genes, proteins, and small molecules dynamically interact to form molecular networks which facilitate sophisticated biological functions. In this paper, we present a survey on recent developments on modelling molecular networks and analyzing synchronization of bio-oscillators in multicellular systems from the viewpoint of systems biology. Attention will be focused on deriving general theoretical results to understand the dynamical behaviors of biological systems based on nonlinear dynamical and control theory. Specifically, we first describe the stochastic and deterministic approaches to model molecular networks and give a brief comparison between them. Then, we explain how to construct a molecular network, in particular, a gene regulatory network with specific functions, e.g., switches and oscillators, in individual cells at the molecular level by using feedback systems, and how to model a general multicellular system with the consideration of external fluctuations and intercellular coupling to study the general cooperative behaviors for a population of bio-oscillators. Finally, as an illustrative example, a synthetic multicellular system is designed to show how synchronization is effectively achieved and how dynamics of individual cells is efficiently controlled. Some recent developments and perspectives of analysis on biological oscillations in future are also discussed. Luonan Chen, Kazuyuki Aihara |
Proc. IEEE | 3 |
| 2007 | Double Three-Level Inverter Based Variable Frequency Drive with Minimal Total Harmonic Distortion Using Particle Swarm Optimization
Huibo Lou, Chengxiong Mao, Jiming Lu, Dan Wang 0008, Luonan Chen |
ICIC (1) | 5 |
| 2007 | Identifying Modules in Complex Networks by a Graph-Theoretical Method and Its Application in Protein Interaction Networks
Rui-Sheng Wang, Xiang-Sun Zhang, Luonan Chen |
ICIC (2) | 4 |
| 2007 | Alignment of molecular networks by integer quadratic programmingabstractMOTIVATION: With more and more data on molecular networks (e.g. protein interaction networks, gene regulatory networks and metabolic networks) available, the discovery of conserved patterns or signaling pathways by comparing various kinds of networks among different species or within a species becomes an increasingly important problem. However, most of the conventional approaches either restrict comparative analysis to special structures, such as pathways, or adopt heuristic algorithms due to computational burden. RESULTS: In this article, to find the conserved substructures, we develop an efficient algorithm for aligning molecular networks based on both molecule similarity and architecture similarity, by using integer quadratic programming (IQP). Such an IQP can be relaxed into the corresponding quadratic programming (QP) which almost always ensures an integer solution, thereby making molecular network alignment tractable without any approximation. The proposed framework is very flexible and can be applied to many kinds of molecular networks including weighted and unweighted, directed and undirected networks with or without loops. AVAILABILITY: Matlab code and data are available from http://zhangroup.aporc.org/bioinfo/MNAligner or http://intelligent.eic.osaka-sandai.ac.jp/chenen/software/MNAligner, or upon request from authors. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhen-Ping Li, Yong Wang 0001, Xiang-Sun Zhang, Luonan Chen |
Bioinform. | 5 |
| 2007 | Inferring transcriptional regulatory networks from high-throughput dataabstractMOTIVATION: Inferring the relationships between transcription factors (TFs) and their targets has utmost importance for understanding the complex regulatory mechanisms in cellular systems. However, the transcription factor activities (TFAs) cannot be measured directly by standard microarray experiment owing to various post-translational modifications. In particular, cooperative mechanism and combinatorial control are common in gene regulation, e.g. TFs usually recruit other proteins cooperatively to facilitate transcriptional reaction processes. RESULTS: In this article, we propose a novel method for inferring transcriptional regulatory networks (TRN) from gene expression data based on protein transcription complexes and mass action law. With gene expression data and TFAs estimated from transcription complex information, the inference of TRN is formulated as a linear programming (LP) problem which has a globally optimal solution in terms of L(1) norm error. The proposed method not only can easily incorporate ChIP-Chip data as prior knowledge, but also can integrate multiple gene expression datasets from different experiments simultaneously. A unique feature of our method is to take into account protein cooperation in transcription process. We tested our method by using both synthetic data and several experimental datasets in yeast. The extensive results illustrate the effectiveness of the proposed method for predicting transcription regulatory relationships between TFs with co-regulators and target genes. Rui-Sheng Wang, Yong Wang 0001, Xiang-Sun Zhang, Luonan Chen |
Bioinform. | 4 |
| 2007 | Predicting gene ontology functions from protein's regional surface structuresabstractBACKGROUND: Annotation of protein functions is an important task in the post-genomic era. Most early approaches for this task exploit only the sequence or global structure information. However, protein surfaces are believed to be crucial to protein functions because they are the main interfaces to facilitate biological interactions. Recently, several databases related to structural surfaces, such as pockets and cavities, have been constructed with a comprehensive library of identified surface structures. For example, CASTp provides identification and measurements of surface accessible pockets as well as interior inaccessible cavities. RESULTS: A novel method was proposed to predict the Gene Ontology (GO) functions of proteins from the pocket similarity network, which is constructed according to the structure similarities of pockets. The statistics of the networks were presented to explore the relationship between the similar pockets and GO functions of proteins. Cross-validation experiments were conducted to evaluate the performance of the proposed method. Results and codes are available at: http://zhangroup.aporc.org/bioinfo/PSN/. CONCLUSION: The computational results demonstrate that the proposed method based on the pocket similarity network is effective and efficient for predicting GO functions of proteins in terms of both computational complexity and prediction accuracy. The proposed method revealed strong relationship between small surface patterns (or pockets) and GO functions, which can be further used to identify active sites or functional motifs. The high quality performance of the prediction method together with the statistics also indicates that pockets play essential roles in biological interactions or the GO functions. Moreover, in addition to pockets, the proposed network framework can also be used for adopting other protein spatial surface patterns to predict the protein functions. Zhi-Ping Liu, Ling-Yun Wu, Yong Wang 0001, Luonan Chen, Xiang-Sun Zhang |
BMC Bioinform. | 4 |
| 2007 | Analysis on multi-domain cooperation for predicting protein-protein interactionsabstractBACKGROUND: Domains are the basic functional units of proteins. It is believed that protein-protein interactions are realized through domain interactions. Revealing multi-domain cooperation can provide deep insights into the essential mechanism of protein-protein interactions at the domain level and be further exploited to improve the accuracy of protein interaction prediction. RESULTS: In this paper, we aim to identify cooperative domains for protein interactions by extending two-domain interactions to multi-domain interactions. Based on the high-throughput experimental data from multiple organisms with different reliabilities, the interactions of domains were inferred by a Linear Programming algorithm with Multi-domain pairs (LPM) and an Association Probabilistic Method with Multi-domain pairs (APMM). Experimental results demonstrate that our approach not only can find cooperative domains effectively but also has a higher accuracy for predicting protein interaction than the existing methods. Cooperative domains, including strongly cooperative domains and superdomains, were detected from major interaction databases MIPS and DIP, and many of them were verified by physical interactions from the crystal structures of protein complexes in PDB which provide intuitive evidences for such cooperation. Comparison experiments in terms of protein/domain interaction prediction justified the benefit of considering multi-domain cooperation. CONCLUSION: From the computational viewpoint, this paper gives a general framework to predict protein interactions in a more accurate manner by considering the information of both multi-domains and multiple organisms, which can also be applied to identify cooperative domains, to reconstruct large complexes and further to annotate functions of domains. Supplementary information and software are provided in http://intelligent.eic.osaka-sandai.ac.jp/chenen/MDCinfer.htm and http://zhangroup.aporc.org/bioinfo/MDCinfer. Rui-Sheng Wang, Yong Wang 0001, Ling-Yun Wu, Xiang-Sun Zhang, Luonan Chen |
BMC Bioinform. | 5 |
| 2006 | Supervised Inference of Gene Regulatory Networks by Linear Programming
Yong Wang 0001, Trupti Joshi, Dong Xu 0002, Xiang-Sun Zhang, Luonan Chen |
ICIC (3) | 5 |
| 2006 | Forecasting Electricity Demand by Hybrid Machine Learning Model
Shu Fan, Chengxiong Mao, Luonan Chen |
ICONIP (2) | 4 |
| 2006 | Designing synthetic biological networksabstractWe aim to design and construct synthetic biological networks, such as gene switch or gene oscillator in this paper from the viewpoint of integrated systems biology. The theoretical results provide insights into the modular structures of gene regulatory networks and their quantitative functions. A hybrid feedback network is adopted to provide a variety of functional designs, which demonstrate the crucial role of interactions between biological modules. Luonan Chen, Xiabo Zhou, S. Wong |
ISCAS | 1 |
| 2006 | Automatic Classification of Protein Structures Based on Convex Hull Representation by Integrated Neural Network
Yong Wang 0001, Ling-Yun Wu, Xiang-Sun Zhang, Luonan Chen |
TAMC | 4 |
| 2006 | Synchronizing a multicellular system by external input: an artificial control strategyabstractMOTIVATION: Although there are significant advances on elucidating the collective behaviors on biological organisms in recent years, the essential mechanisms by which the collective rhythms arise remain to be fully understood, and further how to synchronize multicellular networks by artificial control strategy has not yet been well explored. RESULTS: A control strategy is developed to synchronize gene regulatory networks in a multicellular system when spontaneous synchronization cannot be achieved. We first construct an impulsive control system to model the process of periodically injecting coupling substances with constant or random impulsive control amounts into the common extracellular medium, and further study its effects on the dynamics of individual cells. We derive the threshold of synchronization induced by the periodic substance input. Therefore, we can synchronize the multicellular network to a specific collective behavior by changing the frequency and amplitude of the periodic stimuli. Moreover, a two-stage scheme is proposed to facilitate the synchronization in this paper. We show that the presence of the external input may also initiate different dynamics. The multicellular network of coupled repressilators is used to show the effectiveness of the proposed method. The results not only provide a perspective to understand the interactions between external stimuli and intrinsic physiological rhythms, but also may lead to development of realistic artificial control strategy and medical therapy. CONTACT: [email protected]. Luonan Chen, Kazuyuki Aihara |
Bioinform. | 2 |
| 2006 | Inferring gene regulatory networks from multiple microarray datasetsabstractMOTIVATION: Microarray gene expression data has increasingly become the common data source that can provide insights into biological processes at a system-wide level. One of the major problems with microarrays is that a dataset consists of relatively few time points with respect to a large number of genes, which makes the problem of inferring gene regulatory network an ill-posed one. On the other hand, gene expression data generated by different groups worldwide are increasingly accumulated on many species and can be accessed from public databases or individual websites, although each experiment has only a limited number of time-points. RESULTS: This paper proposes a novel method to combine multiple time-course microarray datasets from different conditions for inferring gene regulatory networks. The proposed method is called GNR (Gene Network Reconstruction tool) which is based on linear programming and a decomposition procedure. The method theoretically ensures the derivation of the most consistent network structure with respect to all of the datasets, thereby not only significantly alleviating the problem of data scarcity but also remarkably improving the prediction reliability. We tested GNR using both simulated data and experimental data in yeast and Arabidopsis. The result demonstrates the effectiveness of GNR in terms of predicting new gene regulatory relationship in yeast and Arabidopsis. AVAILABILITY: The software is available from http://zhangorup.aporc.org/bioinfo/grninfer/, http://digbio.missouri.edu/grninfer/ and http://intelligent.eic.osaka-sandai.ac.jp or upon request from the authors. Yong Wang 0001, Trupti Joshi, Xiang-Sun Zhang, Dong Xu 0002, Luonan Chen |
Bioinform. | 5 |
| 2006 | Transient Resetting: A Novel Mechanism for Synchrony and Its Biological ExamplesabstractThe study of synchronization in biological systems is essential for the understanding of the rhythmic phenomena of living organisms at both molecular and cellular levels. In this paper, by using simple dynamical systems theory, we present a novel mechanism, named transient resetting, for the synchronization of uncoupled biological oscillators with stimuli. This mechanism not only can unify and extend many existing results on (deterministic and stochastic) stimulus-induced synchrony, but also may actually play an important role in biological rhythms. We argue that transient resetting is a possible mechanism for the synchronization in many biological organisms, which might also be further used in the medical therapy of rhythmic disorders. Examples of the synchronization of neural and circadian oscillators as well as a chaotic neuron model are presented to verify our hypothesis. Luonan Chen, Kazuyuki Aihara |
PLoS Comput. Biol. | 2 |
| 2005 | Peak Load Forecasting Using the Self-organizing Map
Shu Fan, Chengxiong Mao, Luonan Chen |
ISNN (3) | 3 |
| 2005 | Noise-induced cooperative behavior in a multicell systemabstractMOTIVATION: All cell components exhibit intracellular noise on account of random births and deaths of individual molecules, and extracellular noise because of environment perturbations. Gene regulation in particular, is an inherently noisy process with transcriptional control, alternative splicing, translation, diffusion and chemical modification reactions, all of which involve stochastic fluctuations. Such stochastic noises may not only affect the dynamics of the entire system but may also be exploited by living organisms to actively facilitate certain functions, such as cooperative behavior and communication. RESULTS: We have provided a general model and an analytic tool to examine the cooperative behavior of a multicell system with both intracellular and extracellular stochastic fluctuations. A multicell system with a synthetic gene network is adopted to demonstrate the effects of noises and coupling on collective dynamics. These results establish not only a theoretical foundation but also a quantitative basis for understanding essential roles of noises on cooperative dynamics, such as synchronization and communication among cells. Luonan Chen, Tianshou Zhou, Kazuyuki Aihara |
Bioinform. | 1 |
| 2005 | Protein structure alignment by deterministic annealingabstractMOTIVATION: Protein structure alignment is one of the most important computational problems in molecular biology and plays a key role in protein structure prediction, fold family classification, motif finding, phylogenetic tree reconstruction and so on. From the viewpoint of computational complexity, a pairwise structure alignment is also a NP-hard problem, in contrast to the polynomial time algorithm for a pairwise sequence alignment. RESULTS: We propose a method for solving the structure alignment problem in an accurate manner at the amino acid level, based on a mean field annealing technique. We define the structure alignment as a mixed integer-programming (MIP) problem. By avoiding complicated combinatorial computation and exploiting the special structure of the continuous partial problem, we transform the MIP into a reduced non-linear continuous optimization problem (NCOP) with a much simpler form. To optimize the reduced NCOP, a mean field annealing procedure is adopted with a modified Potts model, whose solution is generally identical to that of the MIP. There is no 'soft constraint' in our mean field model and all constraints are automatically satisfied throughout the annealing process, thereby not only making the optimization more efficient but also eliminating many unnecessary parameters that depend on problems and usually require careful tuning. A number of benchmark examples are tested by the proposed method with comparisons to several existing approaches. Luonan Chen, Tianshou Zhou |
Bioinform. | 1 |
| 2005 | A parsimonious tree-grow method for haplotype inferenceabstractMOTIVATION: Haplotype information has become increasingly important in analyzing fine-scale molecular genetics data, such as disease genes mapping and drug design. Parsimony haplotyping is one of haplotyping problems belonging to NP-hard class. RESULTS: In this paper, we aim to develop a novel algorithm for the haplotype inference problem with the parsimony criterion, based on a parsimonious tree-grow method (PTG). PTG is a heuristic algorithm that can find the minimum number of distinct haplotypes based on the criterion of keeping all genotypes resolved during tree-grow process. In addition, a block-partitioning method is also proposed to improve the computational efficiency. We show that the proposed approach is not only effective with a high accuracy, but also very efficient with the computational complexity in the order of O(m2n) time for n single nucleotide polymorphism sites in m individual genotypes. AVAILABILITY: The software is available upon request from the authors, or from http://zhangroup.aporc.org/bioinfo/ptg/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supporting materials is available from http://zhangroup.aporc.org/bioinfo/ptg/bti572supplementary.pdf Zhen-Ping Li, Wenfeng Zhou, Xiang-Sun Zhang, Luonan Chen |
Bioinform. | 4 |
| 1998 | Bifurcation analysis of hybrid dynamical systemsabstractDefines a model formulated by differential-difference-algebraic equations (DDA) as a hybrid dynamical system (HDS) or constrained sampled-data model. So far, both continuous-time and discrete-time nonlinear systems have attracted considerable attention and a variety of techniques have been developed. In contrast, however, the researches for the hybrid systems are mainly for the systems defined by both differential equations with continuous states and logical-discrete-event equations with discrete states. On the other hand, the sampled-data models or differential-difference equations where all of the states are continuous are investigated mostly for the linear systems. Less attention has been focused on the nonlinear analysis of the DDA or the hybrid dynamical systems where the differential and the difference equations not only have continuous states but also are constrained by algebraic equations. As part of a continuing series of works which attempt to elucidate the properties of HDS following the analysis of the asymptotical stability in previous papers, this paper aims at analysing bifurcations of HDS and further applying the theoretical results to digital control of power systems. Luonan Chen, Kazuyuki Aihara |
SMC | 1 |
| 1995 | Chaotic simulated annealing by a neural network model with transient chaos
Luonan Chen, Kazuyuki Aihara |
Neural Networks | 1 |