EDBT 2026 Demo / reviewers in the wild / expert
Xing-Ming Zhao
dblp:39/5135
· DBLP profile ↗
61ranked-venue papers
11as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 47 · 7 first-author · 17 since 2021Artificial intelligence and machine learning · 14 · 4 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | E 2 AD: Enhanced and explainable Alzheimer's disease detection framework via anatomy- and relation-aware cross-modal knowledge distillation
Sirong Piao, Tao Chen 0055, Zhaoyang Li 0016, Tongrui Zhang, Xing-Ming Zhao, Hongming Shan |
Medical Image Anal. | 8 |
| 2026 | EIGNN: An Explainable Imaging-Genetic Neural Network for Robust Alzheimer's Disease Risk PredictionabstractAccurate risk prediction and early diagnosis are crucial for the early intervention of Alzheimer's disease (AD). Current prediction models usually have limited power in capturing the complex interplays between the heterogeneous inputs or lack biological explainability required for clinical adoption and new diagnosis biomarkers discovery. Inspired by pioneering works on biologically informed network and multi-modal learning, we presented an Explainable Imaging-Genetic Neural Network (EIGNN), integrating genetic and neuroimaging data to generate accurate, robust, and explainable AD risk prediction. The EIGNN model features a biologically-informed architecture, incorporating a multi-GWAS SNP selection strategy, enhanced explainable neural network design, and a modal attention mechanism. Genetic variants were hierarchically mapped to their target genes and biological pathways, and further integrated with neuroimaging features. We demonstrated that the EIGNN model outperformed existing methods and exhibited improved robustness, explainability, and reproducibility. Finally, by applying a novel biologically informed multi-modal feature interaction map, we prioritized a set of AD risk genes and biological pathways, and explored the intricate interactions between the risk genes and brain regions implicated in AD risk. Zi-Chao Zhang 0001, Zhigao Cai, Xingzhong Zhao, Jixin Cao, Yucheng T. Yang, Xing-Ming Zhao |
IEEE Trans. Medical Imaging | 8 |
| 2025 | Generalized Fetal Brain Segmentation and Identification from Multi-View MRIabstractAutomated fetal brain segmentation from MRI remains challenging due to multi-view inconsistency and limited generalization across fetal types and gestational ages. In this work, we proposed Co-UNext, a novel framework for generalized multi-view fetal brain segmentation and identification. The model employed a two-stage cascade architecture processing multi-view stacks from maternal volumes, integrating depthwise separable convolutions and attention mechanisms. Co-UNext achieved precise segmentation across multi-view stacks while providing fetus-specific segmentation results for multiple pregnancies. We evaluated the framework on 212 fetal MRI scans from 20-36 weeks gestation, including 124 singleton and 88 twin pregnancies scanned on 1.5T Siemens systems with axial, coronal, and sagittal views. Co-UNext achieves 96.9% Dice and 1.23 mm HD95 on singletons, 96.3% Dice and 1.67 mm HD95 on twins, significantly outperforming existing models. The proposed method exhibits superior generalization capabilities and robustness to motion artifacts, facilitating various subsequent quantitative analyses in prenatal care. Zhigao Cai, Qiongjie Zhou, Xing-Ming Zhao |
SMC | 4 |
| 2024 | MicroHDF: predicting host phenotypes with metagenomic data using a deep forest-based frameworkabstractThe gut microbiota plays a vital role in human health, and significant effort has been made to predict human phenotypes, especially diseases, with the microbiota as a promising indicator or predictor with machine learning (ML) methods. However, the accuracy is impacted by a lot of factors when predicting host phenotypes with the metagenomic data, e.g. small sample size, class imbalance, high-dimensional features, etc. To address these challenges, we propose MicroHDF, an interpretable deep learning framework to predict host phenotypes, where a cascade layers of deep forest units is designed for handling sample class imbalance and high dimensional features. The experimental results show that the performance of MicroHDF is competitive with that of existing state-of-the-art methods on 13 publicly available datasets of six different diseases. In particular, it performs best with the area under the receiver operating characteristic curve of 0.9182 ± 0.0098 and 0.9469 ± 0.0076 for inflammatory bowel disease (IBD) and liver cirrhosis, respectively. Our MicroHDF also shows better performance and robustness in cross-study validation. Furthermore, MicroHDF is applied to two high-risk diseases, IBD and autism spectrum disorder, as case studies to identify potential biomarkers. In conclusion, our method provides an effective and reliable prediction of the host phenotype and discovers informative features with biological insights. Kai Shi 0004, Qiaohui Liu, Qingrong Ji, Qisheng He, Xing-Ming Zhao |
Briefings Bioinform. | 5 |
| 2023 | Prioritizing genes associated with brain disorders by leveraging enhancer-promoter interactions in diverse neural cells and tissuesabstractPrioritizing genes that underlie complex brain disorders poses a considerable challenge. By using the CAGE (Cap Analysis of Gene Expression) read alignment files for 439 human cell and tissue types (including primary cells, tissues and cell lines) from FANTOM5 project, we predicted enhancer-promoter interactions (EPIs) of 439 cell and tissue types in human, and examined their reliability. We found that identified EPIs showed activity specificity and network aggregation in cell and tissue types, and discovered that most neurological disorders exhibit heritability enrichment in neural stem cells and astrocytes, while psychiatric disorders and behavioral-cognitive phenotypes exhibit enrichment in neurons. Xingzhong Zhao, Yucheng T. Yang, Xing-Ming Zhao |
BIBM | 3 |
| 2023 | A survey on computational strategies for genome-resolved gut metagenomicsabstractRecovering high-quality metagenome-assembled genomes (HQ-MAGs) is critical for exploring microbial compositions and microbe-phenotype associations. However, multiple sequencing platforms and computational tools for this purpose may confuse researchers and thus call for extensive evaluation. Here, we systematically evaluated a total of 40 combinations of popular computational tools and sequencing platforms (i.e. strategies), involving eight assemblers, eight metagenomic binners and four sequencing technologies, including short-, long-read and metaHiC sequencing. We identified the best tools for the individual tasks (e.g. the assembly and binning) and combinations (e.g. generating more HQ-MAGs) depending on the availability of the sequencing data. We found that the combination of the hybrid assemblies and metaHiC-based binning performed best, followed by the hybrid and long-read assemblies. More importantly, both long-read and metaHiC sequencings link more mobile elements and antibiotic resistance genes to bacterial hosts and improve the quality of public human gut reference genomes with 32% (34/105) HQ-MAGs that were either of better quality than those in the Unified Human Gastrointestinal Genome catalog version 2 or novel. Long-Hao Jia, Yingjian Wu, Yanqi Dong, Jingchao Chen, Wei-Hua Chen, Xing-Ming Zhao |
Briefings Bioinform. | 6 |
| 2023 | Deciphering the genetic architecture of human brain structure and function: a brief survey on recent advances of neuroimaging genomicsabstractBrain imaging genomics is an emerging interdisciplinary field, where integrated analysis of multimodal medical image-derived phenotypes (IDPs) and multi-omics data, bridging the gap between macroscopic brain phenotypes and their cellular and molecular characteristics. This approach aims to better interpret the genetic architecture and molecular mechanisms associated with brain structure, function and clinical outcomes. More recently, the availability of large-scale imaging and multi-omics datasets from the human brain has afforded the opportunity to the discovering of common genetic variants contributing to the structural and functional IDPs of the human brain. By integrative analyses with functional multi-omics data from the human brain, a set of critical genes, functional genomic regions and neuronal cell types have been identified as significantly associated with brain IDPs. Here, we review the recent advances in the methods and applications of multi-omics integration in brain imaging analysis. We highlight the importance of functional genomic datasets in understanding the biological functions of the identified genes and cell types that are associated with brain IDPs. Moreover, we summarize well-known neuroimaging genetics datasets and discuss challenges and future directions in this field. Xingzhong Zhao, Anyi Yang, Zi-Chao Zhang 0001, Yucheng T. Yang, Xing-Ming Zhao |
Briefings Bioinform. | 5 |
| 2023 | SemiBin2: self-supervised contrastive learning leads to better MAGs for short- and long-read sequencingabstractMOTIVATION: Metagenomic binning methods to reconstruct metagenome-assembled genomes (MAGs) from environmental samples have been widely used in large-scale metagenomic studies. The recently proposed semi-supervised binning method, SemiBin, achieved state-of-the-art binning results in several environments. However, this required annotating contigs, a computationally costly and potentially biased process. RESULTS: We propose SemiBin2, which uses self-supervised learning to learn feature embeddings from the contigs. In simulated and real datasets, we show that self-supervised learning achieves better results than the semi-supervised learning used in SemiBin1 and that SemiBin2 outperforms other state-of-the-art binners. Compared to SemiBin1, SemiBin2 can reconstruct 8.3-21.5% more high-quality bins and requires only 25% of the running time and 11% of peak memory usage in real short-read sequencing samples. To extend SemiBin2 to long-read data, we also propose ensemble-based DBSCAN clustering algorithm, resulting in 13.1-26.3% more high-quality genomes than the second best binner for long-read data. AVAILABILITY AND IMPLEMENTATION: SemiBin2 is available as open source software at https://github.com/BigDataBiology/SemiBin/ and the analysis scripts used in the study can be found at https://github.com/BigDataBiology/SemiBin2_benchmark. Shaojun Pan, Xing-Ming Zhao, Luís Pedro Coelho |
Bioinform. | 2 |
| 2023 | Improving Alzheimer's Disease Diagnosis With Multi-Modal PET Embedding Features by a 3D Multi-Task MLP-Mixer Neural NetworkabstractPositron emission tomography (PET) with fluorodeoxyglucose (FDG) or florbetapir (AV45) has been proved effective in the diagnosis of Alzheimer's disease. However, the expensive and radioactive nature of PET has limited its application. Here, employing multi-layer perceptron mixer architecture, we present a deep learning model, namely 3-dimensional multi-task multi-layer perceptron mixer, for simultaneously predicting the standardized uptake value ratios (SUVRs) for FDG-PET and AV45-PET from the cheap and widely used structural magnetic resonance imaging data, and the model can be further used for Alzheimer's disease diagnosis based on embedding features derived from SUVR prediction. Experiment results demonstrate the high prediction accuracy of the proposed method for FDG/AV45-PET SUVRs, where we achieved Pearson's correlation coefficients of 0.66 and 0.61 respectively between the estimated and actual SUVR and the estimated SUVRs also show high sensitivity and distinct longitudinal patterns for different disease status. By taking into account PET embedding features, the proposed method outperforms other competing methods on five independent datasets in the diagnosis of Alzheimer's disease and discriminating between stable and progressive mild cognitive impairments, achieving the area under receiver operating characteristic curves of 0.968 and 0.776 respectively on ADNI dataset, and generalizes better to other external datasets. Moreover, the top-weighted patches extracted from the trained model involve important brain regions related to Alzheimer's disease, suggesting good biological interpretability of our proposed method." Zi-Chao Zhang 0001, Xingzhong Zhao, Guiying Dong, Xing-Ming Zhao |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | MorbidGCN: prediction of multimorbidity with a graph convolutional network based on integration of population phenotypes and disease networkabstractExploring multimorbidity relationships among diseases is of great importance for understanding their shared mechanisms, precise diagnosis and treatment. However, the landscape of multimorbidities is still far from complete due to the complex nature of multimorbidity. Although various types of biological data, such as biomolecules and clinical symptoms, have been used to identify multimorbidities, the population phenotype information (e.g. physical activity and diet) remains less explored for multimorbidity. Here, we present a graph convolutional network (GCN) model, named MorbidGCN, for multimorbidity prediction by integrating population phenotypes and disease network. Specifically, MorbidGCN treats the multimorbidity prediction as a missing link prediction problem in the disease network, where a novel feature selection method is embedded to select important phenotypes. Benchmarking results on two large-scale multimorbidity data sets, i.e. the UK Biobank (UKB) and Human Disease Network (HuDiNe) data sets, demonstrate that MorbidGCN outperforms other competitive methods. With MorbidGCN, 9742 and 14 010 novel multimorbidities are identified in the UKB and HuDiNe data sets, respectively. Moreover, we notice that the selected phenotypes that are generally differentially distributed between multimorbidity patients and single-disease patients can help interpret multimorbidities and show potential for prognosis of multimorbidities. Guiying Dong, Zi-Chao Zhang 0001, Jianfeng Feng, Xing-Ming Zhao |
Briefings Bioinform. | 4 |
| 2022 | An integrated brain-specific network identifies genes associated with neuropathologic and clinical traits of Alzheimer's diseaseabstractAlzheimer's disease (AD) has a strong genetic predisposition. However, its risk genes remain incompletely identified. We developed an Alzheimer's brain gene network-based approach to predict AD-associated genes by leveraging the functional pattern of known AD-associated genes. Our constructed network outperformed existing networks in predicting AD genes. We then systematically validated the predictions using independent genetic, transcriptomic, proteomic data, neuropathological and clinical data. First, top-ranked genes were enriched in AD-associated pathways. Second, using external gene expression data from the Mount Sinai Brain Bank study, we found that the top-ranked genes were significantly associated with neuropathological and clinical traits, including the Consortium to Establish a Registry for Alzheimer's Disease score, Braak stage score and clinical dementia rating. The analysis of Alzheimer's brain single-cell RNA-seq data revealed cell-type-specific association of predicted genes with early pathology of AD. Third, by interrogating proteomic data in the Religious Orders Study and Memory and Aging Project and Baltimore Longitudinal Study of Aging studies, we observed a significant association of protein expression level with cognitive function and AD clinical severity. The network, method and predictions could become a valuable resource to advance the identification of risk genes for AD. Cui-Xiang Lin, Hong-Dong Li, Weisheng Liu, Shannon Erhardt, Fang-Xiang Wu, Xing-Ming Zhao, Yuanfang Guan, Jun Wang 0153, Daifeng Wang, Bin Hu 0001, Jianxin Wang 0001 |
Briefings Bioinform. | 7 |
| 2022 | CITEdb: a manually curated database of cell-cell interactions in humanabstractMOTIVATION: The interactions among various types of cells play critical roles in cell functions and the maintenance of the entire organism. While cell-cell interactions are traditionally revealed from experimental studies, recent developments in single-cell technologies combined with data mining methods have enabled computational prediction of cell-cell interactions, which have broadened our understanding of how cells work together, and have important implications in therapeutic interventions targeting cell-cell interactions for cancers and other diseases. Despite the importance, to our knowledge, there is no database for systematic documentation of high-quality cell-cell interactions at the cell type level, which hinders the development of computational approaches to identify cell-cell interactions. RESULTS: We develop a publicly accessible database, CITEdb (Cell-cell InTEraction database, https://citedb.cn/), which not only facilitates interactive exploration of cell-cell interactions in specific physiological contexts (e.g. a disease or an organ) but also provides a benchmark dataset to interpret and evaluate computationally derived cell-cell interactions from different tools. CITEdb contains 728 pairs of cell-cell interactions in human that are manually curated. Each interaction is equipped with structured annotations including the physiological context, the ligand-receptor pairs that mediate the interaction, etc. Our database provides a web interface to search, visualize and download cell-cell interactions. Users can search for cell-cell interactions by selecting the physiological context of interest or specific cell types involved. CITEdb is the first attempt to catalogue cell-cell interactions at the cell type level, which is beneficial to both experimental, computational and clinical studies of cell-cell interactions. AVAILABILITY AND IMPLEMENTATION: CITEdb is freely available at https://citedb.cn/ and the R package implementing benchmark is available at https://github.com/shanny01/benchmark. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Nayang Shan, Dongyu Li, Jitong Jiang, Linlin Yan, Jiudong Gao, Xing-Ming Zhao, Lin Hou 0003 |
Bioinform. | 9 |
| 2022 | Co-VAE: Drug-Target Binding Affinity Prediction by Co-Regularized Variational AutoencodersabstractIdentifying drug-target interactions has been a key step in drug discovery. Many computational methods have been proposed to directly determine whether drugs and targets can interact or not. Drug-target binding affinity is another type of data which could show the strength of the binding interaction between a drug and a target. However, it is more challenging to predict drug-target binding affinity, and thus a very few studies follow this line. In our work, we propose a novel co-regularized variational autoencoders (Co-VAE) to predict drug-target binding affinity based on drug structures and target sequences. The Co-VAE model consists of two VAEs for generating drug SMILES strings and target sequences, respectively, and a co-regularization part for generating the binding affinities. We theoretically prove that the Co-VAE model is to maximize the lower bound of the joint likelihood of drug, protein and their affinity. The Co-VAE could predict drug-target affinity and generate new drugs which share similar targets with the input drugs. The experimental results on two datasets show that the Co-VAE could predict drug-target affinity better than existing affinity prediction methods such as DeepDTA and DeepAffinity, and could generate more new valid drugs than existing methods such as GAN and VAE. Xing-Ming Zhao |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | nMAGMA: a network-enhanced method for inferring risk genes from GWAS summary statistics and its application to schizophreniaabstractMOTIVATION: Annotating genetic variants from summary statistics of genome-wide association studies (GWAS) is crucial for predicting risk genes of various disorders. The multimarker analysis of genomic annotation (MAGMA) is one of the most popular tools for this purpose, where MAGMA aggregates signals of single nucleotide polymorphisms (SNPs) to their nearby genes. In biology, SNPs may also affect genes that are far away in the genome, thus missed by MAGMA. Although different upgrades of MAGMA have been proposed to extend gene-wise variant annotations with more information (e.g. Hi-C or eQTL), the regulatory relationships among genes and the tissue specificity of signals have not been taken into account. RESULTS: We propose a new approach, namely network-enhanced MAGMA (nMAGMA), for gene-wise annotation of variants from GWAS summary statistics. Compared with MAGMA and H-MAGMA, nMAGMA significantly extends the lists of genes that can be annotated to SNPs by integrating local signals, long-range regulation signals (i.e. interactions between distal DNA elements), and tissue-specific gene networks. When applied to schizophrenia (SCZ), nMAGMA is able to detect more risk genes (217% more than MAGMA and 57% more than H-MAGMA) that are involved in SCZ compared with MAGMA and H-MAGMA, and more of nMAGMA results can be validated with known SCZ risk genes. Some disease-related functions (e.g. the ATPase pathway in Cortex) are also uncovered in nMAGMA but not in MAGMA or H-MAGMA. Moreover, nMAGMA provides tissue-specific risk signals, which are useful for understanding disorders with multitissue origins. Anyi Yang, Jingqi Chen, Xing-Ming Zhao |
Briefings Bioinform. | 3 |
| 2021 | Identifying age-specific gene signatures of the human cerebral cortex with joint analysis of transcriptomes and functional connectomesabstractThe human cerebral cortex undergoes profound structural and functional dynamic variations across the lifespan, whereas the underlying molecular mechanisms remain unclear. Here, with a novel method transcriptome-connectome correlation analysis (TCA), which integrates the brain functional magnetic resonance images and region-specific transcriptomes, we identify age-specific cortex (ASC) gene signatures for adolescence, early adulthood and late adulthood. The ASC gene signatures are significantly correlated with the cortical thickness (P-value <2.00e-3) and myelination (P-value <1.00e-3), two key brain structural features that vary in accordance with brain development. In addition to the molecular underpinning of age-related brain functions, the ASC gene signatures allow delineation of the molecular mechanisms of neuropsychiatric disorders, such as the regulation between ARNT2 and its target gene ETF1 involved in Schizophrenia. We further validate the ASC gene signatures with published gene sets associated with the adult cortex, and confirm the robustness of TCA on other brain image datasets. Availability: All scripts are written in R. Scripts for the TCA method and related statistics result can be freely accessed at https://github.com/Soulnature/TCA. Additional data related to this paper may be requested from the authors. Xingzhong Zhao, Jingqi Chen, Peipei Xiao, Jianfeng Feng, Qing Nie, Xing-Ming Zhao |
Briefings Bioinform. | 6 |
| 2021 | PhosIDN: an integrated deep neural network for improving protein phosphorylation site prediction by combining sequence and protein-protein interaction informationabstractMOTIVATION: Phosphorylation is one of the most studied post-translational modifications, which plays a pivotal role in various cellular processes. Recently, deep learning methods have achieved great success in prediction of phosphorylation sites, but most of them are based on convolutional neural network that may not capture enough information about long-range dependencies between residues in a protein sequence. In addition, existing deep learning methods only make use of sequence information for predicting phosphorylation sites, and it is highly desirable to develop a deep learning architecture that can combine heterogeneous sequence and protein-protein interaction (PPI) information for more accurate phosphorylation site prediction. RESULTS: We present a novel integrated deep neural network named PhosIDN, for phosphorylation site prediction by extracting and combining sequence and PPI information. In PhosIDN, a sequence feature encoding sub-network is proposed to capture not only local patterns but also long-range dependencies from protein sequences. Meanwhile, useful PPI features are also extracted in PhosIDN by a PPI feature encoding sub-network adopting a multi-layer deep neural network. Moreover, to effectively combine sequence and PPI information, a heterogeneous feature combination sub-network is introduced to fully exploit the complex associations between sequence and PPI features, and their combined features are used for final prediction. Comprehensive experiment results demonstrate that the proposed PhosIDN significantly improves the prediction performance of phosphorylation sites and compares favorably with existing general and kinase-specific phosphorylation site prediction methods. AVAILABILITY AND IMPLEMENTATION: PhosIDN is freely available at https://github.com/ustchangyuanyang/PhosIDN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hangyuan Yang, Xing-Ming Zhao, Ao Li 0001 |
Bioinform. | 4 |
| 2021 | MAT2: manifold alignment of single-cell transcriptomes with cell tripletsabstractMOTIVATION: Aligning single-cell transcriptomes is important for the joint analysis of multiple single-cell RNA sequencing datasets, which in turn is vital to establishing a holistic cellular landscape of certain biological processes. Although numbers of approaches have been proposed for this problem, most of which only consider mutual neighbors when aligning the cells without taking into account known cell type annotations. RESULTS: In this work, we present MAT2 that aligns cells in the manifold space with a deep neural network employing contrastive learning strategy. Compared with other manifold-based approaches, MAT2 has two-fold advantages. Firstly, with cell triplets defined based on known cell type annotations, the consensus manifold yielded by the alignment procedure is more robust especially for datasets with limited common cell types. Secondly, the batch-effect-free gene expression reconstructed by MAT2 can better help annotate cell types. Benchmarking results on real scRNA-seq datasets demonstrate that MAT2 outperforms existing popular methods. Moreover, with MAT2, the hematopoietic stem cells are found to differentiate at different paces between human and mouse. AVAILABILITY AND IMPLEMENTATION: MAT2 is publicly available at https://github.com/Zhang-Jinglong/MAT2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ying Wang 0005, Xing-Ming Zhao |
Bioinform. | 5 |
| 2021 | Identifying Molecular Biomarkers for Diseases With Machine Learning Based on Integrative OmicsabstractMolecular biomarkers are certain molecules or set of molecules that can be of help for diagnosis or prognosis of diseases or disorders. In the past decades, thanks to the advances in high-throughput technologies, a huge amount of molecular 'omics' data, e.g., transcriptomics and proteomics, have been accumulated. The availability of these omics data makes it possible to screen biomarkers for diseases or disorders. Accordingly, a number of computational approaches have been developed to identify biomarkers by exploring the omics data. In this review, we present a comprehensive survey on the recent progress of identification of molecular biomarkers with machine learning approaches. Specifically, we categorize the machine learning approaches into supervised, un-supervised and recommendation approaches, where the biomarkers including single genes, gene sets and small gene networks. In addition, we further discuss potential problems underlying bio-medical data that may pose challenges for machine learning, and provide possible directions for future biomarker identification. Kai Shi 0004, Wei Lin 0003, Xing-Ming Zhao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2020 | scTSSR: gene expression recovery for single-cell RNA sequencing using two-side sparse self-representationabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) methods make it possible to reveal gene expression patterns at single-cell resolution. Due to technical defects, dropout events in scRNA-seq will add noise to the gene-cell expression matrix and hinder downstream analysis. Therefore, it is important for recovering the true gene expression levels before carrying out downstream analysis. RESULTS: In this article, we develop an imputation method, called scTSSR, to recover gene expression for scRNA-seq. Unlike most existing methods that impute dropout events by borrowing information across only genes or cells, scTSSR simultaneously leverages information from both similar genes and similar cells using a two-side sparse self-representation model. We demonstrate that scTSSR can effectively capture the Gini coefficients of genes and gene-to-gene correlations observed in single-molecule RNA fluorescence in situ hybridization (smRNA FISH). Down-sampling experiments indicate that scTSSR performs better than existing methods in recovering the true gene expression levels. We also show that scTSSR has a competitive performance in differential expression analysis, cell clustering and cell trajectory inference. AVAILABILITY AND IMPLEMENTATION: The R package is available at https://github.com/Zhangxf-ccnu/scTSSR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Le Ou-Yang, Xing-Ming Zhao, Hong Yan 0001, Xiao-Fei Zhang |
Bioinform. | 3 |
| 2020 | A graph regularized generalized matrix factorization model for predicting links in biomedical bipartite networksabstractMOTIVATION: Predicting potential links in biomedical bipartite networks can provide useful insights into the diagnosis and treatment of complex diseases and the discovery of novel drug targets. Computational methods have been proposed recently to predict potential links for various biomedical bipartite networks. However, existing methods are usually rely on the coverage of known links, which may encounter difficulties when dealing with new nodes without any known link information. RESULTS: In this study, we propose a new link prediction method, named graph regularized generalized matrix factorization (GRGMF), to identify potential links in biomedical bipartite networks. First, we formulate a generalized matrix factorization model to exploit the latent patterns behind observed links. In particular, it can take into account the neighborhood information of each node when learning the latent representation for each node, and the neighborhood information of each node can be learned adaptively. Second, we introduce two graph regularization terms to draw support from affinity information of each node derived from external databases to enhance the learning of latent representations. We conduct extensive experiments on six real datasets. Experiment results show that GRGMF can achieve competitive performance on all these datasets, which demonstrate the effectiveness of GRGMF in prediction potential links in biomedical bipartite networks. AVAILABILITY AND IMPLEMENTATION: The package is available at https://github.com/happyalfred2016/GRGMF. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zi-Chao Zhang 0001, Xiao-Fei Zhang, Min Wu 0008, Le Ou-Yang, Xing-Ming Zhao, Xiaoli Li 0001 |
Bioinform. | 5 |
| 2020 | Deep convolutional neural network for accurate segmentation and quantification of white matter hyperintensities
Liangliang Liu 0001, Shaowu Chen, Xiaofeng Zhu 0001, Xing-Ming Zhao, Fang-Xiang Wu, Jianxin Wang 0001 |
Neurocomputing | 4 |
| 2019 | DeepPhos: prediction of protein phosphorylation sites with deep learningabstractMOTIVATION: Phosphorylation is the most studied post-translational modification, which is crucial for multiple biological processes. Recently, many efforts have been taken to develop computational predictors for phosphorylation site prediction, but most of them are based on feature selection and discriminative classification. Thus, it is useful to develop a novel and highly accurate predictor that can unveil intricate patterns automatically for protein phosphorylation sites. RESULTS: In this study we present DeepPhos, a novel deep learning architecture for prediction of protein phosphorylation. Unlike multi-layer convolutional neural networks, DeepPhos consists of densely connected convolutional neuron network blocks which can capture multiple representations of sequences to make final phosphorylation prediction by intra block concatenation layers and inter block concatenation layers. DeepPhos can also be used for kinase-specific prediction varying from group, family, subfamily and individual kinase level. The experimental results demonstrated that DeepPhos outperforms competitive predictors in general and kinase-specific phosphorylation site prediction. AVAILABILITY AND IMPLEMENTATION: The source code of DeepPhos is publicly deposited at https://github.com/USTCHIlab/DeepPhos. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Fenglin Luo, Yu Liu 0113, Xing-Ming Zhao, Ao Li 0001 |
Bioinform. | 4 |
| 2019 | EnImpute: imputing dropout events in single-cell RNA-sequencing data via ensemble learningabstractSUMMARY: Imputation of dropout events that may mislead downstream analyses is a key step in analyzing single-cell RNA-sequencing (scRNA-seq) data. We develop EnImpute, an R package that introduces an ensemble learning method for imputing dropout events in scRNA-seq data. EnImpute combines the results obtained from multiple imputation methods to generate a more accurate result. A Shiny application is developed to provide easier implementation and visualization. Experiment results show that EnImpute outperforms the individual state-of-the-art methods in almost all situations. EnImpute is useful for correcting the noisy scRNA-seq data before performing downstream analysis. AVAILABILITY AND IMPLEMENTATION: The R package and Shiny application are available through Github at https://github.com/Zhangxf-ccnu/EnImpute. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiao-Fei Zhang, Le Ou-Yang, Xing-Ming Zhao, Xiaohua Hu 0001, Hong Yan 0001 |
Bioinform. | 4 |
| 2019 | DrPOCS: Drug Repositioning Based on Projection Onto Convex SetsabstractDrug repositioning, i.e., identifying new indications for known drugs, has attracted a lot of attentions recently and is becoming an effective strategy in drug development. In literature, several computational approaches have been proposed to identify potential indications of old drugs based on various types of data sources. In this paper, by formulating the drug-disease associations as a low-rank matrix, we propose a novel method, namely DrPOCS, to identify candidate indications of old drugs based on projection onto convex sets (POCS). With the integration of drug structure and disease phenotype information, DrPOCS predicts potential associations between drugs and diseases with matrix completion. Benchmarking results demonstrate that our proposed approach outperforms popular existing approaches with high accuracy. In addition, a number of novel predicted indications are validated with various types of evidences, indicating the predictive power of our proposed approach. Yin-Ying Wang, Chunfeng Cui, Liqun Qi 0001, Hong Yan 0001, Xing-Ming Zhao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2019 | EmDL: Extracting miRNA-Drug Interactions from LiteratureabstractThe microRNAs (miRNAs), regulators of post-transcriptional processes, have been found to affect the efficacy of drugs by regulating the biological processes in which the target proteins of drugs may be involved. For example, some drugs develop resistance when certain miRNAs are overexpressed. Therefore, identifying miRNAs that affect drug effects can help understand the mechanisms of drug actions and design more efficient drugs. Although some computational approaches have been developed to predict miRNA-drug associations, such associations rarely provide explicit information about which miRNAs and how they affect drug efficacy. On the other hand, there are rich information about which miRNAs affect the efficacy of which drugs in the literature. In this paper, we present a novel text mining approach, named as EmDL (Extracting miRNA-Drug interactions from Literature), to extract the relationships of miRNAs affecting drug efficacy from literature. Benchmarking on the drug-miRNA interactions manually extracted from MEDLINE and PubMed Central, EmDL outperforms traditional text mining approaches as well as other popular methods for predicting drug-miRNA associations. Specifically, EmDL can effectively identify the sentences that describe the relationships of miRNAs affecting drug effects. The drug-miRNA interactome presented here can help understand how miRNAs affect drug effects and provide insights into the mechanisms of drug actions. In addition, with the information about drug-miRNA interactions, more effective drugs or combinatorial strategies can be designed in the future. The data used here can be accessed at http://mtd.comp-sysbio.org/. Wen-Bin Xie, Hong Yan 0001, Xing-Ming Zhao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2019 | Joint Learning of Multiple Differential Networks With Latent VariablesabstractGraphical models have been widely used to learn the conditional dependence structures among random variables. In many controlled experiments, such as the studies of disease or drug effectiveness, learning the structural changes of graphical models under two different conditions is of great importance. However, most existing graphical models are developed for estimating a single graph and based on a tacit assumption that there is no missing relevant variables, which wastes the common information provided by multiple heterogeneous data sets and underestimates the influence of latent/unobserved relevant variables. In this paper, we propose a joint differential network analysis (JDNA) model to jointly estimate multiple differential networks with latent variables from multiple data sets. The JDNA model is built on a penalized D-trace loss function, with group lasso or generalized fused lasso penalties. We implement a proximal gradient-based alternating direction method of multipliers to tackle the corresponding convex optimization problems. Extensive simulation experiments demonstrate that JDNA model outperforms state-of-the-art methods in estimating the structural changes of graphical models. Moreover, a series of experiments on several real-world data sets have been performed and experiment results consistently show that our proposed JDNA model is effective in identifying differential networks under different conditions. Le Ou-Yang, Xiao-Fei Zhang, Xing-Ming Zhao, Debby Dan Wang, Fu Lee Wang, Bai Ying Lei, Hong Yan 0001 |
IEEE Trans. Cybern. | 3 |
| 2018 | Prediction of Drug Response with a Topology Based Dual-Layer Network Model
Suyun Huang, Xing-Ming Zhao |
ISBRA | 2 |
| 2017 | PhosD: inferring kinase-substrate interactions based on protein domainsabstractMOTIVATION: Identifying the kinase-substrate relationships is vital to understanding the phosphorylation events and various biological processes, especially signal transductions. Although large amount of phosphorylation sites have been detected, unfortunately, it is rarely known which kinases activate those sites. Despite distinct computational approaches have been proposed to predict the kinase-substrate interactions, the prediction accuracy still needs to be improved. RESULTS: In this paper, we propose a novel probabilistic model named as PhosD to predict kinase-substrate relationships based on protein domains with the assumption that kinase-substrate interactions are accomplished with kinase-domain interactions. By further taking into account protein-protein interactions, our PhosD outperforms other popular approaches on several benchmark datasets with higher precision. In addition, some of our predicted kinase-substrate relationships are validated by signaling pathways, indicating the predictive power of our approach. Furthermore, we notice that given a kinase, the more substrates are known for the kinase the more accurate its predicted substrates will be, and the domains involved in kinase-substrate interactions are found to be more conserved across proteins phosphorylated by multiple kinases. These findings can help develop more efficient computational approaches in the future. AVAILABILITY AND IMPLEMENTATION: The data and results are available at http://comp-sysbio.org/phosd. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gui-Min Qin, Rui-Yi Li, Xing-Ming Zhao |
Bioinform. | 3 |
| 2017 | PCID: A Novel Approach for Predicting Disease Comorbidity by Integrating Multi-Scale DataabstractDisease comorbidity is the presence of one or more diseases along with a primary disorder, which causes additional pain to patients and leads to the failure of standard treatments compared with single diseases. Therefore, the identification of potential comorbidity can help prevent those comorbid diseases when treating a primary disease. Unfortunately, most of current known disease comorbidities are discovered occasionally in clinic, and our knowledge about comorbidity is far from complete. Despite the fact that many efforts have been made to predict disease comorbidity, the prediction accuracy of existing computational approaches needs to be improved. By investigating the factors underlying disease comorbidity, e.g., mutated genes and rewired protein-protein interactions (PPIs), we here present a novel algorithm to predict disease comorbidity by integrating multi-scale data ranging from genes to phenotypes. Benchmark results on real data show that our approach outperforms existing algorithms, and some of our novel predictions are validated with those reported in literature, indicating the effectiveness and predictive power of our approach. In addition, we identify some pathway and PPI patterns that underlie the co-occurrence between a primary disease and certain disease classes, which can help explain how the comorbidity is initiated from molecular perspectives. Feng He 0004, Yin-Ying Wang, Xing-Ming Zhao, De-Shuang Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2016 | The exploration of functional divergence between human and macaque brains based on gene networksabstractAlthough human and macaque share a lot of genes, their brains have different functions. In this study, we compare human and macaque brains from a systematic perspective by assuming that the difference between human and macaque brains is due to the difference of molecular circuits underlying brains instead of certain individual genes. Specially, we construct a gene interaction network for each brain region and investigate the difference between human and macaque brains based on those region-specific gene networks. We found that the genes in human brain networks tend to interact with each other compared with those of macaque brain networks. In addition, the human genes are more likely to form functional modules compared against macaque genes. Moreover, the human specific genes tend to interact with hub genes, while both the hub genes and human specific genes are enriched in brain functions. These findings can help better understand the functional divergence between human and macaque brains. Peipei Xiao, Xing-Ming Zhao |
BIBM | 2 |
| 2016 | A hybrid sequential feature selection approach for the diagnosis of Alzheimer's DiseaseabstractThe timely and accurate diagnosis of Alzheimer's Disease (AD) is important for preventing the progress of the irreversible disease. Recently, various types of imaging techniques, e.g. Magnetic Resonance Imaging (MRI) and Positron Emission Tomography (PET), have been widely used for the diagnosis of AD. Despite their usefulness, the image data generated are high-dimensional and noisy which make it difficult to give accurate diagnosis. In this paper, we propose a novel feature selection approach to detect the informative features from MRI data, which improves both the diagnosis accuracy and reduces the computational cost. Benchmarking results on the AD challenge datasets show that our proposed approach outperforms other popular feature selection methods. Furthermore, compared with the top winners of the challenge, our approach performs best in diagnosis of AD and has comparable performance with others in the detection of mild cognitive impairments (MCIs). The good performance on the real data demonstrate that our proposed approach is effective in prediction of ADs. Xing-Ming Zhao |
IJCNN | 2 |
| 2016 | Understanding tissue-specificity with human tissue-specific regulatory networks
Wei-Li Guo, Lin Zhu 0008, Suping Deng, Xing-Ming Zhao, De-Shuang Huang |
Sci. China Inf. Sci. | 4 |
| 2016 | Data mining in systems biology
Ao Li 0001, Xing-Ming Zhao, Shuigeng Zhou |
Neurocomputing | 2 |
| 2016 | A systematic exploration of the associations between amino acid variants and post-translational modifications
Gui-Min Qin, Yi-Bo Hou, Xing-Ming Zhao |
Neurocomputing | 3 |
| 2016 | Identifying Disease Associated miRNAs Based on Protein DomainsabstractMicroRNAs (miRNAs) are a class of small endogenous non-coding genes, acting as regulators in the post-transcriptional processes. Recently, the miRNAs are found to be widely involved in different types of diseases. Therefore, the identification of disease associated miRNAs can help understand the mechanisms that underlie the disease and identify new biomarkers. However, it is not easy to identify the miRNAs related to diseases due to its extensive involvements in various biological processes. In this work, we present a new approach to identify disease associated miRNAs based on domains, the functional and structural blocks of proteins. The results on real datasets demonstrate that our method can effectively identify disease related miRNAs with high precision. Gui-Min Qin, Rui-Yi Li, Xing-Ming Zhao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2016 | Data Mining in Systems BiologyabstractThe papers in this special issue were presented at the Chinese Control Conference (CCC2015), which was successfully held in Hangzhou, China, during July 28 to 30, 2015. The conference has provided an opportunity for scientists and researchers from different backgrounds to present their latest works on data mining in systems biology. Xing-Ming Zhao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2015 | jNMFMA: a joint non-negative matrix factorization meta-analysis of transcriptomics dataabstractMOTIVATION: Tremendous amount of omics data being accumulated poses a pressing challenge of meta-analyzing the heterogeneous data for mining new biological knowledge. Most existing methods deal with each gene independently, thus often resulting in high false positive rates in detecting differentially expressed genes (DEG). To our knowledge, no or little effort has been devoted to methods that consider dependence structures underlying transcriptomics data for DEG identification in meta-analysis context. RESULTS: This article proposes a new meta-analysis method for identification of DEGs based on joint non-negative matrix factorization (jNMFMA). We mathematically extend non-negative matrix factorization (NMF) to a joint version (jNMF), which is used to simultaneously decompose multiple transcriptomics data matrices into one common submatrix plus multiple individual submatrices. By the jNMF, the dependence structures underlying transcriptomics data can be interrogated and utilized, while the high-dimensional transcriptomics data are mapped into a low-dimensional space spanned by metagenes that represent hidden biological signals. jNMFMA finally identifies DEGs as genes that are associated with differentially expressed metagenes. The ability of extracting dependence structures makes jNMFMA more efficient and robust to identify DEGs in meta-analysis context. Furthermore, jNMFMA is also flexible to identify DEGs that are consistent among various types of omics data, e.g. gene expression and DNA methylation. Experimental results on both simulation data and real-world cancer data demonstrate the effectiveness of jNMFMA and its superior performance over other popular approaches. AVAILABILITY AND IMPLEMENTATION: R code for jNMFMA is available for non-commercial use via http://micblab.iim.ac.cn/Download/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hong-Qiang Wang, Chun-Hou Zheng 0001, Xing-Ming Zhao |
Bioinform. | 3 |
| 2015 | Identifying cancer-related microRNAs based on gene expression dataabstractMOTIVATION: MicroRNAs (miRNAs) are short non-coding RNAs that play important roles in post-transcriptional regulations as well as other important biological processes. Recently, accumulating evidences indicate that miRNAs are extensively involved in cancer. However, it is a big challenge to identify which miRNAs are related to which cancer considering the complex processes involved in tumors, where one miRNA may target hundreds or even thousands of genes and one gene may regulate multiple miRNAs. Despite integrative analysis of matched gene and miRNA expression data can help identify cancer-associated miRNAs, such kind of data is not commonly available. On the other hand, there are huge amount of gene expression data that are publicly accessible. It will significantly improve the efficiency of characterizing miRNA's function in cancer if we can identify cancer miRNAs directly from gene expression data. RESULTS: We present a novel computational framework to identify the cancer-related miRNAs based solely on gene expression profiles without requiring either miRNA expression data or the matched gene and miRNA expression data. The results on multiple cancer datasets show that our proposed method can effectively identify cancer-related miRNAs with higher precision compared with other popular approaches. Furthermore, some of our novel predictions are validated by both differentially expressed miRNAs and evidences from literature, implying the predictive power of our proposed method. In addition, we construct a cancer-miRNA-pathway network, which can help explain how miRNAs are involved in cancer. AVAILABILITY AND IMPLEMENTATION: The R code and data files for the proposed method are available at http://comp-sysbio.org/miR_Path/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: supplementary data are available at Bioinformatics online. Xing-Ming Zhao, Keqin Liu, Feng He 0004, Béatrice Duval, Jean-Michel Richer, De-Shuang Huang, Jin-Kao Hao, Luonan Chen |
Bioinform. | 1 |
| 2014 | Cascleave 2.0, a new approach for predicting caspase and granzyme cleavage targetsabstractMOTIVATION: Caspases and granzyme B (GrB) are important proteases involved in fundamental cellular processes and play essential roles in programmed cell death, necrosis and inflammation. Although a number of substrates for both types have been experimentally identified, the complete repertoire of caspases and granzyme B substrates remained to be fully characterized. Accordingly, systematic bioinformatics studies of known cleavage sites may provide important insights into their substrate specificity and facilitate the discovery of novel substrates. RESULTS: We develop a new bioinformatics tool, termed Cascleave 2.0, which builds on previous success of the Cascleave tool for predicting generic caspase cleavage sites. It can be efficiently used to predict potential caspase-specific cleavage sites for the human caspase-1, 3, 6, 7, 8 and GrB. In particular, we integrate heterogeneous sequence and protein functional information from various sources to improve the prediction accuracy of Cascleave 2.0. During classification, we use both maximum relevance minimum redundancy and forward feature selection techniques to quantify the relative contribution of each feature to prediction and thus remove redundant as well as irrelevant features. A systematic evaluation of Cascleave 2.0 using the benchmark data and comparison with other state-of-the-art tools using independent test data indicate that Cascleave 2.0 outperforms other tools on protease-specific cleavage site prediction of caspase-1, 3, 6, 7 and GrB. Cascleave 2.0 is anticipated to be used as a powerful tool for identifying novel substrates and cleavage sites of caspases and GrB and help understand the functional roles of these important proteases in human proteolytic cascades. AVAILABILITY AND IMPLEMENTATION: http://www.structbioinfor.org/cascleave2/. Xing-Ming Zhao, Tatsuya Akutsu, James C. Whisstock, Jiangning Song |
Bioinform. | 2 |
| 2014 | Pattern recognition in bioinformatics
Xing-Ming Zhao, Alioune Ngom, Jin-Kao Hao |
Neurocomputing | 1 |
| 2014 | Comments on "Human Dominant Disease Genes Are Enriched in Paralogs Originating from Whole Genome Duplication"abstractWe previously showed that monogenic disease genes (MDs) are enriched in duplicates and hypothesized that functional redundancy among duplicates underlies this enrichment [1]. In their comment, Singh et al. refine this enrichment to genes resulting from whole genome duplications (WGDs) [2]; they, furthermore, “could not find any significant enrichment in duplicates in support of possible functional compensation for essential genes” [2] by using gene essentiality data from mouse (transferred to human through orthology).
We appreciate the scientific argument, but we would like to point out that confounding factors and data biases can lead to seemingly opposing conclusions. For example, we carefully considered the duplication age of genes, which is a known confounder in such analyses [3], [4], as well as the use of gene subsets that have known biases such as the mouse essentiality data [3], which, in addition, have issues when conclusions are being transferred to human genes.
First, when using the data of Singh et al. [2] and stratifying small-scale duplicates (SSDs) into old and young groups according to the duplication age relative to WGD, we found that MDs are enriched in old SSDs; limiting this analysis to recessive MDs produced similar results (Figure 1A). In contrast, MDs are depleted in young SSDs (Figure 1B), which is consistent with our hypothesis and with our findings that coexpression decreases with increased duplication age. Thus, when the duplication is old, the ability of the functional copy to compensate for the mutation-carrying malfunctioning copy could be easily disrupted because of random fluctuation in gene expression in a subpopulation; consequently, the gene is associated with a disease, but it will not be purged from the whole population. Therefore, functional compensation can promote the spreading of disease genes in duplicates. However, in young duplicates, the fluctuation in gene expression among duplicates may not be that huge; thus, deleterious mutations could be tolerated, and the corresponding genes are unlikely to associate with any diseases.
Figure 1
Enrichment of MDs in old SSDs and distinct characteristics of the old SSDs as compared with the young ones.
Second, mouse essentiality data are biased [5], e.g., towards developmental genes; i.e., they do not correspond to the full spectrum of MDs. Dividing the tested mouse genes into subgroups, the proportion of essential genes in young SSDs is significantly lower than that of singletons (Figure 1C), consistent with functional redundancy among duplicates; however, the opposite is found in old SSDs (Figure 1D). The latter has led to the somewhat counterintuitive conclusion that “duplicates are as essential as singletons” [6], which has been argued against by several follow-up studies [3]–[5]. These results, again, highlight the importance of taking duplication age into consideration. As previous studies suggested, it is not trivial to correct the biases [3]–[5], and hence, conclusions from this data regarding duplications have to be taken with caution. Furthermore, the essentiality status of mouse genes cannot be reliably transferred to human and vice versa. For example, using data from OGEE [7], an online gene essentiality database, 2,322 mouse essential genes have one-to-one orthologs in human; only 476 out of the 2,322 human genes (approximately 20%) were essential according to a genome-wide small interfering RNA (siRNA) experiment [8].
Finally, only less than 30% of the MDs we collected [1] were used in the analyses by Singh et al.; the intersection with the essentiality dataset is even smaller (approximately 18.6% of the MDs used in [1]) because, so far, only less than one-third (approximately 6,400) of mouse genes has been tested for essentiality [9]. Thus, extrapolating any observations on these data to the whole genome would be difficult; for example, some functional signals might only become statistically significant in larger datasets.
Elucidating the molecular basis of human genetic disorders is one of the most important tasks in medical biology. With the relevant data, such as those from genome-wide association studies (GWAS), accumulated at an astonishing speed, integrative and comparative analyses through bioinformatics are much needed. In this regard, Singh et al. did provide an important contribution by refining the enrichment of dominant MDs in duplicates to those derived from WGD. However, we don't believe that they nullified our functional compensation hypothesis with the analyses performed, but they certainly encouraged further studies on more complete datasets, hopefully to be available in the near future. Wei-Hua Chen, Xing-Ming Zhao, Vera van Noort, Peer Bork |
PLoS Comput. Biol. | 2 |
| 2013 | NARROMI: a noise and redundancy reduction technique improves accuracy of gene regulatory network inferenceabstractMOTIVATION: Reconstruction of gene regulatory networks (GRNs) is of utmost interest to biologists and is vital for understanding the complex regulatory mechanisms within the cell. Despite various methods developed for reconstruction of GRNs from gene expression profiles, they are notorious for high false positive rate owing to the noise inherited in the data, especially for the dataset with a large number of genes but a small number of samples. RESULTS: In this work, we present a novel method, namely NARROMI, to improve the accuracy of GRN inference by combining ordinary differential equation-based recursive optimization (RO) and information theory-based mutual information (MI). In the proposed algorithm, the noisy regulations with low pairwise correlations are first removed by using MI, and the redundant regulations from indirect regulators are further excluded by RO to improve the accuracy of inferred GRNs. In particular, the RO step can help to determine regulatory directions without prior knowledge of regulators. The results on benchmark datasets from Dialogue for Reverse Engineering Assessments and Methods challenge and experimentally determined GRN of Escherichia coli show that NARROMI significantly outperforms other popular methods in terms of false positive rates and accuracy. AVAILABILITY: All the source data and code are available at: http://csb.shu.edu.cn/narromi.htm. Keqin Liu, Zhi-Ping Liu, Béatrice Duval, Jean-Michel Richer, Xing-Ming Zhao, Jin-Kao Hao, Luonan Chen |
Bioinform. | 6 |
| 2013 | Human Monogenic Disease Genes Have Frequently Functionally Redundant ParalogsabstractMendelian disorders are often caused by mutations in genes that are not lethal but induce functional distortions leading to diseases. Here we study the extent of gene duplicates that might compensate genes causing monogenic diseases. We provide evidence for pervasive functional redundancy of human monogenic disease genes (MDs) by duplicates by manifesting 1) genes involved in human genetic disorders are enriched in duplicates and 2) duplicated disease genes tend to have higher functional similarities with their closest paralogs in contrast to duplicated non-disease genes of similar age. We propose that functional compensation by duplication of genes masks the phenotypic effects of deleterious mutations and reduces the probability of purging the defective genes from the human population; this functional compensation could be further enhanced by higher purification selection between disease genes and their duplicates as well as their orthologous counterpart compared to non-disease genes. However, due to the intrinsic expression stochasticity among individuals, the deleterious mutations could still be present as genetic diseases in some subpopulations where the duplicate copies are expressed at low abundances. Consequently the defective genes are linked to genetic disorders while they continue propagating within the population. Our results provide insight into the molecular basis underlying the spreading of duplicated disease genes. Wei-Hua Chen, Xing-Ming Zhao, Vera van Noort, Peer Bork |
PLoS Comput. Biol. | 2 |
| 2012 | Drug-target network in myocardial infarction: A structural analysisabstractThe identification of drug-target interactions is a crucial step in the drug-discovery process. It has been suggested that drug-target interactions are driven by drug-domain interactions. Based on the integration of two recently published datasets, i.e., Drug-target interactions in myocardial infarction (My-DTome) and drug-domain interaction network, this paper reports the association between drugs and protein domains in the context of myocardial infarction (MI). A MI drug-domain interaction network, My-DDome, was constructed. The functional similarity between domains based on their Gene Ontology (GO) annotations was estimated. The association between domains and therapeutic effects was investigated. Lists of GO annotations and Anatomical Therapeutic Chemical classification (ATC) codes highly enriched in My-DDome were identified. We show that drugs acting on blood and blood forming organs (ATC code B) and sensory organs (ATC code S) are significantly enriched in My-DDome (p <; 0.000001). Top enriched GO terms include GO:0003824 (catalytic activity), GO:0008152 (metabolic process) and GO:0030170 (pyridoxal phosphate binding). By incorporating protein domain information into My-DTome, more detailed insights into the interplay between drugs, their known targets and seemingly unrelated proteins are provided. Haiying Wang 0001, Huiru Zheng, Francisco Azuaje, Xing-Ming Zhao |
BIBM | 4 |
| 2012 | Inferring gene regulatory networks from gene expression data by path consistency algorithm based on conditional mutual informationabstractMOTIVATION: Reconstruction of gene regulatory networks (GRNs), which explicitly represent the causality of developmental or regulatory process, is of utmost interest and has become a challenging computational problem for understanding the complex regulatory mechanisms in cellular systems. However, all existing methods of inferring GRNs from gene expression profiles have their strengths and weaknesses. In particular, many properties of GRNs, such as topology sparseness and non-linear dependence, are generally in regulation mechanism but seldom are taken into account simultaneously in one computational method. RESULTS: In this work, we present a novel method for inferring GRNs from gene expression data considering the non-linear dependence and topological structure of GRNs by employing path consistency algorithm (PCA) based on conditional mutual information (CMI). In this algorithm, the conditional dependence between a pair of genes is represented by the CMI between them. With the general hypothesis of Gaussian distribution underlying gene expression data, CMI between a pair of genes is computed by a concise formula involving the covariance matrices of the related gene expression profiles. The method is validated on the benchmark GRNs from the DREAM challenge and the widely used SOS DNA repair network in Escherichia coli. The cross-validation results confirmed the effectiveness of our method (PCA-CMI), which outperforms significantly other previous methods. Besides its high accuracy, our method is able to distinguish direct (or causal) interactions from indirect associations. AVAILABILITY: All the source data and code are available at: http://csb.shu.edu.cn/subweb/grn.htm. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xing-Ming Zhao, Kun He 0007, Le Lu 0001, Yongwei Cao, Jingdong Liu, Jin-Kao Hao, Zhi-Ping Liu, Luonan Chen |
Bioinform. | 2 |
| 2012 | Identifying dysregulated pathways in cancers from pathway interaction networksabstractBACKGROUND: Cancers, a group of multifactorial complex diseases, are generally caused by mutation of multiple genes or dysregulation of pathways. Identifying biomarkers that can characterize cancers would help to understand and diagnose cancers. Traditional computational methods that detect genes differentially expressed between cancer and normal samples fail to work due to small sample size and independent assumption among genes. On the other hand, genes work in concert to perform their functions. Therefore, it is expected that dysregulated pathways will serve as better biomarkers compared with single genes. RESULTS: In this paper, we propose a novel approach to identify dysregulated pathways in cancer based on a pathway interaction network. Our contribution is three-fold. Firstly, we present a new method to construct pathway interaction network based on gene expression, protein-protein interactions and cellular pathways. Secondly, the identification of dysregulated pathways in cancer is treated as a feature selection problem, which is biologically reasonable and easy to interpret. Thirdly, the dysregulated pathways are identified as subnetworks from the pathway interaction networks, where the subnetworks characterize very well the functional dependency or crosstalk between pathways. The benchmarking results on several distinct cancer datasets demonstrate that our method can obtain more reliable and accurate results compared with existing state of the art methods. Further functional analysis and independent literature evidence also confirm that our identified potential pathogenic pathways are biologically reasonable, indicating the effectiveness of our method. CONCLUSIONS: Dysregulated pathways can serve as better biomarkers compared with single genes. In this work, by utilizing pathway interaction networks and gene expression data, we propose a novel approach that effectively identifies dysregulated pathways, which can not only be used as biomarkers to diagnose cancers but also serve as potential drug targets in the future. Keqin Liu, Zhi-Ping Liu, Jin-Kao Hao, Luonan Chen, Xing-Ming Zhao |
BMC Bioinform. | 5 |
| 2012 | Exploring drug combinations in genetic interaction networkabstractBACKGROUND: Drug combination that consists of distinctive agents is an attractive strategy to combat complex diseases and has been widely used clinically with improved therapeutic effects. However, the identification of efficacious drug combinations remains a non-trivial and challenging task due to the huge number of possible combinations among the candidate drugs. As an important factor, the molecular context in which drugs exert their functions can provide crucial insights into the mechanism underlying drug combinations. RESULTS: In this work, we present a network biology approach to investigate drug combinations and their target proteins in the context of genetic interaction networks and the related human pathways, in order to better understand the underlying rules of effective drug combinations. Our results indicate that combinatorial drugs tend to have a smaller effect radius in the genetic interaction networks, which is an important parameter to describe the therapeutic effect of a drug combination from the network perspective. We also find that drug combinations are more likely to modulate functionally related pathways. CONCLUSIONS: This study confirms that the molecular networks where drug combinations exert their functions can indeed provide important insights into the underlying rules of effective drug combinations. We hope that our findings can help shortcut the expedition of the future discovery of novel drug combinations. Yin-Ying Wang, Ke-Jia Xu, Jiangning Song, Xing-Ming Zhao |
BMC Bioinform. | 4 |
| 2012 | Identifying disease genes and module biomarkers by differential interactionsabstractOBJECTIVE: A complex disease is generally caused by the mutation of multiple genes or by the dysfunction of multiple biological processes. Systematic identification of causal disease genes and module biomarkers can provide insights into the mechanisms underlying complex diseases, and help develop efficient therapies or effective drugs. MATERIALS AND METHODS: In this paper, we present a novel approach to predict disease genes and identify dysfunctional networks or modules, based on the analysis of differential interactions between disease and control samples, in contrast to the analysis of differential gene or protein expressions widely adopted in existing methods. RESULTS AND DISCUSSION: As an example, we applied our method to the study of three-stage microarray data for gastric cancer. We identified network modules or module biomarkers that include a set of genes related to gastric cancer, implying the predictive power of our method. The results on holdout validation data sets show that our identified module can serve as an effective module biomarker for accurately detecting or diagnosing gastric cancer, thereby validating the efficiency of our method. CONCLUSION: We proposed a new approach to detect module biomarkers for diseases, and the results on gastric cancer demonstrated that the differential interactions are useful to detect dysfunctional modules in the molecular interaction network, which in turn can be used as robust module biomarkers. Xiaoping Liu 0002, Zhi-Ping Liu, Xing-Ming Zhao, Luonan Chen |
J. Am. Medical Informatics Assoc. | 3 |
| 2011 | Prediction of Drug Combinations by Integrating Molecular and Pharmacological DataabstractCombinatorial therapy is a promising strategy for combating complex disorders due to improved efficacy and reduced side effects. However, screening new drug combinations exhaustively is impractical considering all possible combinations between drugs. Here, we present a novel computational approach to predict drug combinations by integrating molecular and pharmacological data. Specifically, drugs are represented by a set of their properties, such as their targets or indications. By integrating several of these features, we show that feature patterns enriched in approved drug combinations are not only predictive for new drug combinations but also provide insights into mechanisms underlying combinatorial therapy. Further analysis confirmed that among our top ranked predictions of effective combinations, 69% are supported by literature, while the others represent novel potential drug combinations. We believe that our proposed approach can help to limit the search space of drug combinations and provide a new way to effectively utilize existing drugs for new purposes. Xing-Ming Zhao, Murat Iskar, Georg Zeller, Michael Kuhn 0004, Vera van Noort, Peer Bork |
PLoS Comput. Biol. | 1 |
| 2010 | APIS: accurate prediction of hot spots in protein interfaces by combining protrusion index with solvent accessibilityabstractBACKGROUND: It is well known that most of the binding free energy of protein interaction is contributed by a few key hot spot residues. These residues are crucial for understanding the function of proteins and studying their interactions. Experimental hot spots detection methods such as alanine scanning mutagenesis are not applicable on a large scale since they are time consuming and expensive. Therefore, reliable and efficient computational methods for identifying hot spots are greatly desired and urgently required. RESULTS: In this work, we introduce an efficient approach that uses support vector machine (SVM) to predict hot spot residues in protein interfaces. We systematically investigate a wide variety of 62 features from a combination of protein sequence and structure information. Then, to remove redundant and irrelevant features and improve the prediction performance, feature selection is employed using the F-score method. Based on the selected features, nine individual-feature based predictors are developed to identify hot spots using SVMs. Furthermore, a new ensemble classifier, namely APIS (A combined model based on Protrusion Index and Solvent accessibility), is developed to further improve the prediction accuracy. The results on two benchmark datasets, ASEdb and BID, show that this proposed method yields significantly better prediction accuracy than those previously published in the literature. In addition, we also demonstrate the predictive power of our proposed method by modelling two protein complexes: the calmodulin/myosin light chain kinase complex and the heat shock locus gene products U and V complex, which indicate that our method can identify more hot spots in these two complexes compared with other state-of-the-art methods. CONCLUSION: We have developed an accurate prediction model for hot spot residues, given the structure of a protein complex. A major contribution of this study is to propose several new features based on the protrusion index of amino acid residues, which has been shown to significantly improve the prediction performance of hot spots. Moreover, we identify a compact and useful feature subset that has an important implication for identifying hot spot residues. Our results indicate that these features are more effective than the conventional evolutionary conservation, pairwise residue potentials and other traditional features considered previously, and that the combination of our and traditional features may support the creation of a discriminative feature set for efficient prediction of hot spot residues. The data and source code are available on web site http://home.ustc.edu.cn/~jfxia/hotspot.html. Junfeng Xia, Xing-Ming Zhao, Jiangning Song, De-Shuang Huang |
BMC Bioinform. | 2 |
| 2010 | Analysis of Gene Expression Data Using Rpem Algorithm in Normal Mixture Model with Dynamic Adjustment of Learning RateabstractMicroarray technology is a useful tool for monitoring the expression levels of thousands of genes simultaneously. Recently, mixture modeling has been used to extract expression signatures from gene expression profiles. In general, two separate steps are utilized to estimate the number of classes and model parameters, respectively. However, such a method is often time-consuming and leads to suboptimal solutions. In this paper, we therefore apply a one-step approach, namely Rival Penalized Expectation-Maximization (RPEM) algorithm, to analyze the gene expression data. The RPEM algorithm is capable of estimating the parameters of normal mixture model, while determining the number of classes automatically at the same time. Furthermore, we speed up the learning procedure of RPEM by proposing a new mechanism to adjust the learning rate dynamically. The numerical results on real gene expression data demonstrate that our proposed method is indeed effective and efficient. Xing-Ming Zhao, Yiu-Ming Cheung, De-Shuang Huang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2008 | Automatic Modeling of Signal Pathways from Protein-Protein Interaction Networks
Xing-Ming Zhao, Rui-Sheng Wang, Luonan Chen, Kazuyuki Aihara |
APBC | 1 |
| 2008 | Gene function prediction using labeled and unlabeled dataabstractBACKGROUND: In general, gene function prediction can be formalized as a classification problem based on machine learning technique. Usually, both labeled positive and negative samples are needed to train the classifier. For the problem of gene function prediction, however, the available information is only about positive samples. In other words, we know which genes have the function of interested, while it is generally unclear which genes do not have the function, i.e. the negative samples. If all the genes outside of the target functional family are seen as negative samples, the imbalanced problem will arise because there are only a relatively small number of genes annotated in each family. Furthermore, the classifier may be degraded by the false negatives in the heuristically generated negative samples. RESULTS: In this paper, we present a new technique, namely Annotating Genes with Positive Samples (AGPS), for defining negative samples in gene function prediction. With the defined negative samples, it is straightforward to predict the functions of unknown genes. In addition, the AGPS algorithm is able to integrate various kinds of data sources to predict gene functions in a reliable and accurate manner. With the one-class and two-class Support Vector Machines as the core learning algorithm, the AGPS algorithm shows good performances for function prediction on yeast genes. CONCLUSION: We proposed a new method for defining negative samples in gene function prediction. Experimental results on yeast genes show that AGPS yields good performances on both training and test sets. In addition, the overlapping between prediction results and GO annotations on unknown genes also demonstrates the effectiveness of the proposed method. Xing-Ming Zhao, Yong Wang 0001, Luonan Chen, Kazuyuki Aihara |
BMC Bioinform. | 1 |
| 2006 | Classifying G-Protein Coupled Receptors with Hydropathy Blocks and Support Vector Machines
Xing-Ming Zhao, De-Shuang Huang, Shiwu Zhang, Yiu-Ming Cheung |
ICIC (3) | 1 |
| 2006 | A new technique for selecting features from protein sequencesabstractA new method for selecting features from protein sequences is proposed in this paper. First, the protein sequences are converted into fixed-dimensional feature vectors. Then, a subset of features is selected using relative entropy method and used as the inputs for Support Vector Machine (SVM). Finally, the trained SVM classifier is utilized to classify protein sequences into certain known protein families. Experimental results over proteins obtained from PIR database and GPCRs have shown that our proposed approach is really effective and efficient in selecting features from protein sequences. Xing-Ming Zhao, Jixiang Du, Hong-Qiang Wang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2006 | Classifying protein sequences using hydropathy blocks
De-Shuang Huang, Xing-Ming Zhao, Guang-Bin Huang, Yiu-Ming Cheung |
Pattern Recognit. | 2 |
| 2005 | A novel approach to extracting features from motif content and protein composition for protein sequence classification
Xing-Ming Zhao, Yiu-Ming Cheung, De-Shuang Huang |
Neural Networks | 1 |
| 2004 | A Novel Hybrid GA/SVM System for Protein Sequences Classification
Xing-Ming Zhao, De-Shuang Huang, Yiu-Ming Cheung, Hong-Qiang Wang |
IDEAL | 1 |
| 2004 | Representation of DNA sequences with multiple resolutions and BP neural network based classificationabstractIn this paper, we propose a new representation of DNA sequences, which constructs the word frequency vector with multiple resolutions based on the chaos game representation. Compared with the traditional vector, it combines a range of resolutions and reserves higher resolutions, but the dimension is reduced greatly relatively. The algorithm is detailed, which calculates coding format and codes each sequence. To evaluate the significance of our method, we represent Alu sequences by our proposed coding format. After that, the acquired vectors are used to train BP neural networks to recognize the Alu sequences. The experimental results show that this representation of DNA sequences is significant and efficient in biological data processing. De-Shuang Huang, Hong-Qiang Wang, Xing-Ming Zhao |
IJCNN | 4 |
| 2004 | A feature_core and SVM-based algorithm for identification of bioprocess-specific genome featuresabstractThis work presents a SVM and feature/spl I.bar/core-based algorithm for identification of key genome features. The significant difficulty in selecting key features from a high dimensional space of features is the curse of dimensions in searching. For this reason, a feature/spl I.bar/core-based search strategy is proposed in this algorithm. The strategy integrates the forward selection and backward elimination techniques. In this algorithm all key genome features are formed through the agglomeration and expansion of the feature core based on potential information about relevance in a SVM classifier. The application given proves that the algorithm is faster and more efficient than other methods such as clustering, and the single SVM-based method. Hong-Qiang Wang, De-Shuang Huang, Guang-Zheng Zhang, Xing-Ming Zhao |
IJCNN | 4 |
| 2004 | A Novel Clustering Analysis Based on PCA and SOMs for Gene Expression Patterns
Hong-Qiang Wang, De-Shuang Huang, Xing-Ming Zhao |
ISNN (2) | 3 |