Xia Li 0004

dblp:97/30-4 · DBLP profile ↗
← Back
44ranked-venue papers
3as first author
8since 2021 · last 2024
0000-0002-9794-2648ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 41 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-author
YearPublicationVenuePosition
2024 Identifying cell type-specific transcription factor-mediated activity immune modules reveal implications for immunotherapy and molecular classification of pan-cancer
abstract
Systematic investigation of tumor-infiltrating immune (TII) cells is important to the development of immunotherapies, and the clinical response prediction in cancers. There exists complex transcriptional regulation within TII cells, and different immune cell types display specific regulation patterns. To dissect transcriptional regulation in TII cells, we first integrated the gene expression profiles from single-cell datasets, and proposed a computational pipeline to identify TII cell type-specific transcription factor (TF) mediated activity immune modules (TF-AIMs). Our analysis revealed key TFs, such as BACH2 and NFKB1 play important roles in B and NK cells, respectively. We also found some of these TF-AIMs may contribute to tumor pathogenesis. Based on TII cell type-specific TF-AIMs, we identified eight CD8+ T cell subtypes. In particular, we found the PD1 + CD8+ T cell subset and its specific TF-AIMs associated with immunotherapy response. Furthermore, the TII cell type-specific TF-AIMs displayed the potential to be used as predictive markers for immunotherapy response of cancer patients. At the pan-cancer level, we also identified and characterized six molecular subtypes across 9680 samples based on the activation status of TII cell type-specific TF-AIMs. Finally, we constructed a user-friendly web interface CellTF-AIMs (http://bio-bigdata.hrbmu.edu.cn/CellTF-AIMs/) for exploring transcriptional regulatory pattern in various TII cell types. Our study provides valuable implications and a rich resource for understanding the mechanisms involved in cancer microenvironment and immunotherapy.
Mengyue Li, Yongjuan Tang, Liying Pei, Chunlong Zhang, Xia Li 0004, Yanjun Xu
Briefings Bioinform.11
2022 AgingBank: a manually curated knowledgebase and high-throughput analysis platform that provides experimentally supported multi-omics data relevant to aging in multiple species
abstract
Discovering the biological basis of aging is one of the greatest remaining challenges for biomedical field. Work on the biology of aging has discovered a range of interventions and pathways that control aging rate. Thus, we developed AgingBank (http://bio-bigdata.hrbmu.edu.cn/AgingBank) which was a manually curated comprehensive database and high-throughput analysis platform that provided experimentally supported multi-omics data relevant to aging in multiple species. AgingBank contained 3771 experimentally verified aging-related multi-omics entries from studies across more than 50 model organisms, including human, mice, worms, flies and yeast. The records included genome (single nucleotide polymorphism, copy number variation and somatic mutation), transcriptome [mRNA, long non-coding RNA (lncRNA), microRNA (miRNA) and circular RNA (circRNA)], epigenome (DNA methylation and histone modification), other modification and regulation elements (transcription factor, enhancer, promoter, gene silence, alternative splicing and RNA editing). In addition, AgingBank was also an online computational analysis platform containing five useful tools (Aging Landscape, Differential Expression Analyzer, Data Heat Mapper, Co-Expression Network and Functional Annotation Analyzer), nearly 112 high-throughput experiments of genes, miRNAs, lncRNAs, circRNAs and methylation sites related with aging. Cancer & Aging module was developed to explore the relationships between aging and cancer. Submit & Analysis module allows users upload and analyze their experiments data. AginBank is a valuable resource for elucidating aging-related biomarkers and relationships with other diseases.
Yue Gao 0004, Shipeng Shang, Shuang Guo 0001, Hanxiao Zhou, Jing Gan, Yakun Zhang 0003, Xia Li 0004, Shangwei Ning
Briefings Bioinform.9
2022 Identifying and characterizing drug sensitivity-related lncRNA-TF-gene regulatory triplets
abstract
Recently, many studies have shown that lncRNA can mediate the regulation of TF-gene in drug sensitivity. However, there is still a lack of systematic identification of lncRNA-TF-gene regulatory triplets for drug sensitivity. In this study, we propose a novel analytic approach to systematically identify the lncRNA-TF-gene regulatory triplets related to the drug sensitivity by integrating transcriptome data and drug sensitivity data. Totally, 1570 drug sensitivity-related lncRNA-TF-gene triplets were identified, and 16 307 relationships were formed between drugs and triplets. Then, a comprehensive characterization was performed. Drug sensitivity-related triplets affect a variety of biological functions including drug response-related pathways. Phenotypic similarity analysis showed that the drugs with many shared triplets had high similarity in their two-dimensional structures and indications. In addition, Network analysis revealed the diverse regulation mechanism of lncRNAs in different drugs. Also, survival analysis indicated that lncRNA-TF-gene triplets related to the drug sensitivity could be candidate prognostic biomarkers for clinical applications. Next, using the random walk algorithm, the results of which we screen therapeutic drugs for patients across three cancer types showed high accuracy in the drug-cell line heterogeneity network based on the identified triplets. Besides, we developed a user-friendly web interface-DrugSETs (http://bio-bigdata.hrbmu.edu.cn/DrugSETs/) available to explore 1570 lncRNA-TF-gene triplets relevant with 282 drugs. It can also submit a patient's expression profile to predict therapeutic drugs conveniently. In summary, our research may promote the study of lncRNAs in the drug resistance mechanism and improve the effectiveness of treatment.
Congxue Hu, Yingqi Xu, Wanqi Mi, Shuaijun Chen, Xia Li 0004, Yanjun Xu
Briefings Bioinform.9
2022 Comprehensive characterization genetic regulation and chromatin landscape of enhancer-associated long non-coding RNAs and their implication in human cancer
abstract
Long non-coding RNAs (lncRNAs) that emanate from enhancer regions (defined as enhancer-associated lncRNAs, or elncRNAs) are emerging as critical regulators in disease progression. However, their biological characteristics and clinical relevance have not been fully portrayed. Here, based on the traditional expression quantitative loci (eQTL) and our optimized residual eQTL method, we comprehensively described the genetic effect on elncRNA expression in more than 300 lymphoblastoid cell lines. Meanwhile, a chromatin atlas of elncRNAs relative to the genetic regulation state was depicted. By applying the maximum likelihood estimate method, we successfully identified causal elncRNAs for protein-coding gene expression reprogramming and showed their associated single nucleotide polymorphisms (SNPs) favor binding of transcription factors. Further epigenome analysis revealed two immune-associated elncRNAs AL662844.4 and LINC01215 possess high levels of H3K27ac and H3K4me1 in human cancer. Besides, pan-cancer analysis of 3D genome, transcriptome, and regulatome data showed they potentially regulate tumor-immune cell interaction through affecting MHC class I genes and CD47, respectively. Moreover, our study showed there exist associations between elncRNA and patient survival. Finally, we made a user-friendly web interface available for exploring the regulatory relationship of SNP-elncRNA-protein-coding gene triplets (http://bio-bigdata.hrbmu.edu.cn/elncVarReg). Our study provides critical mechanistic insights for elncRNA function and illustrates their implications in human cancer.
Hanxiao Zhou, Peng Wang 0082, Yue Gao 0004, Shipeng Shang, Shuang Guo 0001, Jie Sun 0021, Zhiying Xiong, Shangwei Ning, Xia Li 0004
Briefings Bioinform.12
2022 PRES: a webserver for decoding the functional perturbations of RNA editing sites
abstract
Rapid progresses in RNA-Seq and computational methods have assisted in quantifying A-to-I RNA editing and altered RNA editing sites have been widely observed in various diseases. Nevertheless, functional characterization of the altered RNA editing sites still remains a challenge. Here, we developed perturbations of RNA editing sites (PRES; http://bio-bigdata.hrbmu.edu.cn/PRES/) as the webserver for decoding functional perturbations of RNA editing sites based on editome profiling. After uploading an editome profile among samples of different groups, PRES will first annotate the editing sites to various genomic elements and detect differential editing sites under the user-selected method and thresholds. Next, the downstream functional perturbations of differential editing sites will be characterized from gain or loss miRNA/RNA binding protein regulation, RNA and protein structure changes, and the perturbed biological pathways. A prioritization module was developed to rank genes based on their functional consequences of RNA editing events. PRES provides user-friendly functionalities, ultra-efficient calculation, intuitive table and figure visualization interface to display the annotated RNA editing events, filtering options and elaborate application notebooks. We anticipate PRES will provide an opportunity for better understanding the regulatory mechanisms of RNA editing in human complex diseases.
Dezhong Lv, Kang Xu 0002, Changbo Yang, Ya Luo, Haozhe Zou, Na Ding, Xia Li 0004, Tingting Shao, Yongsheng Li 0001, Juan Xu 0001
Briefings Bioinform.10
2022 Revealing the contribution of somatic gene mutations to shaping tumor immune microenvironment
abstract
Interaction between tumor cells and immune cells determined highly heterogeneous microenvironments across patients, leading to substantial variation in clinical benefits from immunotherapy. Somatic gene mutations were found not only to elicit adaptive immunity but also to influence the composition of tumor immune microenvironment and various processes of antitumor immunity. However, due to an incomplete view of associations between gene mutations and immunophenotypes, how tumor cells shape the immune microenvironment and further determine the clinical benefit of immunotherapy is still unclear. To address this, we proposed a computational approach, inference of mutation effect on immunophenotype by integrated gene set enrichment analysis (MEIGSEA), for tracing back the genomic factor responsible for differences in immunophenotypes. MEIGSEA was demonstrated to accurately identify the previous confirmed immune-associated gene mutations, and systematic evaluation in simulation data further supported its performance. We used MEIGSEA to investigate the influence of driver gene mutations on the infiltration of 22 immune cell types across 19 cancers from The Cancer Genome Atlas. The top associated gene mutations with infiltration of CD8 T cells, such as CASP8, KRAS and EGFR, also showed extensive impact on other immune components; meanwhile, immune effector cells shared critical gene mutations that collaboratively contribute to shaping distinct tumor immune microenvironment. Furthermore, we highlighted the predictive capacity of gene mutations that are positively associated with CD8 T cells for the clinical benefit of immunotherapy. Taken together, we present a computational framework to help illustrate the potential of somatic gene mutations in shaping the tumor immune microenvironment.
Shiwei Zhu, Yujia Lan, Zedong Jiang, Jiali Zhu, Gaoming Liao, Yanyan Ping, Jinyuan Xu, Yun Xiao 0001, Xia Li 0004
Briefings Bioinform.13
2021 CeRNASeek: an R package for identification and analysis of ceRNA regulation
abstract
Competitive endogenous RNA (ceRNA) represents a novel layer of gene regulation that controls both physiological and pathological processes. However, there is still lack of computational tools for quickly identifying ceRNA regulation. To address this problem, we presented an R-package, CeRNASeek, which allows identifying and analyzing ceRNA-ceRNA interactions by integration of multiple-omics data. CeRNASeek integrates six widely used computational methods to identify ceRNA-ceRNA interactions, including two global and four context-specific ceRNA regulation prediction methods. In addition, it provides several downstream analyses for predicted ceRNA-ceRNA pairs, including regulatory network analysis, functional annotation and survival analysis. With examples of cancer-related ceRNA prioritization and cancer subtyping, we demonstrate that CeRNASeek is a valuable tool for investigating the function of ceRNAs in complex diseases. In summary, CeRNASeek provides a comprehensive and efficient tool for identifying and analysis of ceRNA regulation. The package is available on the Comprehensive R Archive Network (CRAN) at https://CRAN.R-project.org/package=CeRNASeek.
Mengying Zhang 0007, Xiyun Jin, Juan Xu 0001, Yongsheng Li 0001, Xia Li 0004
Briefings Bioinform.9
2021 SurvivalMeth: a web server to investigate the effect of DNA methylation-related functional elements on prognosis
abstract
Aberrant DNA methylation is a fundamental characterization of epigenetics for carcinogenesis. Abnormality of DNA methylation-related functional elements (DMFEs) may lead to dysfunction of regulatory genes in the progression of cancers, contributing to prognosis of many cancers. There is an urgent need to construct a tool to comprehensively assess the impact of DMFEs on prognosis. Therefore, we developed SurvivalMeth (http://bio-bigdata.hrbmu.edu.cn/survivalmeth) to explore the prognosis-related DMFEs, which documented many kinds of DMFEs, including 309,465 CpG island-related elements, 104,748 transcript-related elements, 77,634 repeat elements, as well as cell-type specific 1,689,653 super enhancers (SE) and 1,304,902 CTCF binding regions for analysis. SurvivalMeth is a convenient tool which collected DNA methylation profiles of 36 cancers and allowed users to query their genes of interest in different datasets for prognosis. Furthermore, SurvivalMeth not only integrated different combinations, including single DMFE, multiple DMFEs, SEs and clinical data, to perform survival analysis on preupload data but also allowed for uploading customized DNA methylation profile of DMFEs from various diseases to analyze. SurvivalMeth provided a comprehensive resource and automated analysis for prognostic DMFEs, including DMFE methylation level, correlation analysis, clinical analysis, differential analysis, DMFE annotation, survival-related detailed result and visualization of survival analysis. In summary, we believe that SurvivalMeth will facilitate prognostic research of DMFEs in diverse cancers.
Chunlong Zhang, Dezhong Lv, Yongsheng Li 0001, Juan Xu 0001, Xia Li 0004
Briefings Bioinform.10
2020 RNAactDrug: a comprehensive database of RNAs associated with drug sensitivity from multi-omics data
abstract
Drug sensitivity has always been at the core of individualized cancer chemotherapy. However, we have been overwhelmed by large-scale pharmacogenomic data in the era of next-generation sequencing technology, which makes it increasingly challenging for researchers, especially those without bioinformatic experience, to perform data integration, exploration and analysis. To bridge this gap, we developed RNAactDrug, a comprehensive database of RNAs associated with drug sensitivity from multi-omics data, which allows users to explore drug sensitivity and RNA molecule associations directly. It provides association data between drug sensitivity and RNA molecules including mRNAs, long non-coding RNAs (lncRNAs) and microRNAs (miRNAs) at four molecular levels (expression, copy number variation, mutation and methylation) from integrated analysis of three large-scale pharmacogenomic databases (GDSC, CellMiner and CCLE). RNAactDrug currently stores more than 4 924 200 associations of RNA molecules and drug sensitivity at four molecular levels covering more than 19 770 mRNAs, 11 119 lncRNAs, 438 miRNAs and 4155 drugs. A user-friendly interface enriched with various browsing sections augmented with advance search facility for querying the database is offered for users retrieving. RNAactDrug provides a comprehensive resource for RNA molecules acting in drug sensitivity, and it could be used to prioritize drug sensitivity-related RNA molecules, further promoting the identification of clinically actionable biomarkers in drug sensitivity and drug development more cost-efficiently by making this knowledge accessible to both basic researchers and clinical practitioners. Database URL: http://bio-bigdata.hrbmu.edu.cn/RNAactDrug.
Qun Dong, Yanjun Xu, Yingqi Xu, Desi Shang, Chunlong Zhang, Haixiu Yang, Zihan Tian, Kai Mi 0003, Xia Li 0004
Briefings Bioinform.11
2020 A comprehensive overview of oncogenic pathways in human cancer
abstract
Alterations of biological pathways can lead to oncogenesis. An overview of these oncogenic pathways would be highly valuable for researchers to reveal the pathogenic mechanism and develop novel therapeutic approaches for cancers. Here, we reviewed approximately 8500 literatures and documented experimentally validated cancer-pathway associations as benchmarking data set. This data resource includes 4709 manually curated relationships between 1557 paths and 49 cancers with 2427 upstream regulators in 7 species. Based on this resource, we first summarized the cancer-pathway associations and revealed some commonly deregulated pathways across tumor types. Then, we systematically analyzed these oncogenic pathways by integrating TCGA pan-cancer data sets. Multi-omics analysis showed oncogenic pathways may play different roles across tumor types under different omics contexts. We also charted the survival relevance landscape of oncogenic pathways in 26 tumor types, identified dominant omics features and found survival relevance for oncogenic pathways varied in tumor types and omics levels. Moreover, we predicted upstream regulators and constructed a hierarchical network model to understand the pathogenic mechanism of human cancers underlying oncogenic pathway context. Finally, we developed `CPAD' (freely available at http://bio-bigdata.hrbmu.edu.cn/CPAD/), an online resource for exploring oncogenic pathways in human cancers, that integrated manually curated cancer-pathway associations, TCGA pan-cancer multi-omics data sets, drug-target data, drug sensitivity and multi-omics data for cancer cell lines. In summary, our study provides a comprehensive characterization of oncogenic pathways and also presents a valuable resource for investigating the pathogenesis of human cancer.
Tan Wu, Yanjun Xu, Qun Dong, Yingqi Xu, Chunlong Zhang, Jianxia Gao, Liqiu Liu, Xiaoxu Hu, Jian Huang 0004, Xia Li 0004
Briefings Bioinform.13
2020 Identification and comprehensive characterization of lncRNAs with copy number variations and their driving transcriptional perturbed subpathways reveal functional significance for cancer
abstract
Numerous studies have shown that copy number variation (CNV) in lncRNA regions play critical roles in the initiation and progression of cancer. However, our knowledge about their functionalities is still limited. Here, we firstly provided a computational method to identify lncRNAs with copy number variation (lncRNAs-CNV) and their driving transcriptional perturbed subpathways by integrating multidimensional omics data of cancer. The high reliability and accuracy of our method have been demonstrated. Then, the method was applied to 14 cancer types, and a comprehensive characterization and analysis was performed. LncRNAs-CNV had high specificity in cancers, and those with high CNV level may perturb broad biological functions. Some core subpathways and cancer hallmarks widely perturbed by lncRNAs-CNV were revealed. Moreover, subpathways highlighted the functional diversity of lncRNAs-CNV in various cancers. Survival analysis indicated that functional lncRNAs-CNV could be candidate prognostic biomarkers for clinical applications, such as ST7-AS1, CDKN2B-AS1 and EGFR-AS1. In addition, cascade responses and a functional crosstalk model among lncRNAs-CNV, impacted genes, driving subpathways and cancer hallmarks were proposed for understanding the driving mechanism of lncRNAs-CNV. Finally, we developed a user-friendly web interface-LncCASE (http://bio-bigdata.hrbmu.edu.cn/LncCASE/) for exploring lncRNAs-CNV and their driving subpathways in various cancer types. Our study identified and systematically characterized lncRNAs-CNV and their driving subpathways and presented valuable resources for investigating the functionalities of non-coding variations and the mechanisms of tumorigenesis.
Yanjun Xu, Tan Wu, Qun Dong, Desi Shang, Yingqi Xu, Chunlong Zhang, Yiying Dou, Congxue Hu, Haixiu Yang, Lihua Wang 0002, Xia Li 0004
Briefings Bioinform.15
2019 Improving Identification of Essential Proteins by a Novel Ensemble Method
Wei Dai 0012, Xia Li 0004, Wei Peng 0004, Jurong Song, Jiancheng Zhong, Jianxin Wang 0001
ISBRA2
2019 Identifying mutual exclusivity across cancer genomes: computational approaches to discover genetic interaction and reveal tumor vulnerability
abstract
Systematic sequencing of cancer genomes has revealed prevalent heterogeneity, with patients harboring various combinatorial patterns of genetic alteration. In particular, a phenomenon that a group of genes exhibits mutually exclusive patterns has been widespread across cancers, covering a broad spectrum of crucial cancer pathways. Recently, there is considerable evidence showing that, mutual exclusivity reflects alternative functions in tumor initiation and progression, or suggests adverse effects of their concurrence. Given its importance, numerous computational approaches have been proposed to study mutual exclusivity using genomic profiles alone, or by integrating networks and phenotypes. Some of them have been routinely used to explore genetic associations, which lead to a deeper understanding of carcinogenic mechanisms and reveals unexpected tumor vulnerabilities. Here, we present an overview of mutual exclusivity from the perspective of cancer genome. We describe the common hypothesis underlying mutual exclusivity, summarize the strategies for the identification of significant mutually exclusive patterns, compare the performance of representative algorithms from simulated data sets and discuss their common confounders.
Yulan Deng, Shangyi Luo, Chunyu Deng, Wenkang Yin, Hongyi Zhang 0005, Yujia Lan, Yanyan Ping, Yun Xiao 0001, Xia Li 0004
Briefings Bioinform.12
2019 Landscape of the long non-coding RNA transcriptome in human heart
abstract
Long non-coding RNAs (lncRNAs) have been revealed to play essential roles in the human cardiovascular system. However, information about their mechanisms is limited, and a comprehensive view of cardiac lncRNAs is lacking from a multiple tissues perspective to date. Here, the landscape of the lncRNA transcriptome in human heart was summarized. We summarized all lncRNA transcripts from publicly available human transcriptome resources (156 heart samples and 210 samples from 29 other tissues) and systematically analysed all annotated and novel lncRNAs expressed in heart. A total of 7485 lncRNAs whose expression was elevated in heart (HE lncRNAs) and 453 lncRNAs expressed in all 30 analysed tissues (EIA lncRNAs) were extracted. Using various bioinformatics resources, methods and tools, the features of these lncRNAs were discussed from various perspectives, including genomic structure, conservation, dynamic variation during heart development, cis-regulation, differential expression in cardiovascular diseases and cancers as well as regulation at transcriptional and post-transcriptional levels. Afterwards, all the features discussed above were integrated into a user-friendly resource named CARDIO-LNCRNAS (http://bio-bigdata.hrbmu.edu.cn/CARDIO-LNCRNAS/ or http://www.bio-bigdata.net/CARDIO-LNCRNAS/). This study represents the first global view of lncRNAs in the human cardiovascular system based on multiple tissues and sheds light on the role of lncRNAs in developments and heart disorders.
Chunjie Jiang, Na Ding, Xiyun Jin, Caiqin Huo, Yongsheng Li 0001, Juan Xu 0001, Xia Li 0004
Briefings Bioinform.10
2019 Systematic review regulatory principles of non-coding RNAs in cardiovascular diseases
abstract
Cardiovascular diseases (CVDs) continue to be a major cause of morbidity and mortality, and non-coding RNAs (ncRNAs) play critical roles in CVDs. With the recent emergence of high-throughput technologies, including small RNA sequencing, investigations of CVDs have been transformed from candidate-based studies into genome-wide undertakings, and a number of ncRNAs in CVDs were discovered in various studies. A comprehensive review of these ncRNAs would be highly valuable for researchers to get a complete picture of the ncRNAs in CVD. To address these knowledge gaps and clinical needs, in this review, we first discussed dysregulated ncRNAs and their critical roles in cardiovascular development and related diseases. Moreover, we reviewed >28 561 published papers and documented the ncRNA-CVD association benchmarking data sets to summarize the principles of ncRNA regulation in CVDs. This data set included 13 249 curated relationships between 9503 ncRNAs and 139 CVDs in 12 species. Based on this comprehensive resource, we summarized the regulatory principles of dysregulated ncRNAs in CVDs, including the complex associations between ncRNA and CVDs, tissue specificity and ncRNA synergistic regulation. The highlighted principles are that CVD microRNAs (miRNAs) are highly expressed in heart tissue and that they play central roles in miRNA-miRNA functional synergistic network. In addition, CVD-related miRNAs are close to one another in the functional network, indicating the modular characteristic features of CVD miRNAs. We believe that the regulatory principles summarized here will further contribute to our understanding of ncRNA function and dysregulation mechanisms in CVDs.
Yongsheng Li 0001, Caiqin Huo, Xiyun Jin, Jinwen Zhang, Zheng Guo 0002, Juan Xu 0001, Xia Li 0004
Briefings Bioinform.11
2019 Survey of miRNA-miRNA cooperative regulation principles across cancer types
abstract
Cooperative regulation among multiple microRNAs (miRNAs) is a complex type of posttranscriptional regulation in human; however, the global view of the system-level regulatory principles across cancers is still unclear. Here, we investigated miRNA-miRNA cooperative regulatory landscape across 18 cancer types and summarized the regulatory principles of miRNAs. The miRNA-miRNA cooperative pan-cancer network exhibited a scale-free and modular architecture. Cancer types with similar tissue origins had high similarity in cooperative network structure and expression of cooperative miRNA pairs. In addition, cooperative miRNAs showed divergent properties, including higher expression, greater expression variation and a stronger regulatory strength towards targets and were likely to regulate cancer hallmark-related functions. We found a marked rewiring of miRNA-miRNA cooperation between various cancers and revealed conserved and rewired network miRNA hubs. We further identified the common hubs, cancer-specific hubs and other hubs, which tend to target known anticancer drug targets. Finally, miRNA cooperative modules were found to be associated with patient survival in several cancer types. Our study highlights the potential of pan-cancer miRNA-miRNA cooperative regulation as a novel paradigm that may aid in the discovery of tumorigenesis mechanisms and development of anticancer drugs.
Tingting Shao, Guangjuan Wang, Hong Chen 0022, Yunjin Xie, Xiyun Jin, Jing Bai 0014, Juan Xu 0001, Xia Li 0004, Jian Huang 0004, Yan Jin 0007, Yongsheng Li 0001
Briefings Bioinform.8
2019 Breast cancer prognosis signature: linking risk stratification to disease subtypes
abstract
Breast cancer is a very complex and heterogeneous disease with variable molecular mechanisms of carcinogenesis and clinical behaviors. The identification of prognostic risk factors may enable effective diagnosis and treatment of breast cancer. In particular, numerous gene-expression-based prognostic signatures were developed and some of them have already been applied into clinical trials and practice. In this study, we summarized several representative gene-expression-based signatures with significant prognostic value and separately assessed their ability of prognosis prediction in their originally targeted populations of breast cancer. Notably, many of the collected signatures were originally designed to predict the outcomes of estrogen receptor positive (ER+) patients or the whole breast cancer cohort; there are no typical signatures used for the prognostic prediction in a specific population of patients with the intrinsic subtype. We thus attempted to identify subtype-specific prognostic signatures via a computational framework for analyzing multi-omics profiles and patient survival. For both the discovery and an independent data set, we confirmed that subtype-specific signature is a strong and significant independent prognostic factor in the corresponding cohort. These results indicate that the subtype-specific prognostic signature has a much higher resolution in the risk stratification, which may lead to improved therapies and precision medicine for patients with breast cancer.
Fulong Yu, Fei Quan, Jinyuan Xu, Yujia Lan, Huating Yuan, Hongyi Zhang 0005, Shujun Cheng, Yun Xiao 0001, Xia Li 0004
Briefings Bioinform.12
2018 Combinatorial epigenetic regulation of non-coding RNAs has profound effects on oncogenic pathways in breast cancer subtypes
abstract
Although systematic genomic studies have identified a broad spectrum of non-coding RNAs (ncRNAs) that are involved in breast cancer, our understanding of the epigenetic dysregulation of those ncRNAs remains limited. Here, we systematically analysed the epigenetic alterations of microRNAs (miRNAs) and long non-coding RNAs (lncRNAs) in two breast cancer subtypes (luminal and basal). Widespread epigenetic alterations of miRNAs and lncRNAs were observed in both cancer subtypes. In contrast to protein-coding genes, the majority of epigenetically dysregulated ncRNAs were shared between subtypes, but a subset of transcriptomic and corresponding epigenetic changes occurred in a subtype-specific manner. In addition, our findings suggested that various types of epi-modifications might synergistically modulate ncRNA transcription. Our observations further highlighted the complementary dysregulation of epi-modifications, particularly of miRNA members within the same family, which produced the same directed alterations as a result of diverse epi-modifications. Functional enrichment analysis revealed that epigenetically dysregulated ncRNAs were significantly involved in several hallmarks of cancers. Finally, our analysis of epigenetic modification-mediated miRNA regulatory networks revealed that cancer progression was associated with specific miRNA-gene modules in two subtypes. This study enhances understanding of the aberrant epigenetic patterns of ncRNA expression and provides new insights into the functions of ncRNAs in breast cancer subtypes.
Juan Xu 0001, Zishan Wang, Shengli Li 0005, Jinwen Zhang, Chunjie Jiang, Jing Li 0115, Yongsheng Li 0001, Xia Li 0004
Briefings Bioinform.10
2017 A comprehensive overview of lncRNA annotation resources
abstract
Long noncoding RNAs (lncRNAs) are emerging as a class of important regulators participating in various biological functions and disease processes. With the widespread application of next-generation sequencing technologies, large numbers of lncRNAs have been identified, producing plenty of lncRNA annotation resources in different contexts. However, at present, we lack a comprehensive overview of these lncRNA annotation resources. In this study, we reviewed 24 currently available lncRNA annotation resources referring to > 205 000 lncRNAs in over 50 tissues and cell lines. We characterized these annotation resources from different aspects, including exon structure, expression, histone modification and function. We found many distinct properties among these annotation resources. Especially, these resources showed diverse chromatin signatures, remarkable tissue and cell type dependence and functional specificity. Our results suggested the incompleteness and complementarity of current lncRNA annotations and the necessity of integration of multiple resources to comprehensively characterize lncRNAs. Finally, we developed 'LNCat' (lncRNA atlas, freely available at http://biocc.hrbmu.edu.cn/LNCat/), a user-friendly database that provides a genome browser of lncRNA structures, visualization of different resources from multiple angles and download of different combinations of lncRNA annotations, and supports rapid exploration, comparison and integration of lncRNA annotation resources. Overall, our study provides a comprehensive comparison of numerous lncRNA annotations, and can facilitate understanding of lncRNAs in human disease.
Jinyuan Xu, Jing Bai 0014, Yanling Lv, Yonghui Gong, Hongying Zhao, Fulong Yu, Yanyan Ping, Guanxiong Zhang, Yujia Lan, Yun Xiao 0001, Xia Li 0004
Briefings Bioinform.13
2017 miRNA-miRNA crosstalk: from genomics to phenomics
abstract
The discovery of microRNA (miRNA)-miRNA crosstalk has greatly improved our understanding of complex gene regulatory networks in normal and disease-specific physiological conditions. Numerous approaches have been proposed for modeling miRNA-miRNA networks based on genomic sequences, miRNA-mRNA regulation, functional information and phenomics alone, or by integrating heterogeneous data. In addition, it is expected that miRNA-miRNA crosstalk can be reprogrammed in different tissues or specific diseases. Thus, transcriptome data have also been integrated to construct context-specific miRNA-miRNA networks. In this review, we summarize the state-of-the-art miRNA-miRNA network modeling methods, which range from genomics to phenomics, where we focus on the need to integrate heterogeneous types of omics data. Finally, we suggest future directions for studies of crosstalk of noncoding RNAs. This comprehensive summarization and discussion elucidated in this work provide constructive insights into miRNA-miRNA crosstalk.
Juan Xu 0001, Tingting Shao, Na Ding, Yongsheng Li 0001, Xia Li 0004
Briefings Bioinform.5
2015 Identifying novel associations between small molecules and miRNAs based on integrated molecular networks
abstract
MOTIVATION: miRNAs play crucial roles in human diseases and newly discovered could be targeted by small molecule (SM) drug compounds. Thus, the identification of small molecule drug compounds (SM) that target dysregulated miRNAs in cancers will provide new insight into cancer biology and accelerate drug discovery for cancer therapy. RESULTS: In this study, we aimed to develop a novel computational method to comprehensively identify associations between SMs and miRNAs. To this end, exploiting multiple molecular interaction databases, we first established an integrated SM-miRNA association network based on 690 561 SM to SM interactions, 291 600 miRNA to miRNA associations, as well as 664 known SM to miRNA targeting pairs. Then, by performing Random Walk with Restart algorithm on the integrated network, we prioritized the miRNAs associated to each of the SMs. By validating our results utilizing an independent dataset we obtained an area under the ROC curve greater than 0.7. Furthermore, comparisons indicated our integrated approach significantly improved the identification performance of those simple modeled methods. This computational framework as well as the prioritized SM-miRNA targeting relationships will promote the further developments of targeted cancer therapies. CONTACT: [email protected], [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yingli Lv, Shuyuan Wang, Fan-Lin Meng, Jing Wang 0004, Wei Jiang 0023, Xia Li 0004
Bioinform.10
2015 Prediction of potential disease-associated microRNAs based on random walk
abstract
MOTIVATION: Identifying microRNAs associated with diseases (disease miRNAs) is helpful for exploring the pathogenesis of diseases. Because miRNAs fulfill function via the regulation of their target genes and because the current number of experimentally validated targets is insufficient, some existing methods have inferred potential disease miRNAs based on the predicted targets. It is difficult for these methods to achieve excellent performance due to the high false-positive and false-negative rates for the target prediction results. Alternatively, several methods have constructed a network composed of miRNAs based on their associated diseases and have exploited the information within the network to predict the disease miRNAs. However, these methods have failed to take into account the prior information regarding the network nodes and the respective local topological structures of the different categories of nodes. Therefore, it is essential to develop a method that exploits the more useful information to predict reliable disease miRNA candidates. RESULTS: miRNAs with similar functions are normally associated with similar diseases and vice versa. Therefore, the functional similarity between a pair of miRNAs is calculated based on their associated diseases to construct a miRNA network. We present a new prediction method based on random walk on the network. For the diseases with some known related miRNAs, the network nodes are divided into labeled nodes and unlabeled nodes, and the transition matrices are established for the two categories of nodes. Furthermore, different categories of nodes have different transition weights. In this way, the prior information of nodes can be completely exploited. Simultaneously, the various ranges of topologies around the different categories of nodes are integrated. In addition, how far the walker can go away from the labeled nodes is controlled by restarting the walking. This is helpful for relieving the negative effect of noisy data. For the diseases without any known related miRNAs, we extend the walking on a miRNA-disease bilayer network. During the prediction process, the similarity between diseases, the similarity between miRNAs, the known miRNA-disease associations and the topology information of the bilayer network are exploited. Moreover, the importance of information from different layers of network is considered. Our method achieves superior performance for 18 human diseases with AUC values ranging from 0.786 to 0.945. Moreover, case studies on breast neoplasms, lung neoplasms, prostatic neoplasms and 32 diseases further confirm the ability of our method to discover potential disease miRNAs. AVAILABILITY AND IMPLEMENTATION: A web service for the prediction and analysis of disease miRNAs is available at http://bioinfolab.stx.hk/midp/.
Ping Xuan, Yahong Guo, Jin Li 0024, Xia Li 0004, Yingli Zhong, Zhaogong Zhang
Bioinform.5
2014 LincSNP: a database of linking disease-associated SNPs to human large intergenic non-coding RNAs
abstract
BACKGROUND: Genome-wide association studies (GWAS) have successfully identified a large number of single nucleotide polymorphisms (SNPs) that are associated with a wide range of human diseases. However, many of these disease-associated SNPs are located in non-coding regions and have remained largely unexplained. Recent findings indicate that disease-associated SNPs in human large intergenic non-coding RNA (lincRNA) may lead to susceptibility to diseases through their effects on lincRNA expression. There is, therefore, a need to specifically record these SNPs and annotate them as potential candidates for disease. DESCRIPTION: We have built LincSNP, an integrated database, to identify and annotate disease-associated SNPs in human lincRNAs. The current release of LincSNP contains approximately 140,000 disease-associated SNPs (or linkage disequilibrium SNPs), which can be mapped to around 5,000 human lincRNAs, together with their comprehensive functional annotations. The database also contains annotated, experimentally supported SNP-lincRNA-disease associations and disease-associated lincRNAs. It provides flexible search options for data extraction and searches can be performed by disease/phenotype name, SNP ID, lincRNA name and chromosome region. In addition, we provide users with a link to download all the data from LincSNP and have developed a web interface for the submission of novel identified SNP-lincRNA-disease associations. CONCLUSIONS: The LincSNP database aims to integrate disease-associated SNPs and human lincRNAs, which will be an important resource for the investigation of the functions and mechanisms of lincRNAs in human disease. The database is available at http://bioinfo.hrbmu.edu.cn/LincSNP.
Shangwei Ning, Zuxianglan Zhao, Jingrun Ye, Peng Wang 0082, Ronghong Li, Xia Li 0004
BMC Bioinform.8
2014 The detection of risk pathways, regulated by miRNAs, via the integration of sample-matched miRNA-mRNA profiles and pathway structure
Jing Li 0115, Chunquan Li 0002, Junwei Han 0003, Chunlong Zhang, Desi Shang, Qianlan Yao, Yanjun Xu, Wei Liu 0187, Meng Zhou 0003, Haixiu Yang, Xia Li 0004
J. Biomed. Informatics13
2013 Identification of active transcription factor and miRNA regulatory pathways in Alzheimer's disease
abstract
MOTIVATION: Alzheimer's disease (AD) is a severe neurodegenerative disease of the central nervous system that may be caused by perturbation of regulatory pathways rather than the dysfunction of a single gene. However, the pathology of AD has yet to be fully elucidated. RESULTS: In this study, we systematically analyzed AD-related mRNA and miRNA expression profiles as well as curated transcription factor (TF) and miRNA regulation to identify active TF and miRNA regulatory pathways in AD. By mapping differentially expressed genes and miRNAs to the curated TF and miRNA regulatory network as active seed nodes, we obtained a potential active subnetwork in AD. Next, by using the breadth-first-search technique, potential active regulatory pathways, which are the regulatory cascade of TFs, miRNAs and their target genes, were identified. Finally, based on the known AD-related genes and miRNAs, the hypergeometric test was used to identify active pathways in AD. As a result, nine pathways were found to be significantly activated in AD. A comprehensive literature review revealed that eight out of nine genes and miRNAs in these active pathways were associated with AD. In addition, we inferred that the pathway hsa-miR-146a→STAT1→MYC, which is the source of all nine significantly active pathways, may play an important role in AD progression, which should be further validated by biological experiments. Thus, this study provides an effective approach to finding active TF and miRNA regulatory pathways in AD and can be easily applied to other complex diseases.
Wei Jiang 0023, Fan-Lin Meng, Baofeng Lian, Xuexin Yu, Enyu Dai, Shuyuan Wang, Xia Li 0004
Bioinform.12
2013 Topologically inferring risk-active pathways toward precise cancer classification by directed random walk
abstract
MOTIVATION: The accurate prediction of disease status is a central challenge in clinical cancer research. Microarray-based gene biomarkers have been identified to predict outcome and outperform traditional clinical parameters. However, the robustness of the individual gene biomarkers is questioned because of their little reproducibility between different cohorts of patients. Substantial progress in treatment requires advances in methods to identify robust biomarkers. Several methods incorporating pathway information have been proposed to identify robust pathway markers and build classifiers at the level of functional categories rather than of individual genes. However, current methods consider the pathways as simple gene sets but ignore the pathway topological information, which is essential to infer a more robust pathway activity. RESULTS: Here, we propose a directed random walk (DRW)-based method to infer the pathway activity. DRW evaluates the topological importance of each gene by capturing the structure information embedded in the directed pathway network. The strategy of weighting genes by their topological importance greatly improved the reproducibility of pathway activities. Experiments on 18 cancer datasets showed that the proposed method yielded a more accurate and robust overall performance compared with several existing gene-based and pathway-based classification methods. The resulting risk-active pathways are more reliable in guiding therapeutic selection and the development of pathway-specific therapeutic strategies. AVAILABILITY: DRW is freely available at http://210.46.85.180:8080/DRWPClass/
Wei Liu 0187, Chunquan Li 0002, Yanjun Xu, Haixiu Yang, Qianlan Yao, Junwei Han 0003, Desi Shang, Chunlong Zhang, Yun Xiao 0001, Xia Li 0004
Bioinform.14
2013 SM2miR: a database of the experimentally validated small molecules' effects on microRNA expression
abstract
UNLABELLED: The inappropriate expression of microRNAs (miRNAs) is closely related with disease diagnosis, prognosis and therapy response. Recently, many studies have demonstrated that bioactive small molecules (or drugs) can regulate miRNA expression, which indicates that targeting miRNAs with small molecules is a new therapy for human diseases. In this study, we established the SM2miR database, which recorded 2925 relationships between 151 small molecules and 747 miRNAs in 17 species after manual curation from nearly 2000 articles. Each entry contains the detailed information about small molecules, miRNAs and evidences of their relationships, such as species, miRBase Accession number, DrugBank Accession number, PubChem Compound Identifier (CID), expression pattern of miRNA, experimental method, tissues or conditions for detection. SM2miR database has a user-friendly interface to retrieve by miRNA or small molecule. In addition, we offered a submission page. Thus, SM2miR provides a fairly comprehensive repository about the influences of small molecules on miRNA expression, which will promote the development of miRNA therapeutics. AVAILABILITY: SM2miR is freely available at http://bioinfo.hrbmu.edu.cn/SM2miR/.
Shuyuan Wang, Fan-Lin Meng, Jizhe Wang, Enyu Dai, Xuexin Yu, Xia Li 0004, Wei Jiang 0023
Bioinform.8
2012 Dissection of human MiRNA regulatory influence to subpathway
abstract
The global insight into the relationships between miRNAs and their regulatory influences remains poorly understood. And most of complex diseases may be attributed to certain local areas of pathway (subpathway) instead of the entire pathway. Here, we reviewed the studies on miRNA regulations to pathways and constructed a bipartite miRNAs and subpathways network for systematic analyzing the miRNA regulatory influences to subpathways. We found that a small fraction of miRNAs were global regulators, environmental information processing pathways were preferentially regulated by miRNAs, and miRNAs had synergistic effect on regulating group of subpathways with similar function. Integrating the disease states of miRNAs, we also found that disease miRNAs regulated more subpathways than nondisease miRNAs, and for all miRNAs, the number of regulated subpathways was not in proportion to the number of the related diseases. Therefore, the study not only provided a global view on the relationships among disease, miRNA and subpathway, but also uncovered the function aspects of miRNA regulations and potential pathogenesis of complex diseases. A web server to query, visualize and download for all the data can be freely accessed at http://bioinfo.hrbmu.edu.cn/miR2Subpath.
Xia Li 0004, Wei Jiang 0023, Baofeng Lian, Shuyuan Wang, Mingzhi Liao, Yanqiu Wang, Yingli Lv
Briefings Bioinform.1
2011 A sub-pathway-based approach for identifying drug response principal network
abstract
MOTIVATION: The high redundancy of and high degree of cross-talk between biological pathways hint that a sub-pathway may respond more effectively or sensitively than the whole pathway. However, few current pathway enrichment analysis methods account for the sub-pathways or structures of the tested pathways. We present a sub-pathway-based enrichment approach for identifying a drug response principal network, which takes into consideration the quantitative structures of the pathways. RESULT: We validated this new approach on a microarray experiment that captures the transcriptional profile of dexamethasone (DEX)-treated human prostate cancer PC3 cells. Compared with GeneTrail and DAVID, our approach is more sensitive to the DEX response pathways. Specifically, not only pathways but also the principal components of sub-pathways and networks related to prostate cancer and DEX response could be identified and verified by literature retrieval.
Xiujie Chen, Jiankai Xu, Bangqing Huang, Jin Li 0024, Xiusen Bian, Fujian Tan, Xia Li 0004
Bioinform.12
2011 A novel network-based method for measuring the functional relationship between gene sets
abstract
MOTIVATION: In the functional genomic era, a large number of gene sets have been identified via high-throughput genomic and proteomic technologies. These gene sets of interest are often related to the same or similar disorders or phenotypes, and are commonly presented as differentially expressed gene lists, co-expressed gene modules, protein complexes or signaling pathways. However, biologists are still faced by the challenge of comparing gene sets and interpreting the functional relationships between gene sets into an understanding of the underlying biological mechanisms. RESULTS: We introduce a novel network-based method, designated corrected cumulative rank score (CCRS), which analyzes the functional communication and physical interaction between genes, and presents an easy-to-use web-based toolkit called GsNetCom to quantify the functional relationship between two gene sets. To evaluate the performance of our method in assessing the functional similarity between two gene sets, we analyzed the functional coherence of complexes in functional catalog and identified protein complexes in the same functional catalog. The results suggested that CCRS can offer a significant advance in addressing the functional relationship between different gene sets compared with several other available tools or algorithms with similar functionality. We also conducted the case study based on our method, and succeeded in prioritizing candidate leukemia-associated protein complexes and expanding the prioritization and analysis of cancer-related complexes to other cancer types. In addition, GsNetCom provides a new insight into the communication between gene modules, such as exploring gene sets from the perspective of well-annotated protein complexes. AVAILABILITY AND IMPLEMENTATION: GsNetCom is a freely available web accessible toolkit at http://bioinfo.hrbmu.edu.cn/GsNetCom.
Qianghu Wang, Jie Sun 0021, Meng Zhou 0003, Haixiu Yang, Sali Lv, Xia Li 0004
Bioinform.8
2011 DOSim: An R package for similarity between diseases based on Disease Ontology
abstract
BACKGROUND: The construction of the Disease Ontology (DO) has helped promote the investigation of diseases and disease risk factors. DO enables researchers to analyse disease similarity by adopting semantic similarity measures, and has expanded our understanding of the relationships between different diseases and to classify them. Simultaneously, similarities between genes can also be analysed by their associations with similar diseases. As a result, disease heterogeneity is better understood and insights into the molecular pathogenesis of similar diseases have been gained. However, bioinformatics tools that provide easy and straight forward ways to use DO to study disease and gene similarity simultaneously are required. RESULTS: We have developed an R-based software package (DOSim) to compute the similarity between diseases and to measure the similarity between human genes in terms of diseases. DOSim incorporates a DO-based enrichment analysis function that can be used to explore the disease feature of an independent gene set. A multilayered enrichment analysis (GO and KEGG annotation) annotation function that helps users explore the biological meaning implied in a newly detected gene module is also part of the DOSim package. We used the disease similarity application to demonstrate the relationship between 128 different DO cancer terms. The hierarchical clustering of these 128 different cancers showed modular characteristics. In another case study, we used the gene similarity application on 361 obesity-related genes. The results revealed the complex pathogenesis of obesity. In addition, the gene module detection and gene module multilayered annotation functions in DOSim when applied on these 361 obesity-related genes helped extend our understanding of the complex pathogenesis of obesity risk phenotypes and the heterogeneity of obesity-related diseases. CONCLUSIONS: DOSim can be used to detect disease-driven gene modules, and to annotate the modules for functions and pathways. The DOSim package can also be used to visualise DO structure. DOSim can reflect the modular characteristic of disease related genes and promote our understanding of the complex pathogenesis of diseases. DOSim is available on the Comprehensive R Archive Network (CRAN) or http://bioinfo.hrbmu.edu.cn/dosim.
Binsheng Gong, Tao Liu 0031, Chunquan Li 0002, Shaoqi Rao, Xia Li 0004
BMC Bioinform.10
2009 Prioritizing risk pathways: a novel association approach to searching for disease pathways fusing SNPs and pathways
abstract
Abstract Motivation: Complex diseases are generally thought to be under the influence of one or more mutated risk genes as well as genetic and environmental factors. Many traditional methods have been developed to identify susceptibility genes assuming a single-gene disease model (‘single-locus methods’). Pathway-based approaches, combined with traditional methods, consider the joint effects of genetic factor and biologic network context. With the accumulation of high-throughput SNP datasets and human biologic pathways, it becomes feasible to search for risk pathways associated with complex diseases using bioinformatics methods. By analyzing the contribution of genetic factor and biologic network context in KEGG (Kyoto Encyclopedia of Genes and Genomes) pathways, we proposed an approach to prioritize risk pathways for complex diseases: Prioritizing Risk Pathways fusing SNPs and pathways (PRP). A risk-scoring (RS) measurement was used to prioritize risk biologic pathways. This could help to demonstrate the pathogenesis of complex diseases from a new perspective and provide new hypotheses. We introduced this approach to five complex diseases and found that these five diseases not only share common risk pathways, but also have their specific risk pathways, which is verified by literature retrieval. Availability: Genotype frequencies of five case–control samples were downloaded from the WTCCC online system and the address is https://www.wtccc.org.uk/info/access_to_data_samples.shtml Contact: [email protected]; [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Lina Chen, Liangcai Zhang, Liangde Xu, Yukui Shang, Xia Li 0004
Bioinform.9
2008 Towards patterns tree of gene coexpression in eukaryotic species
abstract
MOTIVATION: Cellular pathways behave coordinated regulation activity, and some reported works also have affirmed that genes in the same pathway have similar expression pattern. However, the complexity of biological systems regulation actually causes expression relationships between genes to display multiple patterns, such as linear, non-linear, local, global, linear with time-delayed, non-linear with time-delayed, monotonic and non-monotonic, which should be the explicit representation of cellular inner regulation mechanism in mRNA level. To investigate the relationship between different patterns, our work aims to systematically reveal gene-expression relationship patterns in cellular pathways and to check for the existence of dominating gene-expression pattern. By a large scale analysis of genes expression in three eukaryotic species, Saccharomyces cerevisiae, Caenorhabditis elegans and Human, we constructed gene coexpression patterns tree to systematically and hierarchically illustrate the different patterns and their interrelations. RESULTS: The results show that the linear is the dominating expression pattern in the same pathway. The time-shifted pattern is another important relationship pattern. Many genes from the different pathway also present coexpression patterns. The non-linear, non-monotonic and time-delayed relationship patterns reflect the remote interactions between the genes in cellular processes. Gene coexpression phenomena in the same pathways are diverse in different species. Genes in S.cerevisiae and C.elegans present strong coexpression relationships, especially in C.elegans, coexpression is more universal and stronger due to its special array of genes. However in Human, gene coexpression is not apparent and the human genome involves more complicated functional relationships. In conclusion, different patterns corresponding to different coordinating behaviors coexist. The patterns trees of different species give us comprehensive insight and understanding of genes expression activity in the cellular society.
Xia Li 0004, Bairong Shen, Min Ding 0006, Ziyin Shen
Bioinform.3
2008 Apparently low reproducibility of true differential expression discoveries in microarray studies
abstract
MOTIVATION: Differentially expressed gene (DEG) lists detected from different microarray studies for a same disease are often highly inconsistent. Even in technical replicate tests using identical samples, DEG detection still shows very low reproducibility. It is often believed that current small microarray studies will largely introduce false discoveries. RESULTS: Based on a statistical model, we show that even in technical replicate tests using identical samples, it is highly likely that the selected DEG lists will be very inconsistent in the presence of small measurement variations. Therefore, the apparently low reproducibility of DEG detection from current technical replicate tests does not indicate low quality of microarray technology. We also demonstrate that heterogeneous biological variations existing in real cancer data will further reduce the overall reproducibility of DEG detection. Nevertheless, in small subsamples from both simulated and real data, the actual false discovery rate (FDR) for each DEG list tends to be low, suggesting that each separately determined list may comprise mostly true DEGs. Rather than simply counting the overlaps of the discovery lists from different studies for a complex disease, novel metrics are needed for evaluating the reproducibility of discoveries characterized with correlated molecular changes. Supplementaty information: Supplementary data are available at Bioinformatics online.
Min Zhang 0009, Zheng Guo 0002, Jinfeng Zou, Lin Zhang 0057, Dong Wang 0011, Da Yang 0003, Jing Zhu 0004, Xia Li 0004
Bioinform.12
2006 A Novel Method for Expanding Current Annotations in Gene Ontology
Dapeng Hao, Xia Li 0004, Lei Du 0002, Liangde Xu, Jiankai Xu, Shaoqi Rao
ICIC (3)2
2006 Analysis of Sib-Pair IBD Profiles Using Ensemble Decision Tree Approach: Application to Alcoholism
Qingpu Zhang, Xia Li 0004, Lei Du 0002, Wei Jiang 0023, Jing Li 0115, Shaoqi Rao
ICIC (3)3
2006 Association Research on Potassium Channel Subtypes and Functional Sites
Xia Li 0004, Shaoqi Rao, Wei Jiang 0023, Chuanxing Li
ICIC (3)2
2006 An Analysis of Gene Expression Relationships Between Periodically Expressed Genes in the Hela Cells
Yun Xiao 0001, Xia Li 0004, Shaoqi Rao, Lei Du 0002
ICIC (3)2
2006 Effects of replacing the unreliable cDNA microarray measurements on the disease classification based on gene expression profiles and functional modules
abstract
MOTIVATION: Microarrays datasets frequently contain a large number of missing values (MVs), which need to be estimated and replaced for subsequent data mining. The focus of the paper is to study the effects of different MV treatments for cDNA microarray data on disease classification analysis. RESULTS: By analyzing five datasets, we demonstrate that among three kinds of classifiers evaluated in this study, support vector machine (SVM) classifiers are robust to varied MV imputation methods [e.g. replacing MVs by zero, K nearest-neighbor (KNN) imputation algorithm, local least square imputation and Bayesian principal component analysis], while the classification and regression tree classifiers are sensitive in terms of classification accuracy. The KNNclassifiers built on differentially expressed genes (DEGs) are robust to the varied MV treatments, but the performances of the KNN classifiers based on all measured genes can be significantly deteriorated when imputing MVs for genes with larger missing rate (MR) (e.g. MR > 5%). Generally, while replacing MVs by zero performs relatively poor, the other imputation algorithms have little difference in affecting classification performances of the SVM or KNN classifiers. We further demonstrate the power and feasibility of our recently proposed functional expression profile (FEP) approach as means to handle microarray data with MVs. The FEPs, which are derived from the functional modules that are enriched with sets of DEGs and thus can be consistently identified under varied MV treatments, achieve precise disease classification with better biological interpretation. We conclude that the choice of MV treatments should be determined in context of the later approaches used for disease classification. The suggested exclusion criterion of ignoring the genes with larger MR (e.g. >5%), while justifiable for some classifiers such as KNN classifiers, might not be considered as a general rule for all classifiers.
Dong Wang 0011, Yingli Lv, Zheng Guo 0002, Xia Li 0004, Jing Zhu 0004, Da Yang 0003, Jianzhen Xu, Chenguang Wang 0004, Shaoqi Rao, Baofeng Yang
Bioinform.4
2006 Discovery of time-delayed gene regulatory networks based on temporal gene expression profiling
abstract
BACKGROUND: It is one of the ultimate goals for modern biological research to fully elucidate the intricate interplays and the regulations of the molecular determinants that propel and characterize the progression of versatile life phenomena, to name a few, cell cycling, developmental biology, aging, and the progressive and recurrent pathogenesis of complex diseases. The vast amount of large-scale and genome-wide time-resolved data is becoming increasing available, which provides the golden opportunity to unravel the challenging reverse-engineering problem of time-delayed gene regulatory networks. RESULTS: In particular, this methodological paper aims to reconstruct regulatory networks from temporal gene expression data by using delayed correlations between genes, i.e., pairwise overlaps of expression levels shifted in time relative each other. We have thus developed a novel model-free computational toolbox termed TdGRN (Time-delayed Gene Regulatory Network) to address the underlying regulations of genes that can span any unit(s) of time intervals. This bioinformatics toolbox has provided a unified approach to uncovering time trends of gene regulations through decision analysis of the newly designed time-delayed gene expression matrix. We have applied the proposed method to yeast cell cycling and human HeLa cell cycling and have discovered most of the underlying time-delayed regulations that are supported by multiple lines of experimental evidence and that are remarkably consistent with the current knowledge on phase characteristics for the cell cyclings. CONCLUSION: We established a usable and powerful model-free approach to dissecting high-order dynamic trends of gene-gene interactions. We have carefully validated the proposed algorithm by applying it to two publicly available cell cycling datasets. In addition to uncovering the time trends of gene regulations for cell cycling, this unified approach can also be used to study the complex gene regulations related to the development, aging and progressive pathogenesis of a complex disease where potential dependences between different experiment units might occurs.
Xia Li 0004, Shaoqi Rao, Wei Jiang 0023, Chuanxing Li, Yun Xiao 0001, Zheng Guo 0002, Qingpu Zhang, Lei Du 0002, Jing Li 0115, Li Li 0090, Tianwen Zhang, Qing K. Wang
BMC Bioinform.1
2005 Towards precise classification of cancers based on robust gene functional expression profiles
abstract
BACKGROUND: Development of robust and efficient methods for analyzing and interpreting high dimension gene expression profiles continues to be a focus in computational biology. The accumulated experiment evidence supports the assumption that genes express and perform their functions in modular fashions in cells. Therefore, there is an open space for development of the timely and relevant computational algorithms that use robust functional expression profiles towards precise classification of complex human diseases at the modular level. RESULTS: Inspired by the insight that genes act as a module to carry out a highly integrated cellular function, we thus define a low dimension functional expression profile for data reduction. After annotating each individual gene to functional categories defined in a proper gene function classification system such as Gene Ontology applied in this study, we identify those functional categories enriched with differentially expressed genes. For each functional category or functional module, we compute a summary measure (s) for the raw expression values of the annotated genes to capture the overall activity level of the module. In this way, we can treat the gene expressions within a functional module as an integrative data point to replace the multiple values of individual genes. We compare the classification performance of decision trees based on functional expression profiles with the conventional gene expression profiles using four publicly available datasets, which indicates that precise classification of tumour types and improved interpretation can be achieved with the reduced functional expression profiles. CONCLUSION: This modular approach is demonstrated to be a powerful alternative approach to analyzing high dimension microarray data and is robust to high measurement noise and intrinsic biological variance inherent in microarray data. Furthermore, efficient integration with current biological knowledge has facilitated the interpretation of the underlying molecular mechanisms for complex human diseases at the modular level.
Zheng Guo 0002, Tianwen Zhang, Xia Li 0004, Jianzhen Xu, Jing Zhu 0004, Chenguang Wang 0004, Eric J. Topol, Shaoqi Rao
BMC Bioinform.3
2004 Classification of Cancer Types Based on Decision Tree Analysis of Gene Function Expression Profiles
Zheng Guo 0002, Tianwen Zhang, Shaoqi Rao, Xia Li 0004
SNPD6
2004 Feature Gene Selection Based on a Hybrid between Genetic Algorithm and Support Vector Machine
Li Li 0090, Wei Jiang 0023, Wu Linna, Shaoqi Rao, Zheng Guo 0002, Xia Li 0004
SNPD6
2004 Reverse Engineering of Multiple Time-delayed Gene Regulatory Networks
Xia Li 0004, Tianwen Zhang, Wei Jiang 0023, Shaoqi Rao, Li Li 0090, Zheng Guo 0002
SNPD1