EDBT 2026 Demo / reviewers in the wild / expert
Li Chen 0029
dblp:181/2847-29
· DBLP profile ↗
23ranked-venue papers
8as first author
14since 2021 · last 2025
0000-0001-9372-5606ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 23 · 8 first-author · 14 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Bayesian framework for genome-wide circadian rhythmicity biomarker detectionabstractCircadian rhythms are endogenous $\sim $24-h cycles that significantly influence physiological and behavioral processes. These rhythms are governed by a transcriptional-translational feedback loop of core circadian genes and are essential for maintaining overall health. The study of circadian rhythms has expanded into various omics datasets, necessitating accurate analytical methodology for circadian biomarker detection. Here, we introduce a novel Bayesian framework for the detection of circadian rhythms in genome-wide transcriptomic applications that is capable of incorporating prior biological knowledge and adjusting for multiple testing issue via a false discovery rate (FDR) approach. Our framework leverages a Bayesian hierarchical model and employs a reverse jump Markov chain Monte Carlo technique for model selection. Through extensive simulations, our method, BayesCircRhy, demonstrated favorable FDR control over competing methods, robustness against heavier-tailed error distributions, and better performance compared with existing approaches. The method's efficacy was further validated in two RNA-sequencing data, including a human-restricted feeding data and a mouse aging data, where it successfully identified known and novel circadian genes. Haocheng Ding, Lingsong Meng, Andrew J. Bryant, Chengguo Xing, Karyn A. Esser, Li Chen 0029, Yitong Feng, Zhiguang Huo |
Briefings Bioinform. | 7 |
| 2025 | A novel Bayesian hierarchical model for detecting differential circadian pattern in transcriptomic applicationsabstractCircadian rhythm plays a critical role in regulating various physiological processes, and disruptions in these rhythms have been linked to a wide range of diseases. Identifying molecular biomarkers showing differential circadian (DC) patterns between biological conditions or disease status is important for disease prevention, diagnosis, and treatment. However, circadian pattern is characterized by three key components: amplitude, phase, and MESOR, which poses a great challenge for DC analysis. Existing statistical methods focus on detecting differential shape (amplitude and phase) but often overlook MESOR difference. Additionally, these methods lack flexibility to incorporate external knowledge such as differential circadian information from similar clinical and biological context to improve the current DC analysis. To address these limitation, we introduce a novel Bayesian hierarchical model, BayesDCirc, designed for detecting differential circadian patterns in a two-group experimental design, which offer the advantage of testing MESOR difference and incorporating external knowledge. Benefiting from explicitly testing MESOR within the Bayesian modeling framework, BayesDCirc demonstrates superior FDR control over existing methods, with further performance improvement by leveraging external knowledge of DC genes. Applied to two real datasets, BayesDCirc successfully identify key circadian genes, particularly with external knowledge incorporated. The R package "BayesDCirc" for the method is publicly available on GitHub at https://github.com/lichen-lab/BayesDCirc. Haocheng Ding, Zhiguang Huo, Li Chen 0029 |
Briefings Bioinform. | 4 |
| 2024 | AIGen: an artificial intelligence software for complex genetic data analysisabstractThe recent development of artificial intelligence (AI) technology, especially the advance of deep neural network (DNN) technology, has revolutionized many fields. While DNN plays a central role in modern AI technology, it has rarely been used in genetic data analysis due to analytical and computational challenges brought by high-dimensional genetic data and an increasing number of samples. To facilitate the use of AI in genetic data analysis, we developed a C++ package, AIGen, based on two newly developed neural networks (i.e. kernel neural networks and functional neural networks) that are capable of modeling complex genotype-phenotype relationships (e.g. interactions) while providing robust performance against high-dimensional genetic data. Moreover, computationally efficient algorithms (e.g. a minimum norm quadratic unbiased estimation approach and batch training) are implemented in the package to accelerate the computation, making them computationally efficient for analyzing large-scale datasets with thousands or even millions of samples. By applying AIGen to the UK Biobank dataset, we demonstrate that it can efficiently analyze large-scale genetic data, attain improved accuracy, and maintain robust performance. Availability: AIGen is developed in C++ and its source code, along with reference libraries, is publicly accessible on GitHub at https://github.com/TingtHou/AIGen. Tingting Hou, Xiaoxi Shen, Muxuan Liang, Li Chen 0029, Qing Lu 0004 |
Briefings Bioinform. | 5 |
| 2024 | Multimodal functional deep learning for multiomics dataabstractWith rapidly evolving high-throughput technologies and consistently decreasing costs, collecting multimodal omics data in large-scale studies has become feasible. Although studying multiomics provides a new comprehensive approach in understanding the complex biological mechanisms of human diseases, the high dimensionality of omics data and the complexity of the interactions among various omics levels in contributing to disease phenotypes present tremendous analytical challenges. There is a great need of novel analytical methods to address these challenges and to facilitate multiomics analyses. In this paper, we propose a multimodal functional deep learning (MFDL) method for the analysis of high-dimensional multiomics data. The MFDL method models the complex relationships between multiomics variants and disease phenotypes through the hierarchical structure of deep neural networks and handles high-dimensional omics data using the functional data analysis technique. Furthermore, MFDL leverages the structure of the multimodal model to capture interactions between different types of omics data. Through simulation studies and real-data applications, we demonstrate the advantages of MFDL in terms of prediction accuracy and its robustness to the high dimensionality and noise within the data. Pei Geng, Feifei Xiao, Guoshuai Cai, Li Chen 0029, Qing Lu 0004 |
Briefings Bioinform. | 6 |
| 2024 | Deep5hmC: predicting genome-wide 5-hydroxymethylcytosine landscape via a multimodal deep learning modelabstractMOTIVATION: 5-Hydroxymethylcytosine (5hmC), a crucial epigenetic mark with a significant role in regulating tissue-specific gene expression, is essential for understanding the dynamic functions of the human genome. Despite its importance, predicting 5hmC modification across the genome remains a challenging task, especially when considering the complex interplay between DNA sequences and various epigenetic factors such as histone modifications and chromatin accessibility. RESULTS: Using tissue-specific 5hmC sequencing data, we introduce Deep5hmC, a multimodal deep learning framework that integrates both the DNA sequence and epigenetic features such as histone modification and chromatin accessibility to predict genome-wide 5hmC modification. The multimodal design of Deep5hmC demonstrates remarkable improvement in predicting both qualitative and quantitative 5hmC modification compared to unimodal versions of Deep5hmC and state-of-the-art machine learning methods. This improvement is demonstrated through benchmarking on a comprehensive set of 5hmC sequencing data collected at four developmental stages during forebrain organoid development and across 17 human tissues. Compared to DeepSEA and random forest, Deep5hmC achieves close to 4% and 17% improvement of Area Under the Receiver Operating Characteristic (AUROC) across four forebrain developmental stages, and 6% and 27% across 17 human tissues for predicting binary 5hmC modification sites; and 8% and 22% improvement of Spearman correlation coefficient across four forebrain developmental stages, and 17% and 30% across 17 human tissues for predicting continuous 5hmC modification. Notably, Deep5hmC showcases its practical utility by accurately predicting gene expression and identifying differentially hydroxymethylated regions (DhMRs) in a case-control study of Alzheimer's disease (AD). Deep5hmC significantly improves our understanding of tissue-specific gene regulation and facilitates the development of new biomarkers for complex diseases. AVAILABILITY AND IMPLEMENTATION: Deep5hmC is available via https://github.com/lichen-lab/Deep5hmC. Xin Ma 0028, Sai Ritesh Thela, Fengdi Zhao, Zhexing Wen, Jinying Zhao, Li Chen 0029 |
Bioinform. | 8 |
| 2024 | MPRAVarDB: an online database and web server for exploring regulatory effects of genetic variantsabstractSUMMARY: Massively parallel reporter assay (MPRA) is an important technology for evaluating the impact of genetic variants on gene regulation. Here, we present MPRAVarDB, an online database and web server for exploring regulatory effects of genetic variants. MPRAVarDB harbors 18 MPRA experiments designed to assess the regulatory effects of genetic variants associated with GWAS loci, eQTLs, and genomic features, totaling 242 818 variants tested more than 30 cell lines and 30 human diseases or traits. MPRAVarDB enables users to query MPRA variants by genomic region, disease and cell line, or any combination of these parameters. Notably, MPRAVarDB offers a suite of pretrained machine-learning models tailored to the specific disease and cell line, facilitating the prediction of regulatory variants. The user-friendly interface allows users to receive query and prediction results with just a few clicks. AVAILABILITY AND IMPLEMENTATION: https://mpravardb.rc.ufl.edu. Weijia Jin, Javlon Nizomov, Qing Lu 0004, Li Chen 0029 |
Bioinform. | 7 |
| 2024 | scaDA: A novel statistical method for differential analysis of single-cell chromatin accessibility sequencing dataabstractSingle-cell ATAC-seq sequencing data (scATAC-seq) has been widely used to investigate chromatin accessibility on the single-cell level. One important application of scATAC-seq data analysis is differential chromatin accessibility (DA) analysis. However, the data characteristics of scATAC-seq such as excessive zeros and large variability of chromatin accessibility across cells impose a unique challenge for DA analysis. Existing statistical methods focus on detecting the mean difference of the chromatin accessible regions while overlooking the distribution difference. Motivated by real data exploration that distribution difference exists among cell types, we introduce a novel composite statistical test named "scaDA", which is based on zero-inflated negative binomial model (ZINB), for performing differential distribution analysis of chromatin accessibility by jointly testing the abundance, prevalence and dispersion simultaneously. Benefiting from both dispersion shrinkage and iterative refinement of mean and prevalence parameter estimates, scaDA demonstrates its superiority to both ZINB-based likelihood ratio tests and published methods by achieving the highest power and best FDR control in a comprehensive simulation study. In addition to demonstrating the highest power in three real sc-multiome data analyses, scaDA successfully identifies differentially accessible regions in microglia from sc-multiome data for an Alzheimer's disease (AD) study that are most enriched in GO terms related to neurogenesis and the clinical phenotype of AD, and AD-associated GWAS SNPs. Fengdi Zhao, Xin Ma 0028, Qing Lu 0004, Li Chen 0029 |
PLoS Comput. Biol. | 5 |
| 2023 | DeepPHiC: predicting promoter-centered chromatin interactions using a novel deep learning approachabstractMOTIVATION: Promoter-centered chromatin interactions, which include promoter-enhancer (PE) and promoter-promoter (PP) interactions, are important to decipher gene regulation and disease mechanisms. The development of next-generation sequencing technologies such as promoter capture Hi-C (pcHi-C) leads to the discovery of promoter-centered chromatin interactions. However, pcHi-C experiments are expensive and thus may be unavailable for tissues/cell types of interest. In addition, these experiments may be underpowered due to insufficient sequencing depth or various artifacts, which results in a limited finding of interactions. Most existing computational methods for predicting chromatin interactions are based on in situ Hi-C and can detect chromatin interactions across the entire genome. However, they may not be optimal for predicting promoter-centered chromatin interactions. RESULTS: We develop a supervised multi-modal deep learning model, which utilizes a comprehensive set of features such as genomic sequence, epigenetic signal, anchor distance, evolutionary features and DNA structural features to predict tissue/cell type-specific PE and PP interactions. We further extend the deep learning model in a multi-task learning and a transfer learning framework and demonstrate that the proposed approach outperforms state-of-the-art deep learning methods. Moreover, the proposed approach can achieve comparable prediction performance using predefined biologically relevant tissues/cell types compared to using all tissues/cell types in the pretraining especially for predicting PE interactions. The prediction performance can be further improved by using computationally inferred biologically relevant tissues/cell types in the pretraining, which are defined based on the common genes in the proximity of two anchors in the chromatin interactions. AVAILABILITY AND IMPLEMENTATION: https://github.com/lichen-lab/DeepPHiC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Aman Agarwal, Li Chen 0029 |
Bioinform. | 2 |
| 2023 | TIVAN-indel: a computational framework for annotating and predicting non-coding regulatory small insertions and deletionsabstractMOTIVATION: Small insertion and deletion (sindel) of human genome has an important implication for human disease. One important mechanism for non-coding sindel (nc-sindel) to have an impact on human diseases and phenotypes is through the regulation of gene expression. Nevertheless, current sequencing experiments may lack statistical power and resolution to pinpoint the functional sindel due to lower minor allele frequency or small effect size. As an alternative strategy, a supervised machine learning method can identify the otherwise masked functional sindels by predicting their regulatory potential directly. However, computational methods for annotating and predicting the regulatory sindels, especially in the non-coding regions, are underdeveloped. RESULTS: By leveraging labeled nc-sindels identified by cis-expression quantitative trait loci analyses across 44 tissues in Genotype-Tissue Expression (GTEx), and a compilation of both generic functional annotations and large-scale epigenomic profiles, we develop TIssue-specific Variant Annotation for Non-coding indel (TIVAN-indel), which is a supervised computational framework for predicting non-coding regulatory sindels. As a result, we demonstrate that TIVAN-indel achieves the best prediction performance in both with-tissue prediction and cross-tissue prediction. As an independent evaluation, we train TIVAN-indel from the 'Whole Blood' tissue in GTEx and test the model using 15 immune cell types from an independent study named Database of Immune Cell Expression. Lastly, we perform an enrichment analysis for both true and predicted sindels in key regulatory regions such as chromatin interactions, open chromatin regions and histone modification sites, and find biologically meaningful enrichment patterns. AVAILABILITY AND IMPLEMENTATION: https://github.com/lichen-lab/TIVAN-indel. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Aman Agarwal, Fengdi Zhao, Li Chen 0029 |
Bioinform. | 4 |
| 2022 | Exploiting deep transfer learning for the prediction of functional non-coding variants using genomic sequenceabstractMOTIVATION: Though genome-wide association studies have identified tens of thousands of variants associated with complex traits and most of them fall within the non-coding regions, they may not be the causal ones. The development of high-throughput functional assays leads to the discovery of experimental validated non-coding functional variants. However, these validated variants are rare due to technical difficulty and financial cost. The small sample size of validated variants makes it less reliable to develop a supervised machine learning model for achieving a whole genome-wide prediction of non-coding causal variants. RESULTS: We will exploit a deep transfer learning model, which is based on convolutional neural network, to improve the prediction for functional non-coding variants (NCVs). To address the challenge of small sample size, the transfer learning model leverages both large-scale generic functional NCVs to improve the learning of low-level features and context-specific functional NCVs to learn high-level features toward the context-specific prediction task. By evaluating the deep transfer learning model on three MPRA datasets and 16 GWAS datasets, we demonstrate that the proposed model outperforms deep learning models without pretraining or retraining. In addition, the deep transfer learning model outperforms 18 existing computational methods in both MPRA and GWAS datasets. AVAILABILITY AND IMPLEMENTATION: https://github.com/lichen-lab/TLVar. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Li Chen 0029, Ye Wang 0024, Fengdi Zhao |
Bioinform. | 1 |
| 2022 | DeepPerVar: a multi-modal deep learning framework for functional interpretation of genetic variants in personal genomeabstractMOTIVATION: Understanding the functional consequence of genetic variants, especially the non-coding ones, is important but particularly challenging. Genome-wide association studies (GWAS) or quantitative trait locus analyses may be subject to limited statistical power and linkage disequilibrium, and thus are less optimal to pinpoint the causal variants. Moreover, most existing machine-learning approaches, which exploit the functional annotations to interpret and prioritize putative causal variants, cannot accommodate the heterogeneity of personal genetic variations and traits in a population study, targeting a specific disease. RESULTS: By leveraging paired whole-genome sequencing data and epigenetic functional assays in a population study, we propose a multi-modal deep learning framework to predict genome-wide quantitative epigenetic signals by considering both personal genetic variations and traits. The proposed approach can further evaluate the functional consequence of non-coding variants on an individual level by quantifying the allelic difference of predicted epigenetic signals. By applying the approach to the ROSMAP cohort studying Alzheimer's disease (AD), we demonstrate that the proposed approach can accurately predict quantitative genome-wide epigenetic signals and in key genomic regions of AD causal genes, learn canonical motifs reported to regulate gene expression of AD causal genes, improve the partitioning heritability analysis and prioritize putative causal variants in a GWAS risk locus. Finally, we release the proposed deep learning model as a stand-alone Python toolkit and a web server. AVAILABILITY AND IMPLEMENTATION: https://github.com/lichen-lab/DeepPerVar. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ye Wang 0024, Li Chen 0029 |
Bioinform. | 2 |
| 2021 | Genome-wide circadian rhythm detection methods: systematic evaluations and practical guidelinesabstractCircadian rhythms are oscillations of behavior, physiology and metabolism in many organisms. Recent advancements in omics technology make it possible for genome-wide profiling of circadian rhythms. Here, we conducted a comprehensive analysis of seven existing algorithms commonly used for circadian rhythm detection. Using gold-standard circadian and non-circadian genes, we systematically evaluated the accuracy and reproducibility of the algorithms on empirical datasets generated from various omics platforms under different experimental designs. We also carried out extensive simulation studies to test each algorithm's robustness to key variables, including sampling patterns, replicates, waveforms, signal-to-noise ratios, uneven samplings and missing values. Furthermore, we examined the distributions of the nominal $P$-values under the null and raised issues with multiple testing corrections using traditional approaches. With our assessment, we provide method selection guidelines for circadian rhythm detection, which are applicable to different types of high-throughput omics data. Wenwen Mei, Zhiwen Jiang, Yang Chen 0063, Li Chen 0029, Aziz Sancar |
Briefings Bioinform. | 4 |
| 2021 | A novel deep learning method for predictive modeling of microbiome dataabstractWith the development and decreasing cost of next-generation sequencing technologies, the study of the human microbiome has become a rapid expanding research field, which provides an unprecedented opportunity in various clinical applications such as drug response predictions and disease diagnosis. It is thus essential and desirable to build a prediction model for clinical outcomes based on microbiome data that usually consist of taxon abundance and a phylogenetic tree. Importantly, all microbial species are not uniformly distributed in the phylogenetic tree but tend to be clustered at different phylogenetic depths. Therefore, the phylogenetic tree represents a unique correlation structure of microbiome, which can be an important prior to improve the prediction performance. However, prediction methods that consider the phylogenetic tree in an efficient and rigorous way are under-developed. Here, we develop a novel deep learning prediction method MDeep (microbiome-based deep learning method) to predict both continuous and binary outcomes. Conceptually, MDeep designs convolutional layers to mimic taxonomic ranks with multiple convolutional filters on each convolutional layer to capture the phylogenetic correlation among microbial species in a local receptive field and maintain the correlation structure across different convolutional layers via feature mapping. Taken together, the convolutional layers with its built-in convolutional filters capture microbial signals at different taxonomic levels while encouraging local smoothing and preserving local connectivity induced by the phylogenetic tree. We use both simulation studies and real data applications to demonstrate that MDeep outperforms competing methods in both regression and binary classifications. Availability and Implementation: MDeep software is available at https://github.com/lichen-lab/MDeep Contact:[email protected]. Ye Wang 0024, Tathagata Bhattacharya, Xiao Qin 0001, Andrew J. Saykin, Li Chen 0029 |
Briefings Bioinform. | 8 |
| 2021 | WEVar: a novel statistical learning framework for predicting noncoding regulatory variantsabstractUnderstanding the functional consequence of noncoding variants is of great interest. Though genome-wide association studies or quantitative trait locus analyses have identified variants associated with traits or molecular phenotypes, most of them are located in the noncoding regions, making the identification of causal variants a particular challenge. Existing computational approaches developed for prioritizing noncoding variants produce inconsistent and even conflicting results. To address these challenges, we propose a novel statistical learning framework, which directly integrates the precomputed functional scores from representative scoring methods. It will maximize the usage of integrated methods by automatically learning the relative contribution of each method and produce an ensemble score as the final prediction. The framework consists of two modes. The first 'context-free' mode is trained using curated causal regulatory variants from a wide range of context and is applicable to predict regulatory variants of unknown and diverse context. The second 'context-dependent' mode further improves the prediction when the training and testing variants are from the same context. By evaluating the framework via both simulation and empirical studies, we demonstrate that it outperforms integrated scoring methods and the ensemble score successfully prioritizes experimentally validated regulatory variants in multiple risk loci. Ye Wang 0024, Xiao Qin 0001, Andrew J. Saykin, Li Chen 0029 |
Briefings Bioinform. | 9 |
| 2020 | powmic: an R package for power assessment in microbiome case-control studiesabstractSUMMARY: Power analysis is essential to decide the sample size of metagenomic sequencing experiments in a case-control study for identifying differentially abundant (DA) microbes. However, the complexity of microbial data characteristics, such as excessive zeros, over-dispersion, compositionality, intrinsically microbial correlations and variable sequencing depths, makes the power analysis particularly challenging because the analytical form is usually unavailable. Here, we develop a simulation-based power assessment strategy and R package powmic, which considers the complexity of microbial data characteristics. A real data example demonstrates the usage of powmic. AVAILABILITY AND IMPLEMENTATION: powmic R package and online tutorial are available at https://github.com/lichen-lab/powmic. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Li Chen 0029 |
Bioinform. | 1 |
| 2020 | circMeta: a unified computational framework for genomic feature annotation and differential expression analysis of circular RNAsabstractMOTIVATION: Circular RNAs (circRNAs), a class of non-coding RNAs generated from non-canonical back-splicing events, have emerged to play key roles in many biological processes. Though numerous tools have been developed to detect circRNAs from rRNA-depleted RNA-seq data based on back-splicing junction-spanning reads, computational tools to identify critical genomic features regulating circRNA biogenesis are still lacking. In addition, rigorous statistical methods to perform differential expression (DE) analysis of circRNAs remain under-developed. RESULTS: We present circMeta, a unified computational framework for circRNA analyses. circMeta has three primary functional modules: (i) a pipeline for comprehensive genomic feature annotation related to circRNA biogenesis, including length of introns flanking circularized exons, repetitive elements such as Alu elements and SINEs, competition score for forming circulation and RNA editing in back-splicing flanking introns; (ii) a two-stage DE approach of circRNAs based on circular junction reads to quantitatively compare circRNA levels and (iii) a Bayesian hierarchical model for DE analysis of circRNAs based on the ratio of circular reads to linear reads in back-splicing sites to study spatial and temporal regulation of circRNA production. Both proposed DE methods without and with considering host genes outperform existing methods by obtaining better control of false discovery rate and comparable statistical power. Moreover, the identified DE circRNAs by the proposed two-stage DE approach display potential biological functions in Gene Ontology and circRNA-miRNA-mRNA networks that are not able to be detected using existing mRNA DE methods. Furthermore, top DE circRNAs have been further validated by RT-qPCR using divergent primers spanning back-splicing junctions. AVAILABILITY AND IMPLEMENTATION: The software circMeta is freely available at https://github.com/lichen-lab/circMeta. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Li Chen 0029, Feng Wang 0078, Emily C. Bruggeman |
Bioinform. | 1 |
| 2020 | Application of topic models to a compendium of ChIP-Seq datasets uncovers recurrent transcriptional regulatory modulesabstractMOTIVATION: The availability of thousands of genome-wide coupling chromatin immunoprecipitation (ChIP)-Seq datasets across hundreds of transcription factors (TFs) and cell lines provides an unprecedented opportunity to jointly analyze large-scale TF-binding in vivo, making possible the discovery of the potential interaction and cooperation among different TFs. The interacted and cooperated TFs can potentially form a transcriptional regulatory module (TRM) (e.g. co-binding TFs), which helps decipher the combinatorial regulatory mechanisms. RESULTS: We develop a computational method tfLDA to apply state-of-the-art topic models to multiple ChIP-Seq datasets to decipher the combinatorial binding events of multiple TFs. tfLDA is able to learn high-order combinatorial binding patterns of TFs from multiple ChIP-Seq profiles, interpret and visualize the combinatorial patterns. We apply the tfLDA to two cell lines with a rich collection of TFs and identify combinatorial binding patterns that show well-known TRMs and related TF co-binding events. AVAILABILITY AND IMPLEMENTATION: A software R package tfLDA is freely available at https://github.com/lichen-lab/tfLDA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Aiqun Ma, Zhaohui S. Qin, Li Chen 0029 |
Bioinform. | 4 |
| 2019 | TIVAN: tissue-specific cis-eQTL single nucleotide variant annotation and predictionabstractSUMMARY: Predicting genetic regulatory variants, most of which locate in non-coding genomic regions, still remain a challenge in genetic research. Among all non-coding regulatory variants, cis-eQTL single nucleotide variants (SNVs) are of particular interest for their crucial role in regulating gene expression. Since different gene expression patterns are believed to contribute to the etiologies of different phenotypes, it is desirable to characterize the impact of cis-eQTL SNVs in a context-specific manner. Though computational methods for predicting the potential of variants being pathogenic or deleterious are well-established, methods for annotating and predicting cis-eQTL SNVs are under-developed. Here, we present TIVAN (TIssue-specific Variant ANnotation and prediction), an ensemble method of decision trees, to predict tissue-specific cis-eQTL SNVs. TIVAN is trained based on a comprehensive collection of features, including genome-wide genomic and epigenomic profiling data. As a result, TIVAN has been shown to accurately discriminate cis-eQTL SNVs from non-eQTL SNVs and perform favorably to other methods by obtaining higher five-fold cross-validation AUC values (CV-AUC) and Leave-One-Chromosome-Out predicted AUC values (LOCO-AUC) across 44 different tissues belonging to 27 different tissue classes. Finally, TIVAN consistently maintains top performance on an independent testing dataset, which includes 7 tissues in 11 studies. AVAILABILITY AND IMPLEMENTATION: TIVAN software is available at https://github.com/lichen-lab/TIVAN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Li Chen 0029, Ye Wang 0024, Amit Mitra, Xu Wang 0026, Xiao Qin 0001 |
Bioinform. | 1 |
| 2019 | Destin: toolkit for single-cell analysis of chromatin accessibilityabstractSUMMARY: Single-cell assay of transposase-accessible chromatin followed by sequencing (scATAC-seq) is an emerging new technology for the study of gene regulation with single-cell resolution. The data from scATAC-seq are unique-sparse, binary and highly variable even within the same cell type. As such, neither methods developed for bulk ATAC-seq nor single-cell RNA-seq data are appropriate. Here, we present Destin, a bioinformatic and statistical framework for comprehensive scATAC-seq data analysis. Destin performs cell-type clustering via weighted principle component analysis, weighting accessible chromatin regions by existing genomic annotations and publicly available regulomic datasets. The weights and additional tuning parameters are determined via model-based likelihood. We evaluated the performance of Destin using downsampled bulk ATAC-seq data of purified samples and scATAC-seq data from seven diverse experiments. Compared to existing methods, Destin was shown to outperform across all datasets and platforms. For demonstration, we further applied Destin to 2088 adult mouse forebrain cells and identified cell-type-specific association of previously reported schizophrenia GWAS loci. AVAILABILITY AND IMPLEMENTATION: Destin toolkit is freely available as an R package at https://github.com/urrutiag/destin. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Eugene Urrutia, Li Chen 0029 |
Bioinform. | 2 |
| 2016 | traseR: an R package for performing trait-associated SNP enrichment analysis in genomic intervalsabstractUNLABELLED: Genome-wide association studies (GWASs) have successfully identified many sequence variants that are significantly associated with common diseases and traits. Tens of thousands of such trait-associated SNPs have already been cataloged, which we believe form a great resource for genomic research. Recent studies have demonstrated that the collection of trait-associated SNPs can be exploited to indicate whether a given genomic interval or intervals are likely to be functionally connected with certain phenotypes or diseases. Despite this importance, currently, there is no ready-to-use computational tool able to connect genomic intervals to phenotypes. Here, we present traseR, an easy-to-use R Bioconductor package that performs enrichment analyses of trait-associated SNPs in arbitrary genomic intervals with flexible options, including testing method, type of background and inclusion of SNPs in LD. AVAILABILITY AND IMPLEMENTATION: The traseR R package preloaded with up-to-date collection of trait-associated SNPs are freely available in Bioconductor CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Li Chen 0029, Zhaohui S. Qin |
Bioinform. | 1 |
| 2015 | glmgraph: an R package for variable selection and predictive modeling of structured genomic dataabstractUNLABELLED: One central theme of modern high-throughput genomic data analysis is to identify relevant genomic features as well as build up a predictive model based on selected features for various tasks such as personalized medicine. Correlating the large number of 'omics' features with a certain phenotype is particularly challenging due to small sample size (n) and high dimensionality (p). To address this small n, large p problem, various forms of sparse regression models have been proposed by exploiting the sparsity assumption. Among these, network-constrained sparse regression model is of particular interest due to its ability to utilize the prior graph/network structure in the omics data. Despite its potential usefulness for omics data analysis, no efficient R implementation is publicly available. Here we present an R software package 'glmgraph' that implements the graph-constrained regularization for both sparse linear regression and sparse logistic regression. We implement both the L1 penalty and minimax concave penalty for variable selection and Laplacian penalty for coefficient smoothing. Efficient coordinate descent algorithm is used to solve the optimization problem. We demonstrate the use of the package by applying it to a human microbiome dataset, where phylogeny structure among bacterial taxa is available. AVAILABILITY AND IMPLEMENTATION: 'glmgraph' is implemented in R and C++ Armadillo and publicly available under CRAN. Li Chen 0029, Han Liu 0001, Jean-Pierre A. Kocher, Hongzhe Li |
Bioinform. | 1 |
| 2015 | A novel statistical method for quantitative comparison of multiple ChIP-seq datasetsabstractMOTIVATION: ChIP-seq is a powerful technology to measure the protein binding or histone modification strength in the whole genome scale. Although there are a number of methods available for single ChIP-seq data analysis (e.g. 'peak detection'), rigorous statistical method for quantitative comparison of multiple ChIP-seq datasets with the considerations of data from control experiment, signal to noise ratios, biological variations and multiple-factor experimental designs is under-developed. RESULTS: In this work, we develop a statistical method to perform quantitative comparison of multiple ChIP-seq datasets and detect genomic regions showing differential protein binding or histone modification. We first detect peaks from all datasets and then union them to form a single set of candidate regions. The read counts from IP experiment at the candidate regions are assumed to follow Poisson distribution. The underlying Poisson rates are modeled as an experiment-specific function of artifacts and biological signals. We then obtain the estimated biological signals and compare them through the hypothesis testing procedure in a linear model framework. Simulations and real data analyses demonstrate that the proposed method provides more accurate and robust results compared with existing ones. AVAILABILITY AND IMPLEMENTATION: An R software package ChIPComp is freely available at http://web1.sph.emory.edu/users/hwu30/software/ChIPComp.html. Li Chen 0029, Chi Wang 0003, Zhaohui S. Qin, Hao Wu 0003 |
Bioinform. | 1 |
| 2011 | hmChIP: a database and web server for exploring publicly available human and mouse ChIP-seq and ChIP-chip dataabstractUNLABELLED: hmChIP is a database of genome-wide chromatin immunoprecipitation (ChIP) data in human and mouse. Currently, the database contains 2016 samples from 492 ChIP-seq and ChIP-chip experiments, representing a total of 170 proteins and 11 069 914 protein-DNA interactions. A web server provides interface for database query. Protein-DNA binding intensities can be retrieved from individual samples for user-provided genomic regions. The retrieved intensities can be used to cluster samples and genomic regions to facilitate exploration of combinatorial patterns, cell-type dependencies, and cross-sample variability of protein-DNA interactions. AVAILABILITY: http://jilab.biostat.jhsph.edu/database/cgi-bin/hmChIP.pl. Li Chen 0029, George Wu, Hongkai Ji |
Bioinform. | 1 |