Bingshan Li

dblp:147/6934 · DBLP profile ↗
← Back
18ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0003-2129-168XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 18 · 8 since 2021
YearPublicationVenuePosition
2025 Leveraging scHi-C Data for Integrated Single-Cell Omics Analysis
abstract
The integration of single-cell multi-omics data is essential for deciphering complex gene regulatory programs. While single-cell Hi-C (scHi-C) provides insight into 3D genome architecture, its utilization in multi-omics integration remains underexplored. Here, we use mouse brain single-cell multi-omics integration as a case study to demonstrate the dual utility of scHiC data within a knowledge graph-based integrative framework. First, we use scHi-C as a gene regulatory prior to construct a Hi-C-driven guidance graph. This approach enhances integration of scRNA-seq and scATAC-seq data, resulting in an improved alignment score (FOSCTTM$=0.0356)$. Second, we show that the framework can directly integrate scHi-C as a primary data modality with scRNA-seq. We validate this method on a paired scRNA-seq and scHi-C dataset, achieving 85.0% mapping accuracy, and demonstrate its power in a cross-modal labeltransfer application. This direct integration successfully refines a broader neuronal cluster into finer, distinct hippocampal granule and pyramidal subtypes. Our work presents an effective strategy for leveraging scHi-C data in multi-omics integration, providing a powerful tool to dissect cellular regulatory programs through combining 3D genome organization with gene expression.
Weixin Liu 0001, Rui Chen 0021, Yuting Tan 0005, Xue Zhong, Bingshan Li, Zhijun Yin
BIBM6
2025 Integrating Single Cell RNA Sequencing Data and Protein Embeddings to Infer Cell-Cell Communication in Alzheimer's Disease
abstract
Cell-cell communication (CCC) plays a critical role in the pathogenesis of Alzheimer's disease (AD), yet most computational methods for CCC inference rely exclusively on transcriptomic data, showing low consistency across datasets or methods. In this study, we present a Protein Embedding-Infused Cell Talk (PEICTalk) inference method that integrates singlecell RNA-sequencing (scRNA-seq) data with protein embeddings from the ESM-2 language model to improve the accuracy and interpretability of CCC inference. We applied PEICTalk to two large-scale scRNA-seq datasets, after harmonizing celltype annotations using the SEA-AD taxonomy via MapMyCells. PEICTalk outperformed standard tools such as CellChat and CellPhoneDB, identifying more biologically meaningful ligandreceptor interactions with higher cross-dataset reproducibility. Ablation analysis confirmed the crucial value of protein-level information. This integrative strategy offers a robust foundation for uncovering novel intercellular mechanisms in AD and may serve as a blueprint for future CCC studies.
Yuting Tan 0005, Rui Chen 0021, Anshul Tiwari, Zhexing Wen, Xue Zhong, Zhijun Yin, Bingshan Li
BIBM8
2025 Tensor decomposition of multi-dimensional splicing events across multiple tissues to identify splicing-mediated risk genes associated with complex traits
abstract
Identifying risk genes associated with complex traits remains challenging. Integrating gene expression data with Genome-Wide Association Study (GWAS) through Transcriptome-Wide Association Study (TWAS) methods has discovered candidate risk genes for various complex traits. Splicing, which explains a comparable heritability of complex traits as gene expression, is under-explored due to its multidimensionality. To leverage multiple splicing events in a gene and shared splicing across tissues, we develop Multi-tissue Splicing Gene (MTSG), which employs tensor decomposition and sparse Canonical Correlation Analysis (sCCA) to extract meaningful information from high-dimensional multiple splicing events across multiple tissues. We build MTSG models using GTEx data and apply them to GWAS summary statistics of Alzheimer's disease (AD) (111,326 cases and 677,663 controls) and schizophrenia (SCZ) (36,989 cases and 113,075 controls). We identify 174 and 497 significant splicing-mediated risk genes for AD and SCZ, respectively, at Bonferroni correction. For AD, our results demonstrate significant enrichment of AD related pathways and identify additional AD risk genes not detected in the single-tissue analysis, while preserving most top genes identified in the brain frontal cortex. Consistently, for SCZ, genes identified by our brain-wide MTSG model, built from a cluster of 13 brain tissues, exhibit stronger enrichment in SCZ-relevant genes and MTSG identifies unique SCZ risk genes compared to single-tissue models. These results showcase that our MTSG models capture distinctive splicing events across tissues, which might be overlooked when using single tissue alone. Our MTSG models can be applied to other complex traits to help identify splicing-mediated disease risk genes.
Rui Chen 0021, Hakmook Kang, Yuting Tan 0005, Anshul Tiwari, Zhexing Wen, Xue Zhong, Bingshan Li
PLoS Comput. Biol.9
2024 Improving Genetic Perturbation Response Prediction with an Enhanced Biological Knowledge Graph
abstract
Perturb-seq is a technique that combines scRNA-seq and CRISPR to explore cellular system operations and disease-associated genes, providing profound insights into the mechanisms behind biological processes. Although powerful, such a method is limited by its scalability for its cost-intensive and time-consuming nature, which calls for in silico prediction of genetic perturbation responses. Among all computational methods, GEARS represents the state-of-the-art by explicitly modeling the response of each gene to the perturbed gene, exploiting gene-gene relationships derived from Gene Ontology annotations. However, our evaluation of Gene Ontology annotations indicated that they are insufficient as the sole source of prior knowledge for predicting genetic perturbation responses. Therefore, they cannot fully support predicting genetic perturbation responses. We addressed this gap by constructing an augmented gene ontology network that incorporates extensive knowledge of diseases, drugs, and genes to capture nuanced gene-gene relationships not indicated by Gene Ontology alone. By replacing only the Gene Ontology graph in GEARS, our method outperforms GEARS in both single gene and combinational perturbation predictions. These findings suggest the effectiveness and importance of incorporating finer prior knowledge in predicting genetic perturbation responses, thereby encouraging future works on improving knowledge representation for single-cell perturbation prediction.
Rui Chen 0021, Yuting Tan 0005, Xue Zhong, Bingshan Li, Zhijun Yin
BIBM5
2022 A computational framework to unify orthogonal information in DNA methylation and copy number aberrations in cell-free DNA for early cancer detection
abstract
Cell-free DNA (cfDNA) provides a convenient diagnosis avenue for noninvasive cancer detection. The current methods are focused on identifying circulating tumor DNA (ctDNA)s genomic aberrations, e.g. mutations, copy number aberrations (CNAs) or methylation changes. In this study, we report a new computational method that unifies two orthogonal pieces of information, namely methylation and CNAs, derived from whole-genome bisulfite sequencing (WGBS) data to quantify low tumor content in cfDNA. It implements a Bayes model to enrich ctDNA from WGBS data based on hypomethylation haplotypes, and subsequently, models CNAs for cancer detection. We generated WGBS data in a total of 262 samples, including high-depth (>20×, deduped high mapping quality reads) data in 76 samples with matched triplets (tumor, adjacent normal and cfDNA) and low-depth (~2.5×, deduped high mapping quality reads) data in 186 samples. We identified a total of 54 Mb regions of hypomethylation haplotypes for model building, a vast majority of which are not covered in the HumanMethylation450 arrays. We showed that our model is able to substantially enrich ctDNA reads (tens of folds), with clearly elevated CNAs that faithfully match the CNAs in the paired tumor samples. In the 19 hepatocellular carcinoma cfDNA samples, the estimated enrichment is as high as 16 fold, and in the simulation data, it can achieve over 30-fold enrichment for a ctDNA level of 0.5% with a sequencing depth of 600×. We also found that these hypomethylation regions are also shared among many cancer types, thus demonstrating the potential of our framework for pancancer early detection.
Jiaze An, Jinliang Xing, Bingshan Li
Briefings Bioinform.9
2022 TVAR: assessing tissue-specific functional effects of non-coding variants with deep learning
abstract
MOTIVATION: Analysis of whole-genome sequencing (WGS) for genetics is still a challenge due to the lack of accurate functional annotation of non-coding variants, especially the rare ones. As eQTLs have been extensively implicated in the genetics of human diseases, we hypothesize that rare non-coding variants discovered in WGS play a regulatory role in predisposing disease risk. RESULTS: With thousands of tissue- and cell-type-specific epigenomic features, we propose TVAR. This multi-label learning-based deep neural network predicts the functionality of non-coding variants in the genome based on eQTLs across 49 human tissues in the GTEx project. TVAR learns the relationships between high-dimensional epigenomics and eQTLs across tissues, taking the correlation among tissues into account to understand shared and tissue-specific eQTL effects. As a result, TVAR outputs tissue-specific annotations, with an average AUROC of 0.77 across these tissues. We evaluate TVAR's performance on four complex diseases (coronary artery disease, breast cancer, Type 2 diabetes and Schizophrenia), using TVAR's tissue-specific annotations, and observe its superior performance in predicting functional variants for both common and rare variants, compared with five existing state-of-the-art tools. We further evaluate TVAR's G-score, a scoring scheme across all tissues, on ClinVar, fine-mapped GWAS loci, Massive Parallel Reporter Assay (MPRA) validated variants and observe the consistently better performance of TVAR compared with other competing tools. AVAILABILITY AND IMPLEMENTATION: The TVAR source code and its scores on the ClinVar catalog, fine mapped GWAS Loci, high confidence eQTLs from GTEx dataset, and MPRA validated functional variants are available at https://github.com/haiyang1986/TVAR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hai Yang 0002, Rui Chen 0021, Quan Wang 0004, Ying Ji 0002, Xue Zhong, Bingshan Li
Bioinform.7
2022 A Bayesian framework to integrate multi-level genome-scale data for Autism risk gene prioritization
abstract
BACKGROUND: Autism spectrum disorder (ASD) is a group of complex neurodevelopment disorders with a strong genetic basis. Large scale sequencing studies have identified over one hundred ASD risk genes. Nevertheless, the vast majority of ASD risk genes remain to be discovered, as it is estimated that more than 1000 genes are likely to be involved in ASD risk. Prioritization of risk genes is an effective strategy to increase the power of identifying novel risk genes in genetics studies of ASD. As ASD risk genes are likely to exhibit distinct properties from multiple angles, we reason that integrating multiple levels of genomic data is a powerful approach to pinpoint genuine ASD risk genes. RESULTS: We present BNScore, a Bayesian model selection framework to probabilistically prioritize ASD risk genes through explicitly integrating evidence from sequencing-identified ASD genes, biological annotations, and gene functional network. We demonstrate the validity of our approach and its improved performance over existing methods by examining the resulting top candidate ASD risk genes against sets of high-confidence benchmark genes and large-scale ASD genome-wide association studies. We assess the tissue-, cell type- and development stage-specific expression properties of top prioritized genes, and find strong expression specificity in brain tissues, striatal medium spiny neurons, and fetal developmental stages. CONCLUSIONS: In summary, we show that by integrating sequencing findings, functional annotation profiles, and gene-gene functional network, our proposed BNScore provides competitive performance compared to current state-of-the-art methods in prioritizing ASD genes. Our method offers a general and flexible strategy to risk gene prioritization that can potentially be applied to other complex traits as well.
Ying Ji 0002, Rui Chen 0021, Quan Wang 0004, Bingshan Li
BMC Bioinform.6
2021 DDIWAS: High-throughput electronic health record-based screening of drug-drug interactions
abstract
OBJECTIVE: We developed and evaluated Drug-Drug Interaction Wide Association Study (DDIWAS). This novel method detects potential drug-drug interactions (DDIs) by leveraging data from the electronic health record (EHR) allergy list. MATERIALS AND METHODS: To identify potential DDIs, DDIWAS scans for drug pairs that are frequently documented together on the allergy list. Using deidentified medical records, we tested 616 drugs for potential DDIs with simvastatin (a common lipid-lowering drug) and amlodipine (a common blood-pressure lowering drug). We evaluated the performance to rediscover known DDIs using existing knowledge bases and domain expert review. To validate potential novel DDIs, we manually reviewed patient charts and searched the literature. RESULTS: DDIWAS replicated 34 known DDIs. The positive predictive value to detect known DDIs was 0.85 and 0.86 for simvastatin and amlodipine, respectively. DDIWAS also discovered potential novel interactions between simvastatin-hydrochlorothiazide, amlodipine-omeprazole, and amlodipine-valacyclovir. A software package to conduct DDIWAS is publicly available. CONCLUSIONS: In this proof-of-concept study, we demonstrate the value of incorporating information mined from existing allergy lists to detect DDIs in a real-world clinical setting. Since allergy lists are routinely collected in EHRs, DDIWAS has the potential to detect and validate DDI signals across institutions.
Patrick Wu, Scott D. Nelson, Juan Zhao 0003, Cosby A. Stone Jr., QiPing Feng, Qingxia Chen, Eric A. Larson, Bingshan Li, Nancy J. Cox, C. Michael Stein, Elizabeth Phillips, Dan M. Roden, Joshua C. Denny, Wei-Qi Wei
J. Am. Medical Informatics Assoc.8
2020 Ultrasound Image-Based Diagnosis of Cirrhosis with an End-to-End Deep Learning model
abstract
Cirrhosis is a chronic liver disease that seriously jeopardizes the life and health of patients. Currently, ultrasound (US) imaging is commonly used by the computer-aided diagnosis (CAD) system to diagnose cirrhosis. With the rapid development of artificial intelligence, deep learning methods for cirrhosis diagnosis using ultrasound image data have emerged. However, due to US images' complexity and variability, this input usually requires manual annotation. This study proposes LiverTL, an end-to-end deep learning approach for the automatic cirrhosis ultrasound image classification to overcome these limitations. LiverTL includes an automatic region of interest (ROI) detection module to support various ultrasound images' ROI extraction. Simultaneously, the classification module utilizes ROI areas and obtain the cirrhosis diagnosis results through the transfer learning network. We find that LiverTL achieves high classification accuracy on our evaluation data set. The cirrhosis data experiments suggest that a proper pre-training model for transfer learning is crucial for the classification results. These findings potentially pave the way to advance the diagnosis and therapy of cirrhosis.
Ligang Cui, Bingshan Li
BIBM5
2020 DRAMS: A tool to detect and re-align mixed-up samples for integrative studies of multi-omics data
abstract
Studies of complex disorders benefit from integrative analyses of multiple omics data. Yet, sample mix-ups frequently occur in multi-omics studies, weakening statistical power and risking false findings. Accurately aligning sample information, genotype, and corresponding omics data is critical for integrative analyses. We developed DRAMS (https://github.com/Yi-Jiang/DRAMS) to Detect and Re-Align Mixed-up Samples to address the sample mix-up problem. It uses a logistic regression model followed by a modified topological sorting algorithm to identify the potential true IDs based on data relationships of multi-omics. According to tests using simulated data, the more types of omics data used or the smaller the proportion of mix-ups, the better that DRAMS performs. Applying DRAMS to real data from the PsychENCODE BrainGVEX project, we detected and corrected 201 (12.5% of total data generated) mix-ups. Of the 21 mix-ups involving errors of racial identity, DRAMS re-assigned all data to the correct racial group in the 1000 Genomes project. In doing so, quantitative trait loci (QTL) (FDR<0.01) increased by an average of 1.62-fold. The use of DRAMS in multi-omics studies will strengthen statistical power of the study and improve quality of the results. Even though very limited studies have multi-omics data in place, we expect such data will increase quickly with the needs of DRAMS.
Gina Giase, Kay Grennan, Annie W. Shieh, Lide Han, Quan Wang 0004, Rui Chen 0021, Kevin P. White, Chao Chen 0041, Bingshan Li, Chunyu Liu 0001
PLoS Comput. Biol.13
2019 De novo pattern discovery enables robust assessment of functional consequences of non-coding variants
abstract
MOTIVATION: Given the complexity of genome regions, prioritize the functional effects of non-coding variants remains a challenge. Although several frameworks have been proposed for the evaluation of the functionality of non-coding variants, most of them used 'black boxes' methods that simplify the task as the pathogenicity/benign classification problem, which ignores the distinct regulatory mechanisms of variants and leads to less desirable performance. In this study, we developed DVAR, an unsupervised framework that leverage various biochemical and evolutionary evidence to distinguish the gene regulatory categories of variants and assess their comprehensive functional impact simultaneously. RESULTS: DVAR performed de novo pattern discovery in high-dimensional data and identified five regulatory clusters of non-coding variants. Leveraging the new insights into the multiple functional patterns, it measures both the between-class and the within-class functional implication of the variants to achieve accurate prioritization. Compared to other two-class learning methods, it showed improved performance in identification of clinically significant variants, fine-mapped GWAS variants, eQTLs and expression-modulating variants. Moreover, it has superior performance on disease causal variants verified by genome-editing (like CRISPR-Cas9), which could provide a pre-selection strategy for genome-editing technologies across the whole genome. Finally, evaluated in BioVU and UK Biobank, two large-scale DNA biobanks linked to complete electronic health records, DVAR demonstrated its effectiveness in prioritizing non-coding variants associated with medical phenotypes. AVAILABILITY AND IMPLEMENTATION: The C++ and Python source codes, the pre-computed DVAR-cluster labels and DVAR-scores across the whole genome are available at https://www.vumc.org/cgg/dvar. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hai Yang 0002, Rui Chen 0021, Quan Wang 0004, Ying Ji 0002, Guangze Zheng 0002, Xue Zhong, Nancy J. Cox, Bingshan Li
Bioinform.9
2017 Cancer driver gene discovery through an integrative genomics approach in a non-parametric Bayesian framework
abstract
Motivation: Comprehensive catalogue of genes that drive tumor initiation and progression in cancer is key to advancing diagnostics, therapeutics and treatment. Given the complexity of cancer, the catalogue is far from complete yet. Increasing evidence shows that driver genes exhibit consistent aberration patterns across multiple-omics in tumors. In this study, we aim to leverage complementary information encoded in each of the omics data to identify novel driver genes through an integrative framework. Specifically, we integrated mutations, gene expression, DNA copy numbers, DNA methylation and protein abundance, all available in The Cancer Genome Atlas (TCGA) and developed iDriver, a non-parametric Bayesian framework based on multivariate statistical modeling to identify driver genes in an unsupervised fashion. iDriver captures the inherent clusters of gene aberrations and constructs the background distribution that is used to assess and calibrate the confidence of driver genes identified through multi-dimensional genomic data. Results: We applied the method to 4 cancer types in TCGA and identified candidate driver genes that are highly enriched with known drivers. (e.g.: P < 3.40 × 10 -36 for breast cancer). We are particularly interested in novel genes and observed multiple lines of supporting evidence. Using systematic evaluation from multiple independent aspects, we identified 45 candidate driver genes that were not previously known across these 4 cancer types. The finding has important implications that integrating additional genomic data with multivariate statistics can help identify cancer drivers and guide the next stage of cancer genomics research. Availability and Implementation: The C ++ source code is freely available at https://medschool.vanderbilt.edu/cgg/ . Contacts: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Xue Zhong, Hushan Yang, Bingshan Li
Bioinform.5
2016 Joint detection of copy number variations in parent-offspring trios
abstract
MOTIVATION: Whole genome sequencing (WGS) of parent-offspring trios is a powerful approach for identifying disease-associated genes via detecting copy number variations (CNVs). Existing approaches, which detect CNVs for each individual in a trio independently, usually yield low-detection accuracy. Joint modeling approaches leveraging Mendelian transmission within the parent-offspring trio can be an efficient strategy to improve CNV detection accuracy. RESULTS: In this study, we developed TrioCNV, a novel approach for jointly detecting CNVs in parent-offspring trios from WGS data. Using negative binomial regression, we modeled the read depth signal while considering both GC content bias and mappability bias. Moreover, we incorporated the family relationship and used a hidden Markov model to jointly infer CNVs for three samples of a parent-offspring trio. Through application to both simulated data and a trio from 1000 Genomes Project, we showed that TrioCNV achieved superior performance than existing approaches. AVAILABILITY AND IMPLEMENTATION: The software TrioCNV implemented using a combination of Java and R is freely available from the website at https://github.com/yongzhuang/TrioCNV CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yongzhuang Liu, Jianguo Lu, Jiajie Peng, Liran Juan, Xiaolin Zhu 0002, Bingshan Li, Yadong Wang 0001
Bioinform.7
2016 RVTESTS: an efficient and comprehensive tool for rare variant association analysis using sequence data
abstract
MOTIVATION: Next-generation sequencing technologies have enabled the large-scale assessment of the impact of rare and low-frequency genetic variants for complex human diseases. Gene-level association tests are often performed to analyze rare variants, where multiple rare variants in a gene region are analyzed jointly. Applying gene-level association tests to analyze sequence data often requires integrating multiple heterogeneous sources of information (e.g. annotations, functional prediction scores, allele frequencies, genotypes and phenotypes) to determine the optimal analysis unit and prioritize causal variants. Given the complexity and scale of current sequence datasets and bioinformatics databases, there is a compelling need for more efficient software tools to facilitate these analyses. To answer this challenge, we developed RVTESTS, which implements a broad set of rare variant association statistics and supports the analysis of autosomal and X-linked variants for both unrelated and related individuals. RVTESTS also provides useful companion features for annotating sequence variants, integrating bioinformatics databases, performing data quality control and sample selection. We illustrate the advantages of RVTESTS in functionality and efficiency using the 1000 Genomes Project data. AVAILABILITY AND IMPLEMENTATION: RVTESTS is available on Linux, MacOS and Windows. Source code and executable files can be obtained at https://github.com/zhanxw/rvtests CONTACT: [email protected]; [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiaowei Zhan, Youna Hu, Bingshan Li, Gonçalo R. Abecasis, Dajiang J. Liu
Bioinform.3
2016 A computational method for genotype calling in family-based sequencing data
abstract
BACKGROUND: As sequencing technologies can help researchers detect common and rare variants across the human genome in many individuals, it is known that jointly calling genotypes across multiple individuals based on linkage disequilibrium (LD) can facilitate the analysis of low to modest coverage sequence data. However, genotype-calling methods for family-based sequence data, particularly for complex families beyond parent-offspring trios, are still lacking. RESULTS: In this study, first, we proposed an algorithm that considers both linkage disequilibrium (LD) patterns and familial transmission in nuclear and multi-generational families while retaining the computational efficiency. Second, we extended our method to incorporate external reference panels to analyze family-based sequence data with a small sample size. In simulation studies, we show that modeling multiple offspring can dramatically increase genotype calling accuracy and reduce phasing and Mendelian errors, especially at low to modest coverage. In addition, we show that using external panels can greatly facilitate genotype calling of sequencing data with a small number of individuals. We applied our method to a whole genome sequencing study of 1339 individuals at ~10X coverage from the Minnesota Center for Twin and Family Research. CONCLUSIONS: The aggregated results show that our methods significantly outperform existing ones that ignore family constraints or LD information. We anticipate that our method will be useful for many ongoing family-based sequencing projects. We have implemented our methods efficiently in a C++ program FamLDCaller, which is available from http://www.pitt.edu/~wec47/famldcaller.html.
Lun-Ching Chang, Bingshan Li, Scott Vrieze, Matthew McGue, William G. Iacono, George C. Tseng, Wei Chen 0074
BMC Bioinform.2
2015 A haplotype-based framework for group-wise transmission/disequilibrium tests for rare variant association analysis
abstract
MOTIVATION: A major focus of current sequencing studies for human genetics is to identify rare variants associated with complex diseases. Aside from reduced power of detecting associated rare variants, controlling for population stratification is particularly challenging for rare variants. Transmission/disequilibrium tests (TDT) based on family designs are robust to population stratification and admixture, and therefore provide an effective approach to rare variant association studies to eliminate spurious associations. To increase power of rare variant association analysis, gene-based collapsing methods become standard approaches for analyzing rare variants. Existing methods that extend this strategy to rare variants in families usually combine TDT statistics at individual variants and therefore lack the flexibility of incorporating other genetic models. RESULTS: In this study, we describe a haplotype-based framework for group-wise TDT (gTDT) that is flexible to encompass a variety of genetic models such as additive, dominant and compound heterozygous (CH) (i.e. recessive) models as well as other complex interactions. Unlike existing methods, gTDT constructs haplotypes by transmission when possible and inherently takes into account the linkage disequilibrium among variants. Through extensive simulations we showed that type I error was correctly controlled for rare variants under all models investigated, and this remained true in the presence of population stratification. Under a variety of genetic models, gTDT showed increased power compared with the single marker TDT. Application of gTDT to an autism exome sequencing data of 118 trios identified potentially interesting candidate genes with CH rare variants. AVAILABILITY AND IMPLEMENTATION: We implemented gTDT in C++ and the source code and the detailed usage are available on the authors' website (https://medschool.vanderbilt.edu/cgg). CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Rui Chen 0021, Xiaowei Zhan, Xue Zhong, James S. Sutcliffe, Nancy J. Cox, Edwin H. Cook Jr., Wei Chen 0074, Bingshan Li
Bioinform.10
2015 A Bayesian framework for de novo mutation calling in parents-offspring trios
abstract
MOTIVATION: Spontaneous (de novo) mutations play an important role in the disease etiology of a range of complex diseases. Identifying de novo mutations (DNMs) in sporadic cases provides an effective strategy to find genes or genomic regions implicated in the genetics of disease. High-throughput next-generation sequencing enables genome- or exome-wide detection of DNMs by sequencing parents-proband trios. It is challenging to sift true mutations through massive amount of noise due to sequencing error and alignment artifacts. One of the critical limitations of existing methods is that for all genomic regions the same pre-specified mutation rate is assumed, which has a significant impact on the DNM calling accuracy. RESULTS: In this study, we developed and implemented a novel Bayesian framework for DNM calling in trios (TrioDeNovo), which overcomes these limitations by disentangling prior mutation rates from evaluation of the likelihood of the data so that flexible priors can be adjusted post-hoc at different genomic sites. Through extensively simulations and application to real data we showed that this new method has improved sensitivity and specificity over existing methods, and provides a flexible framework to further improve the efficiency by incorporating proper priors. The accuracy is further improved using effective filtering based on sequence alignment characteristics. AVAILABILITY AND IMPLEMENTATION: The C++ source code implementing TrioDeNovo is freely available at https://medschool.vanderbilt.edu/cgg. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiaowei Zhan, Xue Zhong, Yongzhuang Liu, Yujun Han, Wei Chen 0074, Bingshan Li
Bioinform.7
2014 A gradient-boosting approach for filtering de novo mutations in parent-offspring trios
abstract
MOTIVATION: Whole-genome and -exome sequencing on parent-offspring trios is a powerful approach to identifying disease-associated genes by detecting de novo mutations in patients. Accurate detection of de novo mutations from sequencing data is a critical step in trio-based genetic studies. Existing bioinformatic approaches usually yield high error rates due to sequencing artifacts and alignment issues, which may either miss true de novo mutations or call too many false ones, making downstream validation and analysis difficult. In particular, current approaches have much worse specificity than sensitivity, and developing effective filters to discriminate genuine from spurious de novo mutations remains an unsolved challenge. RESULTS: In this article, we curated 59 sequence features in whole genome and exome alignment context which are considered to be relevant to discriminating true de novo mutations from artifacts, and then employed a machine-learning approach to classify candidates as true or false de novo mutations. Specifically, we built a classifier, named De Novo Mutation Filter (DNMFilter), using gradient boosting as the classification algorithm. We built the training set using experimentally validated true and false de novo mutations as well as collected false de novo mutations from an in-house large-scale exome-sequencing project. We evaluated DNMFilter's theoretical performance and investigated relative importance of different sequence features on the classification accuracy. Finally, we applied DNMFilter on our in-house whole exome trios and one CEU trio from the 1000 Genomes Project and found that DNMFilter could be coupled with commonly used de novo mutation detection approaches as an effective filtering approach to significantly reduce false discovery rate without sacrificing sensitivity. AVAILABILITY: The software DNMFilter implemented using a combination of Java and R is freely available from the website at http://humangenome.duke.edu/software.
Yongzhuang Liu, Bingshan Li, Renjie Tan, Xiaolin Zhu 0002, Yadong Wang 0001
Bioinform.2