VLDB 2026 Research / reviewers in the wild / expert
Wei Chen 0074
dblp:181/2832-74
· DBLP profile ↗
21ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0001-7196-8703ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 5 since 2021Computer networks · 4 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fast and scalable Wasserstein-1 neural optimal transport solver for single-cell perturbation predictionabstractMOTIVATION: Predicting single-cell perturbation responses requires mapping between two unpaired single-cell data distributions. Optimal transport (OT) theory provides a principled framework for constructing such mappings by minimizing transport cost. Recently, Wasserstein-2 (W2) neural optimal transport solvers (e.g. CellOT) have been used for this prediction task. However, W2 OT relies on the general Kantorovich dual formulation, which involves optimizing over two conjugate functions, leading to a complex min-max optimization problem that converges slowly. RESULTS: To address these challenges, we propose a novel solver based on the Wasserstein-1 (W1) dual formulation. Unlike W2, the W1 dual simplifies the optimization to a maximization problem over a single 1-Lipschitz function, thus eliminating the need for time-consuming min-max optimization. While solving the W1 dual only reveals the transport direction and does not directly provide a unique optimal transport map, we incorporate an additional step using adversarial training to determine an appropriate transport step size, effectively recovering the transport map. Our experiments demonstrate that the proposed W1 neural optimal transport solver can mimic the W2 OT solvers in finding a unique and "monotonic" map on 2D datasets. Moreover, the W1 OT solver achieves performance on par with or surpasses W2 OT solvers on real single-cell perturbation datasets. Furthermore, we show that W1 OT solver achieves 25∼45× speedup, scales better on high dimensional transportation task, and can be directly applied on single-cell RNA-seq dataset with highly variable genes. AVAILABILITY AND IMPLEMENTATION: Our implementation and experiments are open-sourced at https://github.com/poseidonchan/w1ot. Yanshuo Chen, Zhengmian Hu, Wei Chen 0074, Heng Huang 0001 |
Bioinform. | 3 |
| 2025 | Enhancing privacy in biosecurity with watermarked protein designabstractMOTIVATION: The biosecurity issue arises as the capability of deep-learning-based protein design has rapidly increased in recent years. Current regulation procedures for DNA synthesizing focus on the biosecurity but ignore the data privacy. RESULTS: We propose a general framework for adding watermarks to protein sequences designed by various autoregressive deep-learning models. Compared to current regulation procedures, watermarks also ensure robust traceability to achieve biosecurity but maintain privacy of designed sequences by local verification. Benchmarked with other watermarking techniques, the watermark detection efficiency of our method is substantially increased to be more practical in real-world scenarios. Moreover, it provides a convenient way for researchers to claim their own intellectual property since the designer's information could be embedded into the sequence with our framework. AVAILABILITY AND IMPLEMENTATION: The implementation of the protein watermark framework is freely available to noncommercial users at https://github.com/poseidonchan/ProteinWatermark. Yanshuo Chen, Zhengmian Hu, Yongrui Jin, Marcus Zhan, Chengjin Xie, Wei Chen 0074, Heng Huang 0001 |
Bioinform. | 8 |
| 2023 | PTEase: Objective Airway Examination for Pulmonary Telemedicine using Commodity SmartphonesabstractRemote monitoring and evaluation of pulmonary diseases via tele-medicine are important to disease diagnosis and management, but current telemedicine solutions have limited capability of objectively examining the airway's internal physiological conditions that are crucial to pulmonary disease evaluation. Existing solutions based on smartphone sensing are also limited to externally monitoring breath rates, respiratory events, or lung function. In this paper, we present PTEase, a new system design that addresses these limitations and uses commodity smartphones to examine the airway's internal physiological conditions. PTEase uses active acoustic sensing to measure the internal changes of lower airway caliber, and then leverages machine learning to analyze the sensory data for pulmonary disease evaluation. We implemented PTEase as a smartphone app, and verified its measurement error in lab-controlled settings as <10%. Clinical studies further showed that PTEase reaches 75% accuracy on disease prediction and 11%-15% errors in estimating lung function indices. Given that such accuracy is comparable with that in clinical practice using spirometry, PTEase can be reliably used as an assistive telemedicine tool for disease evaluation and monitoring. Xiangyu Yin 0002, Kai Huang 0007, Erick Forno, Wei Chen 0074, Heng Huang 0001, Wei Gao 0006 |
MobiSys | 4 |
| 2022 | Multi-modal Genotype and Phenotype Mutual Learning to Enhance Single-Modal Input Based Longitudinal Outcome Prediction
Alireza Ganjdanesh, Wei Chen 0074, Heng Huang 0001 |
RECOMB | 3 |
| 2022 | Out-Clinic Pulmonary Disease Evaluation via Acoustic Sensing and Multi-Task Learning on Commodity SmartphonesabstractPulmonary diseases, such as asthma and Chronic Obstructive Pulmonary Disease (COPD), constitute a major public health challenge. The disease symptoms, including airway obstruction and inflammation, usually result in changes in airway mechanical properties, such as the caliber and impedance of the airway. To measure such airway properties for disease evaluation and diagnosis purposes, pulmonary function tests (PFT) has been widely adopted. However, most existing PFT systems require expensive and cumbersome hardware that are impossible to be used out of clinic. To allow out-clinic continuous pulmonary disease evaluation, in this paper we present AWARE, a new sensing and AI system that supports accurate and reliable PFT using commodity smartphones. AWARE uses a smartphone to transmit acoustic signals and reconstructs the profile of human airway based on the analysis of reflected acoustic waves captured from the smartphone's microphone. The subject's pulmonary condition is then evaluated by a multi-task learning model that integrates both the airway measurements and the subject's lung function records as the ground truth. Evaluations on 75 human subjects demonstrate that AWARE has the capability to achieve 80% accuracy on distinguishing between humans with healthy pulmonary function and with asthma symptoms. Xiangyu Yin 0002, Kai Huang 0007, Erick Forno, Wei Chen 0074, Heng Huang 0001, Wei Gao 0006 |
SenSys | 4 |
| 2022 | Robust and accurate estimation of cellular fraction from tissue omics data via ensemble deconvolutionabstractMOTIVATION: Tissue-level omics data such as transcriptomics and epigenomics are an average across diverse cell types. To extract cell-type-specific (CTS) signals, dozens of cellular deconvolution methods have been proposed to infer cell-type fractions from tissue-level data. However, these methods produce vastly different results under various real data settings. Simulation-based benchmarking studies showed no universally best deconvolution approaches. There have been attempts of ensemble methods, but they only aggregate multiple single-cell references or reference-free deconvolution methods. RESULTS: To achieve a robust estimation of cellular fractions, we proposed EnsDeconv (Ensemble Deconvolution), which adopts CTS robust regression to synthesize the results from 11 single deconvolution methods, 10 reference datasets, 5 marker gene selection procedures, 5 data normalizations and 2 transformations. Unlike most benchmarking studies based on simulations, we compiled four large real datasets of 4937 tissue samples in total with measured cellular fractions and bulk gene expression from different tissues. Comprehensive evaluations demonstrated that EnsDeconv yields more stable, robust and accurate fractions than existing methods. We illustrated that EnsDeconv estimated cellular fractions enable various CTS downstream analyses such as differential fractions associated with clinical variables. We further extended EnsDeconv to analyze bulk DNA methylation data. AVAILABILITY AND IMPLEMENTATION: EnsDeconv is freely available as an R-package from https://github.com/randel/EnsDeconv. The RNA microarray data from the TRAUMA study are available and can be accessed in GEO (GSE36809). The demographic and clinical phenotypes can be shared on reasonable request to the corresponding authors. The RNA-seq data from the EVAPR study cannot be shared publicly due to the privacy of individuals that participated in the clinical research in compliance with the IRB approval at the University of Pittsburgh. The RNA microarray data from the FHS study are available from dbGaP (phs000007.v32.p13). The RNA-seq data from ROS study is downloaded from AD Knowledge Portal. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Manqi Cai, Molin Yue, Tianmeng Chen, Jinling Liu, Erick Forno, Xinghua Lu 0001, Timothy Billiar, Juan C. Celedón, Chris McKennan, Wei Chen 0074, Jiebiao Wang |
Bioinform. | 10 |
| 2021 | CHIT: an allele-specific method for testing the association between molecular quantitative traits and phenotype-genotype interactionabstractMOTIVATION: Allele-specific differences in molecular traits can be obtained from next-generation sequencing data and could potentially improve testing power, but such information is usually overlooked in association studies. Furthermore, the variation of molecular quantitative traits (e.g. gene expression) could result from the interaction effect of genotypes and phenotypes, but it is challenging to identify such interaction signals in complex disease studies in humans due to small genetic effect sizes and/or small sample sizes. RESULTS: We develop a novel statistical method, the combined haplotype interaction test (CHIT), which tests for association between molecular quantitative traits and phenotype-genotype interactions by modeling the total read counts and allele-specific reads in a target region. CHIT can be used as a supplementary analysis to the regular linear interaction regression. In our simulations, CHIT obtains non-inflated type I error rates, and it has higher power than a standard interaction quantitative trait locus approach based on linear regression models. Finally, we illustrate CHIT by testing associations between gene expression obtained by RNA-seq and the interaction of SNPs and atopy status from a study of childhood asthma in Puerto Ricans, and results demonstrate that CHIT could be more powerful than a standard linear interaction expression quantitative trait loci approach. AVAILABILITY AND IMPLEMENTATION: The CHIT algorithm has been implemented in Python. The source code and documentation are available and can be downloaded from https://github.com/QiYanPitt/CHIT. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qi Yan 0008, Erick Forno, Juan C. Celedón, Wei Chen 0074, Daniel E. Weeks |
Bioinform. | 4 |
| 2020 | SpiroSonic: monitoring human lung function via acoustic sensing on commodity smartphonesabstractRespiratory diseases have been a significant public health challenge. Efficient disease evaluation and monitoring call for daily spirometry tests, as an effective way of pulmonary function testing, out of clinic. This requirement, however, is hard to be satisfied due to the large size and high costs of current spirometry equipments. In this paper, we present SpiroSonic, a new system design that uses commodity smartphones to support complete, accurate yet reliable spirometry tests in regular home settings with various environmental and human factors. SpiroSonic measures the humans' chest wall motion via acoustic sensing and interprets such motion into lung function indices, based on the clinically validated correlation between them. We implemented SpiroSonic as a smartphone app, and verified SpiroSonic's monitoring error over healthy humans as <3%. Clinical studies further show that SpiroSonic reaches 5%-10% monitoring error among 83 pediatric patients. Given that the error of in-clinic spirometry is usually around 5%, SpiroSonic can be reliably used for disease tracking and evaluation out of clinic. Xingzhe Song, Boyuan Yang 0001, Ruirong Chen, Erick Forno, Wei Chen 0074, Wei Gao 0006 |
MobiCom | 6 |
| 2020 | Deep Large-Scale Multi-task Learning Network for Gene Expression Inference
Kamran Ghasedi Dizaji, Wei Chen 0074, Heng Huang 0001 |
RECOMB | 2 |
| 2020 | Artificial-cell-type aware cell-type classification in CITE-seqabstractMOTIVATION: Cellular Indexing of Transcriptomes and Epitopes by sequencing (CITE-seq), couples the measurement of surface marker proteins with simultaneous sequencing of mRNA at single cell level, which brings accurate cell surface phenotyping to single-cell transcriptomics. Unfortunately, multiplets in CITE-seq datasets create artificial cell types (ACT) and complicate the automation of cell surface phenotyping. RESULTS: We propose CITE-sort, an artificial-cell-type aware surface marker clustering method for CITE-seq. CITE-sort is aware of and is robust to multiplet-induced ACT. We benchmarked CITE-sort with real and simulated CITE-seq datasets and compared CITE-sort against canonical clustering methods. We show that CITE-sort produces the best clustering performance across the board. CITE-sort not only accurately identifies real biological cell types (BCT) but also consistently and reliably separates multiplet-induced artificial-cell-type droplet clusters from real BCT droplet clusters. In addition, CITE-sort organizes its clustering process with a binary tree, which facilitates easy interpretation and verification of its clustering result and simplifies cell-type annotation with domain knowledge in CITE-seq. AVAILABILITY AND IMPLEMENTATION: http://github.com/QiuyuLian/CITE-sort. SUPPLEMENTARY INFORMATION: Supplementary data is available at Bioinformatics online. Qiuyu Lian, Hongyi Xin, Jianzhu Ma, Liza Konnikova, Wei Chen 0074, Jin Gu, Kong Chen |
Bioinform. | 5 |
| 2020 | CSMD: a computational subtraction-based microbiome discovery pipeline for species-level characterization of clinical metagenomic samplesabstractMOTIVATION: Microbiome analyses of clinical samples with low microbial biomass are challenging because of the very small quantities of microbial DNA relative to the human host, ubiquitous contaminating DNA in sequencing experiments and the large and rapidly growing microbial reference databases. RESULTS: We present computational subtraction-based microbiome discovery (CSMD), a bioinformatics pipeline specifically developed to generate accurate species-level microbiome profiles for clinical samples with low microbial loads. CSMD applies strategies for the maximal elimination of host sequences with minimal loss of microbial signal and effectively detects microorganisms present in the sample with minimal false positives using a stepwise convergent solution. CSMD was benchmarked in a comparative evaluation with other classic tools on previously published well-characterized datasets. It showed higher sensitivity and specificity in host sequence removal and higher specificity in microbial identification, which led to more accurate abundance estimation. All these features are integrated into a free and easy-to-use tool. Additionally, CSMD applied to cell-free plasma DNA showed that microbial diversity within these samples is substantially broader than previously believed. AVAILABILITY AND IMPLEMENTATION: CSMD is freely available at https://github.com/liuyu8721/csmd. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Paul W. Bible, Qiaoxing Liang, Cong Dong, Xiaofeng Wen, Xiaofei Ge, Xifang Li, Xiuli Deng, Shixin Guo, Juanran Liang, Wenliang Pan, Wei Chen 0074 |
Bioinform. | 17 |
| 2019 | Device-Free Acoustic Motion Tracking over Targets with Large SizesabstractDevice-free acoustic motion tracking allows a commodity mobile device to precisely track the human user's motion, without applying any extra hardware tracker on the human body. Most of current device-free acoustic motion tracking systems, however, are limited to tracking the motion of small parts of the human body with negligible sizes, such as human fingers. Their accuracy of motion tracking will significantly degrade when being applied to targets with large sizes, such as humans' hands, arms or body trunk. We envision the key reason to such degradation as the target size's significant impact on the pattern of the reflected acoustic signal, and develop analytical modeling of such reflected acoustic signal from large targets. Based on such modeling, we present a new system called Acoustic Tracking over targets with LArge Sizes (ATLAS), which ensures precise motion tracking over large targets by correctly interpreting the reflected acoustic signal and extracting the phase from the signal. Experiment results over commodity Android smartphones show that ATLAS can reduce the error of motion tracking by more than 75%, when being applied to targets with heterogeneous sizes in practice. Ruirong Chen, Xingzhe Song, Wei Gao 0006, Wei Chen 0074, Erick Forno |
MASS | 5 |
| 2018 | Bayesian integrative model for multi-omics data with missingnessabstractMotivation: Integrative analysis of multi-omics data from different high-throughput experimental platforms provides valuable insight into regulatory mechanisms associated with complex diseases, and gains statistical power to detect markers that are otherwise overlooked by single-platform omics analysis. In practice, a significant portion of samples may not be measured completely due to insufficient tissues or restricted budget (e.g. gene expression profile are measured but not methylation). Current multi-omics integrative methods require complete data. A common practice is to ignore samples with any missing platform and perform complete case analysis, which leads to substantial loss of statistical power. Methods: In this article, inspired by the popular Integrative Bayesian Analysis of Genomics data (iBAG), we propose a full Bayesian model that allows incorporation of samples with missing omics data. Results: Simulation results show improvement of the new full Bayesian approach in terms of outcome prediction accuracy and feature selection performance when sample size is limited and proportion of missingness is large. When sample size is large or the proportion of missingness is low, incorporating samples with missingness may introduce extra inference uncertainty and generate worse prediction and feature selection performance. To determine whether and how to incorporate samples with missingness, we propose a self-learning cross-validation (CV) decision scheme. Simulations and a real application on child asthma dataset demonstrate superior performance of the CV decision scheme when various types of missing mechanisms are evaluated. Availability and implementation: Freely available on the GitHub at https://github.com/CHPGenetics/FBM. Supplementary information: Supplementary data are available at Bioinformatics online. Tianzhou Ma, Gong Tang, Qi Yan 0008, Ting Wang 0003, Juan C. Celedón, Wei Chen 0074, George C. Tseng |
Bioinform. | 8 |
| 2018 | DIMM-SC: a Dirichlet mixture model for clustering droplet-based single cell transcriptomic dataabstractMotivation: Single cell transcriptome sequencing (scRNA-Seq) has become a revolutionary tool to study cellular and molecular processes at single cell resolution. Among existing technologies, the recently developed droplet-based platform enables efficient parallel processing of thousands of single cells with direct counting of transcript copies using Unique Molecular Identifier (UMI). Despite the technology advances, statistical methods and computational tools are still lacking for analyzing droplet-based scRNA-Seq data. Particularly, model-based approaches for clustering large-scale single cell transcriptomic data are still under-explored. Results: We developed DIMM-SC, a Dirichlet Mixture Model for clustering droplet-based Single Cell transcriptomic data. This approach explicitly models UMI count data from scRNA-Seq experiments and characterizes variations across different cell clusters via a Dirichlet mixture prior. We performed comprehensive simulations to evaluate DIMM-SC and compared it with existing clustering methods such as K-means, CellTree and Seurat. In addition, we analyzed public scRNA-Seq datasets with known cluster labels and in-house scRNA-Seq datasets from a study of systemic sclerosis with prior biological knowledge to benchmark and validate DIMM-SC. Both simulation studies and real data applications demonstrated that overall, DIMM-SC achieves substantially improved clustering accuracy and much lower clustering variability compared to other existing clustering methods. More importantly, as a model-based approach, DIMM-SC is able to quantify the clustering uncertainty for each single cell, facilitating rigorous statistical inference and biological interpretations, which are typically unavailable from existing clustering methods. Availability and implementation: DIMM-SC has been implemented in a user-friendly R package with a detailed tutorial available on www.pitt.edu/∼wec47/singlecell.html. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Zhe Sun 0007, Ting Wang 0003, Robert Lafyatis, Ying Ding 0003, Ming Hu 0001, Wei Chen 0074 |
Bioinform. | 8 |
| 2018 | KMgene: a unified R package for gene-based association analysis for complex traitsabstractSummary: In this report, we introduce an R package KMgene for performing gene-based association tests for familial, multivariate or longitudinal traits using kernel machine (KM) regression under a generalized linear mixed model framework. Extensive simulations were performed to evaluate the validity of the approaches implemented in KMgene. Availability and implementation: http://cran.r-project.org/web/packages/KMgene. Supplementary information: Supplementary data are available at Bioinformatics online. Qi Yan 0008, Wei Chen 0074 |
Bioinform. | 3 |
| 2018 | SILGGM: An extensive R package for efficient statistical inference in large-scale gene networksabstractGene co-expression network analysis is extremely useful in interpreting a complex biological process. The recent droplet-based single-cell technology is able to generate much larger gene expression data routinely with thousands of samples and tens of thousands of genes. To analyze such a large-scale gene-gene network, remarkable progress has been made in rigorous statistical inference of high-dimensional Gaussian graphical model (GGM). These approaches provide a formal confidence interval or a p-value rather than only a single point estimator for conditional dependence of a gene pair and are more desirable for identifying reliable gene networks. To promote their widespread use, we herein introduce an extensive and efficient R package named SILGGM (Statistical Inference of Large-scale Gaussian Graphical Model) that includes four main approaches in statistical inference of high-dimensional GGM. Unlike the existing tools, SILGGM provides statistically efficient inference on both individual gene pair and whole-scale gene pairs. It has a novel and consistent false discovery rate (FDR) procedure in all four methodologies. Based on the user-friendly design, it provides outputs compatible with multiple platforms for interactive network visualization. Furthermore, comparisons in simulation illustrate that SILGGM can accelerate the existing MATLAB implementation to several orders of magnitudes and further improve the speed of the already very efficient R package FastGGM. Testing results from the simulated data confirm the validity of all the approaches in SILGGM even in a very large-scale setting with the number of variables or genes to a ten thousand level. We have also applied our package to a novel single-cell RNA-seq data set with pan T cells. The results show that the approaches in SILGGM significantly outperform the conventional ones in a biological sense. The package is freely available via CRAN at https://cran.r-project.org/package=SILGGM. Zhao Ren, Wei Chen 0074 |
PLoS Comput. Biol. | 3 |
| 2016 | A computational method for genotype calling in family-based sequencing dataabstractBACKGROUND: As sequencing technologies can help researchers detect common and rare variants across the human genome in many individuals, it is known that jointly calling genotypes across multiple individuals based on linkage disequilibrium (LD) can facilitate the analysis of low to modest coverage sequence data. However, genotype-calling methods for family-based sequence data, particularly for complex families beyond parent-offspring trios, are still lacking. RESULTS: In this study, first, we proposed an algorithm that considers both linkage disequilibrium (LD) patterns and familial transmission in nuclear and multi-generational families while retaining the computational efficiency. Second, we extended our method to incorporate external reference panels to analyze family-based sequence data with a small sample size. In simulation studies, we show that modeling multiple offspring can dramatically increase genotype calling accuracy and reduce phasing and Mendelian errors, especially at low to modest coverage. In addition, we show that using external panels can greatly facilitate genotype calling of sequencing data with a small number of individuals. We applied our method to a whole genome sequencing study of 1339 individuals at ~10X coverage from the Minnesota Center for Twin and Family Research. CONCLUSIONS: The aggregated results show that our methods significantly outperform existing ones that ignore family constraints or LD information. We anticipate that our method will be useful for many ongoing family-based sequencing projects. We have implemented our methods efficiently in a C++ program FamLDCaller, which is available from http://www.pitt.edu/~wec47/famldcaller.html. Lun-Ching Chang, Bingshan Li, Scott Vrieze, Matthew McGue, William G. Iacono, George C. Tseng, Wei Chen 0074 |
BMC Bioinform. | 8 |
| 2016 | FastGGM: An Efficient Algorithm for the Inference of Gaussian Graphical Model in Biological NetworksabstractBiological networks provide additional information for the analysis of human diseases, beyond the traditional analysis that focuses on single variables. Gaussian graphical model (GGM), a probability model that characterizes the conditional dependence structure of a set of random variables by a graph, has wide applications in the analysis of biological networks, such as inferring interaction or comparing differential networks. However, existing approaches are either not statistically rigorous or are inefficient for high-dimensional data that include tens of thousands of variables for making inference. In this study, we propose an efficient algorithm to implement the estimation of GGM and obtain p-value and confidence interval for each edge in the graph, based on a recent proposal by Ren et al., 2015. Through simulation studies, we demonstrate that the algorithm is faster by several orders of magnitude than the current implemented algorithm for Ren et al. without losing any accuracy. Then, we apply our algorithm to two real data sets: transcriptomic data from a study of childhood asthma and proteomic data from a study of Alzheimer's disease. We estimate the global gene or protein interaction networks for the disease and healthy samples. The resulting networks reveal interesting interactions and the differential networks between cases and controls show functional relevance to the diseases. In conclusion, we provide a computationally fast algorithm to implement a statistically sound procedure for constructing Gaussian graphical model and making inference with high-dimensional biological data. The algorithm has been implemented in an R package named "FastGGM". Ting Wang 0003, Zhao Ren, Ying Ding 0003, Zhe Sun 0007, Matthew L. MacDonald, Robert A. Sweet, Jieru Wang, Wei Chen 0074 |
PLoS Comput. Biol. | 9 |
| 2015 | A haplotype-based framework for group-wise transmission/disequilibrium tests for rare variant association analysisabstractMOTIVATION: A major focus of current sequencing studies for human genetics is to identify rare variants associated with complex diseases. Aside from reduced power of detecting associated rare variants, controlling for population stratification is particularly challenging for rare variants. Transmission/disequilibrium tests (TDT) based on family designs are robust to population stratification and admixture, and therefore provide an effective approach to rare variant association studies to eliminate spurious associations. To increase power of rare variant association analysis, gene-based collapsing methods become standard approaches for analyzing rare variants. Existing methods that extend this strategy to rare variants in families usually combine TDT statistics at individual variants and therefore lack the flexibility of incorporating other genetic models. RESULTS: In this study, we describe a haplotype-based framework for group-wise TDT (gTDT) that is flexible to encompass a variety of genetic models such as additive, dominant and compound heterozygous (CH) (i.e. recessive) models as well as other complex interactions. Unlike existing methods, gTDT constructs haplotypes by transmission when possible and inherently takes into account the linkage disequilibrium among variants. Through extensive simulations we showed that type I error was correctly controlled for rare variants under all models investigated, and this remained true in the presence of population stratification. Under a variety of genetic models, gTDT showed increased power compared with the single marker TDT. Application of gTDT to an autism exome sequencing data of 118 trios identified potentially interesting candidate genes with CH rare variants. AVAILABILITY AND IMPLEMENTATION: We implemented gTDT in C++ and the source code and the detailed usage are available on the authors' website (https://medschool.vanderbilt.edu/cgg). CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rui Chen 0021, Xiaowei Zhan, Xue Zhong, James S. Sutcliffe, Nancy J. Cox, Edwin H. Cook Jr., Wei Chen 0074, Bingshan Li |
Bioinform. | 9 |
| 2015 | A Bayesian framework for de novo mutation calling in parents-offspring triosabstractMOTIVATION: Spontaneous (de novo) mutations play an important role in the disease etiology of a range of complex diseases. Identifying de novo mutations (DNMs) in sporadic cases provides an effective strategy to find genes or genomic regions implicated in the genetics of disease. High-throughput next-generation sequencing enables genome- or exome-wide detection of DNMs by sequencing parents-proband trios. It is challenging to sift true mutations through massive amount of noise due to sequencing error and alignment artifacts. One of the critical limitations of existing methods is that for all genomic regions the same pre-specified mutation rate is assumed, which has a significant impact on the DNM calling accuracy. RESULTS: In this study, we developed and implemented a novel Bayesian framework for DNM calling in trios (TrioDeNovo), which overcomes these limitations by disentangling prior mutation rates from evaluation of the likelihood of the data so that flexible priors can be adjusted post-hoc at different genomic sites. Through extensively simulations and application to real data we showed that this new method has improved sensitivity and specificity over existing methods, and provides a flexible framework to further improve the efficiency by incorporating proper priors. The accuracy is further improved using effective filtering based on sequence alignment characteristics. AVAILABILITY AND IMPLEMENTATION: The C++ source code implementing TrioDeNovo is freely available at https://medschool.vanderbilt.edu/cgg. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaowei Zhan, Xue Zhong, Yongzhuang Liu, Yujun Han, Wei Chen 0074, Bingshan Li |
Bioinform. | 6 |
| 2015 | DISSCO: direct imputation of summary statistics allowing covariatesabstractBACKGROUND: Imputation of individual level genotypes at untyped markers using an external reference panel of genotyped or sequenced individuals has become standard practice in genetic association studies. Direct imputation of summary statistics can also be valuable, for example in meta-analyses where individual level genotype data are not available. Two methods (DIST and ImpG-Summary/LD), that assume a multivariate Gaussian distribution for the association summary statistics, have been proposed for imputing association summary statistics. However, both methods assume that the correlations between association summary statistics are the same as the correlations between the corresponding genotypes. This assumption can be violated in the presence of confounding covariates. METHODS: We analytically show that in the absence of covariates, correlation among association summary statistics is indeed the same as that among the corresponding genotypes, thus serving as a theoretical justification for the recently proposed methods. We continue to prove that in the presence of covariates, correlation among association summary statistics becomes the partial correlation of the corresponding genotypes controlling for covariates. We therefore develop direct imputation of summary statistics allowing covariates (DISSCO). RESULTS: We consider two real-life scenarios where the correlation and partial correlation likely make practical difference: (i) association studies in admixed populations; (ii) association studies in presence of other confounding covariate(s). Application of DISSCO to real datasets under both scenarios shows at least comparable, if not better, performance compared with existing correlation-based methods, particularly for lower frequency variants. For example, DISSCO can reduce the absolute deviation from the truth by 3.9-15.2% for variants with minor allele frequency <5%. Zheng Xu 0010, Qing Duan, Wei Chen 0074, Mingyao Li, Ethan M. Lange |
Bioinform. | 4 |