VLDB 2026 Research / reviewers in the wild / expert
Jun S. Liu
dblp:83/1842
· DBLP profile ↗
55ranked-venue papers
1as first author
7since 2021 · last 2024
0000-0002-4450-7239ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 45 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 2 since 2021Systems, architecture and hardware · 2Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Power of knockoff: The impact of ranking algorithm, augmented design, and symmetric statisticabstractThe knockoff filter is a recent false discovery rate (FDR) control method for high-dimensional linear models. We point out that knockoff has three key components: ranking algorithm, augmented design, and symmetric statistic, and each component admits multiple choices. By considering various combinations of the three components, we obtain a collection of variants of knockoff. All these variants guarantee finite-sample FDR control, and our goal is to compare their power. We assume a Rare and Weak signal model on regression coeffi- cients and compare the power of different variants of knockoff by deriving explicit formulas of false positive rate and false negative rate. Our results provide new insights on how to improve power when controlling FDR at a targeted level. We also compare the power of knockoff with its propotype - a method that uses the same ranking algorithm but has access to an ideal threshold. The comparison reveals the additional price one pays by finding a data-driven threshold to control FDR. Zheng Tracy Ke, Jun S. Liu, Yucong Ma |
J. Mach. Learn. Res. | 2 |
| 2024 | A phylogenetic method linking nucleotide substitution rates to rates of continuous trait evolutionabstractGenomes contain conserved non-coding sequences that perform important biological functions, such as gene regulation. We present a phylogenetic method, PhyloAcc-C, that associates nucleotide substitution rates with changes in a continuous trait of interest. The method takes as input a multiple sequence alignment of conserved elements, continuous trait data observed in extant species, and a background phylogeny and substitution process. Gibbs sampling is used to assign rate categories (background, conserved, accelerated) to lineages and explore whether the assigned rate categories are associated with increases or decreases in the rate of trait evolution. We test our method using simulations and then illustrate its application using mammalian body size and lifespan data previously analyzed with respect to protein coding genes. Like other studies, we find processes such as tumor suppression, telomere maintenance, and p53 regulation to be related to changes in longevity and body size. In addition, we also find that skeletal genes, and developmental processes, such as sprouting angiogenesis, are relevant. Patrick Gemmell, Timothy B. Sackton, Scott V. Edwards, Jun S. Liu |
PLoS Comput. Biol. | 4 |
| 2022 | A data-adaptive Bayesian regression approach for polygenic risk predictionabstractMOTIVATION: Polygenic risk score (PRS) has been widely exploited for genetic risk prediction due to its accuracy and conceptual simplicity. We introduce a unified Bayesian regression framework, NeuPred, for PRS construction, which accommodates varying genetic architectures and improves overall prediction accuracy for complex diseases by allowing for a wide class of prior choices. To take full advantage of the framework, we propose a summary-statistics-based cross-validation strategy to automatically select suitable chromosome-level priors, which demonstrates a striking variability of the prior preference of each chromosome, for the same complex disease, and further significantly improves the prediction accuracy. RESULTS: Simulation studies and real data applications with seven disease datasets from the Wellcome Trust Case Control Consortium cohort and eight groups of large-scale genome-wide association studies demonstrate that NeuPred achieves substantial and consistent improvements in terms of predictive r2 over existing methods. In addition, NeuPred has similar or advantageous computational efficiency compared with the state-of-the-art Bayesian methods. AVAILABILITY AND IMPLEMENTATION: The R package implementing NeuPred is available at https://github.com/shuangsong0110/NeuPred. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shuang Song 0006, Lin Hou 0003, Jun S. Liu |
Bioinform. | 3 |
| 2022 | Erratum to: A data-adaptive Bayesian regression approach for polygenic risk predictionabstractBioinformatics (2022), https://doi.org/10.1093/bioinformatics/btac024 In the originally published version of this manuscript, there was an address error in affiliations 1, 2, and 3. This error has been corrected. Shuang Song 0006, Lin Hou 0003, Jun S. Liu |
Bioinform. | 3 |
| 2021 | MiRACLe: an individual-specific approach to improve microRNA-target prediction based on a random contact modelabstractDeciphering microRNA (miRNA) targets is important for understanding the function of miRNAs as well as miRNA-based diagnostics and therapeutics. Given the highly cell-specific nature of miRNA regulation, recent computational approaches typically exploit expression data to identify the most physiologically relevant target messenger RNAs (mRNAs). Although effective, those methods usually require a large sample size to infer miRNA-mRNA interactions, thus limiting their applications in personalized medicine. In this study, we developed a novel miRNA target prediction algorithm called miRACLe (miRNA Analysis by a Contact modeL). It integrates sequence characteristics and RNA expression profiles into a random contact model, and determines the target preferences by relative probability of effective contacts in an individual-specific manner. Evaluation by a variety of measures shows that fitting TargetScan, a frequently used prediction tool, into the framework of miRACLe can improve its predictive power with a significant margin and consistently outperform other state-of-the-art methods in prediction accuracy, regulatory potential and biological relevance. Notably, the superiority of miRACLe is robust to various biological contexts, types of expression data and validation datasets, and the computation process is fast and efficient. Additionally, we show that the model can be readily applied to other sequence-based algorithms to improve their predictive power, such as DIANA-microT-CDS, miRanda-mirSVR and MirTarget4. MiRACLe is publicly available at https://github.com/PANWANG2014/miRACLe. Yibo Gao, Jun S. Liu |
Briefings Bioinform. | 5 |
| 2021 | Openness weighted association studies: leveraging personal genome information to prioritize non-coding variantsabstractMOTIVATION: Identification and interpretation of non-coding variations that affect disease risk remain a paramount challenge in genome-wide association studies (GWAS) of complex diseases. Experimental efforts have provided comprehensive annotations of functional elements in the human genome. On the other hand, advances in computational biology, especially machine learning approaches, have facilitated accurate predictions of cell-type-specific functional annotations. Integrating functional annotations with GWAS signals has advanced the understanding of disease mechanisms. In previous studies, functional annotations were treated as static of a genomic region, ignoring potential functional differences imposed by different genotypes across individuals. RESULTS: We develop a computational approach, Openness Weighted Association Studies (OWAS), to leverage and aggregate predictions of chromosome accessibility in personal genomes for prioritizing GWAS signals. The approach relies on an analytical expression we derived for identifying disease associated genomic segments whose effects in the etiology of complex diseases are evaluated. In extensive simulations and real data analysis, OWAS identifies genes/segments that explain more heritability than existing methods, and has a better replication rate in independent cohorts than GWAS. Moreover, the identified genes/segments show tissue-specific patterns and are enriched in disease relevant pathways. We use rheumatic arthritis and asthma as examples to demonstrate how OWAS can be exploited to provide novel insights on complex diseases. AVAILABILITY AND IMPLEMENTATION: The R package OWAS that implements our method is available at https://github.com/shuangsong0110/OWAS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shuang Song 0006, Nayang Shan, Xiting Yan, Jun S. Liu, Lin Hou 0003 |
Bioinform. | 5 |
| 2021 | Bayesian Text Classification and Summarization via A Class-Specified Topic ModelabstractWe propose the class-specified topic model (CSTM) to deal with the tasks of text classification and class-specific text summarization. The model assumes that in addition to a set of latent topics that are shared across classes, there is a set of class-specific latent topics for each class. Each document is a probabilistic mixture of the class-specific topics associated with its class and the shared topics. Each class-specific or shared topic has its own probability distribution over a given dictionary. We develop a Bayesian inference of CSTM in the semisupervised scenario, with the supervised scenario as a special case. We analyze in detail the 20 Newsgroups dataset, a benchmark dataset for text classification, and demonstrate that CSTM has better performance than a two stage approach based on latent Dirichlet allocation (LDA), several existing supervised extensions of LDA, and an $L^1$ penalized logistic regression. The favorable performance of CSTM is also demonstrated through Monte Carlo simulations and an analysis of the Reuters dataset. Junni L. Zhang, Jun S. Liu |
J. Mach. Learn. Res. | 5 |
| 2020 | Probabilistic Connection Importance Inference and Lossless Compression of Deep Neural Networks
Long Sha, Pengyu Hong, Zuofeng Shang, Jun S. Liu |
ICLR | 5 |
| 2020 | NGM: Neural Gaussian Mirror for Controlled Feature Selection in Neural NetworksabstractDeep neural networks (DNNs) have become increasingly popular and achieved outstanding performance in predictive tasks. However, the DNN framework itself cannot inform the user which features are more or less relevant for making the prediction, which limits its applicability in many scientific fields. We introduce neural Gaussian mirrors (NGMs), in which mirrored features are created, via a structured perturbation based on a kernel-based conditional dependence measure, to help evaluate feature importance. We design two modifications of the DNN architecture for incorporating mirrored features and providing mirror statistics to measure feature importance. As shown in simulated and real data examples, the proposed method controls the feature selection error rate at a predefined level and maintains a high selection power even with the presence of highly correlated features. Yu Gui, Chenguang Dai, Jun S. Liu |
ICMLA | 4 |
| 2020 | New Algorithms in RNA Structure Prediction Based on BHGabstractThere are some NP-hard problems in the prediction of RNA structures. Prediction of RNA folding structure in RNA nucleotide sequence remains an unsolved challenge. We investigate the computing algorithm in RNA folding structural prediction based on extended structure and basin hopping graph, it is a computing mode of basin hopping graph in RNA folding structural prediction including pseudoknots. This study presents the predicting algorithm based on extended structure, it also proposes an improved computing algorithm based on barrier tree and basin hopping graph, which are the attractive approaches in RNA folding structural prediction. Many experiments have been implemented in Rfam14.1 database and PseudoBase database, the experimental results show that our two algorithms are efficient and accurate than the other existing algorithms. Jun S. Liu |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2020 | Inferring Spatial Organization of Individual Topologically Associated Domains via Piecewise Helical ModelabstractThe recently developed Hi-C technology enables a genome-wide view of chromosome spatial organizations, and has shed deep insights into genome structure and genome function. However, multiple sources of uncertainties make downstream data analysis and interpretation challenging. Specifically, statistical models for inferring three-dimensional (3D) chromosomal structure from Hi-C data are far from their maturity. Most existing methods are highly over-parameterized, lacking clear interpretations, and sensitive to outliers. In this study, we propose a parsimonious, easy to interpret, and robust piecewise helical model for the inference of 3D chromosomal structure of individual topologically associated domain from Hi-C data. When applied to a real Hi-C dataset, the piecewise helical model not only achieves much better model fitting than existing models, but also reveals that geometric properties of chromatin spatial organization are closely related to genome function. Ming Hu 0001, Zhaohui S. Qin, Jun S. Liu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2019 | Characterizing Mutually Exclusive Driver Mutations in Pan-CancerabstractEfforts on cancer genomic study have revealed that mutations in cancer-driving oncogenes tend to mutually exclusive. To understand this tumorigenic mechanism and common cancer features, large-scale genomic study and molecular analyses across many cancer types are essential. We compiled somatic mutation and transcriptome profiles for a pan-cancer dataset that contains tumor tissue samples from 32 cancer types for over 10,000 patients. The genomic and molecular analysis was applied to four pan-cancer subgroups including pan-gastrointestinal, pan-gynecological, pan-kidney and pan-squamous, as well as the entire patient group. We identified seven genes, TP53, PIK3CA, PTEN, PIK3R1, LRP1B, CDH1, and TRRAP exhibited mutually exclusive mutation patterns across all cancer types (adjusted ). The mutations in these genes occurred in 55.9% of all cancer patients and over 70% of patients with pan-gastrointestinal, pan-gynecological and pan-squamous. Excepting LRP1B, the other six genes are included in the known cancer pathways. Notably, TP53, PIK3CA, and PIK3R1 are involved in over 20 cancer-related pathways. The co-expression network analysis suggested that LRP1B was associated with multiple genes that showed significant enrichment in Wnt signal pathway (adjusted p=0.0023) and ERBB signal pathway (adjusted p=0.0097). We further constructed a regulatory network to infer novel regulations of LRP1B in cancer. In addition, based on the mutations in the seven genes, we divided the patients into different groups for survival analysis. We found that the survival rate of the patient group with TP53 mutations was significantly shorter than in other groups (log rank p = 5.88e-10) in pan-gynecological. Our analysis systematically revealed the mutually exclusive mutation patterns of seven cancer genes across 32 cancer types. It helps us to better understanding the relations among different cancer types and pan-cancer subgroups, providing new insights into common mechanisms underlying diverse cancer types. William Yang, Jun S. Liu, Mary Yang |
BIBM | 3 |
| 2017 | Relation discovery and hotspots analysis on diabetes mellitus and obesity with representation modelabstractDiabetes mellitus and obesity are becoming some of the most serious public health challenges in the world. To help researchers more quickly reveal the complex relationships existing between diabetes mellitus, obesity, and related diseases in the literature, and give them an inspiration to search the effective treatments for these diseases, we propose a novel model named as representative latent Dirichlet allocation topic model (RLDA). We conducted the representation learning model on more than 337,000 pieces of diabetes and obesity related literature published in the recent decade. Then, an explicit analysis of the final result using a series of visualization tools to discover meaningful relations among diabetes mellitus, obesity, and other diseases was performed. In order to show the credibility of our discoveries, we used clinical reports, such as Standards of Medical Care in Diabetes, which were not used in our training data, to verify our results. Fortunately, a sufficient number of the reports were direct matches. With the help of our model, we achieved satisfactory results for diabetes mellitus and obesity. For example, we discovered that 22 other diseases are closely related to diabetes mellitus, 10 with obesity and 8 with both. In addition, the tumor, adolescent/child, inflammation, and hypertension will be the hottest research topics relating to diabetes and obesity in the near future. We believe that the representational learning model we have built can help biomedical researchers direct the focus and adjust the direction of their work. Guannan He, Yanchun Liang 0001, William Yang, Jun S. Liu, Mary Yang, Renchu Guan |
BIBM | 5 |
| 2017 | CLIC, a tool for expanding biological pathways based on co-expression across thousands of datasetsabstractIn recent years, there has been a huge rise in the number of publicly available transcriptional profiling datasets. These massive compendia comprise billions of measurements and provide a special opportunity to predict the function of unstudied genes based on co-expression to well-studied pathways. Such analyses can be very challenging, however, since biological pathways are modular and may exhibit co-expression only in specific contexts. To overcome these challenges we introduce CLIC, CLustering by Inferred Co-expression. CLIC accepts as input a pathway consisting of two or more genes. It then uses a Bayesian partition model to simultaneously partition the input gene set into coherent co-expressed modules (CEMs), while assigning the posterior probability for each dataset in support of each CEM. CLIC then expands each CEM by scanning the transcriptome for additional co-expressed genes, quantified by an integrated log-likelihood ratio (LLR) score weighted for each dataset. As a byproduct, CLIC automatically learns the conditions (datasets) within which a CEM is operative. We implemented CLIC using a compendium of 1774 mouse microarray datasets (28628 microarrays) or 1887 human microarray datasets (45158 microarrays). CLIC analysis reveals that of 910 canonical biological pathways, 30% consist of strongly co-expressed gene modules for which new members are predicted. For example, CLIC predicts a functional connection between protein C7orf55 (FMC1) and the mitochondrial ATP synthase complex that we have experimentally validated. CLIC is freely available at www.gene-clic.org. We anticipate that CLIC will be valuable both for revealing new components of biological pathways as well as the conditions in which they are active. Alexis A. Jourdain, Sarah E. Calvo, Jun S. Liu, Vamsi K. Mootha |
PLoS Comput. Biol. | 4 |
| 2016 | Predicting regulatory variants with composite statisticabstractMOTIVATION: Prediction and prioritization of human non-coding regulatory variants is critical for understanding the regulatory mechanisms of disease pathogenesis and promoting personalized medicine. Existing tools utilize functional genomics data and evolutionary information to evaluate the pathogenicity or regulatory functions of non-coding variants. However, different algorithms lead to inconsistent and even conflicting predictions. Combining multiple methods may increase accuracy in regulatory variant prediction. RESULTS: Here, we compiled an integrative resource for predictions from eight different tools on functional annotation of non-coding variants. We further developed a composite strategy to integrate multiple predictions and computed the composite likelihood of a given variant being regulatory variant. Benchmarked by multiple independent causal variants datasets, we demonstrated that our composite model significantly improves the prediction performance. AVAILABILITY AND IMPLEMENTATION: We implemented our model and scoring procedure as a tool, named PRVCS, which is freely available to academic and non-profit usage at http://jjwanglab.org/PRVCS CONTACT: [email protected], [email protected], or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mulin Jun Li, Zipeng Liu, Jiexing Wu, Panwen Wang, Zhengyuan Xia, Pak Chung Sham, Jean-Pierre A. Kocher, Miao-Xin Li, Jun S. Liu, Junwen Wang |
Bioinform. | 12 |
| 2016 | Statistical inference for time course RNA-Seq data using a negative binomial mixed-effect modelabstractBACKGROUND: Accurate identification of differentially expressed (DE) genes in time course RNA-Seq data is crucial for understanding the dynamics of transcriptional regulatory network. However, most of the available methods treat gene expressions at different time points as replicates and test the significance of the mean expression difference between treatments or conditions irrespective of time. They thus fail to identify many DE genes with different profiles across time. In this article, we propose a negative binomial mixed-effect model (NBMM) to identify DE genes in time course RNA-Seq data. In the NBMM, mean gene expression is characterized by a fixed effect, and time dependency is described by random effects. The NBMM is very flexible and can be fitted to both unreplicated and replicated time course RNA-Seq data via a penalized likelihood method. By comparing gene expression profiles over time, we further classify the DE genes into two subtypes to enhance the understanding of expression dynamics. A significance test for detecting DE genes is derived using a Kullback-Leibler distance ratio. Additionally, a significance test for gene sets is developed using a gene set score. RESULTS: Simulation analysis shows that the NBMM outperforms currently available methods for detecting DE genes and gene sets. Moreover, our real data analysis of fruit fly developmental time course RNA-Seq data demonstrates the NBMM identifies biologically relevant genes which are well justified by gene ontology analysis. CONCLUSIONS: The proposed method is powerful and efficient to detect biologically relevant DE genes and gene sets in time course RNA-Seq data. David Dalpiaz, Jun S. Liu, Wenxuan Zhong, Ping Ma 0001 |
BMC Bioinform. | 4 |
| 2016 | On the Characterization of a Class of Fisher-Consistent Loss Functions and its Application to BoostingabstractAccurate classification of categorical outcomes is essential in a wide range of applications. Due to computational issues with minimizing the empirical 0/1 loss, Fisher consistent losses have been proposed as viable proxies. However, even with smooth losses, direct minimization remains a daunting task. To approximate such a minimizer, various boosting algorithms have been suggested. For example, with exponential loss, the AdaBoost algorithm (Freund and Schapire, 1995) is widely used for two- class problems and has been extended to the multi-class setting (Zhu et al., 2009). Alternative loss functions, such as the logistic and the hinge losses, and their corresponding boosting algorithms have also been proposed (Zou et al., 2008; Wang, 2012). In this paper we demonstrate that a broad class of losses, including non-convex functions, achieve Fisher consistency, and in addition can be used for explicit estimation of the conditional class probabilities. Furthermore, we provide a generic boosting algorithm that is not loss-specific. Extensive simulation results suggest that the proposed boosting algorithms could outperform existing methods with properly chosen losses and bags of weak learners. Matey Neykov, Jun S. Liu, Tianxi Cai |
J. Mach. Learn. Res. | 2 |
| 2016 | L1-Regularized Least Squares for Support Recovery of High Dimensional Single Index Models with Gaussian DesignsabstractIt is known that for a certain class of single index models (SIMs) $Y = f(X_{p \times 1}^\top\beta_0, \varepsilon)$, support recovery is impossible when $X \sim \mathcal{N}(0, I_{p \times p})$ and a model complexity adjusted sample size is below a critical threshold. Recently, optimal algorithms based on Sliced Inverse Regression (SIR) were suggested. These algorithms work provably under the assumption that the design $X$ comes from an i.i.d. Gaussian distribution. In the present paper we analyze algorithms based on covariance screening and least squares with $L_1$ penalization (i.e. LASSO) and demonstrate that they can also enjoy optimal (up to a scalar) rescaled sample size in terms of support recovery, albeit under slightly different assumptions on $f$ and $\varepsilon$ compared to the SIR based algorithms. Furthermore, we show more generally, that LASSO succeeds in recovering the signed support of $\beta_0$ if $X \sim \mathcal{N}(0, \Sigma)$, and the covariance $\Sigma$ satisfies the irrepresentable condition. Our work extends existing results on the support recovery of LASSO for the linear model, to a more general class of SIMs. Matey Neykov, Jun S. Liu, Tianxi Cai |
J. Mach. Learn. Res. | 2 |
| 2015 | Conformational sampling and structure prediction of multiple interacting loops in soluble and β-barrel membrane proteins using multi-loop distance-guided chain-growth Monte Carlo methodabstractMOTIVATION: Loops in proteins are often involved in biochemical functions. Their irregularity and flexibility make experimental structure determination and computational modeling challenging. Most current loop modeling methods focus on modeling single loops. In protein structure prediction, multiple loops often need to be modeled simultaneously. As interactions among loops in spatial proximity can be rather complex, sampling the conformations of multiple interacting loops is a challenging task. RESULTS: In this study, we report a new method called multi-loop Distance-guided Sequential chain-Growth Monte Carlo (M-DiSGro) for prediction of the conformations of multiple interacting loops in proteins. Our method achieves an average RMSD of 1.93 Å for lowest energy conformations of 36 pairs of interacting protein loops with the total length ranging from 12 to 24 residues. We further constructed a data set containing proteins with 2, 3 and 4 interacting loops. For the most challenging target proteins with four loops, the average RMSD of the lowest energy conformations is 2.35 Å. Our method is also tested for predicting multiple loops in β-barrel membrane proteins. For outer-membrane protein G, the lowest energy conformation has a RMSD of 2.62 Å for the three extracellular interacting loops with a total length of 34 residues (12, 12 and 10 residues in each loop). AVAILABILITY AND IMPLEMENTATION: The software is freely available at: tanto.bioe.uic.edu/m-DiSGro. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Samuel W. K. Wong, Jun S. Liu |
Bioinform. | 3 |
| 2015 | dslice: an R package for nonparametric testing of associations with application in QTL and gene set analysisabstractUNLABELLED: Many statistical problems in bioinformatics and genetics can be formulated as the testing of associations between a categorical variable and a continuous variable. A dynamic slicing method was proposed for non-parametric dependence testing, which has been demonstrated to have higher powers compared with traditional methods such as Kolmogorov-Smirnov test. We introduce an R package dslice to facilitate the use of dynamic slicing method in bioinformatic applications such as quantitative trait loci study and gene set enrichment analysis. AVAILABILITY AND IMPLEMENTATION: dslice is implemented in Rcpp and available in the Comprehensive R Archive Network. The package is distributed under the GNU General Public License (version 2 or later). Bo Jiang 0005, Xuegong Zhang, Jun S. Liu |
Bioinform. | 4 |
| 2014 | Advances in translational bioinformatics facilitate revealing the landscape of complex disease mechanismsabstractAdvances of high-throughput technologies have rapidly produced more and more data from DNAs and RNAs to proteins, especially large volumes of genome-scale data. However, connection of the genomic information to cellular functions and biological behaviours relies on the development of effective approaches at higher systems level. In particular, advances in RNA-Seq technology has helped the studies of transcriptome, RNA expressed from the genome, while systems biology on the other hand provides more comprehensive pictures, from which genes and proteins actively interact to lead to cellular behaviours and physiological phenotypes. As biological interactions mediate many biological processes that are essential for cellular function or disease development, it is important to systematically identify genomic information including genetic mutations from GWAS (genome-wide association study), differentially expressed genes, bidirectional promoters, intrinsic disordered proteins (IDP) and protein interactions to gain deep insights into the underlying mechanisms of gene regulations and networks. Furthermore, bidirectional promoters can co-regulate many biological pathways, where the roles of bidirectional promoters can be studied systematically for identifying co-regulating genes at interactive network level. Combining information from different but related studies can ultimately help revealing the landscape of molecular mechanisms underlying complex diseases such as cancer. Jack Y. Yang, A. Keith Dunker, Jun S. Liu, Xiang Qin, Hamid R. Arabnia, William Yang, Andrzej Niemierko, Zhongxue Chen, Zuojie Luo, Liangjiang Wang, Youping Deng, Weida Tong, Mary Yang |
BMC Bioinform. | 3 |
| 2014 | Identification of genes and pathways involved in kidney renal clear cell carcinomaabstractBACKGROUND: Kidney Renal Clear Cell Carcinoma (KIRC) is one of fatal genitourinary diseases and accounts for most malignant kidney tumours. KIRC has been shown resistance to radiotherapy and chemotherapy. Like many types of cancers, there is no curative treatment for metastatic KIRC. Using advanced sequencing technologies, The Cancer Genome Atlas (TCGA) project of NIH/NCI-NHGRI has produced large-scale sequencing data, which provide unprecedented opportunities to reveal new molecular mechanisms of cancer. We combined differentially expressed genes, pathways and network analyses to gain new insights into the underlying molecular mechanisms of the disease development. RESULTS: Followed by the experimental design for obtaining significant genes and pathways, comprehensive analysis of 537 KIRC patients' sequencing data provided by TCGA was performed. Differentially expressed genes were obtained from the RNA-Seq data. Pathway and network analyses were performed. We identified 186 differentially expressed genes with significant p-value and large fold changes (P < 0.01, |log(FC)| > 5). The study not only confirmed a number of identified differentially expressed genes in literature reports, but also provided new findings. We performed hierarchical clustering analysis utilizing the whole genome-wide gene expressions and differentially expressed genes that were identified in this study. We revealed distinct groups of differentially expressed genes that can aid to the identification of subtypes of the cancer. The hierarchical clustering analysis based on gene expression profile and differentially expressed genes suggested four subtypes of the cancer. We found enriched distinct Gene Ontology (GO) terms associated with these groups of genes. Based on these findings, we built a support vector machine based supervised-learning classifier to predict unknown samples, and the classifier achieved high accuracy and robust classification results. In addition, we identified a number of pathways (P < 0.04) that were significantly influenced by the disease. We found that some of the identified pathways have been implicated in cancers from literatures, while others have not been reported in the cancer before. The network analysis leads to the identification of significantly disrupted pathways and associated genes involved in the disease development. Furthermore, this study can provide a viable alternative in identifying effective drug targets. CONCLUSIONS: Our study identified a set of differentially expressed genes and pathways in kidney renal clear cell carcinoma, and represents a comprehensive computational approach to analysis large-scale next-generation sequencing data. The pathway and network analyses suggested that information from distinctly expressed genes can be utilized in the identification of aberrant upstream regulators. Identification of distinctly expressed genes and altered pathways are important in effective biomarker identification for early cancer diagnosis and treatment planning. Combining differentially expressed genes with pathway and network analyses using intelligent computational approaches provide an unprecedented opportunity to identify upstream disease causal genes and effective drug targets. William Yang, Kenji Yoshigoe, Xiang Qin, Jun S. Liu, Jack Y. Yang, Andrzej Niemierko, Youping Deng, A. Keith Dunker, Zhongxue Chen, Liangjiang Wang, Hamid R. Arabnia, Weida Tong, Mary Yang |
BMC Bioinform. | 4 |
| 2013 | Bayesian hierarchical model of protein-binding microarray k-mer data reduces noise and identifies transcription factor subclasses and preferred k-mersabstractMOTIVATION: Sequence-specific transcription factors (TFs) regulate the expression of their target genes through interactions with specific DNA-binding sites in the genome. Data on TF-DNA binding specificities are essential for understanding how regulatory specificity is achieved. RESULTS: Numerous studies have used universal protein-binding microarray (PBM) technology to determine the in vitro binding specificities of hundreds of TFs for all possible 8 bp sequences (8mers). We have developed a Bayesian analysis of variance (ANOVA) model that decomposes these 8mer data into background noise, TF familywise effects and effects due to the particular TF. Adjusting for background noise improves PBM data quality and concordance with in vivo TF binding data. Moreover, our model provides simultaneous identification of TF subclasses and their shared sequence preferences, and also of 8mers bound preferentially by individual members of TF subclasses. Such results may aid in deciphering cis-regulatory codes and determinants of protein-DNA binding specificity. AVAILABILITY AND IMPLEMENTATION: Source code, compiled code and R and Python scripts are available from http://thebrain.bwh.harvard.edu/hierarchicalANOVA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bo Jiang 0005, Jun S. Liu, Martha L. Bulyk |
Bioinform. | 2 |
| 2013 | Bayesian Inference of Spatial Organizations of ChromosomesabstractKnowledge of spatial chromosomal organizations is critical for the study of transcriptional regulation and other nuclear processes in the cell. Recently, chromosome conformation capture (3C) based technologies, such as Hi-C and TCC, have been developed to provide a genome-wide, three-dimensional (3D) view of chromatin organization. Appropriate methods for analyzing these data and fully characterizing the 3D chromosomal structure and its structural variations are still under development. Here we describe a novel Bayesian probabilistic approach, denoted as "Bayesian 3D constructor for Hi-C data" (BACH), to infer the consensus 3D chromosomal structure. In addition, we describe a variant algorithm BACH-MIX to study the structural variations of chromatin in a cell population. Applying BACH and BACH-MIX to a high resolution Hi-C dataset generated from mouse embryonic stem cells, we found that most local genomic regions exhibit homogeneous 3D chromosomal structures. We further constructed a model for the spatial arrangement of chromatin, which reveals structural properties associated with euchromatic and heterochromatic regions in the genome. We observed strong associations between structural properties and several genomic and epigenetic features of the chromosome. Using BACH-MIX, we further found that the structural variations of chromatin are correlated with these genomic and epigenetic features. Our results demonstrate that BACH and BACH-MIX have the potential to provide new insights into the chromosomal architecture of mammalian cells. Ming Hu 0001, Zhaohui S. Qin, Jesse R. Dixon, Siddarth Selvaraj, Jennifer Fang, Jun S. Liu |
PLoS Comput. Biol. | 8 |
| 2012 | IMID: integrated molecular interaction databaseabstractMOTIVATION: Molecular interaction information, such as protein-protein interactions and protein-small molecule interactions, is indispensable for understanding the mechanism of biological processes and discovering treatments for diseases. Many databases have been built by manual annotation of literature to organize such information into structured form. However, most databases focus on only one type of interactions, which are often not well annotated and integrated with related functional information. RESULTS: In this study, we integrate molecular interaction information from literature by automatic information extraction and from manually annotated databases. We further integrate the relationships between protein/gene and other bio-entity terms including gene ontology terms, pathways, species and diseases to build an integrated molecular interaction database (IMID). Interactions can be selected by their associated probabilities. IMID allows complex and versatile queries for context-specific molecular interactions, which are not available currently in other molecular interaction databases. AVAILABILITY: The database is located at www.integrativebiology.org. Sentil Balaji, Charles Mcclendon, Rajesh Chowdhary, Jun S. Liu |
Bioinform. | 4 |
| 2012 | GFOLD: a generalized fold change for ranking differentially expressed genes from RNA-seq dataabstractMOTIVATION: RNA-seq has been widely used in transcriptome analysis to effectively measure gene expression levels. Although sequencing costs are rapidly decreasing, almost 70% of all the human RNA-seq samples in the gene expression omnibus do not have biological replicates and more unreplicated RNA-seq data were published than replicated RNA-seq data in 2011. Despite the large amount of single replicate studies, there is currently no satisfactory method for detecting differentially expressed genes when only a single biological replicate is available. RESULTS: We present the GFOLD (generalized fold change) algorithm to produce biologically meaningful rankings of differentially expressed genes from RNA-seq data. GFOLD assigns reliable statistics for expression changes based on the posterior distribution of log fold change. In this way, GFOLD overcomes the shortcomings of P-value and fold change calculated by existing RNA-seq analysis methods and gives more stable and biological meaningful gene rankings when only a single biological replicate is available. AVAILABILITY: The open source C/C++ program is available at http://www.tongji.edu.cn/∼zhanglab/GFOLD/index.html Jianxing Feng, Clifford A. Meyer, Jun S. Liu, Xiaole Shirley Liu, Yong Zhang 0006 |
Bioinform. | 4 |
| 2012 | HiCNorm: removing biases in Hi-C data via Poisson regressionabstractSUMMARY: We propose a parametric model, HiCNorm, to remove systematic biases in the raw Hi-C contact maps, resulting in a simple, fast, yet accurate normalization procedure. Compared with the existing Hi-C normalization method developed by Yaffe and Tanay, HiCNorm has fewer parameters, runs >1000 times faster and achieves higher reproducibility. AVAILABILITY: Freely available on the web at: http://www.people.fas.harvard.edu/∼junliu/HiCNorm/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ming Hu 0001, Siddarth Selvaraj, Zhaohui S. Qin, Jun S. Liu |
Bioinform. | 6 |
| 2012 | Using Poisson mixed-effects model to quantify transcript-level gene expression in RNA-SeqabstractMOTIVATION: RNA sequencing (RNA-Seq) is a powerful new technology for mapping and quantifying transcriptomes using ultra high-throughput next-generation sequencing technologies. Using deep sequencing, gene expression levels of all transcripts including novel ones can be quantified digitally. Although extremely promising, the massive amounts of data generated by RNA-Seq, substantial biases and uncertainty in short read alignment pose challenges for data analysis. In particular, large base-specific variation and between-base dependence make simple approaches, such as those that use averaging to normalize RNA-Seq data and quantify gene expressions, ineffective. RESULTS: In this study, we propose a Poisson mixed-effects (POME) model to characterize base-level read coverage within each transcript. The underlying expression level is included as a key parameter in this model. Since the proposed model is capable of incorporating base-specific variation as well as between-base dependence that affect read coverage profile throughout the transcript, it can lead to improved quantification of the true underlying expression level. AVAILABILITY AND IMPLEMENTATION: POME can be freely downloaded at http://www.stat.purdue.edu/~yuzhu/pome.html. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ming Hu 0001, Jeremy M. G. Taylor, Jun S. Liu, Zhaohui S. Qin |
Bioinform. | 4 |
| 2010 | Tmod: toolbox of motif discoveryabstractSUMMARY: Motif discovery is an important topic in computational transcriptional regulation studies. In the past decade, many researchers have contributed to the field and many de novo motif-finding tools have been developed, each may have a different strength. However, most of these tools do not have a user-friendly interface and their results are not easily comparable. We present a software called Toolbox of Motif Discovery (Tmod) for Windows operating systems. The current version of Tmod integrates 12 widely used motif discovery programs: MDscan, BioProspector, AlignACE, Gibbs Motif Sampler, MEME, CONSENSUS, MotifRegressor, GLAM, MotifSampler, SeSiMCMC, Weeder and YMF. Tmod provides a unified interface to ease the use of these programs and help users to understand the tuning parameters. It allows plug-in motif-finding programs to run either separately or in a batch mode with predetermined parameters, and provides a summary comprising of outputs from multiple programs. Tmod is developed in C++ with the support of Microsoft Foundation Classes and Cygwin. Tmod can also be easily expanded to include future algorithms. AVAILABILITY: Tmod is available for download at http://www.fas.harvard.edu/~junliu/Tmod/. Hanchang Sun, Jun S. Liu, Hongwei Xie |
Bioinform. | 5 |
| 2010 | A Bayesian Partition Method for Detecting Pleiotropic and Epistatic eQTL ModulesabstractStudies of the relationship between DNA variation and gene expression variation, often referred to as "expression quantitative trait loci (eQTL) mapping", have been conducted in many species and resulted in many significant findings. Because of the large number of genes and genetic markers in such analyses, it is extremely challenging to discover how a small number of eQTLs interact with each other to affect mRNA expression levels for a set of co-regulated genes. We present a Bayesian method to facilitate the task, in which co-expressed genes mapped to a common set of markers are treated as a module characterized by latent indicator variables. A Markov chain Monte Carlo algorithm is designed to search simultaneously for the module genes and their linked markers. We show by simulations that this method is more powerful for detecting true eQTLs and their target genes than traditional QTL mapping methods. We applied the procedure to a data set consisting of gene expression and genotypes for 112 segregants of S. cerevisiae. Our method identified modules containing genes mapped to previously reported eQTL hot spots, and dissected these large eQTL hot spots into several modules corresponding to possibly different biological functions or primary and secondary responses to regulatory perturbations. In addition, we identified nine modules associated with pairs of eQTLs, of which two have been previously reported. We demonstrated that one of the novel modules containing many daughter-cell expressed genes is regulated by AMN1 and BPH1. In conclusion, the Bayesian partition method which simultaneously considers all traits and all markers is more powerful for detecting both pleiotropic and epistatic effects based on both simulated and empirical data. Eric E. Schadt, Jun S. Liu |
PLoS Comput. Biol. | 4 |
| 2009 | Bayesian inference of protein-protein interactions from biological literatureabstractMOTIVATION: Protein-protein interaction (PPI) extraction from published biological articles has attracted much attention because of the importance of protein interactions in biological processes. Despite significant progress, mining PPIs from literatures still rely heavily on time- and resource-consuming manual annotations. RESULTS: In this study, we developed a novel methodology based on Bayesian networks (BNs) for extracting PPI triplets (a PPI triplet consists of two protein names and the corresponding interaction word) from unstructured text. The method achieved an overall accuracy of 87% on a cross-validation test using manually annotated dataset. We also showed, through extracting PPI triplets from a large number of PubMed abstracts, that our method was able to complement human annotations to extract large number of new PPIs from literature. AVAILABILITY: Programs/scripts we developed/used in the study are available at http://stat.fsu.edu/~jinfeng/datasets/Bio-SI-programs-Bayesian-chowdhary-zhang-liu.zip. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rajesh Chowdhary, Jun S. Liu |
Bioinform. | 3 |
| 2009 | Information Flow Analysis of Interactome NetworksabstractRecent studies of cellular networks have revealed modular organizations of genes and proteins. For example, in interactome networks, a module refers to a group of interacting proteins that form molecular complexes and/or biochemical pathways and together mediate a biological process. However, it is still poorly understood how biological information is transmitted between different modules. We have developed information flow analysis, a new computational approach that identifies proteins central to the transmission of biological information throughout the network. In the information flow analysis, we represent an interactome network as an electrical circuit, where interactions are modeled as resistors and proteins as interconnecting junctions. Construing the propagation of biological signals as flow of electrical current, our method calculates an information flow score for every protein. Unlike previous metrics of network centrality such as degree or betweenness that only consider topological features, our approach incorporates confidence scores of protein-protein interactions and automatically considers all possible paths in a network when evaluating the importance of each protein. We apply our method to the interactome networks of Saccharomyces cerevisiae and Caenorhabditis elegans. We find that the likelihood of observing lethality and pleiotropy when a protein is eliminated is positively correlated with the protein's information flow score. Even among proteins of low degree or low betweenness, high information scores serve as a strong predictor of loss-of-function lethality or pleiotropy. The correlation between information flow scores and phenotypes supports our hypothesis that the proteins of high information flow reside in central positions in interactome networks. We also show that the ranks of information flow scores are more consistent than that of betweenness when a large amount of noisy data is added to an interactome. Finally, we combine gene expression data with interaction data in C. elegans and construct an interactome network for muscle-specific genes. We find that genes that rank high in terms of information flow in the muscle interactome network but not in the entire network tend to play important roles in muscle function. This framework for studying tissue-specific networks by the information flow model can be applied to other tissues and other organisms as well. Patrycja Vasilyev Missiuro, Kesheng Liu, Lihua Zou, Brian C. Ross, Guoyan Zhao 0002, Jun S. Liu |
PLoS Comput. Biol. | 6 |
| 2008 | PFP: A Computational Framework for Phylogenetic Footprinting in Prokaryotic Genomes
Dongsheng Che, Shane T. Jensen, Jun S. Liu, Ying Xu 0001 |
ISBRA | 4 |
| 2008 | Genomic Sequence Is Highly Predictive of Local Nucleosome DepletionabstractThe regulation of DNA accessibility through nucleosome positioning is important for transcription control. Computational models have been developed to predict genome-wide nucleosome positions from DNA sequences, but these models consider only nucleosome sequences, which may have limited their power. We developed a statistical multi-resolution approach to identify a sequence signature, called the N-score, that distinguishes nucleosome binding DNA from non-nucleosome DNA. This new approach has significantly improved the prediction accuracy. The sequence information is highly predictive for local nucleosome enrichment or depletion, whereas predictions of the exact positions are only modestly more accurate than a null model, suggesting the importance of other regulatory factors in fine-tuning the nucleosome positions. The N-score in promoter regions is negatively correlated with gene expression levels. Regulatory elements are enriched in low N-score regions. While our model is derived from yeast data, the N-score pattern computed from this model agrees well with recent high-resolution protein-binding data in human. Guo-Cheng Yuan, Jun S. Liu |
PLoS Comput. Biol. | 2 |
| 2008 | Systematic Analysis of Pleiotropy in C. elegans Early EmbryogenesisabstractPleiotropy refers to the phenomenon in which a single gene controls several distinct, and seemingly unrelated, phenotypic effects. We use C. elegans early embryogenesis as a model to conduct systematic studies of pleiotropy. We analyze high-throughput RNA interference (RNAi) data from C. elegans and identify "phenotypic signatures", which are sets of cellular defects indicative of certain biological functions. By matching phenotypic profiles to our identified signatures, we assign genes with complex phenotypic profiles to multiple functional classes. Overall, we observe that pleiotropy occurs extensively among genes involved in early embryogenesis, and a small proportion of these genes are highly pleiotropic. We hypothesize that genes involved in early embryogenesis are organized into partially overlapping functional modules, and that pleiotropic genes represent "connectors" between these modules. In support of this hypothesis, we find that highly pleiotropic genes tend to reside in central positions in protein-protein interaction networks, suggesting that pleiotropic genes act as connecting points between different protein complexes or pathways. Lihua Zou, Sira Sriswasdi, Brian C. Ross, Patrycja Vasilyev Missiuro, Jun S. Liu |
PLoS Comput. Biol. | 5 |
| 2007 | Statistial Analysis of nucleosome occupancy and histone modification dataabstractIn eukaryotic cells, genomic DNAs wrap around beadlike molecules, called nucleosomes, so as to pack more compactly in the nucleus of the cell. The nucleosome is made up of four pairs of histone proteins (H2A, H2B, H3, and H4) who share a very similar structural motif. The positioning of nucleosomes as well as the modifications of various sites of histone proteins (such as acetylation) plays important but incompletely understood roles in gene regulation. We propose statistical models for predicting nucleosome positioning and histone modification patterns using only genomic sequence information. Computation models have been developed to predict genome-wide nucleosome positions from DNA sequences, but these models consider only nucleosome sequences, which may have limited their power. We developed a statistical multi-resolution approach to identify a sequence signature, called the N-score, that distinguishes nucleosome binding DNA from non-nucleosome DNA. The N-score is not sensitive to deletion of short DNA elements and can also be estimated reasonably accurately from coarse nucleosome positioning data. We found that the sequence information is highly predictive for local nucleosome enrichment or depletion, whereas the exact positions may be further fine-tuned by other regulatory factors. We observed that many characteristics of nucleosome positioning, such as the nucleosome depletion in the promoter regions, can be predicted accurately by the sequence information through N-scores. In addition to nucleosome positioning, histone acetylations are also important in directing gene regulation. A comprehensive understanding of the regulatory role of histone acetylation is difficult because many different histone acetylation patterns exist and their effects are confounded by other factors, such as the transcription factor binding sequence motif information and nucleosome occupancy. We analyzed recent genomewide histone acetylation data using a few complementary statistical models and tested the validity of a cumulative model in approximating the global regulatory effect of histone acetylation. Confounding effects due to transcription factor binding sequence information were estimated by using two independent motif-based algorithms followed by a variable selection method. Our analysis confirms that histone acetylation has a significant effect on transcription rates in addition to that attributable to upstream sequence motifs. Our model fits well with observed genome-wide data. Strikingly, including more complicated combinatorial effects does not improve the model's performance. Through an analysis of conditional independence, we found that H4 acetylation may not have significant direct impact on global gene expression. Guo-Cheng Yuan, Jun S. Liu |
BIBE | 2 |
| 2007 | Statistical power of phylo-HMM for evolutionarily conserved element detectionabstractBACKGROUND: An important goal of comparative genomics is the identification of functional elements through conservation analysis. Phylo-HMM was recently introduced to detect conserved elements based on multiple genome alignments, but the method has not been rigorously evaluated. RESULTS: We report here a simulation study to investigate the power of phylo-HMM. We show that the power of the phylo-HMM approach depends on many factors, the most important being the number of species-specific genomes used and evolutionary distances between pairs of species. This finding is consistent with results reported by other groups for simpler comparative genomics models. In addition, the conservation ratio of conserved elements and the expected length of the conserved elements are also major factors. In contrast, the influence of the topology and the nucleotide substitution model are relatively minor factors. CONCLUSION: Our results provide for general guidelines on how to select the number of genomes and their evolutionary distance in comparative genomics studies, as well as the level of power we can expect under different parameter settings. Xiaodan Fan, Eric E. Schadt, Jun S. Liu |
BMC Bioinform. | 4 |
| 2007 | Predicting Gene Expression from Sequence: A ReexaminationabstractAlthough much of the information regarding genes' expressions is encoded in the genome, deciphering such information has been very challenging. We reexamined Beer and Tavazoie's (BT) approach to predict mRNA expression patterns of 2,587 genes in Saccharomyces cerevisiae from the information in their respective promoter sequences. Instead of fitting complex Bayesian network models, we trained naïve Bayes classifiers using only the sequence-motif matching scores provided by BT. Our simple models correctly predict expression patterns for 79% of the genes, based on the same criterion and the same cross-validation (CV) procedure as BT, which compares favorably to the 73% accuracy of BT. The fact that our approach did not use position and orientation information of the predicted binding sites but achieved a higher prediction accuracy, motivated us to investigate a few biological predictions made by BT. We found that some of their predictions, especially those related to motif orientations and positions, are at best circumstantial. For example, the combinatorial rules suggested by BT for the PAC and RRPE motifs are not unique to the cluster of genes from which the predictive model was inferred, and there are simpler rules that are statistically more significant than BT's ones. We also show that CV procedure used by BT to estimate their method's prediction accuracy is inappropriate and may have overestimated the prediction accuracy by about 10%. Lei Guo 0013, Jun S. Liu |
PLoS Comput. Biol. | 4 |
| 2006 | Bayesian models for pooling microarray studies with multiple sources of replicationsabstractBACKGROUND: Biologists often conduct multiple but different cDNA microarray studies that all target the same biological system or pathway. Within each study, replicate slides within repeated identical experiments are often produced. Pooling information across studies can help more accurately identify true target genes. Here, we introduce a method to integrate multiple independent studies efficiently. RESULTS: We introduce a Bayesian hierarchical model to pool cDNA microarray data across multiple independent studies to identify highly expressed genes. Each study has multiple sources of variation, i.e. replicate slides within repeated identical experiments. Our model produces the gene-specific posterior probability of differential expression, which provides a direct method for ranking genes, and provides Bayesian estimates of false discovery rates (FDR). In simulations combining two and five independent studies, with fixed FDR levels, we observed large increases in the number of discovered genes in pooled versus individual analyses. When the number of output genes is fixed (e.g., top 100), the pooled model found appreciably more truly differentially expressed genes than the individual studies. We were also able to identify more differentially expressed genes from pooling two independent studies in Bacillus subtilis than from each individual data set. Finally, we observed that in our simulation studies our Bayesian FDR estimates tracked the true FDRs very well. CONCLUSION: Our method provides a cohesive framework for combining multiple but not identical microarray studies with several sources of replication, with data produced from the same platform. We assume that each study contains only two conditions: an experimental and a control sample. We demonstrated our model's suitability for a small number of studies that have been either pre-scaled or have no outliers. Erin M. Conlon, Joon J. Song, Jun S. Liu |
BMC Bioinform. | 3 |
| 2006 | Recursive SVM feature selection and sample classification for mass-spectrometry and microarray dataabstractBACKGROUND: Like microarray-based investigations, high-throughput proteomics techniques require machine learning algorithms to identify biomarkers that are informative for biological classification problems. Feature selection and classification algorithms need to be robust to noise and outliers in the data. RESULTS: We developed a recursive support vector machine (R-SVM) algorithm to select important genes/biomarkers for the classification of noisy data. We compared its performance to a similar, state-of-the-art method (SVM recursive feature elimination or SVM-RFE), paying special attention to the ability of recovering the true informative genes/biomarkers and the robustness to outliers in the data. Simulation experiments show that a 5%- approximately 20% improvement over SVM-RFE can be achieved regard to these properties. The SVM-based methods are also compared with a conventional univariate method and their respective strengths and weaknesses are discussed. R-SVM was applied to two sets of SELDI-TOF-MS proteomics data, one from a human breast cancer study and the other from a study on rat liver cirrhosis. Important biomarkers found by the algorithm were validated by follow-up biological experiments. CONCLUSION: The proposed R-SVM method is suitable for analyzing noisy high-throughput proteomics and microarray data and it outperforms SVM-RFE in the robustness to noise and in the ability to recover informative features. The multivariate SVM-based method outperforms the univariate method in the classification performance, but univariate methods can reveal more of the differentially expressed features especially when there are correlations between the features. Xuegong Zhang, Xiu-qin Xu, Hon-chiu E. Leung, Lyndsay N. Harris, James D. Iglehart, Alexander Miron, Jun S. Liu, Wing Hung Wong |
BMC Bioinform. | 9 |
| 2006 | On Side-Chain Conformational Entropy of ProteinsabstractThe role of side-chain entropy (SCE) in protein folding has long been speculated about but is still not fully understood. Utilizing a newly developed Monte Carlo method, we conducted a systematic investigation of how the SCE relates to the size of the protein and how it differs among a protein's X-ray, NMR, and decoy structures. We estimated the SCE for a set of 675 nonhomologous proteins, and observed that there is a significant SCE for both exposed and buried residues for all these proteins-the contribution of buried residues approaches approximately 40% of the overall SCE. Furthermore, the SCE can be quite different for structures with similar compactness or even similar conformations. As a striking example, we found that proteins' X-ray structures appear to pack more "cleverly" than their NMR or decoy counterparts in the sense of retaining higher SCE while achieving comparable compactness, which suggests that the SCE plays an important role in favouring native protein structures. By including a SCE term in a simple free energy function, we can significantly improve the discrimination of native protein structures from decoys. Jun S. Liu |
PLoS Comput. Biol. | 2 |
| 2005 | BEST: Binding-site Estimation Suite of ToolsabstractSUMMARY: The purpose of our Binding-site Estimation Suite of Tools (BEST) is two-fold: to provide a platform for using and comparing different motif-finding programs for transcription factor binding site prediction, and to improve the accuracy of these predictions by further optimization. Our software package BEST includes four commonly used motif-finding programs: AlignACE, BioProspector, CONSENSUS and MEME, as well as the optimization program BioOptimizer. BEST allows the user to run programs either separately or sequentially and manages all programs by automating the common inputs and the optimization procedure. The BEST system was implemented in Qt, a C++ application development framework, and was compiled and executed on Linux operating systems. AVAILABILITY: BEST is available for download at http://www.cs.uga.edu/~che/BEST and http://www.fas.harvard.edu/~junliu/BEST CONTACT: [email protected], [email protected]. Dongsheng Che, Shane T. Jensen, Liming Cai, Jun S. Liu |
Bioinform. | 4 |
| 2005 | A boosting approach for motif modeling using ChIP-chip dataabstractMotivation: Building an accurate binding model for a transcription factor (TF) is essential to differentiate its true binding targets from those spurious ones. This is an important step toward understanding gene regulation. Results: This paper describes a boosting approach to modeling TF–DNA binding. Different from the widely used weight matrix model, which predicts TF–DNA binding based on a linear combination of position-specific contributions, our approach builds a TF binding classifier by combining a set of weight matrix based classifiers, thus yielding a non-linear binding decision rule. The proposed approach was applied to the ChIP-chip data of Saccharomyces cerevisiae. When compared with the weight matrix method, our new approach showed significant improvements on the specificity in a majority of cases. Contact: [email protected] Supplementary information: The software and the Supplementary data are available at http://biogibbs.stanford.edu/~hong2004/MotifBooster/. Pengyu Hong, Xiaole Shirley Liu, Jun S. Liu, Wing Hung Wong |
Bioinform. | 5 |
| 2005 | Combining phylogenetic motif discovery and motif clustering to predict co-regulated genesabstractMOTIVATION: We present a sequence-based framework and algorithm PHYLOCLUS for predicting co-regulated genes. In our approach, de novo discovery methods are used to find motifs conserved by evolution and then a Bayesian hierarchical clustering model is used to cluster these motifs, thereby grouping together genes that are putatively co-regulated. Our clustering procedure allows both the number of clusters and the motif width within each cluster to be unknown. RESULTS: We use our framework to predict co-regulated genes in the bacterium Bacillus subtilis using six other closely related bacterial species. Our predicted motifs and gene clusters are validated using several external sources and significant clusters are examined in detail. An extension to the discovery and clustering of two-block motifs can be used for inference about synergistic binding relationships between transcription factors. AVAILABILITY: Software and Supplementary Materials can be downloaded at http://stat.wharton.upenn.edu/~stjensen/research/phyloclus.html or http://www.fas.harvard.edu/~junliu/phyloclus.html CONTACT: [email protected]. Shane T. Jensen, Jun S. Liu |
Bioinform. | 3 |
| 2005 | HapBlock: haplotype block partitioning and tag SNP selection software using a set of dynamic programming algorithmsabstractUNLABELLED: Recent studies have revealed that linkage disequilibrium (LD) patterns vary across the human genome with some regions of high LD interspersed with regions of low LD. Such LD patterns make it possible to select a set of single nucleotide polymorphism (SNPs; tag SNPs) for genome-wide association studies. We have developed a suite of computer programs to analyze the block-like LD patterns and to select the corresponding tag SNPs. Compared to other programs for haplotype block partitioning and tag SNP selection, our program has several notable features. First, the dynamic programming algorithms implemented are guaranteed to find the block partition with minimum number of tag SNPs for the given criteria of blocks and tag SNPs. Second, both haplotype data and genotype data from unrelated individuals and/or from general pedigrees can be analyzed. Third, several existing measures/criteria for haplotype block partitioning and tag SNP selection have been implemented in the program. Finally, the programs provide flexibility to include specific SNPs (e.g. non-synonymous SNPs) as tag SNPs. AVAILABILITY: The HapBlock program and its supplemental documents can be downloaded from the website http://www.cmb.usc.edu/~msms/HapBlock. Zhaohui S. Qin, Ting Chen 0006, Jun S. Liu, Michael S. Waterman, Fengzhu Sun |
Bioinform. | 4 |
| 2005 | RSIR: regularized sliced inverse regression for motif discoveryabstractMOTIVATION: Identification of transcription factor binding motifs (TFBMs) is a crucial first step towards the understanding of regulatory circuitries controlling the expression of genes. In this paper, we propose a novel procedure called regularized sliced inverse regression (RSIR) for identifying TFBMs. RSIR follows a recent trend to combine information contained in both gene expression measurements and genes' promoter sequences. Compared with existing methods, RSIR is efficient in computation, very stable for data with high dimensionality and high collinearity, and improves motif detection sensitivities and specificities by avoiding inappropriate model specification. RESULTS: We compare RSIR with SIR and stepwise regression based on simulated data and find that RSIR has a lower false positive rate. We also demonstrate an excellent performance of RSIR by applying it to the yeast amino acid starvation data and cell cycle data. AVAILABILITY: Matlab programs are available upon request from the authors. Wenxuan Zhong, Ping Ma 0001, Jun S. Liu, Michael Yu Zhu |
Bioinform. | 4 |
| 2004 | BioOptimizer: a Bayesian scoring function approach to motif discoveryabstractMOTIVATION: Transcription factors (TFs) bind directly to short segments on the genome, often within hundreds to thousands of base pairs upstream of gene transcription start sites, to regulate gene expression. The experimental determination of TFs binding sites is expensive and time-consuming. Many motif-finding programs have been developed, but no program is clearly superior in all situations. Practitioners often find it difficult to judge which of the motifs predicted by these algorithms are more likely to be biologically relevant. RESULTS: We derive a comprehensive scoring function based on a full Bayesian model that can handle unknown site abundance, unknown motif width and two-block motifs with variable-length gaps. An algorithm called BioOptimizer is proposed to optimize this scoring function so as to reduce noise in the motif signal found by any motif-finding program. The accuracy of BioOptimizer, which can be used in conjunction with several existing programs, is shown to be superior to using any of these motif-finding programs alone when evaluated by both simulation studies and application to sets of co-regulated genes in bacteria. In addition, this scoring function formulation enables us to compare objectively different predicted motifs and select the optimal ones, effectively combining the strengths of existing programs. AVAILABILITY: BioOptimizer is available for download at www.fas.harvard.edu/~junliu/BioOptimizer/ Shane T. Jensen, Jun S. Liu |
Bioinform. | 2 |
| 2004 | Modeling within-motif dependence for transcription factor binding site predictionsabstractMOTIVATION: The position-specific weight matrix (PWM) model, which assumes that each position in the DNA site contributes independently to the overall protein-DNA interaction, has been the primary means to describe transcription factor binding site motifs. Recent biological experiments, however, suggest that there exists interdependence among positions in the binding sites. In order to exploit this interdependence to aid motif discovery, we extend the PWM model to include pairs of correlated positions and design a Markov chain Monte Carlo algorithm to sample in the model space. We then combine the model sampling step with the Gibbs sampling framework for de novo motif discoveries. RESULTS: Testing on experimentally validated binding sites, we find that about 25% of the transcription factor binding motifs show significant within-site position correlations, and 80% of these motif models can be improved by considering the correlated positions. Using both simulated data and real promoter sequences, we show that the new de novo motif-finding algorithm can infer the true correlated position pairs accurately and is more precise in finding putative transcription factor binding sites than the standard Gibbs sampling algorithms. Jun S. Liu |
Bioinform. | 2 |
| 2004 | Gapped alignment of protein sequence motifs through Monte Carlo optimization of a hidden Markov modelabstractBACKGROUND: Certain protein families are highly conserved across distantly related organisms and belong to large and functionally diverse superfamilies. The patterns of conservation present in these protein sequences presumably are due to selective constraints maintaining important but unknown structural mechanisms with some constraints specific to each family and others shared by a larger subset or by the entire superfamily. To exploit these patterns as a source of functional information, we recently devised a statistically based approach called contrast hierarchical alignment and interaction network (CHAIN) analysis, which infers the strengths of various categories of selective constraints from co-conserved patterns in a multiple alignment. The power of this approach strongly depends on the quality of the multiple alignments, which thus motivated development of theoretical concepts and strategies to improve alignment of conserved motifs within large sets of distantly related sequences. RESULTS: Here we describe a hidden Markov model (HMM), an algebraic system, and Markov chain Monte Carlo (MCMC) sampling strategies for alignment of multiple sequence motifs. The MCMC sampling strategies are useful both for alignment optimization and for adjusting position specific background amino acid frequencies for alignment uncertainties. Associated statistical formulations provide an objective measure of alignment quality as well as automatic gap penalty optimization. Improved alignments obtained in this way are compared with PSI-BLAST based alignments within the context of CHAIN analysis of three protein families: Gialpha subunits, prolyl oligopeptidases, and transitional endoplasmic reticulum (p97) AAA+ ATPases. CONCLUSION: While not entirely replacing PSI-BLAST based alignments, which likewise may be optimized for CHAIN analysis using this approach, these motif-based methods often more accurately align very distantly related sequences and thus can provide a better measure of selective constraints. In some instances, these new approaches also provide a better understanding of family-specific constraints, as we illustrate for p97 ATPases. Programs implementing these procedures and supplementary information are available from the authors. Andrew F. Neuwald, Jun S. Liu |
BMC Bioinform. | 2 |
| 2000 | Dynamic weighting Monte Carlo for constrained floorplan designs in mixed signal applicationabstractSimulated annealing has been one of the most popular stochastic optimization methods used in the VLSI CAD eld in the past tw odecades.Recently, a new Monte Carlo and optimization method, named dynamic weighting Monte Carlo [WL97], has been introduced and successfully applied to the traveling salesman problem, neural net w orktraining [WL97], and spin-glasses simulation [LW99].In this paper, we h a v e successfully applied dynamic w eighting Monte Carlo algorithm to the constrained oorplan design with consideration of both area and wirelength minimization.Our application scenario is the constrained oorplan design for mixed signal MCMs, where w eneed to place all the analog modules together in groups so that they can share common pow er and ground planes, which are separate from those used b y the digital modules.Our experiments indicate that the dynamic weighting Monte Carlo algorithm is very effectiv e for constrained oorplan optimization.It outperforms the simulated annealing for a real mixed signal MCM design b y 19:5% in wirelength, with sligh t area improvement.This is the rst work adopting the dynamic weighting Monte Carlo optimization method for solving VLSI CAD problems.We believe that this method has applications to many other VLSI CAD optimization problems. Jason Cong, Tianming Kong, Faming Liang, Jun S. Liu, Wing Hung Wong, Dongmin Xu |
ASP-DAC | 4 |
| 2000 | Adaptive joint detection and decoding in flat-fading channels via mixture Kalman filteringabstractA novel adaptive Bayesian receiver for signal detection and decoding in fading channels with known channel statistics is developed; it is based on the sequential Monte Carlo methodology that has emerged in the field of statistics. The basic idea is to treat the transmitted signals as "missing data" and to sequentially impute multiple samples of them based on the observed signals. The imputed signal sequences, together with their importance weights, provide a way to approximate the Bayesian estimate of the transmitted signals and the channel states. Adaptive receiver algorithms for both uncoded and convolutionally coded systems are developed. The proposed techniques can easily handle the non-Gaussian ambient channel noise. It is shown through simulations that the proposed sequential Monte Carlo receivers achieve near-bound performance in fading channels for both uncoded and coded systems, without the use of any training/pilot symbols or decision feedback. Moreover, the proposed receiver structure exhibits massive parallelism and is ideally suited for high-speed parallel implementation using the very large scale integration (VLSI) systolic array technology. Jun S. Liu |
IEEE Trans. Inf. Theory | 3 |
| 1999 | Relaxed Simulated Tempering for VLSI Floorplan DesignsabstractIn the past two decades, the simulated annealing technique has been considered as a powerful approach to handle many NP-hard optimization problems in VLSI designs. Recently, a new Monte Carlo and optimization technique, named simulated tempering, was invented and has been successfully applied to many scientific problems, from random field Ising modeling to the traveling salesman problem. It is designed to overcome the drawback in simulated annealing when the problem has a rough energy landscape with many local minima separated by high energy barriers. In this paper, we have successfully applied a version of relaxed simulated tempering to slicing floorplan design with consideration of both area and wirelength optimization. Good experimental results were obtained. Jason Cong, Tianming Kong, Dongmin Xu, Faming Liang, Jun S. Liu, Wing Hung Wong |
ASP-DAC | 5 |
| 1999 | Bayesian inference on biopolymer modelsabstractMOTIVATION: Most existing bioinformatics methods are limited to making point estimates of one variable, e.g. the optimal alignment, with fixed input values for all other variables, e.g. gap penalties and scoring matrices. While the requirement to specify parameters remains one of the more vexing issues in bioinformatics, it is a reflection of a larger issue: the need to broaden the view on statistical inference in bioinformatics. RESULTS: The assignment of probabilities for all possible values of all unknown variables in a problem in the form of a posterior distribution is the goal of Bayesian inference. Here we show how this goal can be achieved for most bioinformatics methods that use dynamic programming. Specifically, a tutorial style description of a Bayesian inference procedure for segmentation of a sequence based on the heterogeneity in its composition is given. In addition, full Bayesian inference algorithms for sequence alignment are described. AVAILABILITY: Software and a set of transparencies for a tutorial describing these ideas are available at http://www.wadsworth.org/res&res/bioinfo/ Jun S. Liu, Charles E. Lawrence |
Bioinform. | 1 |
| 1998 | Bayesian adaptive sequence alignment algorithmsabstractThe selection of a scoring matrix and gap penalty parameters continues to be an important problem in sequence alignment. We describe here an algorithm, the 'Bayes block aligner, which bypasses this requirement. Instead of requiring a fixed set of parameter settings, this algorithm returns the Bayesian posterior probability for the number of gaps and for the scoring matrices in any series of interest. Furthermore, instead of returning the single best alignment for the chosen parameter settings, this algorithm returns the posterior distribution of all alignments considering the full range of gapping and scoring matrices selected, weighing each in proportion to its probability based on the data. We compared the Bayes aligner with the popular Smith-Waterman algorithm with parameter settings from the literature which had been optimized for the identification of structural neighbors, and found that the Bayes aligner correctly identified more structural neighbors. In a detailed examination of the alignment of a pair of kinase and a pair of GTPase sequences, we illustrate the algorithm's potential to identify subsequences that are conserved to different degrees. In addition, this example shows that the Bayes aligner returns an alignment-free assessment of the distance between a pair of sequences. Jun S. Liu, Charles E. Lawrence |
Bioinform. | 2 |
| 1997 | Bayesian Adaptive Alignment and Inference
Jun S. Liu, Charles E. Lawrence |
ISMB | 2 |