Wei Pan 0011

dblp:69/4450-11 · DBLP profile ↗
← Back
34ranked-venue papers
8as first author
5since 2021 · last 2024
0000-0002-1159-0582ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 24 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Probabilistic and Bayesian machine learning · 76% Learning theory · 17% Kernel, tree and ensemble methods · 3%
Interdisciplinary, comprehensive, and emerging computing
20 papers
Bioinformatics and computational biology · 95% Medical and health informatics · 3% Computational science and engineering · 2%
Theoretical computer science
3 papers
Mathematical optimization · 99% Information theory · 1%
Databases, data mining, and information retrieval
3 papers
Data mining · 100%

Topics — the 30 heaviest of 46, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery
0.812024
Causal Discovery with Generalized Linear Models through Peeling Algorithms · J. Mach. Learn. Res. 2024
Machine learning › Probabilistic and Bayesian machine learning › causal inference
instrumental variable
0.812024
Causal Discovery with Generalized Linear Models through Peeling Algorithms · J. Mach. Learn. Res. 2024
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal model
structural equation models
0.812024
Causal Discovery with Generalized Linear Models through Peeling Algorithms · J. Mach. Learn. Res. 2024
Bioinformatics and computational biology › genomics
genome-wide association study
0.732021
Integrative analysis of multi-omics data for discovering low-frequency variants associated with low-density lipoprotein cholesterol levels · Bioinform. 2021
A simple convolutional neural network for prediction of enhancer-promoter interactions with DNA sequence data · Bioinform. 2019
Integration of methylation QTL and enhancer-target gene maps with schizophrenia GWAS summary results identifies novel genes · Bioinform. 2019
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.712023
Inference for a Large Directed Acyclic Graph with Unspecified Interventions · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
directed acyclic graph
0.712023
Inference for a Large Directed Acyclic Graph with Unspecified Interventions · J. Mach. Learn. Res. 2023
Machine learning › Learning theory
hypothesis testing
0.712023
Inference for a Large Directed Acyclic Graph with Unspecified Interventions · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
likelihood ratio test
0.712023
Inference for a Large Directed Acyclic Graph with Unspecified Interventions · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning
statistical inference
0.712023
Inference for a Large Directed Acyclic Graph with Unspecified Interventions · J. Mach. Learn. Res. 2023
Data mining
clustering
0.532016
A New Algorithm and Theory for Penalized Regression-based Clustering · J. Mach. Learn. Res. 2016
Cluster analysis: unsupervised learning via supervised learning with a non-convex penalty · J. Mach. Learn. Res. 2013
Penalized Model-Based Clustering with Application to Variable Selection · J. Mach. Learn. Res. 2007
Machine learning › Learning theory
high-dimensional statistics
0.412020
A Regularization-Based Adaptive Test for High-Dimensional GLMs · J. Mach. Learn. Res. 2020
Mathematical optimization › statistical estimation › regression
regularized regression
0.412020
A Regularization-Based Adaptive Test for High-Dimensional GLMs · J. Mach. Learn. Res. 2020
Bioinformatics and computational biology › gene regulation › enhancer analysis
enhancer-promoter interaction prediction
0.412019
A simple convolutional neural network for prediction of enhancer-promoter interactions with DNA sequence data · Bioinform. 2019
Bioinformatics and computational biology › statistical genetics › association testing
gene-level association testing
0.412019
Integration of methylation QTL and enhancer-target gene maps with schizophrenia GWAS summary results identifies novel genes · Bioinform. 2019
Bioinformatics and computational biology
gene regulation
0.412019
A simple convolutional neural network for prediction of enhancer-promoter interactions with DNA sequence data · Bioinform. 2019
Bioinformatics and computational biology
gene expression analysis
0.482006
Incorporating gene functions as priors in model-based clustering of microarray gene expression data · Bioinform. 2006
Incorporating biological knowledge into distance-based clustering analysis of microarray gene expression data · Bioinform. 2006
Modeling the relationship between LVAD support time and gene expression changes in the human heart by penalized partial least squares · Bioinform. 2004
Bioinformatics and computational biology › statistical genetics
genetic association study
0.312017
Gene- and pathway-based association tests for multiple traits with GWAS summary statistics · Bioinform. 2017
Mathematical optimization › continuous optimization
convex optimization
0.212016
A New Algorithm and Theory for Penalized Regression-based Clustering · J. Mach. Learn. Res. 2016
Mathematical optimization › nonconvex optimization
difference of convex programming
0.212016
A New Algorithm and Theory for Penalized Regression-based Clustering · J. Mach. Learn. Res. 2016
Machine learning › Kernel, tree and ensemble methods
large margin methods
0.222011
Large Margin Hierarchical Classification with Mutually Exclusive Class Membership · J. Mach. Learn. Res. 2011
On Efficient Large Margin Semisupervised Learning: Method and Theory · J. Mach. Learn. Res. 2009
Bioinformatics and computational biology › gene expression analysis
differential expression analysis
0.242008
Incorporating gene networks into statistical tests for genomic data via a spatially correlated mixture model · Bioinform. 2008
Modified Nonparametric Approaches to Detecting Differentially Expressed Genes in Replicated Microarray Experiments · Bioinform. 2003
On the Use of Permutation in and the Performance of A Class of Nonparametric Methods to Detect Differential Gene Expression · Bioinform. 2003
Bioinformatics and computational biology › gene expression analysis
microarray data analysis
0.252010
Penalized mixtures of factor analyzers with application to clustering high-dimensional microarray data · Bioinform. 2010
A note on using permutation-based false discovery rate estimates to compare different analysis methods for microarray data · Bioinform. 2005
Modified Nonparametric Approaches to Detecting Differentially Expressed Genes in Replicated Microarray Experiments · Bioinform. 2003
Bioinformatics and computational biology › biological network › network biology › network inference
gene regulatory network inference
0.212023
Inference for a Large Directed Acyclic Graph with Unspecified Interventions · J. Mach. Learn. Res. 2023
Bioinformatics and computational biology › gene expression analysis
model-based clustering
0.222010
Penalized mixtures of factor analyzers with application to clustering high-dimensional microarray data · Bioinform. 2010
Incorporating gene functions as priors in model-based clustering of microarray gene expression data · Bioinform. 2006
Bioinformatics and computational biology
multi-omics data integration
0.112021
Integrative analysis of multi-omics data for discovering low-frequency variants associated with low-density lipoprotein cholesterol levels · Bioinform. 2021
Bioinformatics and computational biology › gene expression analysis
gene expression classification
0.122007
Incorporating prior knowledge of gene functional groups into regularized discriminant analysis of microarray data · Bioinform. 2007
Incorporating prior knowledge of predictors into penalized classifiers with multiple penalty terms · Bioinform. 2007
Medical and health informatics
alzheimer's disease genetics
0.112020
A Regularization-Based Adaptive Test for High-Dimensional GLMs · J. Mach. Learn. Res. 2020
Computer vision › Image recognition and object detection › image classification
hierarchical classification
0.112011
Large Margin Hierarchical Classification with Mutually Exclusive Class Membership · J. Mach. Learn. Res. 2011
Machine learning › Learning paradigms
semi-supervised learning
0.112009
On Efficient Large Margin Semisupervised Learning: Method and Theory · J. Mach. Learn. Res. 2009
Bioinformatics and computational biology › statistical genetics
quantitative trait locus analysis
0.112009
Network-based multiple locus linkage analysis of expression traits · Bioinform. 2009

Methods — techniques the papers use, named apart from their topics

peeling algorithm · 2.1non-convex penalty · 1.5nodewise regression · 1.3data perturbation · 1.3asymptotic null distribution · 1.3adaptive interaction sum of powered score test · 1.3generalized linear model · 0.8l0 regularization · 0.5adaptive gene-based test · 0.5DC programming · 0.5ADMM · 0.5transfer learning · 0.4summary statistics integration · 0.4convolutional neural network · 0.4supervised learning · 0.2mutually exclusive class constraints · 0.1margin analysis · 0.1penalization · 0.1
YearPublicationVenuePosition
2024 Causal Discovery with Generalized Linear Models through Peeling Algorithms
abstract
This article presents a novel method for causal discovery with generalized structural equation models suited for analyzing diverse types of outcomes, including discrete, continuous, and mixed data. Causal discovery often faces challenges due to unmeasured confounders that hinder the identification of causal relationships. The proposed approach addresses this issue by developing two peeling algorithms (bottom-up and top-down) to ascertain causal relationships and valid instruments. This approach first reconstructs a super-graph to represent ancestral relationships between variables, using a peeling algorithm based on nodewise GLM regressions that exploit relationships between primary and instrumental variables. Then, it estimates parent-child effects from the ancestral relationships using another peeling algorithm while deconfounding a child's model with information borrowed from its parents' models. The article offers a theoretical analysis of the proposed approach, establishing conditions for model identifiability and providing statistical guarantees for accurately discovering parent-child relationships via the peeling algorithms. Furthermore, the article presents numerical experiments showcasing the effectiveness of our approach in comparison to state-of-the-art structure learning methods without confounders. Lastly, it demonstrates an application to Alzheimer's disease (AD), highlighting the method's utility in constructing gene-to-gene and gene-to-disease regulatory networks involving Single Nucleotide Polymorphisms (SNPs) for healthy and AD subjects.
Xiaotong Shen, Wei Pan 0011
J. Mach. Learn. Res.3
2024 Significance Tests of Feature Relevance for a Black-Box Learner
abstract
An exciting recent development is the uptake of deep neural networks in many scientific fields, where the main objective is outcome prediction with a black-box nature. Significance testing is promising to address the black-box issue and explore novel scientific insights and interpretations of the decision-making process based on a deep learning model. However, testing for a neural network poses a challenge because of its black-box nature and unknown limiting distributions of parameter estimates while existing methods require strong assumptions or excessive computation. In this article, we derive one-split and two-split tests relaxing the assumptions and computational complexity of existing black-box tests and extending to examine the significance of a collection of features of interest in a dataset of possibly a complex type, such as an image. The one-split test estimates and evaluates a black-box model based on estimation and inference subsets through sample splitting and data perturbation. The two-split test further splits the inference subset into two but requires no perturbation. Also, we develop their combined versions by aggregating the p -values based on repeated sample splitting. By deflating the bias-sd-ratio, we establish asymptotic null distributions of the test statistics and the consistency in terms of Type 2 error. Numerically, we demonstrate the utility of the proposed tests on seven simulated examples and six real datasets. Accompanying this article is our python library dnn-inference (https://dnn-inference.readthedocs.io/en/latest/) that implements the proposed tests.
Ben Dai, Xiaotong Shen, Wei Pan 0011
IEEE Trans. Neural Networks Learn. Syst.3
2023 Inference for a Large Directed Acyclic Graph with Unspecified Interventions
abstract
Statistical inference of directed relations given some unspecified interventions (i.e., the intervention targets are unknown) is challenging. In this article, we test hypothesized directed relations with unspecified interventions. First, we derive conditions to yield an identifiable model. Unlike classical inference, testing directed relations requires identifying the ancestors and relevant interventions of hypothesis-specific primary variables. To this end, we propose a peeling algorithm based on nodewise regressions to establish a topological order of primary variables. Moreover, we prove that the peeling algorithm yields a consistent estimator in low-order polynomial time. Second, we propose a likelihood ratio test integrated with a data perturbation scheme to account for the uncertainty of identifying ancestors and interventions. Also, we show that the distribution of a data perturbation test statistic converges to the target distribution. Numerical examples demonstrate the utility and effectiveness of the proposed methods, including an application to infer gene regulatory networks.
Chunlin Li 0007, Xiaotong Shen, Wei Pan 0011
J. Mach. Learn. Res.3
2021 Integrative analysis of multi-omics data for discovering low-frequency variants associated with low-density lipoprotein cholesterol levels
abstract
MOTIVATION: The abundance of omics data has facilitated integrative analyses of single and multiple molecular layers with genome-wide association studies focusing on common variants. Built on its successes, we propose a general analysis framework to leverage multi-omics data with sequencing data to improve the statistical power of discovering new associations and understanding of the disease susceptibility due to low-frequency variants. The proposed test features its robustness to model misspecification, high power across a wide range of scenarios and the potential of offering insights into the underlying genetic architecture and disease mechanisms. RESULTS: Using the Framingham Heart Study data, we show that low-frequency variants are predictive of DNA methylation, even after conditioning on the nearby common variants. In addition, DNA methylation and gene expression provide complementary information to functional genomics. In the Avon Longitudinal Study of Parents and Children with a sample size of 1497, one gene CLPTM1 is identified to be associated with low-density lipoprotein cholesterol levels by the proposed powerful adaptive gene-based test integrating information from gene expression, methylation and enhancer-promoter interactions. It is further replicated in the TwinsUK study with 1706 samples. The signal is driven by both low-frequency and common variants. AVAILABILITY AND IMPLEMENTATION: Models are available at https://github.com/ytzhong/DNAm. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tianzhong Yang, Peng Wei 0005, Wei Pan 0011
Bioinform.3
2021 Model checking via testing for direct effects in Mendelian Randomization and transcriptome-wide association studies
abstract
It is of great interest and potential to discover causal relationships between pairs of exposures and outcomes using genetic variants as instrumental variables (IVs) to deal with hidden confounding in observational studies. Two most popular approaches are Mendelian randomization (MR), which usually use independent genetic variants/SNPs across the genome, and transcriptome-wide association studies (TWAS) (or their generalizations) using cis-SNPs local to a gene (or some genome-wide and likely dependent SNPs), as IVs. In spite of their many promising applications, both approaches face a major challenge: the validity of their causal conclusions depends on three critical assumptions on valid IVs, and more generally on other modeling assumptions, which however may not hold in practice. The most likely as well as challenging situation is due to the wide-spread horizontal pleiotropy, leading to two of the three IV assumptions being violated and thus to biased statistical inference. More generally, we'd like to conduct a goodness-of-fit (GOF) test to check the model being used. Although some methods have been proposed as being robust to various degrees to the violation of some modeling assumptions, they often give different and even conflicting results due to their own modeling assumptions and possibly lower statistical efficiency, imposing difficulties to the practitioner in choosing and interpreting varying results across different methods. Hence, it would help to directly test whether any assumption is violated or not. In particular, there is a lack of such tests for TWAS. We propose a new and general GOF test, called TEDE (TEsting Direct Effects), applicable to both correlated and independent SNPs/IVs (as commonly used in TWAS and MR respectively). Through simulation studies and real data examples, we demonstrate high statistical power and advantages of our new method, while confirming the frequent violation of modeling (including valid IV) assumptions in practice and thus the importance of model checking by applying such a test in MR/TWAS analysis.
Yangqing Deng, Wei Pan 0011
PLoS Comput. Biol.2
2020 A Regularization-Based Adaptive Test for High-Dimensional GLMs
abstract
In spite of its urgent importance in the era of big data, testing high-dimensional parameters in generalized linear models (GLMs) in the presence of high-dimensional nuisance parameters has been largely under-studied, especially with regard to constructing powerful tests for general (and unknown) alternatives. Most existing tests are powerful only against certain alternatives and may yield incorrect Type 1 error rates under high-dimensional nuisance parameter situations. In this paper, we propose the adaptive interaction sum of powered score (aiSPU) test in the framework of penalized regression with a non-convex penalty, called truncated Lasso penalty (TLP), which can maintain correct Type 1 error rates while yielding high statistical power across a wide range of alternatives. To calculate its p-values analytically, we derive its asymptotic null distribution. Via simulations, its superior finite-sample performance is demonstrated over several representative existing methods. In addition, we apply it and other representative tests to an Alzheimer's Disease Neuroimaging Initiative (ADNI) data set, detecting possible gene-gender interactions for Alzheimer's disease. We also put R package “aispu” implementing the proposed test on GitHub.
Gongjun Xu, Xiaotong Shen, Wei Pan 0011
J. Mach. Learn. Res.4
2020 A powerful and versatile colocalization test
abstract
Transcriptome-wide association studies (TWAS and PrediXcan) have been increasingly applied to detect associations between genetically predicted gene expressions and GWAS traits, which may suggest, however do not completely determine, causal genes for GWAS traits, due to the likely violation of their imposed strong assumptions for causal inference. Testing colocalization moves it closer to establishing causal relationships: if a GWAS trait and a gene's expression share the same associated SNP, it may suggest a regulatory (and thus putative causal) role of the SNP mediated through the gene on the GWAS trait. Accordingly, it is of interest to develop and apply various colocalization testing approaches. The existing approaches may each have some severe limitations. For instance, some methods test the null hypothesis that there is colocalization, which is not ideal because often the null hypothesis cannot be rejected simply due to limited statistical power (with too small sample sizes). Some other methods arbitrarily restrict the maximum number of causal SNPs in a locus, which may lead to loss of power in the presence of wide-spread allelic heterogeneity. Importantly, most methods cannot be applied to either GWAS/eQTL summary statistics or cases with more than two possibly correlated traits. Here we present a simple and general approach based on conditional analysis of a locus on multiple traits, overcoming the above and other shortcomings of the existing methods. We demonstrate that, compared with other methods, our new method can be applied to a wider range of scenarios and often perform better. We showcase its applications to both simulated and real data, including a large-scale Alzheimer's disease GWAS summary dataset and a gene expression dataset, and a large-scale blood lipid GWAS summary association dataset. An R package "jointsum" implementing the proposed method is publicly available at github.
Yangqing Deng, Wei Pan 0011
PLoS Comput. Biol.2
2020 Penalized regression and model selection methods for polygenic scores on summary statistics
abstract
Polygenic scores quantify the genetic risk associated with a given phenotype and are widely used to predict the risk of complex diseases. There has been recent interest in developing methods to construct polygenic risk scores using summary statistic data. We propose a method to construct polygenic risk scores via penalized regression using summary statistic data and publicly available reference data. Our method bears similarity to existing method LassoSum, extending their framework to the Truncated Lasso Penalty (TLP) and the elastic net. We show via simulation and real data application that the TLP improves predictive accuracy as compared to the LASSO while imposing additional sparsity where appropriate. To facilitate model selection in the absence of validation data, we propose methods for estimating model fitting criteria AIC and BIC. These methods approximate the AIC and BIC in the case where we have a polygenic risk score estimated on summary statistic data and no validation data. Additionally, we propose the so-called quasi-correlation metric, which quantifies the predictive accuracy of a polygenic risk score applied to out-of-sample data for which we have only summary statistic information. In total, these methods facilitate estimation and model selection of polygenic risk scores on summary statistic data, and the application of these polygenic risk scores to out-of-sample data for which we have only summary statistic information. We demonstrate the utility of these methods by applying them to GWA studies of lipids, height, and lung cancer.
Jack Pattee, Wei Pan 0011
PLoS Comput. Biol.2
2019 Integration of methylation QTL and enhancer-target gene maps with schizophrenia GWAS summary results identifies novel genes
abstract
MOTIVATION: Most trait-associated genetic variants identified in genome-wide association studies (GWASs) are located in non-coding regions of the genome and thought to act through their regulatory roles. RESULTS: To account for enriched association signals in DNA regulatory elements, we propose a novel and general gene-based association testing strategy that integrates enhancer-target gene pairs and methylation quantitative trait locus data with GWAS summary results; it aims to both boost statistical power for new discoveries and enhance mechanistic interpretability of any new discovery. By reanalyzing two large-scale schizophrenia GWAS summary datasets, we demonstrate that the proposed method could identify some significant and novel genes (containing no genome-wide significant SNPs nearby) that would have been missed by other competing approaches, including the standard and some integrative gene-based association methods, such as one incorporating enhancer-target gene pairs and one integrating expression quantitative trait loci. AVAILABILITY AND IMPLEMENTATION: Software: wuchong.org/egmethyl.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wei Pan 0011
Bioinform.2
2019 A simple convolutional neural network for prediction of enhancer-promoter interactions with DNA sequence data
abstract
MOTIVATION: Enhancer-promoter interactions (EPIs) in the genome play an important role in transcriptional regulation. EPIs can be useful in boosting statistical power and enhancing mechanistic interpretation for disease- or trait-associated genetic variants in genome-wide association studies. Instead of expensive and time-consuming biological experiments, computational prediction of EPIs with DNA sequence and other genomic data is a fast and viable alternative. In particular, deep learning and other machine learning methods have been demonstrated with promising performance. RESULTS: First, using a published human cell line dataset, we demonstrate that a simple convolutional neural network (CNN) performs as well as, if no better than, a more complicated and state-of-the-art architecture, a hybrid of a CNN and a recurrent neural network. More importantly, in spite of the well-known cell line-specific EPIs (and corresponding gene expression), in contrast to the standard practice of training and predicting for each cell line separately, we propose two transfer learning approaches to training a model using all cell lines to various extents, leading to substantially improved predictive performance. AVAILABILITY AND IMPLEMENTATION: Computer code is available at https://github.com/zzUMN/Combine-CNN-Enhancer-and-Promoters. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhong Zhuang, Xiaotong Shen, Wei Pan 0011
Bioinform.3
2017 Gene- and pathway-based association tests for multiple traits with GWAS summary statistics
abstract
To identify novel genetic variants associated with complex traits and to shed new insights on underlying biology, in addition to the most popular single SNP-single trait association analysis, it would be useful to explore multiple correlated (intermediate) traits at the gene- or pathway-level by mining existing single GWAS or meta-analyzed GWAS data. For this purpose, we present an adaptive gene-based test and a pathway-based test for association analysis of multiple traits with GWAS summary statistics. The proposed tests are adaptive at both the SNP- and trait-levels; that is, they account for possibly varying association patterns (e.g. signal sparsity levels) across SNPs and traits, thus maintaining high power across a wide range of situations. Furthermore, the proposed methods are general: they can be applied to mixed types of traits, and to Z-statistics or P-values as summary statistics obtained from either a single GWAS or a meta-analysis of multiple GWAS. Our numerical studies with simulated and real data demonstrated the promising performance of the proposed methods. AVAILABILITY AND IMPLEMENTATION: The methods are implemented in R package aSPU, freely and publicly available at: https://cran.r-project.org/web/packages/aSPU/ CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online.
Il-Youp Kwak, Wei Pan 0011
Bioinform.2
2016 A New Algorithm and Theory for Penalized Regression-based Clustering
abstract
Clustering is unsupervised and exploratory in nature. Yet, it can be performed through penalized regression with grouping pursuit, as demonstrated in Pan et al. (2013). In this paper, we develop a more efficient algorithm for scalable computation and a new theory of clustering consistency for the method. This algorithm, called DC-ADMM, combines difference of convex (DC) programming with the alternating direction method of multipliers (ADMM). This algorithm is shown to be more computationally efficient than the quadratic penalty based algorithm of Pan et al. (2013) because of the former's closed-form updating formulas. Numerically, we compare the DC- ADMM algorithm with the quadratic penalty algorithm to demonstrate its utility and scalability. Theoretically, we establish a finite-sample mis- clustering error bound for penalized regression based clustering with the $L_0$ constrained regularization in a general setting. On this ground, we provide conditions for clustering consistency of the penalized clustering method. As an end product, we put R package prclust implementing PRclust with various loss and grouping penalty functions available on GitHub and CRAN.
Sunghoon Kwon, Xiaotong Shen, Wei Pan 0011
J. Mach. Learn. Res.4
2013 Cluster analysis: unsupervised learning via supervised learning with a non-convex penalty
Wei Pan 0011, Xiaotong Shen, Binghui Liu
J. Mach. Learn. Res.1
2011 Large Margin Hierarchical Classification with Mutually Exclusive Class Membership
Huixin Wang, Xiaotong Shen, Wei Pan 0011
J. Mach. Learn. Res.3
2010 Penalized mixtures of factor analyzers with application to clustering high-dimensional microarray data
abstract
MOTIVATION: Model-based clustering has been widely used, e.g. in microarray data analysis. Since for high-dimensional data variable selection is necessary, several penalized model-based clustering methods have been proposed tørealize simultaneous variable selection and clustering. However, the existing methods all assume that the variables are independent with the use of diagonal covariance matrices. RESULTS: To model non-independence of variables (e.g. correlated gene expressions) while alleviating the problem with the large number of unknown parameters associated with a general non-diagonal covariance matrix, we generalize the mixture of factor analyzers to that with penalization, which, among others, can effectively realize variable selection. We use simulated data and real microarray data to illustrate the utility and advantages of the proposed method over several existing ones.
Benhuai Xie, Wei Pan 0011, Xiaotong Shen
Bioinform.2
2009 Network-based multiple locus linkage analysis of expression traits
abstract
Abstract Motivation: We consider the problem of multiple locus linkage analysis for expression traits of genes in a pathway or a network. To capitalize on co-expression of functionally related genes, we propose a penalized regression method that maps multiple expression quantitative trait loci (eQTLs) for all related genes simultaneously while accounting for their shared functions as specified a priori by a gene pathway or network. Results: An analysis of a mouse dataset and simulation studies clearly demonstrate the advantage of the proposed method over a standard approach that ignores biological knowledge of gene networks. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Wei Pan 0011
Bioinform.1
2009 Network-based support vector machine for classification of microarray samples
abstract
BACKGROUND: The importance of network-based approach to identifying biological markers for diagnostic classification and prognostic assessment in the context of microarray data has been increasingly recognized. To our knowledge, there have been few, if any, statistical tools that explicitly incorporate the prior information of gene networks into classifier building. The main idea of this paper is to take full advantage of the biological observation that neighboring genes in a network tend to function together in biological processes and to embed this information into a formal statistical framework. RESULTS: We propose a network-based support vector machine for binary classification problems by constructing a penalty term from the Finfinity-norm being applied to pairwise gene neighbors with the hope to improve predictive performance and gene selection. Simulation studies in both low- and high-dimensional data settings as well as two real microarray applications indicate that the proposed method is able to identify more clinically relevant genes while maintaining a sparse model with either similar or higher prediction accuracy compared with the standard and the L1 penalized support vector machines. CONCLUSION: The proposed network-based support vector machine has the potential to be a practically useful classification tool for microarrays and other high-dimensional data.
Yanni Zhu, Xiaotong Shen, Wei Pan 0011
BMC Bioinform.3
2009 On Efficient Large Margin Semisupervised Learning: Method and Theory
Xiaotong Shen, Wei Pan 0011
J. Mach. Learn. Res.3
2008 Incorporating gene networks into statistical tests for genomic data via a spatially correlated mixture model
abstract
MOTIVATION: It is a common task in genomic studies to identify a subset of the genes satisfying certain conditions, such as differentially expressed genes or regulatory target genes of a transcription factor (TF). This can be formulated as a statistical hypothesis testing problem. Most existing approaches treat the genes as having an identical and independent distribution a priori, testing each gene independently or testing some subsets of the genes one by one. On the other hand, it is known that the genes work coordinately as dictated by gene networks. Treating genes equally and independently ignores the important information contained in gene networks, leading to inefficient analysis and reduced power. RESULTS: We propose incorporating gene network information into statistical analysis of genomic data. Specifically, rather than treating the genes equally and independently a priori in a standard mixture model, we assume that gene-specific prior probabilities are correlated as induced by a gene network: while the genes are allowed to have different prior probabilities, those neighboring ones in the network have similar prior probabilities, reflecting their shared biological functions. We applied the two approaches to a real ChIP-chip dataset (and simulated data) to identify the transcriptional target genes of TF GCN4. The new method was found to be more powerful in discovering the target genes.
Peng Wei 0005, Wei Pan 0011
Bioinform.2
2008 Incorporating Gene Functions into Regression Analysis of DNA-Protein Binding Data and Gene Expression Data to Construct Transcriptional Networks
abstract
Useful information on transcriptional networks has been extracted by regression analyses of gene expression data and DNA-protein binding data. However, a potential limitation of these approaches is their assumption on the common and constant activity level of a transcription factor (TF) on all the genes in any given experimental condition; for example, any TF is assumed to be either an activator or a repressor, but not both, while it is known that some TFs can be dual regulators. Rather than assuming a common linear regression model for all the genes, we propose using separate regression models for various gene groups; the genes can be grouped based on their functions or some clustering results. Furthermore, to take advantage of the hierarchical structure of many existing gene function annotation systems, such as Gene Ontology (GO), we propose a shrinkage method that borrows information from relevant gene groups. Applications to a yeast dataset and simulations lend support for our proposed methods. In particular, we find that the shrinkage method consistently works well under various scenarios. We recommend the use of the shrinkage method as a useful alternative to the existing methods.
Peng Wei 0005, Wei Pan 0011
IEEE ACM Trans. Comput. Biol. Bioinform.2
2007 Incorporating prior knowledge of predictors into penalized classifiers with multiple penalty terms
abstract
MOTIVATION: In the context of sample (e.g. tumor) classifications with microarray gene expression data, many methods have been proposed. However, almost all the methods ignore existing biological knowledge and treat all the genes equally a priori. On the other hand, because some genes have been identified by previous studies to have biological functions or to be involved in pathways related to the outcome (e.g. cancer), incorporating this type of prior knowledge into a classifier can potentially improve both the predictive performance and interpretability of the resulting model. RESULTS: We propose a simple and general framework to incorporate such prior knowledge into building a penalized classifier. As two concrete examples, we apply the idea to two penalized classifiers, nearest shrunken centroids (also called PAM) and penalized partial least squares (PPLS). Instead of treating all the genes equally a priori as in standard penalized methods, we group the genes according to their functional associations based on existing biological knowledge or data, and adopt group-specific penalty terms and penalization parameters. Simulated and real data examples demonstrate that, if prior knowledge on gene grouping is indeed informative, our new methods perform better than the two standard penalized methods, yielding higher predictive accuracy and screening out more irrelevant genes.
Feng Tai, Wei Pan 0011
Bioinform.2
2007 Incorporating prior knowledge of gene functional groups into regularized discriminant analysis of microarray data
abstract
MOTIVATION: Discriminant analysis for high-dimensional and low-sample-sized data has become a hot research topic in bioinformatics, mainly motivated by its importance and challenge in applications to tumor classifications for high-dimensional microarray data. Two of the popular methods are the nearest shrunken centroids, also called predictive analysis of microarray (PAM), and shrunken centroids regularized discriminant analysis (SCRDA). Both methods are modifications to the classic linear discriminant analysis (LDA) in two aspects tailored to high-dimensional and low-sample-sized data: one is the regularization of the covariance matrix, and the other is variable selection through shrinkage. In spite of their usefulness, there are potential limitations with each method. The main concern is that both PAM and SCRDA are possibly too extreme: the covariance matrix in the former is restricted to be diagonal while in the latter there is barely any restriction. Based on the biology of gene functions and given the feature of the data, it may be beneficial to estimate the covariance matrix as an intermediate between the two; furthermore, more effective shrinkage schemes may be possible. RESULTS: We propose modified LDA methods to integrate biological knowledge of gene functions (or variable groups) into classification of microarray data. Instead of simply treating all the genes independently or imposing no restriction on the correlations among the genes, we group the genes according to their biological functions extracted from existing biological knowledge or data, and propose regularized covariance estimators that encourages between-group gene independence and within-group gene correlations while maintaining the flexibility of any general covariance structure. Furthermore, we propose a shrinkage scheme on groups of genes that tends to retain or remove a whole group of the genes altogether, in contrast to the standard shrinkage on individual genes. We show that one of the proposed methods performed better than PAM and SCRDA in a simulation study and several real data examples.
Feng Tai, Wei Pan 0011
Bioinform.2
2007 Penalized Model-Based Clustering with Application to Variable Selection
Wei Pan 0011, Xiaotong Shen
J. Mach. Learn. Res.1
2006 Incorporating biological knowledge into distance-based clustering analysis of microarray gene expression data
abstract
MOTIVATION: Because co-expressed genes are likely to share the same biological function, cluster analysis of gene expression profiles has been applied for gene function discovery. Most existing clustering methods ignore known gene functions in the process of clustering. RESULTS: To take advantage of accumulating gene functional annotations, we propose incorporating known gene functions into a new distance metric, which shrinks a gene expression-based distance towards 0 if and only if the two genes share a common gene function. A two-step procedure is used. First, the shrinkage distance metric is used in any distance-based clustering method, e.g. K-medoids or hierarchical clustering, to cluster the genes with known functions. Second, while keeping the clustering results from the first step for the genes with known functions, the expression-based distance metric is used to cluster the remaining genes of unknown function, assigning each of them to either one of the clusters obtained in the first step or some new clusters. A simulation study and an application to gene function prediction for the yeast demonstrate the advantage of our proposal over the standard method.
Desheng Huang, Wei Pan 0011
Bioinform.2
2006 Incorporating gene functions as priors in model-based clustering of microarray gene expression data
abstract
MOTIVATION: Cluster analysis of gene expression profiles has been widely applied to clustering genes for gene function discovery. Many approaches have been proposed. The rationale is that the genes with the same biological function or involved in the same biological process are more likely to co-express, hence they are more likely to form a cluster with similar gene expression patterns. However, most existing methods, including model-based clustering, ignore known gene functions in clustering. RESULTS: To take advantage of accumulating gene functional annotations, we propose incorporating known gene functions as prior probabilities in model-based clustering. In contrast to a global mixture model applicable to all the genes in the standard model-based clustering, we use a stratified mixture model: one stratum corresponds to the genes of unknown function while each of the other ones corresponding to the genes sharing the same biological function or pathway; the genes from the same stratum are assumed to have the same prior probability of coming from a cluster while those from different strata are allowed to have different prior probabilities of coming from the same cluster. We derive a simple EM algorithm that can be used to fit the stratified model. A simulation study and an application to gene function prediction demonstrate the advantage of our proposal over the standard method. CONTACT: [email protected]
Wei Pan 0011
Bioinform.1
2006 Semi-supervised learning via penalized mixture model with application to microarray sample classification
abstract
MOTIVATION: It is biologically interesting to address whether human blood outgrowth endothelial cells (BOECs) belong to or are closer to large vessel endothelial cells (LVECs) or microvascular endothelial cells (MVECs) based on global expression profiling. An earlier analysis using a hierarchical clustering and a small set of genes suggested that BOECs seemed to be closer to MVECs. By taking advantage of the two known classes, LVEC and MVEC, while allowing BOEC samples to belong to either of the two classes or to form their own new class, we take a semi-supervised learning approach; for high-dimensional data as encountered here, we propose a penalized mixture model with a weighted L1 penalty to realize automatic feature selection while fitting the model. RESULTS: We applied our penalized mixture model to a combined dataset containing 27 BOEC, 28 LVEC and 25 MVEC samples. Analysis results indicated that the BOEC samples appeared to form their own new class. A simulation study confirmed that, compared with the standard mixture model with or without initial variable selection, the penalized mixture model performed much better in identifying relevant genes and forming corresponding clusters. The penalized mixture model seems to be promising for high-dimensional data with the capability of novel class discovery and automatic feature selection.
Wei Pan 0011, Xiaotong Shen, Aixiang Jiang, Robert P. Hebbel
Bioinform.1
2005 A note on using permutation-based false discovery rate estimates to compare different analysis methods for microarray data
abstract
MOTIVATION: False discovery rate (FDR) is defined as the expected percentage of false positives among all the claimed positives. In practice, with the true FDR unknown, an estimated FDR can serve as a criterion to evaluate the performance of various statistical methods under the condition that the estimated FDR approximates the true FDR well, or at least, it does not improperly favor or disfavor any particular method. Permutation methods have become popular to estimate FDR in genomic studies. The purpose of this paper is 2-fold. First, we investigate theoretically and empirically whether the standard permutation-based FDR estimator is biased, and if so, whether the bias inappropriately favors or disfavors any method. Second, we propose a simple modification of the standard permutation to yield a better FDR estimator, which can in turn serve as a more fair criterion to evaluate various statistical methods. RESULTS: Both simulated and real data examples are used for illustration and comparison. Three commonly used test statistics, the sample mean, SAM statistic and Student's t-statistic, are considered. The results show that the standard permutation method overestimates FDR. The overestimation is the most severe for the sample mean statistic while the least for the t-statistic with the SAM-statistic lying between the two extremes, suggesting that one has to be cautious when using the standard permutation-based FDR estimates to evaluate various statistical methods. In addition, our proposed FDR estimation method is simple and outperforms the standard method.
Wei Pan 0011, Arkady B. Khodursky
Bioinform.2
2005 A comparative study of discriminating human heart failure etiology using gene expression profiles
abstract
BACKGROUND: Human heart failure is a complex disease that manifests from multiple genetic and environmental factors. Although ischemic and non-ischemic heart disease present clinically with many similar decreases in ventricular function, emerging work suggests that they are distinct diseases with different responses to therapy. The ability to distinguish between ischemic and non-ischemic heart failure may be essential to guide appropriate therapy and determine prognosis for successful treatment. In this paper we consider discriminating the etiologies of heart failure using gene expression libraries from two separate institutions. RESULTS: We apply five new statistical methods, including partial least squares, penalized partial least squares, LASSO, nearest shrunken centroids and random forest, to two real datasets and compare their performance for multiclass classification. It is found that the five statistical methods perform similarly on each of the two datasets: it is difficult to correctly distinguish the etiologies of heart failure in one dataset whereas it is easy for the other one. In a simulation study, it is confirmed that the five methods tend to have close performance, though the random forest seems to have a slight edge. CONCLUSIONS: For some gene expression data, several recently developed discriminant methods may perform similarly. More importantly, one must remain cautious when assessing the discriminating performance using gene expression profiles based on a small dataset; our analysis suggests the importance of utilizing multiple or larger datasets.
Wei Pan 0011, Suzanne Grindle, Xinqiang Han, Soon J. Park, Leslie W. Miller, Jennifer Hall
BMC Bioinform.2
2004 Modeling the relationship between LVAD support time and gene expression changes in the human heart by penalized partial least squares
abstract
MOTIVATION: Heart failure affects more than 20 million people in the world. Heart transplantation is the most effective therapy, but the number of eligible patients far outweighs the number of available donor hearts. The left mechanical ventricular assist device (LVAD) has been developed as a successful substitution therapy that aids the failing ventricle while a patient is waiting for the donor heart. We obtained genomics data from paired human heart samples harvested at the time of LVAD implant and explant. The heart failure patients in our study were supported by the LVAD for various periods of time. The goal of this study is to model the relationship between the time of LVAD support and gene expression changes. RESULTS: To serve the purpose, we propose a novel penalized partial least squares (PPLS) method to build a regression model. Compared with partial least squares and Breiman's random forest method, PPLS gives the best prediction results for the LVAD data.
Wei Pan 0011, Soon J. Park, Xinqiang Han, Leslie W. Miller, Jennifer Hall
Bioinform.2
2003 Linear regression and two-class classification with gene expression data
abstract
MOTIVATION: Using gene expression data to classify (or predict) tumor types has received much research attention recently. Due to some special features of gene expression data, several new methods have been proposed, including the weighted voting scheme of Golub et al., the compound covariate method of Hedenfalk et al. (originally proposed by Tukey), and the shrunken centroids method of Tibshirani et al. These methods look different and are more or less ad hoc. RESULTS: We point out a close connection of the three methods with a linear regression model. Casting the classification problem in the general framework of linear regression naturally leads to new alternatives, such as partial least squares (PLS) methods and penalized PLS (PPLS) methods. Using two real data sets, we show the competitive performance of our new methods when compared with the other three methods.
Wei Pan 0011
Bioinform.2
2003 On the Use of Permutation in and the Performance of A Class of Nonparametric Methods to Detect Differential Gene Expression
abstract
MOTIVATION: Recently a class of nonparametric statistical methods, including the empirical Bayes (EB) method, the significance analysis of microarray (SAM) method and the mixture model method (MMM), have been proposed to detect differential gene expression for replicated microarray experiments conducted under two conditions. All the methods depend on constructing a test statistic Z and a so-called null statistic z. The null statistic z is used to provide some reference distribution for Z such that statistical inference can be accomplished. A common way of constructing z is to apply Z to randomly permuted data. Here we point our that the distribution of z may not approximate the null distribution of Z well, leading to possibly too conservative inference. This observation may apply to other permutation-based nonparametric methods. We propose a new method of constructing a null statistic that aims to estimate the null distribution of a test statistic directly. RESULTS: Using simulated data and real data, we assess and compare the performance of the existing method and our new method when applied in EB, SAM and MMM. Some interesting findings on operating characteristics of EB, SAM and MMM are also reported. Finally, by combining the idea of SAM and MMM, we outline a simple nonparametric method based on the direct use of a test statistic and a null statistic.
Wei Pan 0011
Bioinform.1
2003 Modified Nonparametric Approaches to Detecting Differentially Expressed Genes in Replicated Microarray Experiments
abstract
MOTIVATION: An important goal in analyzing microarray data is to determine which genes are differentially expressed across two kinds of tissue samples or samples obtained under two experimental conditions. Various parametric tests, such as the two-sample t-test, have been used, but their possibly too strong parametric assumptions or large sample justifications may not hold in practice. As alternatives, a class of three nonparametric statistical methods, including the empirical Bayes method of Efron et al. (2001), the significance analysis of microarray (SAM) method of Tusher et al. (2001) and the mixture model method (MMM) of Pan et al. (2001), have been proposed. All the three methods depend on constructing a test statistic and a so-called null statistic such that the null statistic's distribution can be used to approximate the null distribution of the test statistic. However, relatively little effort has been directed toward assessment of the performance or the underlying assumptions of the methods in constructing such test and null statistics. RESULTS: We point out a problem of a current method to construct the test and null statistics, which may lead to largely inflated Type I errors (i.e. false positives). We also propose two modifications that overcome the problem. In the context of MMM, the improved performance of the modified methods is demonstrated using simulated data. In addition, our numerical results also provide evidence to support the utility and effectiveness of MMM.
Yanli Zhao, Wei Pan 0011
Bioinform.2
2002 A comparative review of statistical methods for discovering differentially expressed genes in replicated microarray experiments
abstract
MOTIVATION: A common task in analyzing microarray data is to determine which genes are differentially expressed across two kinds of tissue samples or samples obtained under two experimental conditions. Recently several statistical methods have been proposed to accomplish this goal when there are replicated samples under each condition. However, it may not be clear how these methods compare with each other. Our main goal here is to compare three methods, the t-test, a regression modeling approach (Thomas et al., Genome Res., 11, 1227-1236, 2001) and a mixture model approach (Pan et al., http://www.biostat.umn.edu/cgi-bin/rrs?print+2001,2001a,b) with particular attention to their different modeling assumptions. RESULTS: It is pointed out that all the three methods are based on using the two-sample t-statistic or its minor variation, but they differ in how to associate a statistical significance level to the corresponding statistic, leading to possibly large difference in the resulting significance levels and the numbers of genes detected. In particular, we give an explicit formula for the test statistic used in the regression approach. Using the leukemia data of Golub et al. (Science, 285, 531-537, 1999), we illustrate these points. We also briefly compare the results with those of several other methods, including the empirical Bayesian method of Efron et al. (J. Am. Stat. Assoc., to appear, 2001) and the Significance Analysis of Microarray (SAM) method of Tusher et al. (PROC: Natl Acad. Sci. USA, 98, 5116-5121, 2001).
Wei Pan 0011
Bioinform.1
1999 Shrinking classification trees for bootstrap aggregation
Wei Pan 0011
Pattern Recognit. Lett.1