Dan Nettleton

dblp:57/2915 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
2since 2021 · last 2023
0000-0002-6045-1036ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
7 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
1 paper
Trustworthy machine learning · 46% Kernel, tree and ensemble methods · 23% Motion planning and robot control · 23%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › gene expression analysis
differential expression analysis
1.452023
Adjusting for gene-specific covariates to improve RNA-seq analysis · Bioinform. 2023
rmRNAseq: differential expression analysis for repeated-measures RNA-seq data · Bioinform. 2020
An improved method for computing q-values when the distribution of effect sizes is asymmetric · Bioinform. 2014
Bioinformatics and computational biology › transcriptomics
RNA-seq analysis
0.922023
Adjusting for gene-specific covariates to improve RNA-seq analysis · Bioinform. 2023
SimSeq: a nonparametric approach to simulation of RNA-sequence datasets · Bioinform. 2015
Bioinformatics and computational biology
gene expression analysis
0.842020
rmRNAseq: differential expression analysis for repeated-measures RNA-seq data · Bioinform. 2020
An improved method for computing q-values when the distribution of effect sizes is asymmetric · Bioinform. 2014
Pooling mRNA in microarray experiments and its effect on power · Bioinform. 2007
Machine learning › Trustworthy machine learning › uncertainty estimation › conformal prediction
coverage guarantee
0.712023
Stability of Random Forests and Coverage of Random-Forest Prediction Intervals · NeurIPS 2023
Machine learning › Trustworthy machine learning › uncertainty estimation
prediction intervals
0.712023
Stability of Random Forests and Coverage of Random-Forest Prediction Intervals · NeurIPS 2023
Machine learning › Kernel, tree and ensemble methods › ensemble learning › tree ensembles
random forest
0.712023
Stability of Random Forests and Coverage of Random-Forest Prediction Intervals · NeurIPS 2023
Robotics › Motion planning and robot control
stability analysis
0.712023
Stability of Random Forests and Coverage of Random-Forest Prediction Intervals · NeurIPS 2023
Bioinformatics and computational biology › biostatistics › statistical bioinformatics
statistical genomics
0.712023
Adjusting for gene-specific covariates to improve RNA-seq analysis · Bioinform. 2023
Bioinformatics and computational biology › transcriptomics › RNA-seq analysis
RNA-seq simulation
0.212015
SimSeq: a nonparametric approach to simulation of RNA-sequence datasets · Bioinform. 2015
Bioinformatics and computational biology › gene expression analysis
microarray data analysis
0.232014
Identification of differentially expressed gene categories in microarray studies using nonparametric multivariate analysis · Bioinform. 2008
Pooling mRNA in microarray experiments and its effect on power · Bioinform. 2007
An improved method for computing q-values when the distribution of effect sizes is asymmetric · Bioinform. 2014
Machine learning › Optimization for machine learning › convergence analysis
non-asymptotic analysis
0.212023
Stability of Random Forests and Coverage of Random-Forest Prediction Intervals · NeurIPS 2023
Bioinformatics and computational biology
statistical genetics
0.112020
rmRNAseq: differential expression analysis for repeated-measures RNA-seq data · Bioinform. 2020

Methods — techniques the papers use, named apart from their topics

q-value · 0.8positive false discovery rate · 0.7out-of-bag error · 0.7jackknife · 0.7cross-validation · 0.7parametric bootstrap · 0.4general linear model · 0.4continuous autoregressive structure · 0.4nonparametric simulation · 0.2negative binomial modeling · 0.2multiple testing · 0.2nonparametric multivariate test · 0.1
YearPublicationVenuePosition
2023 Stability of Random Forests and Coverage of Random-Forest Prediction Intervals
abstract
We establish stability of random forests under the mild condition that the squared response ($Y^2$) does not have a heavy tail. In particular, our analysis holds for the practical version of random forests that is implemented in popular packages like \texttt{randomForest} in \texttt{R}. Empirical results show that stability may persist even beyond our assumption and hold for heavy-tailed $Y^2$. Using the stability property, we prove a non-asymptotic lower bound for the coverage probability of prediction intervals constructed from the out-of-bag error of random forests. With another mild condition that is typically satisfied when $Y$ is continuous, we also establish a complementary upper bound, which can be similarly established for the jackknife prediction interval constructed from an arbitrary stable algorithm. We also discuss the asymptotic coverage probability under assumptions weaker than those considered in previous literature. Our work implies that random forests, with its stability property, is an effective machine learning method that can provide not only satisfactory point prediction but also justified interval prediction at almost no extra computational cost.
Huaiqing Wu, Dan Nettleton
NeurIPS3
2023 Adjusting for gene-specific covariates to improve RNA-seq analysis
abstract
SUMMARY: This article suggests a novel positive false discovery rate (pFDR) controlling method for testing gene-specific hypotheses using a gene-specific covariate variable, such as gene length. We suppose the null probability depends on the covariate variable. In this context, we propose a rejection rule that accounts for heterogeneity among tests by using two distinct types of null probabilities. We establish a pFDR estimator for a given rejection rule by following Storey's q-value framework. A condition on a type 1 error posterior probability is provided that equivalently characterizes our rejection rule. We also present a suitable procedure for selecting a tuning parameter through cross-validation that maximizes the expected number of hypotheses declared significant. A simulation study demonstrates that our method is comparable to or better than existing methods across realistic scenarios. In data analysis, we find support for our method's premise that the null probability varies with a gene-specific covariate variable. AVAILABILITY AND IMPLEMENTATION: The source code repository is publicly available at https://github.com/hsjeon1217/conditional_method.
Hyeongseon Jeon, Kyu-Sang Lim, Yet Nguyen, Dan Nettleton
Bioinform.4
2020 rmRNAseq: differential expression analysis for repeated-measures RNA-seq data
abstract
MOTIVATION: With the reduction in price of next-generation sequencing technologies, gene expression profiling using RNA-seq has increased the scope of sequencing experiments to include more complex designs, such as designs involving repeated measures. In such designs, RNA samples are extracted from each experimental unit at multiple time points. The read counts that result from RNA sequencing of the samples extracted from the same experimental unit tend to be temporally correlated. Although there are many methods for RNA-seq differential expression analysis, existing methods do not properly account for within-unit correlations that arise in repeated-measures designs. RESULTS: We address this shortcoming by using normalized log-transformed counts and associated precision weights in a general linear model pipeline with continuous autoregressive structure to account for the correlation among observations within each experimental unit. We then utilize parametric bootstrap to conduct differential expression inference. Simulation studies show the advantages of our method over alternatives that do not account for the correlation among observations within experimental units. AVAILABILITY AND IMPLEMENTATION: We provide an R package rmRNAseq implementing our proposed method (function TC_CAR1) at https://cran.r-project.org/web/packages/rmRNAseq/index.html. Reproducible R codes for data analysis and simulation are available at https://github.com/ntyet/rmRNAseq/tree/master/simulation.
Yet Nguyen, Dan Nettleton
Bioinform.2
2018 A hidden Markov tree model for testing multiple hypotheses corresponding to Gene Ontology gene sets
abstract
BACKGROUND: Testing predefined gene categories has become a common practice for scientists analyzing high throughput transcriptome data. A systematic way of testing gene categories leads to testing hundreds of null hypotheses that correspond to nodes in a directed acyclic graph. The relationships among gene categories induce logical restrictions among the corresponding null hypotheses. An existing fully Bayesian method is powerful but computationally demanding. RESULTS: We develop a computationally efficient method based on a hidden Markov tree model (HMTM). Our method is several orders of magnitude faster than the existing fully Bayesian method. Through simulation and an expression quantitative trait loci study, we show that the HMTM method provides more powerful results than other existing methods that honor the logical restrictions. CONCLUSIONS: The HMTM method provides an individual estimate of posterior probability of being differentially expressed for each gene set, which can be useful for result interpretation. The R package can be found on https://github.com/k22liang/HMTGO .
Kun Liang 0004, Chuanlong Du, Hankun You, Dan Nettleton
BMC Bioinform.4
2018 Crowdsourcing image analysis for plant phenomics to generate ground truth data for machine learning
abstract
The accuracy of machine learning tasks critically depends on high quality ground truth data. Therefore, in many cases, producing good ground truth data typically involves trained professionals; however, this can be costly in time, effort, and money. Here we explore the use of crowdsourcing to generate a large number of training data of good quality. We explore an image analysis task involving the segmentation of corn tassels from images taken in a field setting. We investigate the accuracy, speed and other quality metrics when this task is performed by students for academic credit, Amazon MTurk workers, and Master Amazon MTurk workers. We conclude that the Amazon MTurk and Master Mturk workers perform significantly better than the for-credit students, but with no significant difference between the two MTurk worker types. Furthermore, the quality of the segmentation produced by Amazon MTurk workers rivals that of an expert worker. We provide best practices to assess the quality of ground truth data, and to compare data quality produced by different sources. We conclude that properly managed crowdsourcing can be used to establish large volumes of viable ground truth data at a low cost and high quality, especially in the context of high throughput plant phenotyping. We also provide several metrics for assessing the quality of the generated datasets.
Naihui Zhou, Zachary D. Siegel, Scott Zarecor, Nigel Lee, Darwin A. Campbell, Carson M. Andorf, Dan Nettleton, Carolyn J. Lawrence-Dill, Baskar Ganapathysubramanian, Jonathan W. Kelly, Iddo Friedberg
PLoS Comput. Biol.7
2015 SimSeq: a nonparametric approach to simulation of RNA-sequence datasets
abstract
MOTIVATION: RNA sequencing analysis methods are often derived by relying on hypothetical parametric models for read counts that are not likely to be precisely satisfied in practice. Methods are often tested by analyzing data that have been simulated according to the assumed model. This testing strategy can result in an overly optimistic view of the performance of an RNA-seq analysis method. RESULTS: We develop a data-based simulation algorithm for RNA-seq data. The vector of read counts simulated for a given experimental unit has a joint distribution that closely matches the distribution of a source RNA-seq dataset provided by the user. We conduct simulation experiments based on the negative binomial distribution and our proposed nonparametric simulation algorithm. We compare performance between the two simulation experiments over a small subset of statistical methods for RNA-seq analysis available in the literature. We use as a benchmark the ability of a method to control the false discovery rate. Not surprisingly, methods based on parametric modeling assumptions seem to perform better with respect to false discovery rate control when data are simulated from parametric models rather than using our more realistic nonparametric simulation strategy. AVAILABILITY AND IMPLEMENTATION: The nonparametric simulation algorithm developed in this article is implemented in the R package SimSeq, which is freely available under the GNU General Public License (version 2 or later) from the Comprehensive R Archive Network (http://cran.rproject.org/). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sam Benidt, Dan Nettleton
Bioinform.2
2014 An improved method for computing q-values when the distribution of effect sizes is asymmetric
abstract
MOTIVATION: Asymmetry is frequently observed in the empirical distribution of test statistics that results from the analysis of gene expression experiments. This asymmetry indicates an asymmetry in the distribution of effect sizes. A common method for identifying differentially expressed (DE) genes in a gene expression experiment while controlling false discovery rate (FDR) is Storey's q-value method. This method ranks genes based solely on the P-values from each gene in the experiment. RESULTS: We propose a method that alters and improves upon the q-value method by taking the sign of the test statistics, in addition to the P-values, into account. Through two simulation studies (one involving independent normal data and one involving microarray data), we show that the proposed method, when compared with the traditional q-value method, generally provides a better ranking for genes as well as a higher number of truly DE genes declared to be DE, while still adequately controlling FDR. We illustrate the proposed method by analyzing two microarray datasets, one from an experiment of thale cress seedlings and the other from an experiment of maize leaves. AVAILABILITY AND IMPLEMENTATION: The R code and data files for the proposed method and examples are available at Bioinformatics online.
Megan Orr, Peng Liu 0012, Dan Nettleton
Bioinform.3
2014 Copy number variation detection using next generation sequencing read counts
abstract
BACKGROUND: A copy number variation (CNV) is a difference between genotypes in the number of copies of a genomic region. Next generation sequencing (NGS) technologies provide sensitive and accurate tools for detecting genomic variations that include CNVs. However, statistical approaches for CNV identification using NGS are limited. We propose a new methodology for detecting CNVs using NGS data. This method (henceforth denoted by m-HMM) is based on a hidden Markov model with emission probabilities that are governed by mixture distributions. We use the Expectation-Maximization (EM) algorithm to estimate the parameters in the model. RESULTS: A simulation study demonstrates that our proposed m-HMM approach has greater power for detecting copy number gains and losses relative to existing methods. Furthermore, application of our m-HMM to DNA sequencing data from the two maize inbred lines B73 and Mo17 to identify CNVs that may play a role in creating phenotypic differences between these inbred lines provides results concordant with previous array-based efforts to identify CNVs. CONCLUSIONS: The new m-HMM method is a powerful and practical approach for identifying CNVs from NGS data.
Dan Nettleton, Kai Ying
BMC Bioinform.2
2008 Identification of differentially expressed gene categories in microarray studies using nonparametric multivariate analysis
abstract
MOTIVATION: The field of microarray data analysis is shifting emphasis from methods for identifying differentially expressed genes to methods for identifying differentially expressed gene categories. The latter approaches utilize a priori information about genes to group genes into categories and enhance the interpretation of experiments aimed at identifying expression differences across treatments. While almost all of the existing approaches for identifying differentially expressed gene categories are practically useful, they suffer from a variety of drawbacks. Perhaps most notably, many popular tools are based exclusively on gene-specific statistics that cannot detect many types of multivariate expression change. RESULTS: We have developed a nonparametric multivariate method for identifying gene categories whose multivariate expression distribution differs across two or more conditions. We illustrate our approach and compare its performance to several existing procedures via the analysis of a real data set and a unique data-based simulation study designed to capture the challenges and complexities of practical data analysis. We show that our method has good power for differentiating between differentially expressed and non-differentially expressed gene categories, and we utilize a resampling based strategy for controlling the false discovery rate when testing multiple categories. AVAILABILITY: R code (www.r-project.org) for implementing our approach is available from the first author by request.
Dan Nettleton, Justin Recknor, James M. Reecy
Bioinform.1
2007 Pooling mRNA in microarray experiments and its effect on power
abstract
MOTIVATION: Microarrays can simultaneously measure the expression levels of many genes and are widely applied to study complex biological problems at the genetic level. To contain costs, instead of obtaining a microarray on each individual, mRNA from several subjects can be first pooled and then measured with a single array. mRNA pooling is also necessary when there is not enough mRNA from each subject. Several studies have investigated the impact of pooling mRNA on inferences about gene expression, but have typically modeled the process of pooling as if it occurred in some transformed scale. This assumption is unrealistic. RESULTS: We propose modeling the gene expression levels in a pool as a weighted average of mRNA expression of all individuals in the pool on the original measurement scale, where the weights correspond to individual sample contributions to the pool. Based on these improved statistical models, we develop the appropriate F statistics to test for differentially expressed genes. We present formulae to calculate the power of various statistical tests under different strategies for pooling mRNA and compare resulting power estimates to those that would be obtained by following the approach proposed by Kendziorski et al. (2003). We find that the Kendziorski estimate tends to exceed true power and that the estimate we propose, while somewhat conservative, is less biased. We argue that it is possible to design a study that includes mRNA pooling at a significantly reduced cost but with little loss of information.
Wuyan Zhang, Alicia L. Carriquiry, Dan Nettleton, Jack C. M. Dekkers
Bioinform.3
2006 Scanning microarrays at multiple intensities enhances discovery of differentially expressed genes
abstract
MOTIVATION: Scanning parameters are often overlooked when optimizing microarray experiments. A scanning approach that extends the dynamic data range by acquiring multiple scans of different intensities has been developed. RESULTS: Data from each of three scan intensities (low, medium, high) were analyzed separately using multiple scan and linear regression approaches to identify and compare the sets of genes that exhibit statistically significant differential expression. In the multiple scan approach only one-third of the differentially expressed genes were shared among the three intensities, and each scan intensity identified unique sets of differentially expressed genes. The set of differentially expressed genes from any one scan amounted to < 70% of the total number of genes identified in at least one scan. The average signal intensity of genes that exhibited statistically significant changes in expression was highest for the low-intensity scan and lowest for the high-intensity scan, suggesting that low-intensity scans may be best for detecting expression differences in high-signal genes, while high-intensity scans may be best for detecting expression differences in low-signal genes. Comparison of the differentially expressed genes identified in the multiple scan and linear regression approaches revealed that the multiple scan approach effectively identifies a subset of statistically significant genes that linear regression approach is unable to identify. Quantitative RT-PCR (qRT-PCR) tests demonstrated that statistically significant differences identified at all three scan intensities can be verified. AVAILABILITY: The data presented can be viewed at http://www.ncbi.nlm.nih.gov/geo/ under GEO accession no. GSE3017.
David S. Skibbe, Lisa A. Borsuk, Dan Nettleton, Patrick S. Schnable
Bioinform.5
2002 Expression pattern of yeast sporulation: a case study for regulatory changes after yeast genome duplication
Wei Huang 0032, Dan Nettleton, Xun Gu 0002
Inf. Sci.2