Robert Tibshirani

dblp:t/RobertTibshirani · DBLP profile ↗
← Back
31ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0003-0553-5090ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Structure-preserving multivariate hypothesis testing for mass spectrometry imaging and single-cell data
abstract
MOTIVATION: Mass spectrometry imaging (MSI) and single-cell RNA sequencing (scRNA-seq) offer high-resolution views into tissue and cellular heterogeneity. However, conventional statistical analyses often treat sub-measurements (pixels or cells) as independent, ignoring their nested origin from individual samples. This assumption inflates the effective sample size, increases false discoveries, and undermines biological interpretation. Pixel- or cell-level averaging avoids this but sacrifices spatial or cellular resolution. RESULTS: We evaluated block-SAM on both simulated and real-world datasets. In DESI-MSI data from kidney, lung, and ovarian tumors, block-SAM consistently identified fewer-but more reliable-differential features compared to traditional-SAM. For example, in the kidney dataset, traditional-SAM identified 569 features between tumor and normal tissue, while block-SAM identified 186-all overlapping but excluding 383 likely false positives. Applied to a metastatic RCC scRNA-seq dataset comparing immune checkpoint blockade (ICB)-treated versus untreated patients, traditional-SAM identified over 19,000 differentially expressed genes in malignant cells; block-SAM reduced this to 19. These results demonstrate block-SAM's ability to reduce false discoveries while retaining biologically meaningful signals across diverse high-dimensional datasets. AVAILABILITY AND IMPLEMENTATION: All code and data associated with this study are deposited on Zenodo (https://doi.org/10.5281/zenodo.18273497). The samr package is freely available on the Comprehensive R Archive Network (CRAN).
Keziah E. Liebenberg, Erin Craig, Robert Tibshirani, Livia Eberlin
Bioinform.3
2026 Adaptive Forward Stepwise: A Method for High Sparsity Regression
abstract
This paper proposes a sparse regression method that continuously interpolates between Forward Stepwise selection (FS) and the LASSO. When tuned appropriately, our solutions are much sparser than typical LASSO fits but, unlike FS fits, benefit from the stabilizing effect of shrinkage. Our method, Adaptive Forward Stepwise Regression (AFS) addresses the need for sparser models with shrinkage. We show its connection with boosting via a soft-thresholding viewpoint and demonstrate the ease of adapting the method to classification tasks. In both simulations and real data, our method has lower mean squared error and fewer selected features across multiple settings compared to popular sparse modeling procedures.
Ivy Zhang, Robert Tibshirani
J. Mach. Learn. Res.2
2025 BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
abstract
The development of vision-language models (VLMs) is driven by large-scale and diverse multi-modal datasets. However, progress toward generalist biomedical VLMs is limited by the lack of annotated, publicly accessible datasets across biology and medicine. Existing efforts are limited to narrow domains, missing the full diversity of biomedical knowledge encoded in scientific literature. To address this gap, we introduce BIOMEDICA: a scalable, open-source framework to extract, annotate, and serialize the entirety of the PubMed Central Open Access subset into an easy-to-use, publicly accessible dataset. Our framework produces a comprehensive archive with over 24 million unique image-text pairs from over 6 million articles. Metadata and expert-guided annotations are additionally provided.We demonstrate the utility and accessibility of our resource by releasing BMC-CLIP, a suite of CLIP-style models continuously pre-trained on BIOMEDICA dataset via streaming (eliminating the need to download 27 TB of data locally). On average, our models achieve state-of-the-art performance across 40 tasks — spanning pathology, radiology, ophthalmology, dermatology, surgery, molecular biology, parasitology, and cell biology — excelling in zero-shot classification with 6.56% average improvement (as high as 29.8% and 17.5% in dermatology and ophthalmology, respectively) and stronger image-text retrieval while using 10x less compute. To foster reproducibility and collaboration, we release our codebase1,2and dataset3to the broader research community.
Alejandro Lozano, Min Woo Sun, James Burgess, Liangyu Chen 0005, Jeffrey J. Nirschl, Jeffrey Gu, Iván López 0001, Josiah Aklilu, Anita Rau, Austin Wolfgang Katzer, Collin Chiu, Alfred Seunghoon Song, Robert Tibshirani, Serena Yeung-Levy
CVPR15
2021 LassoNet: Neural Networks with Feature Sparsity
abstract
Much work has been done recently to make neural networks more interpretable, and one approach is to arrange for the network to use only a subset of the available features. In linear models, Lasso (or $\ell_1$-regularized) regression assigns zero weights to the most irrelevant or redundant features, and is widely used in data science. However the Lasso only applies to linear models. Here we introduce LassoNet, a neural network framework with global feature selection. Our approach achieves feature sparsity by allowing a feature to participate in a hidden unit only if its linear representative is active. Unlike other approaches to feature selection for neural nets, our method uses a modified objective function with constraints, and so integrates feature selection with the parameter learning directly. As a result, it delivers an entire regularization path of solutions with a range of feature sparsity. In experiments with real and simulated data, LassoNet significantly outperforms state-of-the-art methods for feature selection and regression. The LassoNet method uses projected proximal gradient descent, and generalizes directly to deep networks. It can be implemented by adding just a few lines of code to a standard neural network.
Ismael Lemhadri, Feng Ruan, Robert Tibshirani
AISTATS3
2021 Fast numerical optimization for genome sequencing data in population biobanks
abstract
MOTIVATION: Large-scale and high-dimensional genome sequencing data poses computational challenges. General-purpose optimization tools are usually not optimal in terms of computational and memory performance for genetic data. RESULTS: We develop two efficient solvers for optimization problems arising from large-scale regularized regressions on millions of genetic variants sequenced from hundreds of thousands of individuals. These genetic variants are encoded by the values in the set {0,1,2,NA}. We take advantage of this fact and use two bits to represent each entry in a genetic matrix, which reduces memory requirement by a factor of 32 compared to a double precision floating point representation. Using this representation, we implemented an iteratively reweighted least square algorithm to solve Lasso regressions on genetic matrices, which we name snpnet-2.0. When the dataset contains many rare variants, the predictors can be encoded in a sparse matrix. We utilize the sparsity in the predictor matrix to further reduce memory requirement and computational speed. Our sparse genetic matrix implementation uses both the compact two-bit representation and a simplified version of compressed sparse block format so that matrix-vector multiplications can be effectively parallelized on multiple CPU cores. To demonstrate the effectiveness of this representation, we implement an accelerated proximal gradient method to solve group Lasso on these sparse genetic matrices. This solver is named sparse-snpnet, and will also be included as part of snpnet R package. Our implementation is able to solve Lasso and group Lasso, linear, logistic and Cox regression problems on sparse genetic matrices that contain 1 000 000 variants and almost 100 000 individuals within 10 min and using less than 32GB of memory. AVAILABILITY AND IMPLEMENTATION: https://github.com/rivas-lab/snpnet/tree/compact.
Christopher Chang, Yosuke Tanigawa, Balasubramanian Narasimhan, Trevor J. Hastie, Robert Tibshirani, Manuel A. Rivas
Bioinform.6
2021 Survival analysis on rare events using group-regularized multi-response Cox regression
abstract
MOTIVATION: The prediction performance of Cox proportional hazard model suffers when there are only few uncensored events in the training data. RESULTS: We propose a Sparse-Group regularized Cox regression method to improve the prediction performance of large-scale and high-dimensional survival data with few observed events. Our approach is applicable when there is one or more other survival responses that 1. has a large number of observed events; 2. share a common set of associated predictors with the rare event response. This scenario is common in the UK Biobank dataset where records for a large number of common and less prevalent diseases of the same set of individuals are available. By analyzing these responses together, we hope to achieve higher prediction performance than when they are analyzed individually. To make this approach practical for large-scale data, we developed an accelerated proximal gradient optimization algorithm as well as a screening procedure inspired by Qian et al. AVAILABILITYANDIMPLEMENTATION: https://github.com/rivas-lab/multisnpnet-Cox.
Yosuke Tanigawa, Johanne M. Justesen, Trevor J. Hastie, Robert Tibshirani, Manuel A. Rivas
Bioinform.6
2021 MassExplorer: a computational tool for analyzing desorption electrospray ionization mass spectrometry data
abstract
Summary: In the last few years, desorption electrospray ionization mass spectrometry imaging (DESI-MSI) has been increasingly used for simultaneous detection of thousands of metabolites and lipids from human tissues and biofluids. To successfully find the most significant differences between two sets of DESI-MSI data (e.g., healthy vs disease) requires the application of accurate computational and statistical methods that can pre-process the data under various normalization settings and help identify these changes among thousands of detected metabolites. Here, we report MassExplorer, a novel computational tool, to help pre-process DESI-MSI data, visualize raw data, build predictive models using the statistical lasso approach to select for a sparse set of significant molecular changes, and interpret selected metabolites. This tool, which is available for both online and offline use, is flexible for both chemists and biologists and statisticians as it helps in visualizing structure of DESI-MSI data and in analyzing the statistically significant metabolites that are differentially expressed across both sample types. Based on the modules in MassExplorer, we expect it to be immediately useful for various biological and chemical applications in mass spectrometry. Availability and implementation: MassExplorer is available as an online R-Shiny application or Mac OS X compatible standalone application. The application, sample performance, source code and corresponding guide can be found at: https://zarelab.com/research/massexplorer-a-tool-to-help-guide-analysis-of-mass-spectrometry-samples/. Supplementary informationMATION: Supplementary data are available at Bioinformatics online.
Vishnu Shankar, Robert Tibshirani, Richard N. Zare
Bioinform.2
2021 LassoNet: A Neural Network with Feature Sparsity
abstract
Much work has been done recently to make neural networks more interpretable, and one approach is to arrange for the network to use only a subset of the available features. In linear models, Lasso (or $\ell_1$-regularized) regression assigns zero weights to the most irrelevant or redundant features, and is widely used in data science. However the Lasso only applies to linear models. Here we introduce LassoNet, a neural network framework with global feature selection. Our approach achieves feature sparsity by adding a skip (residual) layer and allowing a feature to participate in any hidden layer only if its skip-layer representative is active. Unlike other approaches to feature selection for neural nets, our method uses a modified objective function with constraints, and so integrates feature selection with the parameter learning directly. As a result, it delivers an entire regularization path of solutions with a range of feature sparsity. We apply LassoNet to a number of real-data problems and find that it significantly outperforms state-of-the-art methods for feature selection and regression. LassoNet uses projected proximal gradient descent, and generalizes directly to deep networks. It can be implemented by adding just a few lines of code to a standard neural network.
Ismael Lemhadri, Feng Ruan, Louis Abraham, Robert Tibshirani
J. Mach. Learn. Res.4
2021 De novo mutational signature discovery in tumor genomes using SparseSignatures
abstract
Cancer is the result of mutagenic processes that can be inferred from tumor genomes by analyzing rate spectra of point mutations, or "mutational signatures". Here we present SparseSignatures, a novel framework to extract signatures from somatic point mutation data. Our approach incorporates a user-specified background signature, employs regularization to reduce noise in non-background signatures, uses cross-validation to identify the number of signatures, and is scalable to large datasets. We show that SparseSignatures outperforms current state-of-the-art methods on simulated data using a variety of standard metrics. We then apply SparseSignatures to whole genome sequences of pancreatic and breast tumors, discovering well-differentiated signatures that are linked to known mutagenic mechanisms and are strongly associated with patient clinical features.
Avantika Lal, Keli Liu, Robert Tibshirani, Arend Sidow, Daniele Ramazzotti
PLoS Comput. Biol.3
2019 Multiomics modeling of the immunome, transcriptome, microbiome, proteome and metabolome adaptations during human pregnancy
abstract
Motivation: Multiple biological clocks govern a healthy pregnancy. These biological mechanisms produce immunologic, metabolomic, proteomic, genomic and microbiomic adaptations during the course of pregnancy. Modeling the chronology of these adaptations during full-term pregnancy provides the frameworks for future studies examining deviations implicated in pregnancy-related pathologies including preterm birth and preeclampsia. Results: We performed a multiomics analysis of 51 samples from 17 pregnant women, delivering at term. The datasets included measurements from the immunome, transcriptome, microbiome, proteome and metabolome of samples obtained simultaneously from the same patients. Multivariate predictive modeling using the Elastic Net (EN) algorithm was used to measure the ability of each dataset to predict gestational age. Using stacked generalization, these datasets were combined into a single model. This model not only significantly increased predictive power by combining all datasets, but also revealed novel interactions between different biological modalities. Future work includes expansion of the cohort to preterm-enriched populations and in vivo analysis of immune-modulating interventions based on the mechanisms identified. Availability and implementation: Datasets and scripts for reproduction of results are available through: https://nalab.stanford.edu/multiomics-pregnancy/. Supplementary information: Supplementary data are available at Bioinformatics online.
Mohammad Sajjad Ghaemi, Daniel B. DiGiulio, Kévin Contrepois, Benjamin J. Callahan, Thuy T. M. Ngo, Brittany Lee-McMullen, Benoit Lehallier, Anna Robaczewska, David Mcilwain, Yael Rosenberg-Hasson, Ronald J. Wong, Cecele Quaintance, Anthony Culos, Natalie Stanley, Athena Tanada, Amy Tsai, Dyani Gaudilliere, Edward Ganio, Xiaoyuan Han, Kazuo Ando, Leslie McNeil, Martha Tingle, Paul H. Wise, Ivana Maric, Marina Sirota, Tony Wyss-Coray, Virginia D. Winn, Maurice L. Druzin, Ronald Gibbs, Gary L. Darmstadt, David B. Lewis, Vahid Partovi Nia, Bruno Agard, Robert Tibshirani, Garry P. Nolan, Michael Snyder 0001, David A. Relman, Stephen R. Quake, Gary M. Shaw, David K. Stevenson, Martin S. Angst, Brice Gaudilliere, Nima Aghaeepour
Bioinform.34
2013 Scientific research in the age of omics: the good, the bad, and the sloppy
abstract
It has been claimed that most research findings are false, and it is known that large-scale studies involving omics data are especially prone to errors in design, execution, and analysis. The situation is alarming because taxpayer dollars fund a substantial amount of biomedical research, and because the publication of a research article that is later determined to be flawed can erode the credibility of an entire field, resulting in a severe and negative impact for years to come. Here, we urge the development of an online, open-access, postpublication, peer review system that will increase the accountability of scientists for the quality of their research and the ability of readers to distinguish good from sloppy science.
Daniela M. Witten, Robert Tibshirani
J. Am. Medical Informatics Assoc.2
2010 DR-Integrator: a new analytic tool for integrating DNA copy number and gene expression data
abstract
SUMMARY: DNA copy number alterations (CNA) frequently underlie gene expression changes by increasing or decreasing gene dosage. However, only a subset of genes with altered dosage exhibit concordant changes in gene expression. This subset is likely to be enriched for oncogenes and tumor suppressor genes, and can be identified by integrating these two layers of genome-scale data. We introduce DNA/RNA-Integrator (DR-Integrator), a statistical software tool to perform integrative analyses on paired DNA copy number and gene expression data. DR-Integrator identifies genes with significant correlations between DNA copy number and gene expression, and implements a supervised analysis that captures genes with significant alterations in both DNA copy number and gene expression between two sample classes. AVAILABILITY: DR-Integrator is freely available for non-commercial use from the Pollack Lab at http://pollacklab.stanford.edu/ and can be downloaded as a plug-in application to Microsoft Excel and as a package for the R statistical computing environment. The R package is available under the name 'DRI' at http://cran.r-project.org/. An example analysis using DR-Integrator is included as supplemental material. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Keyan Salari, Robert Tibshirani, Jonathan R. Pollack
Bioinform.2
2010 Spectral Regularization Algorithms for Learning Large Incomplete Matrices
Rahul Mazumder, Trevor J. Hastie, Robert Tibshirani
J. Mach. Learn. Res.3
2009 Estimation of Sparse Binary Pairwise Markov Networks using Pseudo-likelihoods
Holger Höfling, Robert Tibshirani
J. Mach. Learn. Res.2
2008 Regularization paths and coordinate descent
abstract
In a statistical world faced with an explosion of data, regularization has become an important ingredient. In a wide variety of problems we have many more input features than observations, and the lasso penalty and its hybrids have become increasingly useful for both feature selection and regularization. This talk presents some effective algorithms based on coordinate descent for fitting large scale regularization paths for a variety of problems.
Trevor J. Hastie, Jerome H. Friedman, Robert Tibshirani
KDD3
2007 Disease-specific genomic analysis: identifying the signature of pathologic biology
abstract
MOTIVATION: Genomic high-throughput technology generates massive data, providing opportunities to understand countless facets of the functioning genome. It also raises profound issues in identifying data relevant to the biology being studied. RESULTS: We introduce a method for the analysis of pathologic biology that unravels the disease characteristics of high dimensional data. The method, disease-specific genomic analysis (DSGA), is intended to precede standard techniques like clustering or class prediction, and enhance their performance and ability to detect disease. DSGA measures the extent to which the disease deviates from a continuous range of normal phenotypes, and isolates the aberrant component of data. In several microarray cancer datasets, we show that DSGA outperforms standard methods. We then use DSGA to highlight a novel subdivision of an important class of genes in breast cancer, the estrogen receptor (ER) cluster. We also identify new markers distinguishing ductal and lobular breast cancers. Although our examples focus on microarrays, DSGA generalizes to any high dimensional genomic/proteomic data.
Monica Nicolau, Robert Tibshirani, Anne-Lise Børresen-Dale, Stefanie S. Jeffrey
Bioinform.2
2007 Margin Trees for High-dimensional Classification
Robert Tibshirani, Trevor J. Hastie
J. Mach. Learn. Res.1
2006 A simple method for assessing sample sizes in microarray experiments
abstract
BACKGROUND: In this short article, we discuss a simple method for assessing sample size requirements in microarray experiments. RESULTS: Our method starts with the output from a permutation-based analysis for a set of pilot data, e.g. from the SAM package. Then for a given hypothesized mean difference and various samples sizes, we estimate the false discovery rate and false negative rate of a list of genes; these are also interpretable as per gene power and type I error. We also discuss application of our method to other kinds of response variables, for example survival outcomes. CONCLUSION: Our method seems to be useful for sample size assessment in microarray experiments.
Robert Tibshirani
BMC Bioinform.1
2004 The Entire Regularization Path for the Support Vector Machine
abstract
In this paper we argue that the choice of the SVM cost parameter can be critical. We then derive an algorithm that can fit the entire path of SVM solutions for every value of the cost parameter, with essentially the same computational cost as fitting one SVM model.
Trevor J. Hastie, Saharon Rosset, Robert Tibshirani, Ji Zhu 0001
NIPS3
2004 Sample classification from protein mass spectrometry, by 'peak probability contrasts'
abstract
MOTIVATION: Early cancer detection has always been a major research focus in solid tumor oncology. Early tumor detection can theoretically result in lower stage tumors, more treatable diseases and ultimately higher cure rates with less treatment-related morbidities. Protein mass spectrometry is a potentially powerful tool for early cancer detection. We propose a novel method for sample classification from protein mass spectrometry data. When applied to spectra from both diseased and healthy patients, the 'peak probability contrast' technique provides a list of all common peaks among the spectra, their statistical significance and their relative importance in discriminating between the two groups. We illustrate the method on matrix-assisted laser desorption and ionization mass spectrometry data from a study of ovarian cancers. RESULTS: Compared to other statistical approaches for class prediction, the peak probability contrast method performs as well or better than several methods that require the full spectra, rather than just labelled peaks. It is also much more interpretable biologically. The peak probability contrast method is a potentially useful tool for sample classification from protein mass spectrometry data.
Robert Tibshirani, Trevor J. Hastie, Balasubramanian Narasimhan, Scott G. Soltys, Gongyi Shi, Albert C. Koong, Quynh-Thu Le
Bioinform.1
2004 Cancer characterization and feature set extraction by discriminative margin clustering
abstract
BACKGROUND: A central challenge in the molecular diagnosis and treatment of cancer is to define a set of molecular features that, taken together, distinguish a given cancer, or type of cancer, from all normal cells and tissues. RESULTS: Discriminative margin clustering is a new technique for analyzing high dimensional quantitative datasets, specially applicable to gene expression data from microarray experiments related to cancer. The goal of the analysis is find highly specialized sub-types of a tumor type which are similar in having a small combination of genes which together provide a unique molecular portrait for distinguishing the sub-type from any normal cell or tissue. Detection of the products of these genes can then, in principle, provide a basis for detection and diagnosis of a cancer, and a therapy directed specifically at the distinguishing constellation of molecular features can, in principle, provide a way to eliminate the cancer cells, while minimizing toxicity to any normal cell. CONCLUSIONS: The new methodology yields highly specialized tumor subtypes which are similar in terms of potential diagnostic markers.
Kamesh Munagala, Robert Tibshirani, Patrick O. Brown
BMC Bioinform.2
2004 The Entire Regularization Path for the Support Vector Machine
Trevor J. Hastie, Saharon Rosset, Robert Tibshirani, Ji Zhu 0001
J. Mach. Learn. Res.3
2003 1-norm Support Vector Machines
abstract
The standard 2-norm SVM is known for its good performance in two- In this paper, we consider the 1-norm SVM. We class classi£cation. argue that the 1-norm SVM may have some advantage over the standard 2-norm SVM, especially when there are redundant noise features. We also propose an ef£cient algorithm that computes the whole solution path of the 1-norm SVM, hence facilitates adaptive selection of the tuning parameter for the 1-norm SVM.
Ji Zhu 0001, Saharon Rosset, Trevor J. Hastie, Robert Tibshirani
NIPS4
2003 Note on "Comparison of Model Selection for Regression" by Vladimir Cherkassky and Yunqian Ma
abstract
While Cherkassky and Ma (2003) raise some interesting issues in comparing techniques for model selection, their article appears to be written largely in protest of comparisons made in our book, Elements of Statistical Learning (2001). Cherkassky and Ma feel that we falsely represented the structural risk minimization (SRM) method, which they defend strongly here. In a two-page section of our book (pp. 212-213), we made an honest attempt to compare the SRM method with two related techniques, Aikaike information criterion (AIC) and Bayesian information criterion (BIC). Apparently, we did not apply SRM in the optimal way. We are also accused of using contrived examples, designed to make SRM look bad. Alas, we did introduce some careless errors in our original simulation--errors that were corrected in the second and subsequent printings. Some of these errors were pointed out to us by Cherkassky and Ma (we supplied them with our source code), and as a result we replaced the assessment "SRM performs poorly overall" with a more moderate "the performance of SRM is mixed" (p. 212).
Trevor J. Hastie, Robert Tibshirani, Jerome H. Friedman
Neural Comput.2
2002 Independent Components Analysis through Product Density Estimation
abstract
We present a simple direct approach for solving the ICA problem, using density estimation and maximum likelihood. Given a candi(cid:173) date orthogonal frame, we model each of the coordinates using a semi-parametric density estimate based on cubic splines. Since our estimates have two continuous derivatives, we can easily run a sec(cid:173) ond order search for the frame parameters. Our method performs very favorably when compared to state-of-the-art techniques.
Trevor J. Hastie, Robert Tibshirani
NIPS2
2001 Missing value estimation methods for DNA microarrays
abstract
MOTIVATION: Gene expression microarray experiments can generate data sets with multiple missing expression values. Unfortunately, many algorithms for gene expression analysis require a complete matrix of gene array values as input. For example, methods such as hierarchical clustering and K-means clustering are not robust to missing data, and may lose effectiveness even with a few missing values. Methods for imputing missing data are needed, therefore, to minimize the effect of incomplete data sets on analyses, and to increase the range of data sets to which these algorithms can be applied. In this report, we investigate automated methods for estimating missing data. RESULTS: We present a comparative study of several methods for the estimation of missing values in gene microarray data. We implemented and evaluated three methods: a Singular Value Decomposition (SVD) based method (SVDimpute), weighted K-nearest neighbors (KNNimpute), and row average. We evaluated the methods using a variety of parameter settings and over different real data sets, and assessed the robustness of the imputation methods to the amount of missing data over the range of 1--20% missing values. We show that KNNimpute appears to provide a more robust and sensitive method for missing value estimation than SVDimpute, and both SVDimpute and KNNimpute surpass the commonly used row average method (as well as filling missing values with zeros). We report results of the comparative experiments and provide recommendations and tools for accurate estimation of missing microarray data under a variety of conditions.
Olga G. Troyanskaya, Michael N. Cantor, Gavin Sherlock, Patrick O. Brown, Trevor J. Hastie, Robert Tibshirani, David Botstein, Russ B. Altman
Bioinform.6
1997 Classification by Pairwise Coupling
Trevor J. Hastie, Robert Tibshirani
NIPS2
1996 A Comparison of Some Error Estimates for Neural Network Models
abstract
We discuss a number of methods for estimating the standard error of predicted values from a multilayer perceptron. These methods include the delta method based on the Hessian, bootstrap estimators, and the “sandwich” estimator. The methods are described and compared in a number of examples. We find that the bootstrap methods perform best, partly because they capture variability due to the choice of starting weights.
Robert Tibshirani
Neural Comput.1
1996 Discriminant Adaptive Nearest Neighbor Classification
abstract
Nearest neighbour classification expects the class conditional probabilities to be locally constant, and suffers from bias in high dimensions. We propose a locally adaptive form of nearest neighbour classification to try to ameliorate this curse of dimensionality. We use a local linear discriminant analysis to estimate an effective metric for computing neighbourhoods. We determine the local decision boundaries from centroid information, and then shrink neighbourhoods in directions orthogonal to these local decision boundaries, and elongate them parallel to the boundaries. Thereafter, any neighbourhood-based classifier can be employed, using the modified neighbourhoods. The posterior probabilities tend to be more homogeneous in the modified neighbourhoods. We also propose a method for global dimension reduction, that combines local dimension information. In a number of examples, the methods demonstrate the potential for substantial improvements over nearest neighbour classification.
Trevor J. Hastie, Robert Tibshirani
IEEE Trans. Pattern Anal. Mach. Intell.2
1995 Discriminant Adaptive Nearest Neighbor Classification
Trevor J. Hastie, Robert Tibshirani
KDD2
1995 Discriminant Adaptive Nearest Neighbor Classification and Regression
Trevor J. Hastie, Robert Tibshirani
NIPS2