VLDB 2026 Research / reviewers in the wild / expert
David M. Rocke
dblp:73/1677
· DBLP profile ↗
20ranked-venue papers
4as first author
0since 2021 · last 2020
0000-0002-3958-7318ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 3 first-authorArtificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
14 papers |
Bioinformatics and computational biology · 96% Computational science and engineering · 3% Medical and health informatics · 1% | |
| Artificial intelligence
3 papers |
Optimization for machine learning · 26% Efficient and distributed learning · 26% Image recognition and object detection · 26% |
Topics — the 23 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › proteomics
peptide identification |
0.6 | 2 | 2018 | Learning Concave Conditional Likelihood Models for Improved Analysis of Tandem Mass Spectra · NeurIPS 2018 Gradients of Generative Models for Improved Discriminative Analysis of Tandem Mass Spectra · NIPS 2017 |
Bioinformatics and computational biology › proteomics
protein identification |
0.6 | 2 | 2018 | Learning Concave Conditional Likelihood Models for Improved Analysis of Tandem Mass Spectra · NeurIPS 2018 Gradients of Generative Models for Improved Discriminative Analysis of Tandem Mass Spectra · NIPS 2017 |
Bioinformatics and computational biology › proteomics
tandem mass spectrometry |
0.6 | 2 | 2018 | Learning Concave Conditional Likelihood Models for Improved Analysis of Tandem Mass Spectra · NeurIPS 2018 Gradients of Generative Models for Improved Discriminative Analysis of Tandem Mass Spectra · NIPS 2017 |
Machine learning › Optimization for machine learning
convex optimization |
0.4 | 1 | 2020 | GPU-Accelerated Primal Learning for Extremely Fast Large-Scale Classification · NeurIPS 2020 |
Machine learning › Efficient and distributed learning › hardware acceleration
GPU training |
0.4 | 1 | 2020 | GPU-Accelerated Primal Learning for Extremely Fast Large-Scale Classification · NeurIPS 2020 |
Computer vision › Image recognition and object detection › image classification › large-scale image classification
large-scale classification |
0.4 | 1 | 2020 | GPU-Accelerated Primal Learning for Extremely Fast Large-Scale Classification · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › bayesian network
dynamic bayesian network |
0.2 | 2 | 2018 | Learning Concave Conditional Likelihood Models for Improved Analysis of Tandem Mass Spectra · NeurIPS 2018 Gradients of Generative Models for Improved Discriminative Analysis of Tandem Mass Spectra · NIPS 2017 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.2 | 2 | 2018 | Learning Concave Conditional Likelihood Models for Improved Analysis of Tandem Mass Spectra · NeurIPS 2018 Gradients of Generative Models for Improved Discriminative Analysis of Tandem Mass Spectra · NIPS 2017 |
Bioinformatics and computational biology › gene expression analysis
microarray data analysis |
0.2 | 4 | 2005 | An expression index for Affymetrix GeneChips based on the generalized logarithm · Bioinform. 2005 Variance-stabilizing transformations for two-color microarrays · Bioinform. 2004 Transformation and normalization of oligonucleotide microarray data · Bioinform. 2003 |
Bioinformatics and computational biology
gene expression analysis |
0.2 | 4 | 2005 | A method for detection of differential gene expression in the presence of inter-individual variability in response · Bioinform. 2005 Variance-stabilizing transformations for two-color microarrays · Bioinform. 2004 Estimation of Transformation Parameters for Microarray Data · Bioinform. 2003 |
Bioinformatics and computational biology
proteomics |
0.1 | 1 | 2020 | GPU-Accelerated Primal Learning for Extremely Fast Large-Scale Classification · NeurIPS 2020 |
Bioinformatics and computational biology › genomics
computational genomics |
0.1 | 3 | 2002 | Partial least squares proportional hazard regression for application to DNA microarray survival data · Bioinform. 2002 Multi-class cancer classification via partial least squares with gene expression profiles · Bioinform. 2002 Tumor classification by partial least squares using microarray gene expression data · Bioinform. 2002 |
Bioinformatics and computational biology › gene expression analysis › gene expression data mining
dimension reduction for gene expression |
0.1 | 3 | 2002 | Partial least squares proportional hazard regression for application to DNA microarray survival data · Bioinform. 2002 Multi-class cancer classification via partial least squares with gene expression profiles · Bioinform. 2002 Tumor classification by partial least squares using microarray gene expression data · Bioinform. 2002 |
Bioinformatics and computational biology › gene expression analysis › microarray data analysis
microarray gene expression analysis |
0.1 | 3 | 2002 | Partial least squares proportional hazard regression for application to DNA microarray survival data · Bioinform. 2002 Multi-class cancer classification via partial least squares with gene expression profiles · Bioinform. 2002 Tumor classification by partial least squares using microarray gene expression data · Bioinform. 2002 |
Computational science and engineering › regression modeling
partial least squares |
0.1 | 3 | 2002 | Partial least squares proportional hazard regression for application to DNA microarray survival data · Bioinform. 2002 Multi-class cancer classification via partial least squares with gene expression profiles · Bioinform. 2002 Tumor classification by partial least squares using microarray gene expression data · Bioinform. 2002 |
Bioinformatics and computational biology › biomarker discovery
cancer biomarker discovery |
0.1 | 1 | 2009 | Detecting glycan cancer biomarkers in serum samples using MALDI FT-ICR mass spectrometry data · Bioinform. 2009 |
Bioinformatics and computational biology › gene expression analysis
microarray data preprocessing |
0.1 | 2 | 2003 | Approximate Variance-stabilizing Transformations for Gene-expression Microarray Data · Bioinform. 2003 A variance-stabilizing transformation for gene-expression microarray data · ISMB 2002 |
Bioinformatics and computational biology › gene expression analysis › microarray data preprocessing
variance stabilization |
0.1 | 2 | 2003 | Transformation and normalization of oligonucleotide microarray data · Bioinform. 2003 A variance-stabilizing transformation for gene-expression microarray data · ISMB 2002 |
Bioinformatics and computational biology › gene expression analysis
differential expression analysis |
0.1 | 2 | 2005 | An expression index for Affymetrix GeneChips based on the generalized logarithm · Bioinform. 2005 Variance-stabilizing transformations for two-color microarrays · Bioinform. 2004 |
Bioinformatics and computational biology › gene expression analysis › differential expression analysis
differential gene expression detection |
0.1 | 1 | 2005 | A method for detection of differential gene expression in the presence of inter-individual variability in response · Bioinform. 2005 |
Bioinformatics and computational biology › cancer genomics
cancer classification |
0.0 | 1 | 2002 | Tumor classification by partial least squares using microarray gene expression data · Bioinform. 2002 |
Bioinformatics and computational biology › survival analysis
survival prediction |
0.0 | 1 | 2002 | Partial least squares proportional hazard regression for application to DNA microarray survival data · Bioinform. 2002 |
Medical and health informatics › oncology
cancer diagnosis |
0.0 | 1 | 2009 | Detecting glycan cancer biomarkers in serum samples using MALDI FT-ICR mass spectrometry data · Bioinform. 2009 |
Methods — techniques the papers use, named apart from their topics
trust region newton · 0.9multithreading · 0.9GPU optimization · 0.9message passing · 0.7maximum likelihood estimation · 0.7convex virtual emissions · 0.7max-product inference · 0.6gradient ascent · 0.6fisher kernel · 0.6dynamic bayesian network · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | GPU-Accelerated Primal Learning for Extremely Fast Large-Scale ClassificationabstractOne of the most efficient methods to solve L2 -regularized primal problems, such as logistic regression and linear support vector machine (SVM) classification, is the widely used trust region Newton algorithm, TRON. While TRON has recently been shown to enjoy substantial speedups on shared-memory multi-core systems, exploiting graphical processing units (GPUs) to speed up the method is significantly more difficult, owing to the highly complex and heavily sequential nature of the algorithm. In this work, we show that using judicious GPU-optimization principles, TRON training time for different losses and feature representations may be drastically reduced. For sparse feature sets, we show that using GPUs to train logistic regression classifiers in LIBLINEAR is up to an order-of-magnitude faster than solely using multithreading. For dense feature sets–which impose far more stringent memory constraints–we show that GPUs substantially reduce the lengthy SVM learning times required for state-of-the-art proteomics analysis, leading to dramatic improvements over recently proposed speedups. Furthermore, we show how GPU speedups may be mixed with multithreading to enable such speedups when the dataset is too large for GPU memory requirements; on a massive dense proteomics dataset of nearly a quarter-billion data instances, these mixed-architecture speedups reduce SVM analysis time from over half a week to less than a single day while using limited GPU memory. John T. Halloran, David M. Rocke |
NeurIPS | 2 |
| 2018 | Learning Concave Conditional Likelihood Models for Improved Analysis of Tandem Mass SpectraabstractThe most widely used technology to identify the proteins present in a complex biological sample is tandem mass spectrometry, which quickly produces a large collection of spectra representative of the peptides (i.e., protein subsequences) present in the original sample. In this work, we greatly expand the parameter learning capabilities of a dynamic Bayesian network (DBN) peptide-scoring algorithm, Didea, by deriving emission distributions for which its conditional log-likelihood scoring function remains concave. We show that this class of emission distributions, called Convex Virtual Emissions (CVEs), naturally generalizes the log-sum-exp function while rendering both maximum likelihood estimation and conditional maximum likelihood estimation concave for a wide range of Bayesian networks. Utilizing CVEs in Didea allows efficient learning of a large number of parameters while ensuring global convergence, in stark contrast to Didea’s previous parameter learning framework (which could only learn a single parameter using a costly grid search) and other trainable models (which only ensure convergence to local optima). The newly trained scoring function substantially outperforms the state-of-the-art in both scoring function accuracy and downstream Fisher kernel analysis. Furthermore, we significantly improve Didea’s runtime performance through successive optimizations to its message passing schedule and derive explicit connections between Didea’s new concave score and related MS/MS scoring functions. John T. Halloran, David M. Rocke |
NeurIPS | 2 |
| 2017 | Gradients of Generative Models for Improved Discriminative Analysis of Tandem Mass SpectraabstractTandem mass spectrometry (MS/MS) is a high-throughput technology used to identify the proteins in a complex biological sample, such as a drop of blood. A collection of spectra is generated at the output of the process, each spectrum of which is representative of a peptide (protein subsequence) present in the original complex sample. In this work, we leverage the log-likelihood gradients of generative models to improve the identification of such spectra. In particular, we show that the gradient of a recently proposed dynamic Bayesian network (DBN) may be naturally employed by a kernel-based discriminative classifier. The resulting Fisher kernel substantially improves upon recent attempts to combine generative and discriminative models for post-processing analysis, outperforming all other methods on the evaluated datasets. We extend the improved accuracy offered by the Fisher kernel framework to other search algorithms by introducing Theseus, a DBN representating a large number of widely used MS/MS scoring functions. Furthermore, with gradient ascent and max-product inference at hand, we use Theseus to learn model parameters without any supervision. John T. Halloran, David M. Rocke |
NIPS | 2 |
| 2009 | Detecting glycan cancer biomarkers in serum samples using MALDI FT-ICR mass spectrometry dataabstractAbstract Motivation: The development of better tests to detect cancer in its earliest stages is one of the most sought-after goals in medicine. Especially important are minimally invasive tests that require only blood or urine samples. By profiling oligosaccharides cleaved from glycosylated proteins shed by tumor cells into the blood stream, we hope to determine glycan profiles that will help identify cancer patients using a simple blood test. The data in this article were generated using matrix-assisted laser desorption/ionization Fourier transform ion cyclotron resonance mass spectrometry (MALDI FT-ICR MS). We have developed novel methods for analyzing this type of mass spectrometry data and applied it to eight datasets from three different types of cancer (breast, ovarian and prostate). Results: The techniques we have developed appear to be effective in the analysis of MALDI FT-ICR MS data. We found significant differences between control and cancer groups in all eight datasets, including two structurally related compounds that were found to be significantly different between control and cancer groups in all three types of cancer studied. Availability: The software used to perform the analysis described in this article is available in the form of an R package called fticrms, version 0.6, either from the Comprehensive R Archive Network (http://www.r-project.org/) or from the first author. Contact: [email protected] Donald A. Barkauskas, Hyun Joo An, Scott R. Kronewitter, Maria Lorna de Leoz, Helen K. Chew, Ralph W. de Vere White, Gary S. Leiserowitz, Suzanne Miyamoto, Carlito B. Lebrilla, David M. Rocke |
Bioinform. | 10 |
| 2009 | Papers on normalization, variable selection, classification or clustering of microarray dataabstractOver the last decade or so, there have been large numbers of methods published on approaches for normalization, variable (gene) selection, classification and clustering of microarray data. As indicated in the scope document for Bioinformatics, this requires papers describing new methods for these problems to meet a very high standard, showing important improvement in results for real biological data, as well as novelty. In this editorial, we describe some standards that need to be met for papers in these areas to be seriously considered. We ask that prospective authors consider these points carefully before submission of their papers to Bioinformatics. The role of simulation: Simulation can be useful in investigating the properties of various methods of data analysis. Yet, there are important barriers to credible use of simulation in microarray studies, largely due to what we do not know about the statistical distribution of measured gene expression levels. First, the distribution across transcripts of true expression values is dependent on the biological state of the tissue or cell, and for a given state this is unknown, even in distributional form, and may further exhibit gene- and platform-specific effects. Second, the correlation within biological replicates of true expression is unknown, and is likely unknowable in detail given that it is expressed by a correlation matrix with on the order of a billion entries. Third, the distribution of changes from one biological state to another is unknown. Fourth, the correlation in observational errors in gene expression across genes is unknown and similarly probably unknowable in detail. On the other hand, the measurement error for a given transcript has been well described by several authors (Ideker et al., 2001; Rocke and Durbin, 2001). Given this gap between knowledge and simulation specification, it is likely that any new method can be shown to be superior to some other method(s) by careful choice of simulation parameters, since simulations often include biases in the distributions selected and in other assumptions of the models. Thus, while simulation may still be worthwhile, and a useful tool for exploring robustness and parameter space of a new method, it is insufficient evidence for superiority of a new method without substantial support from significant improvement in results from analysis of real data. Normalization: Normalization necessarily involves a trade-off between its positive role in reducing variability, and its potentially negative role in increasing bias. There are a number of good image analysis, preprocessing, transformation and normalization methods extant for single- and dual-color DNA microarrays. To show that a new method is better requires comparison demonstrating that results in differential expression analysis, classification or clustering are better with the new normalization method than with previous methods. Not one but several previous methods should be chosen for comparison including the most widely used approaches. Several datasets should be used, including spike-in and dilution studies when feasible, as well as ‘real’ biological datasets. Showing that more genes are differentially expressed using a normalization method is not compelling evidence of superiority without a good estimate of the false-positive rate or a compelling biological analysis of the resulting differentially expressed genes. Variable selection: Typically, new variable selection methods are proposed as part of a classification or clustering strategy, and demonstrating superiority of the variable selection method usually means demonstrating superiority of the combined methodology. It is quite important that metrics for evaluation be used that are robust to intra-array correlations and variable selection artifacts. For example, in cross-validation studies in which variable selection is followed by a classification method, selection of variables using all the data and then cross-validating the classification accuracy introduces substantial bias, making classification methods appear more accurate than they really are (Ambroise and McLachlan, 2002). It is important that any method be compared with several of the most widely used existing methods, including baseline approaches such as filtering by t-score or forward stepwise analysis. Such comparisons should be performed on more than one biological dataset. Further, the method must demonstrate significant improvement over existing methods; incremental improvements will not be considered of sufficient interest to warrant review. Classification and prediction: New classification or prediction methods for microarray data enter a crowded arena. From long-standing techniques such as logistic regression and linear discriminant analysis to the more modern support vector machines and neural networks, most known classification methods have already been applied to microarray data. To show that a newly proposed classification method is a real advance, a substantial improvement in performance needs to be shown over a reasonable selection of existing datasets and methods, including commonly used or simple methods. This is because, consciously or subconsciously, the developer of a new method optimizes its characteristics against the datasets to be used for evaluation. Variable selection and parameter choice for all methods needs to be done strictly in the training set (whether there is one training set or many as in cross validation). Resampling methods like permuting the class labels on the arrays or the bootstrap can be used to provide robust estimates of the significance of differential expression, but do not in themselves give estimates of classification performance except to show that the performance is better than chance. Experience shows that there is considerable noise in classification accuracy experiments, so modest increases in achieved accuracy are usually not convincing. Experience also shows that classification performance in a microarray problem depends strongly on the dataset, and less on the variable selection and classification methods. More than modest differences are required to excite interest in a new method. Authors should keep in mind the ‘No Free Lunch Theorems’ of Wolpert and Macready (1997) which demonstrated that there is no optimization/classification method that outperforms all others in all circumstances (Wolpert, 1996). Clustering: Demonstrating superiority of a clustering method is in many ways more difficult than demonstrating superiority in a classification method. Usually, there is no ground truth against which to compare the clustering results. Defining a criterion (e.g. the Rand index) and showing that a clustering method achieves better scores on this criterion is often not compelling, since such criteria are easily optimized (again, consciously or subconsciously) to ensure superiority. For reasons discussed above, simulation is also not usually sufficient. Ideally, a new clustering method would demonstrate novel biological insights or some attractive statistical properties not available from previous methods, including several commonly used methods. Requiring new biological findings is a difficult standard, but a necessary one to insure that new published methods are useful and likely to be used. To conclude, microarrays remain a useful technology to address a wide array of biological problems and the optimal analysis of these data to extract meaningful results still pose many bioinformatics challenges. However, with a number of successful methods already addressing the well-established microarray data analysis problems, publication of new methods in this area requires either identification of a new challenge and formulation of a new problem or development of a substantially better methodology then those existing that can be benchmarked on a variety of datasets. We hope that suggestions provided above for evaluation and validation of such new methods would increase the likelihood of them supporting biological discoveries in the future. Funding: DMR to NIH grants P42-ES04699 and R01-HG003352. David M. Rocke, Trey Ideker, Olga G. Troyanskaya, John Quackenbush, Joaquín Dopazo |
Bioinform. | 1 |
| 2008 | Assessing probe-specific dye and slide biases in two-color microarray dataabstractBACKGROUND: A primary reason for using two-color microarrays is that the use of two samples labeled with different dyes on the same slide, that bind to probes on the same spot, is supposed to adjust for many factors that introduce noise and errors into the analysis. Most users assume that any differences between the dyes can be adjusted out by standard methods of normalization, so that measures such as log ratios on the same slide are reliable measures of comparative expression. However, even after the normalization, there are still probe specific dye and slide variation among the data. We define a method to quantify the amount of the dye-by-probe and slide-by-probe interaction. This serves as a diagnostic, both visual and numeric, of the existence of probe-specific dye bias. We show how this improved the performance of two-color array analysis for arrays for genomic analysis of biological samples ranging from rice to human tissue. RESULTS: We develop a procedure for quantifying the extent of probe-specific dye and slide bias in two-color microarrays. The primary output is a graphical diagnostic of the extent of the bias which called ECDF (Empirical Cumulative Distribution Function), though numerical results are also obtained. CONCLUSION: We show that the dye and slide biases were high for human and rice genomic arrays in two gene expression facilities, even after the standard intensity-based normalization, and describe how this diagnostic allowed the problems causing the probe-specific bias to be addressed, and resulted in important improvements in performance. The R package LMGene which contains the method described in this paper has been available to download from Bioconductor. Ruixiao Lu, Geun-Cheol Lee, Michael Shultz, Chris Dardick, Kihong Jung, Jirapa Phetsom, Yi Jia, Robert H. Rice, Zelanna Goldberg, Patrick S. Schnable, Pamela C. Ronald, David M. Rocke |
BMC Bioinform. | 12 |
| 2008 | Baseline Correction for NMR Spectroscopic Metabolomics Data AnalysisabstractBACKGROUND: We propose a statistically principled baseline correction method, derived from a parametric smoothing model. It uses a score function to describe the key features of baseline distortion and constructs an optimal baseline curve to maximize it. The parameters are determined automatically by using LOWESS (locally weighted scatterplot smoothing) regression to estimate the noise variance. RESULTS: We tested this method on 1D NMR spectra with different forms of baseline distortions, and demonstrated that it is effective for both regular 1D NMR spectra and metabolomics spectra with over-crowded peaks. CONCLUSION: Compared with the automatic baseline correction function in XWINNMR 3.5, the penalized smoothing method provides more accurate baseline correction for high-signal density metabolomics spectra. Yuanxin Xi, David M. Rocke |
BMC Bioinform. | 2 |
| 2007 | On the analysis of glycomics mass spectrometry data via the regularized area under the ROC curveabstractBACKGROUND: Novel molecular and statistical methods are in rising demand for disease diagnosis and prognosis with the help of recent advanced biotechnology. High-resolution mass spectrometry (MS) is one of those biotechnologies that are highly promising to improve health outcome. Previous literatures have identified some proteomics biomarkers that can distinguish healthy patients from cancer patients using MS data. In this paper, an MS study is demonstrated which uses glycomics to identify ovarian cancer. Glycomics is the study of glycans and glycoproteins. The glycans on the proteins may deviate between a cancer cell and a normal cell and may be visible in the blood. High-resolution MS has been applied to measure relative abundances of potential glycan biomarkers in human serum. Multiple potential glycan biomarkers are measured in MS spectra. With the objection of maximizing the empirical area under the ROC curve (AUC), an analysis method was considered which combines potential glycan biomarkers for the diagnosis of cancer. RESULTS: Maximizing the empirical AUC of glycomics MS data is a large-dimensional optimization problem. The technical difficulty is that the empirical AUC function is not continuous. Instead, it is in fact an empirical 0-1 loss function with a large number of linear predictors. An approach was investigated that regularizes the area under the ROC curve while replacing the 0-1 loss function with a smooth surrogate function. The constrained threshold gradient descent regularization algorithm was applied, where the regularization parameters were chosen by the cross-validation method, and the confidence intervals of the regression parameters were estimated by the bootstrap method. The method is called TGDR-AUC algorithm. The properties of the approach were studied through a numerical simulation study, which incorporates the positive values of mass spectrometry data with the correlations between measurements within person. The simulation proved asymptotic properties that estimated AUC approaches the true AUC. Finally, mass spectrometry data of serum glycan for ovarian cancer diagnosis was analyzed. The optimal combination based on TGDR-AUC algorithm yields plausible result and the detected biomarkers are confirmed based on biological evidence. CONCLUSION: The TGDR-AUC algorithm relaxes the normality and independence assumptions from previous literatures. In addition to its flexibility and easy interpretability, the algorithm yields good performance in combining potential biomarkers and is computationally feasible. Thus, the approach of TGDR-AUC is a plausible algorithm to classify disease status on the basis of multiple biomarkers. Jingjing Ye, Crystal Kirmiz, Carlito B. Lebrilla, David M. Rocke |
BMC Bioinform. | 5 |
| 2005 | A method for detection of differential gene expression in the presence of inter-individual variability in responseabstractMOTIVATION: Many stimuli to biological systems result in transcriptional responses that vary across the individual organism either in type or in timing. This creates substantial difficulties in detecting these responses. This is especially the case when the data for any one individual are limited and when the number of genes, probes or probe sets is large. RESULTS: We have developed a procedure that allows for sensitive detection of transcriptional responses that differ between individuals in type or in timing. This consists of four steps: one is to identify a group of genes, probes or probe sets that detect genes that belong to a molecular class or to a common pathway. The second is to conduct a statistical test of the hypothesis that the gene is differentially expressed for each individual and for each gene in the set. The third is to examine the collection of these statistics to see if there is a detectable signal in the aggregate of them. The final step is to assess the significance of this by resampling to avoid correlational bias. AVAILABILITY: Software in the form of R code to perform the required test is available from the first author or from his website http://www.idav.ucdavis.edu/~dmrocke/software; however the procedures are also easily performed using any standard statistical software. David M. Rocke, Zelanna Goldberg, Chad Schweitert, Alison Santana |
Bioinform. | 1 |
| 2005 | An expression index for Affymetrix GeneChips based on the generalized logarithmabstractMOTIVATION: Affymetrix GeneChip high-density oligonucleotide arrays interrogate a single transcript using multiple short 25mer probes. Usually, a necessary step in the analysis of experiments using these GeneChips is to summarize each of these probe sets into a single expression index that can then be used for determining differential expression, for classification, for clustering, and for other analyses. In this paper, we propose a new expression index that is competitive with the best existing methods, and superior in many cases. We call this expression index method GLA, for GLog Average, since after normalization at the probe level, we take the mean generalized logarithm of perfect match probes. RESULTS: In this paper, we use Affycomp as the primary tool to assess the weaknesses and strengths of GLA. Comparisons are made between GLA and most widely used summary methods (RMA, MAS5.0 and MBEI) in great detail. The substantial reduction in variability and increased ability to detect differential expression, together with the simplicity of implementation, make GLA a plausible candidate for analysis of Affymetrix GeneChip data. David M. Rocke |
Bioinform. | 2 |
| 2004 | Variance-stabilizing transformations for two-color microarraysabstractMOTIVATION: Authors of several recent papers have independently introduced a family of transformations (the generalized-log family), which stabilizes the variance of microarray data up to the first order. However, for data from two-color arrays, tests for differential expression may require that the variance of the difference of transformed observations be constant, rather than that of the transformed observations themselves. RESULTS: We introduce a transformation within the generalized-log family which stabilizes, to the first order, the variance of the difference of transformed observations. We also introduce transformations from the 'started-log' and log-linear-hybrid families which provide good approximate variance stabilization of differences. Examples using control-control data show that any of these transformations may provide sufficient variance stabilization for practical applications, and all perform well compared to log ratios. Blythe Durbin, David M. Rocke |
Bioinform. | 2 |
| 2004 | Classification of contamination in salt marsh plants using hyperspectral reflectanceabstractIn this paper, we compare the classification effectiveness of two relatively new techniques on data consisting of leaf-level reflectance from five species of salt marsh and two species of crop plants (in four experiments) that have been exposed to varying levels of different heavy metal or petroleum toxicity, with a control treatment for each experiment. If these methodologies work well on leaf-level data, then there is hope that they will also work well on data from air- and spaceborne platforms. The classification methods compared were support vector classification (SVC) of exposed and nonexposed plants based on the spectral reflectance data, and partial least squares compression of the spectral reflectance data followed by classification using logistic discrimination (PLS/LD). The statistic we used to compare the effectiveness of the methodologies was the leave-one-out cross-validation estimate of the prediction error. Our results suggest that both techniques perform reasonably well, but that SVC was superior to PLS/LD for use on hyperspectral data and it is worth exploring as a technique for classifying heavy-metal or petroleum exposed plants for the more complicated data from air- and spaceborne sensors. Machelle D. Wilson, Susan L. Ustin, David M. Rocke |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2003 | Estimation of Transformation Parameters for Microarray DataabstractMOTIVATION AND RESULTS: Durbin et al. (2002), Huber et al. (2002) and Munson (2001) independently introduced a family of transformations (the generalized-log family) which stabilizes the variance of microarray data up to the first order. We introduce a method for estimating the transformation parameter in tandem with a linear model based on the procedure outlined in Box and Cox (1964). We also discuss means of finding transformations within the generalized-log family which are optimal under other criteria, such as minimum residual skewness and minimum mean-variance dependency. AVAILABILITY: R and Matlab code and test data are available from the authors on request. Blythe Durbin, David M. Rocke |
Bioinform. | 2 |
| 2003 | Transformation and normalization of oligonucleotide microarray dataabstractMOTIVATION: Most methods of analyzing microarray data or doing power calculations have an underlying assumption of constant variance across all levels of gene expression. The most common transformation, the logarithm, results in data that have constant variance at high levels but not at low levels. Rocke and Durbin showed that data from spotted arrays fit a two-component model and Durbin, Hardin, Hawkins, and Rocke, Huber et al. and Munson provided a transformation that stabilizes the variance as well as symmetrizes and normalizes the error structure. We wish to evaluate the applicability of this transformation to the error structure of GeneChip microarrays. RESULTS: We demonstrate in an example study a simple way to use the two-component model of Rocke and Durbin and the data transformation of Durbin, Hardin, Hawkins and Rocke, Huber et al. and Munson on Affymetrix GeneChip data. In addition we provide a method for normalization of Affymetrix GeneChips simultaneous with the determination of the transformation, producing a data set without chip or slide effects but with constant variance and with symmetric errors. This transformation/normalization process can be thought of as a machine calibration in that it requires a few biologically constant replicates of one sample to determine the constant needed to specify the transformation and normalize. It is hypothesized that this constant needs to be found only once for a given technology in a lab, perhaps with periodic updates. It does not require extensive replication in each study. Furthermore, the variance of the transformed pilot data can be used to do power calculations using standard power analysis programs. AVAILABILITY: SPLUS code for the transformation/normalization for four replicates is available from the first author upon request. A program written in C is available from the last author. Sue C. Geller, Jeff P. Gregg, Paul Hagerman, David M. Rocke |
Bioinform. | 4 |
| 2003 | Approximate Variance-stabilizing Transformations for Gene-expression Microarray DataabstractMOTIVATION: A variance stabilizing transformation for microarray data was recently introduced independently by several research groups. This transformation has sometimes been called the generalized logarithm or glog transformation. In this paper, we derive several alternative approximate variance stabilizing transformations that may be easier to use in some applications. RESULTS: We demonstrate that the started-log and the log-linear-hybrid transformation families can produce approximate variance stabilizing transformations for microarray data that are nearly as good as the generalized logarithm (glog) transformation. These transformations may be more convenient in some applications. David M. Rocke, Blythe Durbin |
Bioinform. | 1 |
| 2003 | Sampling and Subsampling for Cluster Analysis in Data Mining: With Applications to Sky Survey Data
David M. Rocke, Jian J. Dai |
Data Min. Knowl. Discov. | 1 |
| 2002 | A variance-stabilizing transformation for gene-expression microarray dataabstractMOTIVATION: Standard statistical techniques often assume that data are normally distributed, with constant variance not depending on the mean of the data. Data that violate these assumptions can often be brought in line with the assumptions by application of a transformation. Gene-expression microarray data have a complicated error structure, with a variance that changes with the mean in a non-linear fashion. Log transformations, which are often applied to microarray data, can inflate the variance of observations near background. RESULTS: We introduce a transformation that stabilizes the variance of microarray data across the full range of expression. Simulation studies also suggest that this transformation approximately symmetrizes microarray data. Blythe Durbin, Johanna S. Hardin, Douglas M. Hawkins, David M. Rocke |
ISMB | 4 |
| 2002 | Tumor classification by partial least squares using microarray gene expression dataabstractMOTIVATION: One important application of gene expression microarray data is classification of samples into categories, such as the type of tumor. The use of microarrays allows simultaneous monitoring of thousands of genes expressions per sample. This ability to measure gene expression en masse has resulted in data with the number of variables p(genes) far exceeding the number of samples N. Standard statistical methodologies in classification and prediction do not work well or even at all when N < p. Modification of existing statistical methodologies or development of new methodologies is needed for the analysis of microarray data. RESULTS: We propose a novel analysis procedure for classifying (predicting) human tumor samples based on microarray gene expressions. This procedure involves dimension reduction using Partial Least Squares (PLS) and classification using Logistic Discrimination (LD) and Quadratic Discriminant Analysis (QDA). We compare PLS to the well known dimension reduction method of Principal Components Analysis (PCA). Under many circumstances PLS proves superior; we illustrate a condition when PCA particularly fails to predict well relative to PLS. The proposed methods were applied to five different microarray data sets involving various human tumor samples: (1) normal versus ovarian tumor; (2) Acute Myeloid Leukemia (AML) versus Acute Lymphoblastic Leukemia (ALL); (3) Diffuse Large B-cell Lymphoma (DLBCLL) versus B-cell Chronic Lymphocytic Leukemia (BCLL); (4) normal versus colon tumor; and (5) Non-Small-Cell-Lung-Carcinoma (NSCLC) versus renal samples. Stability of classification results and methods were further assessed by re-randomization studies. Danh V. Nguyen, David M. Rocke |
Bioinform. | 2 |
| 2002 | Multi-class cancer classification via partial least squares with gene expression profilesabstractMOTIVATION: Discrimination between two classes such as normal and cancer samples and between two types of cancers based on gene expression profiles is an important problem which has practical implications as well as the potential to further our understanding of gene expression of various cancer cells. Classification or discrimination of more than two groups or classes (multi-class) is also needed. The need for multi-class discrimination methodologies is apparent in many microarray experiments where various cancer types are considered simultaneously. RESULTS: Thus, in this paper we present the extension to the classification methodology proposed earlier Nguyen and Rocke (2002b; Bioinformatics, 18, 39-50) to classify cancer samples from multiple classes. The methodologies proposed in this paper are applied to four gene expression data sets with multiple classes: (a) a hereditary breast cancer data set with (1) BRCA1-mutation, (2) BRCA2-mutation and (3) sporadic breast cancer samples, (b) an acute leukemia data set with (1) acute myeloid leukemia (AML), (2) T-cell acute lymphoblastic leukemia (T-ALL) and (3) B-cell acute lymphoblastic leukemia (B-ALL) samples, (c) a lymphoma data set with (1) diffuse large B-cell lymphoma (DLBCL), (2) B-cell chronic lymphocytic leukemia (BCLL) and (3) follicular lymphoma (FL) samples, and (d) the NCI60 data set with cell lines derived from cancers of various sites of origin. In addition, we evaluated the classification algorithms and examined the variability of the error rates using simulations based on randomization of the real data sets. We note that there are other methods for addressing multi-class prediction recently and our approach is along the line of Nguyen and Rocke (2002b; Bioinformatics, 18, 39-50). CONTACT: [email protected]; [email protected] Danh V. Nguyen, David M. Rocke |
Bioinform. | 2 |
| 2002 | Partial least squares proportional hazard regression for application to DNA microarray survival dataabstractAbstract Motivation: Microarrays are increasingly used in cancer research. When gene transcription data from microarray experiments also contains patient survival information, it is often of interest to predict the survival times based on the gene expression. In this paper we consider the well-known proportional hazard (PH) regression model for survival analysis. Ordinarily, the PH model is used with a few covariates and many observations (subjects). We consider here the case that the number of covariates, p, exceeds the number of samples, N, a setting typical of gene expression data from DNA microarrays. Results: For a given vector of response values which are survival times and p gene expressions (covariates) we examine the problem of how to predict the survival probabilities, when N ≪ p. The approach taken to cope with the high dimensionality is to reduce the dimension using partial least squares with the response variable as the vector of survival times. After dimension reduction, the extracted PLS gene components are then used as covariates in a PH regression to predict the survival probabilities. We demonstrate the use of the methodology on two cDNA gene expression data sets, both containing survival data. The first data set contains 40 diffuse large B-cell lymphoma (DLBCL) tissue samples and the second data set contains 49 tissue samples from patients with locally advanced breast cancer in a prospective study. Availability: The methodology can be implemented using a combination of standard statistical methods, available, for example, in SAS. Sample SAS macro codes to implement the methods will be available at http://stat.tamu.edu/~dnguyen/supplemental.html Contact: [email protected]@ucdavis.edu * To whom correspondence should be addressed. Danh V. Nguyen, David M. Rocke |
Bioinform. | 2 |