Richard Simon

dblp:47/6736 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
0since 2021 · last 2014
0000-0003-3558-4598ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
7 papers
Bioinformatics and computational biology · 100%
Databases, data mining, and information retrieval
1 paper
Data mining · 50% Machine learning and data management · 50%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
cancer genomics
0.332012
Estimating the order of mutations during tumorigenesis from tumor genome sequencing data · Bioinform. 2012
Identifying cancer driver genes in tumor genome sequencing studies · Bioinform. 2011
A bioinformatics tool to select sequences for microarray studies of mouse models of oncogenesis · Bioinform. 2002
Bioinformatics and computational biology › cancer genomics
cancer driver gene identification
0.112011
Identifying cancer driver genes in tumor genome sequencing studies · Bioinform. 2011
Bioinformatics and computational biology › cancer genomics
somatic mutation analysis
0.112011
Identifying cancer driver genes in tumor genome sequencing studies · Bioinform. 2011
Bioinformatics and computational biology › gene expression analysis
microarray experimental design
0.122003
Statistical Design of Reverse Dye Microarrays · Bioinform. 2003
Comparison of microarray designs for class comparison and class discovery · Bioinform. 2002
Bioinformatics and computational biology › gene expression analysis
gene expression classification
0.112005
Prediction error estimation: a comparison of resampling methods · Bioinform. 2005
Bioinformatics and computational biology › gene expression analysis
differential expression analysis
0.012003
A random variance model for detection of differential gene expression in small microarray experiments · Bioinform. 2003
Bioinformatics and computational biology
gene expression analysis
0.012003
A random variance model for detection of differential gene expression in small microarray experiments · Bioinform. 2003
Bioinformatics and computational biology › gene expression analysis
microarray data analysis
0.012003
A random variance model for detection of differential gene expression in small microarray experiments · Bioinform. 2003
Bioinformatics and computational biology
genomics
0.012002
A bioinformatics tool to select sequences for microarray studies of mouse models of oncogenesis · Bioinform. 2002
Bioinformatics and computational biology
microarray probe design
0.012002
A bioinformatics tool to select sequences for microarray studies of mouse models of oncogenesis · Bioinform. 2002
Data mining
model selection
0.012005
Prediction error estimation: a comparison of resampling methods · Bioinform. 2005
Machine learning and data management
resampling
0.012005
Prediction error estimation: a comparison of resampling methods · Bioinform. 2005
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › tree search
AND/OR tree search
0.011971
On the Optimal Solutions to AND/OR Series-Parallel Graphs · J. ACM 1971
Graph algorithms and graph theory
graph algorithms
0.011971
On the Optimal Solutions to AND/OR Series-Parallel Graphs · J. ACM 1971

Methods — techniques the papers use, named apart from their topics

probabilistic modeling · 0.1genome sequencing analysis · 0.1statistical testing · 0.1background mutation rate estimation · 0.1cross-validation · 0.1bootstrap · 0.1probability modeling · 0.0optimal estimation · 0.0inverse gamma model · 0.0empirical bayes · 0.0tree searching · 0.0graph reduction · 0.0
YearPublicationVenuePosition
2014 Using single cell sequencing data to model the evolutionary history of a tumor
abstract
BACKGROUND: The introduction of next-generation sequencing (NGS) technology has made it possible to detect genomic alterations within tumor cells on a large scale. However, most applications of NGS show the genetic content of mixtures of cells. Recently developed single cell sequencing technology can identify variation within a single cell. Characterization of multiple samples from a tumor using single cell sequencing can potentially provide information on the evolutionary history of that tumor. This may facilitate understanding how key mutations accumulate and evolve in lineages to form a heterogeneous tumor. RESULTS: We provide a computational method to infer an evolutionary mutation tree based on single cell sequencing data. Our approach differs from traditional phylogenetic tree approaches in that our mutation tree directly describes temporal order relationships among mutation sites. Our method also accommodates sequencing errors. Furthermore, we provide a method for estimating the proportion of time from the earliest mutation event of the sample to the most recent common ancestor of the sample of cells. Finally, we discuss current limitations on modeling with single cell sequencing data and possible improvements under those limitations. CONCLUSIONS: Inferring the temporal ordering of mutational sites using current single cell sequencing data is a challenge. Our proposed method may help elucidate relationships among key mutations and their role in tumor progression.
Kyung In Kim, Richard Simon
BMC Bioinform.2
2013 Using passenger mutations to estimate the timing of driver mutations and identify mutator alterations
abstract
BACKGROUND: Recent developments in high-throughput genomic technologies make it possible to have a comprehensive view of genomic alterations in tumors on a whole genome scale. Only a small number of somatic alterations detected in tumor genomes are driver alterations which drive tumorigenesis. Most of the somatic alterations are passengers that are neutral to tumor cell selection. Although most research efforts are focused on analyzing driver alterations, the passenger alterations also provide valuable information about the history of tumor development. RESULTS: In this paper, we develop a method for estimating the age of the tumor lineage and the timing of the driver alterations based on the number of passenger alterations. This method also identifies mutator genes which increase genomic instability when they are altered and provides estimates of the increased rate of alterations caused by each mutator gene. We applied this method to copy number data and DNA sequencing data for ovarian and lung tumors. We identified well known mutators such as TP53, PRKDC, BRCA1/2 as well as new mutator candidates PPP2R2A and the chromosomal region 22q13.33. We found that most mutator genes alter early during tumorigenesis and were able to estimate the age of individual tumor lineage in cell generations. CONCLUSIONS: This is the first computational method to identify mutator genes and to take into account the increase of the alteration rate by mutator genes, providing more accurate estimates of the tumor age and the timing of driver alterations.
Ahrim Youn, Richard Simon
BMC Bioinform.2
2012 Estimating the order of mutations during tumorigenesis from tumor genome sequencing data
abstract
MOTIVATION: Tumors are thought to develop and evolve through a sequence of genetic and epigenetic somatic alterations to progenitor cells. Early stages of human tumorigenesis are hidden from view. Here, we develop a method for inferring some aspects of the order of mutational events during tumorigenesis based on genome sequencing data for a set of tumors. This method does not assume that the sequence of driver alterations is the same for each tumor, but enables the degree of similarity or difference in the sequence to be evaluated. RESULTS: To evaluate the new method, we applied it to colon cancer tumor sequencing data and the results are consistent with the multi-step tumorigenesis model previously developed based on comparing stages of cancer. We then applied the new method to DNA sequencing data for a set of lung cancers. The model may be a useful tool for better understanding the process of tumorigenesis. AVAILABILITY: The software is available at: http://linus.nci.nih.gov/Data/YounA/OrderMutation.zip.
Ahrim Youn, Richard Simon
Bioinform.2
2011 Identifying cancer driver genes in tumor genome sequencing studies
abstract
MOTIVATION: Major tumor sequencing projects have been conducted in the past few years to identify genes that contain 'driver' somatic mutations in tumor samples. These genes have been defined as those for which the non-silent mutation rate is significantly greater than a background mutation rate estimated from silent mutations. Several methods have been used for estimating the background mutation rate. RESULTS: We propose a new method for identifying cancer driver genes, which we believe provides improved accuracy. The new method accounts for the functional impact of mutations on proteins, variation in background mutation rate among tumors and the redundancy of the genetic code. We reanalyzed sequence data for 623 candidate genes in 188 non-small cell lung tumors using the new method. We found several important genes like PTEN, which were not deemed significant by the previous method. At the same time, we determined that some genes previously reported as drivers were not significant by the new analysis because mutations in these genes occurred mainly in tumors with large background mutation rates. AVAILABILITY: The software is available at: http://linus.nci.nih.gov/Data/YounA/software.zip.
Ahrim Youn, Richard Simon
Bioinform.2
2011 Construct and Compare Gene Coexpression Networks with DAPfinder and DAPview
abstract
BACKGROUND: DAPfinder and DAPview are novel BRB-ArrayTools plug-ins to construct gene coexpression networks and identify significant differences in pairwise gene-gene coexpression between two phenotypes. RESULTS: Each significant difference in gene-gene association represents a Differentially Associated Pair (DAP). Our tools include several choices of filtering methods, gene-gene association metrics, statistical testing methods and multiple comparison adjustments. Network results are easily displayed in Cytoscape. Analyses of glioma experiments and microarray simulations demonstrate the utility of these tools. CONCLUSIONS: DAPfinder is a new friendly-user tool for reconstruction and comparison of biological networks.
Jeff Skinner, Yuri Kotliarov, Sudhir Varma, Karina L. Mine, Anatoly Yambartsev, Richard Simon, Yentram Huyen, Andrey Morgun
BMC Bioinform.6
2011 Microarray-based Cancer Prediction Using Single Genes
abstract
BACKGROUND: Although numerous methods of using microarray data analysis for cancer classification have been proposed, most utilize many genes to achieve accurate classification. This can hamper interpretability of the models and ease of translation to other assay platforms. We explored the use of single genes to construct classification models. We first identified the genes with the most powerful univariate class discrimination ability and then constructed simple classification rules for class prediction using the single genes. RESULTS: We applied our model development algorithm to eleven cancer gene expression datasets and compared classification accuracy to that for standard methods including Diagonal Linear Discriminant Analysis, k-Nearest Neighbor, Support Vector Machine and Random Forest. The single gene classifiers provided classification accuracy comparable to or better than those obtained by existing methods in most cases. We analyzed the factors that determined when simple single gene classification is effective and when more complex modeling is warranted. CONCLUSIONS: For most of the datasets examined, the single-gene classification methods appear to work as well as more standard methods, suggesting that simple models could perform well in microarray-based cancer prediction.
Richard Simon
BMC Bioinform.2
2006 Bias in error estimation when using cross-validation for model selection
abstract
BACKGROUND: Cross-validation (CV) is an effective method for estimating the prediction error of a classifier. Some recent articles have proposed methods for optimizing classifiers by choosing classifier parameter values that minimize the CV error estimate. We have evaluated the validity of using the CV error estimate of the optimized classifier as an estimate of the true error expected on independent data. RESULTS: We used CV to optimize the classification parameters for two kinds of classifiers; Shrunken Centroids and Support Vector Machines (SVM). Random training datasets were created, with no difference in the distribution of the features between the two classes. Using these "null" datasets, we selected classifier parameter values that minimized the CV error estimate. 10-fold CV was used for Shrunken Centroids while Leave-One-Out-CV (LOOCV) was used for the SVM. Independent test data was created to estimate the true error. With "null" and "non null" (with differential expression between the classes) data, we also tested a nested CV procedure, where an inner CV loop is used to perform the tuning of the parameters while an outer CV is used to compute an estimate of the error. The CV error estimate for the classifier with the optimal parameters was found to be a substantially biased estimate of the true error that the classifier would incur on independent data. Even though there is no real difference between the two classes for the "null" datasets, the CV error estimate for the Shrunken Centroid with the optimal parameters was less than 30% on 18.5% of simulated training data-sets. For SVM with optimal parameters the estimated error rate was less than 30% on 38% of "null" data-sets. Performance of the optimized classifiers on the independent test set was no better than chance. The nested CV procedure reduces the bias considerably and gives an estimate of the error that is very close to that obtained on the independent testing set for both Shrunken Centroids and SVM classifiers for "null" and "non-null" data distributions. CONCLUSION: We show that using CV to compute an error estimate for a classifier that has itself been tuned using CV gives a significantly biased estimate of the true error. Proper use of CV for estimating true error of a classifier developed using a well defined algorithm requires that all steps of the algorithm, including classifier parameter tuning, be repeated in each CV loop. A nested CV procedure provides an almost unbiased estimate of the true error.
Sudhir Varma, Richard Simon
BMC Bioinform.2
2005 Comment on 'Evaluation of the gene-specific dye bias in cDNA microarray experiments'
Kevin Dobbin, Joanna H. Shih, Richard Simon
Bioinform.3
2005 Prediction error estimation: a comparison of resampling methods
abstract
MOTIVATION: In genomic studies, thousands of features are collected on relatively few samples. One of the goals of these studies is to build classifiers to predict the outcome of future observations. There are three inherent steps to this process: feature selection, model selection and prediction assessment. With a focus on prediction assessment, we compare several methods for estimating the 'true' prediction error of a prediction model in the presence of feature selection. RESULTS: For small studies where features are selected from thousands of candidates, the resubstitution and simple split-sample estimates are seriously biased. In these small samples, leave-one-out cross-validation (LOOCV), 10-fold cross-validation (CV) and the .632+ bootstrap have the smallest bias for diagonal discriminant analysis, nearest neighbor and classification trees. LOOCV and 10-fold CV have the smallest bias for linear discriminant analysis. Additionally, LOOCV, 5- and 10-fold CV, and the .632+ bootstrap have the lowest mean square error. The .632+ bootstrap is quite biased in small sample sizes with strong signal-to-noise ratios. Differences in performance among resampling methods are reduced as the number of specimens available increase. SUPPLEMENTARY INFORMATION: A complete compilation of results and R code for simulations and analyses are available in Molinaro et al. (2005) (http://linus.nci.nih.gov/brb/TechReport.htm).
Annette M. Molinaro, Richard Simon, Ruth M. Pfeiffer
Bioinform.2
2004 Iterative class discovery and feature selection using Minimal Spanning Trees
abstract
BACKGROUND: Clustering is one of the most commonly used methods for discovering hidden structure in microarray gene expression data. Most current methods for clustering samples are based on distance metrics utilizing all genes. This has the effect of obscuring clustering in samples that may be evident only when looking at a subset of genes, because noise from irrelevant genes dominates the signal from the relevant genes in the distance calculation. RESULTS: We describe an algorithm for automatically detecting clusters of samples that are discernable only in a subset of genes. We use iteration between Minimal Spanning Tree based clustering and feature selection to remove noise genes in a step-wise manner while simultaneously sharpening the clustering. Evaluation of this algorithm on synthetic data shows that it resolves planted clusters with high accuracy in spite of noise and the presence of other clusters. It also shows a low probability of detecting spurious clusters. Testing the algorithm on some well known micro-array data-sets reveals known biological classes as well as novel clusters. CONCLUSIONS: The iterative clustering method offers considerable improvement over clustering in all genes. This method can be used to discover partitions and their biological significance can be determined by comparing with clinical correlates and gene annotations. The MATLAB programs for the iterative clustering algorithm are available from http://linus.nci.nih.gov/supplement.html
Sudhir Varma, Richard Simon
BMC Bioinform.2
2003 Statistical Design of Reverse Dye Microarrays
abstract
MOTIVATION: In cDNA microarray experiments all samples are labelled with either Cy3 dye or Cy5 dye. Certain genes exhibit dye bias-a tendency to bind more efficiently to one of the dyes. The common reference design avoids the problem of dye bias by running all arrays 'forward', so that the samples being compared are always labelled with the same dye. But comparison of samples labelled with different dyes is sometimes of interest. In these situations, it is necessary to run some arrays 'reverse'-with the dye labelling reversed-in order to correct for the dye bias. The design of these experiments will impact one's ability to identify genes that are differentially expressed in different tissues or conditions. We address the design issue of how many specimens are needed, how many forward and reverse labelled arrays to perform, and how to optimally assign Cy3 and Cy5 labels to the specimens. RESULTS: We consider three types of experiments for which some reverse labelling is needed: paired samples, samples from two predefined groups, and reference design data when comparison with the reference is of interest. We present simple probability models for the data, derive optimal estimators for relative gene expression, and compare the efficiency of the estimators for a range of designs. In each case, we present the optimal design and sample size formulas. We show that reverse labelling of individual arrays is generally not required.
Kevin Dobbin, Joanna H. Shih, Richard Simon
Bioinform.3
2003 A random variance model for detection of differential gene expression in small microarray experiments
abstract
MOTIVATION: Microarray techniques provide a valuable way of characterizing the molecular nature of disease. Unfortunately expense and limited specimen availability often lead to studies with small sample sizes. This makes accurate estimation of variability difficult, since variance estimates made on a gene by gene basis will have few degrees of freedom, and the assumption that all genes share equal variance is unlikely to be true. RESULTS: We propose a model by which the within gene variances are drawn from an inverse gamma distribution, whose parameters are estimated across all genes. This results in a test statistic that is a minor variation of those used in standard linear models. We demonstrate that the model assumptions are valid on experimental data, and that the model has more power than standard tests to pick up large changes in expression, while not increasing the rate of false positives. AVAILABILITY: This method is incorporated into BRB-ArrayTools version 3.0 (http://linus.nci.nih.gov/BRB-ArrayTools.html). SUPPLEMENTARY MATERIAL: ftp://linus.nci.nih.gov/pub/techreport/RVM_supplement.pdf
George W. Wright, Richard Simon
Bioinform.2
2003 Evaluation of normalization methods for microarray data
abstract
BACKGROUND: Microarray technology allows the monitoring of expression levels for thousands of genes simultaneously. This novel technique helps us to understand gene regulation as well as gene by gene interactions more systematically. In the microarray experiment, however, many undesirable systematic variations are observed. Even in replicated experiment, some variations are commonly observed. Normalization is the process of removing some sources of variation which affect the measured gene expression levels. Although a number of normalization methods have been proposed, it has been difficult to decide which methods perform best. Normalization plays an important role in the earlier stage of microarray data analysis. The subsequent analysis results are highly dependent on normalization. RESULTS: In this paper, we use the variability among the replicated slides to compare performance of normalization methods. We also compare normalization methods with regard to bias and mean square error using simulated data. CONCLUSIONS: Our results show that intensity-dependent normalization often performs better than global normalization methods, and that linear and nonlinear normalization methods perform similarly. These conclusions are based on analysis of 36 cDNA microarrays of 3,840 genes obtained in an experiment to search for changes in gene expression profiles during neuronal differentiation of cortical stem cells. Simulation studies confirm our findings.
Taesung Park, Sung-Gon Yi, Sung-Hyun Kang, Seung Yeoun Lee, Yong-Sung Lee, Richard Simon
BMC Bioinform.6
2002 Comparison of microarray designs for class comparison and class discovery
abstract
MOTIVATION: Two-color microarray experiments in which an aliquot derived from a common RNA sample is placed on each array are called reference designs. Traditionally, microarray experiments have used reference designs, but designs without a reference have recently been proposed as alternatives. RESULTS: We develop a statistical model that distinguishes the different levels of variation typically present in cancer data, including biological variation among RNA samples, experimental error and variation attributable to phenotype. Within the context of this model, we examine the reference design and two designs which do not use a reference, the balanced block design and the loop design, focusing particularly on efficiency of estimates and the performance of cluster analysis. We calculate the relative efficiency of designs when there are a fixed number of arrays available, and when there are a fixed number of samples available. Monte Carlo simulation is used to compare the designs when the objective is class discovery based on cluster analysis of the samples. The number of discrepancies between the estimated clusters and the true clusters were significantly smaller for the reference design than for the loop design. The efficiency of the reference design relative to the loop and block designs depends on the relation between inter- and intra-sample variance. These results suggest that if cluster analysis is a major goal of the experiment, then a reference design is preferable. If identification of differentially expressed genes is the main concern, then design selection may involve a consideration of several factors.
Kevin Dobbin, Richard Simon
Bioinform.2
2002 A bioinformatics tool to select sequences for microarray studies of mouse models of oncogenesis
abstract
Abstract One of the challenges to the effective utilization of cDNA microarray analysis in mouse models of oncogenesis is the choice of a critical set of probes that are informative for human disease. Given the thousands of genes with a potential role in human oncogenesis and the hundreds of thousands of mouse sequences available for use as probes, selection of an informative set of mouse probes can be an overwhelming task. We have developed a web based sequence mining tool using DataBase Independent (DBI) Perl to annotate publicly available sequences. The Mouse Oncochip Design Tool uses the Mouse Genome Database (MGD) developed and maintained by the Jackson Laboratories for mouse DNA sequences. There are over 380 000 sequences in their database. The output list has been ordered to present the genes more likely to be informative in a mouse model of human cancer using a candidate set of oncogenes to order the list. Mouse sequences that represent genes that are homologous with a member of a human oncogene set are listed first. In addition it provides a set of links for information on clone source gene function. Contact: http://nciarray.nci.nih.gov/cgi-bin/me/mouse_design.cgi * To whom correspondence should be addressed. 5 Current address: Department of Pathology and Biomedical Informatics, Vanderbilt University Medical Center, MCN C-3321, Nashville TN 37232, USA. 6 Current address: Department of Pharmacology, School of Medicine, University of Colorado Health Sciences Center, 4200 E Ninth Avenue, Denver CO 80262, USA. 7 Current address: Genome Institute of Singapore, 1 Research Link, IMA Building #04-01, National University of Singapore, Singapore 117604.
Mary E. Edgerton, Ronald C. Taylor, John I. Powell, Lawrence Hunter, Richard Simon, Edison T. Liu
Bioinform.5
1971 On the Optimal Solutions to AND/OR Series-Parallel Graphs
abstract
This paper is concerned with efficient ways to find optimal solutions to AND/OR graphs.Although the general methods are still at large, we have found an efficient way to obtain optimal solutions to AND/OR series-parallel graphs.This is achieved by reducing an AND/OR series-parallel graph to an AND/OR tree.Once a graph is reduced to a tree, all the known exact and heuristic methods of tree searching can be applied.
Richard Simon, Richard C. T. Lee
J. ACM1