Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Bertrand S. Clarke

dblp:89/1839 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Theory of computation · 4 · 2 first-authorArtificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Probabilistic and Bayesian machine learning · 52% Learning theory · 46% Kernel, tree and ensemble methods · 2%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Theoretical computer science
4 papers
Coding theory · 57% Information theory · 43%

Topics — the 20 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian model selection
model averaging
0.612022
Model Averaging Is Asymptotically Better Than Model Selection For Prediction · J. Mach. Learn. Res. 2022
Machine learning › Learning theory
model selection
0.422019
Model Selection via the VC Dimension · J. Mach. Learn. Res. 2019
Comparing Bayes Model Averaging and Stacking When Model Approximation Error Cannot be Ignored · J. Mach. Learn. Res. 2003
Machine learning › Learning theory
statistical learning theory
0.412019
Model Selection via the VC Dimension · J. Mach. Learn. Res. 2019
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian model selection › model averaging
bayesian model averaging
0.222022
Model Averaging Is Asymptotically Better Than Model Selection For Prediction · J. Mach. Learn. Res. 2022
Comparing Bayes Model Averaging and Stacking When Model Approximation Error Cannot be Ignored · J. Mach. Learn. Res. 2003
Bioinformatics and computational biology
gene expression analysis
0.112010
Statistical expression deconvolution from mixed tissue samples · Bioinform. 2010
Bioinformatics and computational biology › gene expression analysis › gene expression quantification
gene expression deconvolution
0.112010
Statistical expression deconvolution from mixed tissue samples · Bioinform. 2010
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
approximate bayesian inference
0.112007
Information Conversion, Effective Samples, and Parameter Size · IEEE Trans. Inf. Theory 2007
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.112007
Information Conversion, Effective Samples, and Parameter Size · IEEE Trans. Inf. Theory 2007
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian model selection
0.112007
Information Conversion, Effective Samples, and Parameter Size · IEEE Trans. Inf. Theory 2007
Machine learning › Learning theory › statistical learning theory › statistical complexity
effective number of parameters
0.112007
Information Conversion, Effective Samples, and Parameter Size · IEEE Trans. Inf. Theory 2007
Information theory › information measures › divergence measures
kullback-leibler divergence
0.132007
Asymptotic Normality of the Posterior in Relative Entropy · IEEE Trans. Inf. Theory 1999
Information Conversion, Effective Samples, and Parameter Size · IEEE Trans. Inf. Theory 2007
Information-theoretic asymptotics of Bayes methods · IEEE Trans. Inf. Theory 1990
Machine learning › Kernel, tree and ensemble methods › ensemble learning
stacking
0.012003
Comparing Bayes Model Averaging and Stacking When Model Approximation Error Cannot be Ignored · J. Mach. Learn. Res. 2003
Bioinformatics and computational biology
cancer genomics
0.012010
Statistical expression deconvolution from mixed tissue samples · Bioinform. 2010
Coding theory › source coding
code length
0.021999
Asymptotic Normality of the Posterior in Relative Entropy · IEEE Trans. Inf. Theory 1999
Information-theoretic asymptotics of Bayes methods · IEEE Trans. Inf. Theory 1990
Coding theory
channel coding
0.011999
An Information Criterion for Likelihood Selection · IEEE Trans. Inf. Theory 1999
Coding theory › source coding › rate-distortion theory
rate-distortion function
0.011999
An Information Criterion for Likelihood Selection · IEEE Trans. Inf. Theory 1999
Coding theory › source coding
rate-distortion theory
0.011999
An Information Criterion for Likelihood Selection · IEEE Trans. Inf. Theory 1999
Information theory › algorithmic information theory
stochastic complexity
0.011999
Asymptotic Normality of the Posterior in Relative Entropy · IEEE Trans. Inf. Theory 1999
Information theory › estimation theory
density estimation
0.011990
Information-theoretic asymptotics of Bayes methods · IEEE Trans. Inf. Theory 1990
Coding theory › source coding
universal coding
0.011990
Information-theoretic asymptotics of Bayes methods · IEEE Trans. Inf. Theory 1990

Methods — techniques the papers use, named apart from their topics

stacking · 0.6random forest · 0.6boosting · 0.6bagging · 0.6empirical risk minimization · 0.4relative entropy minimization · 0.1bayesian hierarchical modeling · 0.1statistical deconvolution · 0.1maximum likelihood estimation · 0.1bayesian model averaging · 0.0bayesian asymptotics · 0.0bayes estimation · 0.0fisher information · 0.0bayesian modeling · 0.0
YearPublicationVenuePosition
2022 Model Averaging Is Asymptotically Better Than Model Selection For Prediction
abstract
We compare the performance of six model average predictors---Mallows' model averaging, stacking, Bayes model averaging, bagging, random forests, and boosting---to the components used to form them.In all six cases we identify conditions under which the model average predictor is consistent for its intended limit and performs as well or better than any of its components asymptotically. This is well known empirically, especially for complex problems, although theoretical results do not seem to have been formally established. We have focused our attention on the regression context since that is wheremodel averaging techniques differ most often from current practice.
Tri M. Le, Bertrand S. Clarke
J. Mach. Learn. Res.2
2019 Model Selection via the VC Dimension
abstract
We derive an objective function that can be optimized to give an estimator for the Vapnik-Chervonenkis dimension for use in model selection in regression problems. We verify our estimator is consistent. Then, we verify it performs well compared to seven other model selection techniques. We do this for a variety of types of data sets.
Merlin Mpoudeu, Bertrand S. Clarke
J. Mach. Learn. Res.2
2016 EnsCat: clustering of categorical data via ensembling
abstract
BACKGROUND: Clustering is a widely used collection of unsupervised learning techniques for identifying natural classes within a data set. It is often used in bioinformatics to infer population substructure. Genomic data are often categorical and high dimensional, e.g., long sequences of nucleotides. This makes inference challenging: The distance metric is often not well-defined on categorical data; running time for computations using high dimensional data can be considerable; and the Curse of Dimensionality often impedes the interpretation of the results. Up to the present, however, the literature and software addressing clustering for categorical data has not yet led to a standard approach. RESULTS: We present software for an ensemble method that performs well in comparison with other methods regardless of the dimensionality of the data. In an ensemble method a variety of instantiations of a statistical object are found and then combined into a consensus value. It has been known for decades that ensembling generally outperforms the components that comprise it in many settings. Here, we apply this ensembling principle to clustering. We begin by generating many hierarchical clusterings with different clustering sizes. When the dimension of the data is high, we also randomly select subspaces also of variable size, to generate clusterings. Then, we combine these clusterings into a single membership matrix and use this to obtain a new, ensembled dissimilarity matrix using Hamming distance. CONCLUSIONS: Ensemble clustering, as implemented in R and called EnsCat, gives more clearly separated clusters than other clustering techniques for categorical data. The latest version with manual and examples is available at https://github.com/jlp2duke/EnsCat .
Bertrand S. Clarke, Saeid Amiri, Jennifer Clarke
BMC Bioinform.1
2010 Statistical expression deconvolution from mixed tissue samples
abstract
MOTIVATION: Global expression patterns within cells are used for purposes ranging from the identification of disease biomarkers to basic understanding of cellular processes. Unfortunately, tissue samples used in cancer studies are usually composed of multiple cell types and the non-cancerous portions can significantly affect expression profiles. This severely limits the conclusions that can be made about the specificity of gene expression in the cell-type of interest. However, statistical analysis can be used to identify differentially expressed genes that are related to the biological question being studied. RESULTS: We propose a statistical approach to expression deconvolution from mixed tissue samples in which the proportion of each component cell type is unknown. Our method estimates the proportion of each component in a mixed tissue sample; this estimate can be used to provide estimates of gene expression from each component. We demonstrate our technique on xenograft samples from breast cancer research and publicly available experimental datasets found in the National Center for Biotechnology Information Gene Expression Omnibus repository. AVAILABILITY: R code (http://www.r-project.org/) for estimating sample proportions is freely available to non-commercial users and available at http://www.med.miami.edu/medicine/x2691.xml.
Jennifer Clarke, Pearl Seo, Bertrand S. Clarke
Bioinform.3
2007 Information Conversion, Effective Samples, and Parameter Size
abstract
Consider the relative entropy between a posterior density for a parameter given a sample and a second posterior density for the same parameter, based on a different model and a different data set. Then the relative entropy can be minimized over the second sample to get a virtual sample that would make the second posterior as close as possible to the first in an informational sense. If the first posterior is based on a dependent dataset and the second posterior uses an independence model, the effective inferential power of the dependent sample is transferred into the independent sample by the optimization. Examples of this optimization are presented for models with nuisance parameters, finite mixture models, and models for correlated data. Our approach is also used to choose the effective parameter size in a Bayesian hierarchical model.
Xiaodong Lin 0004, Jennifer Pittman, Bertrand S. Clarke
IEEE Trans. Inf. Theory3
2003 Comparing Bayes Model Averaging and Stacking When Model Approximation Error Cannot be Ignored
Bertrand S. Clarke
J. Mach. Learn. Res.1
1999 Discussion of the Papers by Rissanen, and by Wallace and Dowe
abstract
Commenting on the papers by Dr Rissanen and Professors Wallace and Dowe is a daunting sort of pleasure. Superficially, the papers are not closely related: Dr Rissanen has focused on the Shtarkov optimality criterion in a minimum description length (MDL) context giving some new implications from normalization (see his equation (1)) for model selection and hypothesis testing. By contrast, Wallace and Dowe have given a first, and welcome, effort to axiomatize a setting in which it is reasonable to hope that the algorithmic complexity approaches of Kolmogorov and Solomonoff—streams one and two, using UTMs—may coincide with the Shannon theory approaches—stream three in its minimum message length (MML) and MDL versions. Despite these differences, there are several senses in which these two papers are closely related.
Bertrand S. Clarke
Comput. J.1
1999 Asymptotic Normality of the Posterior in Relative Entropy
abstract
We show that the relative entropy between a posterior density formed from a smooth likelihood and prior and a limiting normal form tends to zero in the independent and identically distributed case. The mode of convergence is in probability and in mean. Applications to code lengths in stochastic complexity and to sample size selection are discussed.
Bertrand S. Clarke
IEEE Trans. Inf. Theory1
1999 An Information Criterion for Likelihood Selection
abstract
For a given source distribution, we establish properties of the conditional density achieving the rate distortion function lower bound as the distortion parameter varies. In the limit as the distortion tolerated goes to zero, the conditional density achieving the rate distortion function lower bound becomes degenerate in the sense that the channel it defines becomes error-free. As the permitted distortion increases to its limit, the conditional density achieving the rate distortion function lower bound defines a channel which no longer depends on the source distribution. In addition to the data compression motivation, we establish two results-one asymptotic, one nonasymptotic-showing that the the conditional densities achieving the rate distortion function lower bound make relatively weak assumptions on the dependence between the source and its representation. This corresponds, in Bayes estimation, to choosing a likelihood which makes relatively weak assumptions on the data generating mechanism if the source is regarded as a prior. Taken together, these results suggest one can use the conditional density obtained from the rate distortion function in data analysis. That is, when it is impossible to identify a "true" parametric family on the basis of physical modeling, our results provide both data compression and channel coding justification for using the conditional density achieving the rate distortion function lower bound as a likelihood.
A. Yuan, Bertrand S. Clarke
IEEE Trans. Inf. Theory2
1990 Information-theoretic asymptotics of Bayes methods
abstract
In the absence of knowledge of the true density function, Bayesian models take the joint density function for a sequence of n random variables to be an average of densities with respect to a prior. The authors examine the relative entropy distance D/sub n/ between the true density and the Bayesian density and show that the asymptotic distance is (d/2)(log n)+c, where d is the dimension of the parameter vector. Therefore, the relative entropy rate D/sub n//n converges to zero at rate (log n)/n. The constant c, which the authors explicitly identify, depends only on the prior density function and the Fisher information matrix evaluated at the true parameter value. Consequences are given for density estimation, universal data compression, composite hypothesis testing, and stock-market portfolio selection.>
Bertrand S. Clarke, Andrew R. Barron
IEEE Trans. Inf. Theory1