Ryan Theisen

dblp:251/5575 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2024
0009-0005-7542-5921ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Learning theory · 57% Kernel, tree and ensemble methods · 16% Generative modeling · 12%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
generalization bounds
1.422024
How many classifiers do we need? · NeurIPS 2024
When are ensembles really effective? · NeurIPS 2023
Machine learning › Learning theory
ensemble learning theory
0.812024
How many classifiers do we need? · NeurIPS 2024
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.712023
When are ensembles really effective? · NeurIPS 2023
Machine learning › Kernel, tree and ensemble methods › ensemble learning
majority voting
0.712023
When are ensembles really effective? · NeurIPS 2023
Machine learning › Learning theory
model selection
0.712023
Test Accuracy vs. Generalization Gap: Model Selection in NLP without Accessing Training or Testing Data · KDD 2023
Natural language and speech › Language models and text generation
pre-trained language model
0.712023
Test Accuracy vs. Generalization Gap: Model Selection in NLP without Accessing Training or Testing Data · KDD 2023
Machine learning › Learning theory › classification › classification theory
bayes error rate
0.512021
Evaluating State-of-the-Art Classification Models Against Bayes Optimality · NeurIPS 2021
Machine learning › Learning theory › statistical learning theory › bayesian learning theory
bayes optimality
0.512021
Evaluating State-of-the-Art Classification Models Against Bayes Optimality · NeurIPS 2021
Machine learning › Learning theory
generalization
0.512021
Taxonomizing local versus global structure in neural network loss landscapes · NeurIPS 2021
Machine learning › Generative modeling › normalizing flow
invertible transformation
0.512021
Evaluating State-of-the-Art Classification Models Against Bayes Optimality · NeurIPS 2021
Machine learning › Deep learning architectures and training
loss landscape
0.512021
Taxonomizing local versus global structure in neural network loss landscapes · NeurIPS 2021
Machine learning › Generative modeling
normalizing flow
0.512021
Evaluating State-of-the-Art Classification Models Against Bayes Optimality · NeurIPS 2021
Machine learning › Learning theory › generalization error
generalization gap
0.212023
Test Accuracy vs. Generalization Gap: Model Selection in NLP without Accessing Training or Testing Data · KDD 2023
Machine learning › Transfer learning and domain adaptation › model adaptation
model interpolation
0.212023
When are ensembles really effective? · NeurIPS 2023
Machine learning › Learning theory › statistical learning theory
statistical physics of learning
0.112021
Taxonomizing local versus global structure in neural network loss landscapes · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

neural polarization law · 0.8entropy-constrained error bounds · 0.8power law spectral distribution · 0.7heavy-tail analysis · 0.7disagreement-error ratio analysis · 0.7competence condition · 0.7temperature variation · 0.5normalizing flow · 0.5holmes-diaconis-ross integration · 0.5empirical analysis · 0.5
YearPublicationVenuePosition
2024 How many classifiers do we need?
abstract
As performance gains through scaling data and/or model size experience diminishing returns, it is becoming increasingly popular to turn to ensembling, where the predictions of multiple models are combined to improve accuracy. In this paper, we provide a detailed analysis of how the disagreement and the polarization (a notion we introduce and define in this paper) among classifiers relate to the performance gain achieved by aggregating individual classifiers, for majority vote strategies in classification tasks. We address these questions in the following ways. (1) An upper bound for polarization is derived, and we propose what we call a neural polarization law: most interpolating neural network models are 4/3-polarized. Our empirical results not only support this conjecture but also show that polarization is nearly constant for a dataset, regardless of hyperparameters or architectures of classifiers. (2) The error rate of the majority vote classifier is considered under restricted entropy conditions, and we present a tight upper bound that indicates that the disagreement is linearly correlated with the error rate, and that the slope is linear in the polarization. (3) We prove results for the asymptotic behavior of the disagreement in terms of the number of classifiers, which we show can help in predicting the performance for a larger number of classifiers from that of a smaller number. Our theoretical findings are supported by empirical results on several image classification tasks with various types of neural networks.
Liam Hodgkinson, Ryan Theisen, Michael W. Mahoney
NeurIPS3
2023 Test Accuracy vs. Generalization Gap: Model Selection in NLP without Accessing Training or Testing Data
abstract
Selecting suitable architecture parameters and training hyperparameters is essential for enhancing machine learning (ML) model performance. Several recent empirical studies conduct large-scale correlational analysis on neural networks (NNs) to search for effective generalization metrics that can guide this type of model selection. Effective metrics are typically expected to correlate strongly with test performance. In this paper, we expand on prior analyses by examining generalization-metric-based model selection with the following objectives: (i) focusing on natural language processing (NLP) tasks, as prior work primarily concentrates on computer vision (CV) tasks; (ii) considering metrics that directly predict test error instead of the generalization gap; (iii) exploring metrics that do not need access to data to compute. From these objectives, we are able to provide the first model selection results on large pretrained Transformers from Huggingface using generalization metrics. Our analyses consider (I) hundreds of Transformers trained in different settings, in which we systematically vary the amount of data, the model size and the optimization hyperparameters, (II) a total of 51 pretrained Transformers from eight families of Huggingface NLP models, including GPT2, BERT, etc., and (III) a total of 28 existing and novel generalization metrics. Despite their niche status, we find that metrics derived from the heavy-tail (HT) perspective are particularly useful in NLP tasks, exhibiting stronger correlations than other, more popular metrics. To further examine these metrics, we extend prior formulations relying on power law (PL) spectral distributions to exponential (EXP) and exponentially-truncated power law (E-TPL) families.
Yaoqing Yang 0002, Ryan Theisen, Liam Hodgkinson, Joseph Gonzalez 0001, Kannan Ramchandran, Charles H. Martin, Michael W. Mahoney
KDD2
2023 When are ensembles really effective?
abstract
Ensembling has a long history in statistical data analysis, with many impactful applications. However, in many modern machine learning settings, the benefits of ensembling are less ubiquitous and less obvious. We study, both theoretically and empirically, the fundamental question of when ensembling yields significant performance improvements in classification tasks. Theoretically, we prove new results relating the \emph{ensemble improvement rate} (a measure of how much ensembling decreases the error rate versus a single model, on a relative scale) to the \emph{disagreement-error ratio}. We show that ensembling improves performance significantly whenever the disagreement rate is large relative to the average error rate; and that, conversely, one classifier is often enough whenever the disagreement rate is low relative to the average error rate. On the way to proving these results, we derive, under a mild condition called \emph{competence}, improved upper and lower bounds on the average test error rate of the majority vote classifier. To complement this theory, we study ensembling empirically in a variety of settings, verifying the predictions made by our theory, and identifying practical scenarios where ensembling does and does not result in large performance improvements. Perhaps most notably, we demonstrate a distinct difference in behavior between interpolating models (popular in current practice) and non-interpolating models (such as tree-based methods, where ensembling is popular), demonstrating that ensembling helps considerably more in the latter case than in the former.
Ryan Theisen, Yaoqing Yang 0002, Liam Hodgkinson, Michael W. Mahoney
NeurIPS1
2021 Good Classifiers are Abundant in the Interpolating Regime
abstract
Within the machine learning community, the widely-used uniform convergence framework has been used to answer the question of how complex, over-parameterized models can generalize well to new data. This approach bounds the test error of the \emph{worst-case} model one could have fit to the data, but it has fundamental limitations. Inspired by the statistical mechanics approach to learning, we formally define and develop a methodology to compute precisely the full distribution of test errors among interpolating classifiers from several model classes. We apply our method to compute this distribution for several real and synthetic datasets, with both linear and random feature classification models. We find that test errors tend to concentrate around a small \emph{typical} value $\varepsilon^*$, which deviates substantially from the test error of the worst-case interpolating model on the same datasets, indicating that “bad” classifiers are extremely rare. We provide theoretical results in a simple setting in which we characterize the full asymptotic distribution of test errors, and we show that these indeed concentrate around a value $\varepsilon^*$, which we also identify exactly. We then formalize a more general conjecture supported by our empirical findings. Our results show that the usual style of analysis in statistical learning theory may not be fine-grained enough to capture the good generalization performance observed in practice, and that approaches based on the statistical mechanics of learning may offer a promising alternative.
Ryan Theisen, Jason M. Klusowski, Michael W. Mahoney
AISTATS1
2021 Evaluating State-of-the-Art Classification Models Against Bayes Optimality
abstract
Evaluating the inherent difficulty of a given data-driven classification problem is important for establishing absolute benchmarks and evaluating progress in the field. To this end, a natural quantity to consider is the \emph{Bayes error}, which measures the optimal classification error theoretically achievable for a given data distribution. While generally an intractable quantity, we show that we can compute the exact Bayes error of generative models learned using normalizing flows. Our technique relies on a fundamental result, which states that the Bayes error is invariant under invertible transformation. Therefore, we can compute the exact Bayes error of the learned flow models by computing it for Gaussian base distributions, which can be done efficiently using Holmes-Diaconis-Ross integration. Moreover, we show that by varying the temperature of the learned flow models, we can generate synthetic datasets that closely resemble standard benchmark datasets, but with almost any desired Bayes error. We use our approach to conduct a thorough investigation of state-of-the-art classification models, and find that in some --- but not all --- cases, these models are capable of obtaining accuracy very near optimal. Finally, we use our method to evaluate the intrinsic "hardness" of standard benchmark datasets.
Ryan Theisen, Huan Wang 0016, Lav R. Varshney, Caiming Xiong, Richard Socher
NeurIPS1
2021 Taxonomizing local versus global structure in neural network loss landscapes
abstract
Viewing neural network models in terms of their loss landscapes has a long history in the statistical mechanics approach to learning, and in recent years it has received attention within machine learning proper. Among other things, local metrics (such as the smoothness of the loss landscape) have been shown to correlate with global properties of the model (such as good generalization performance). Here, we perform a detailed empirical analysis of the loss landscape structure of thousands of neural network models, systematically varying learning tasks, model architectures, and/or quantity/quality of data. By considering a range of metrics that attempt to capture different aspects of the loss landscape, we demonstrate that the best test accuracy is obtained when: the loss landscape is globally well-connected; ensembles of trained models are more similar to each other; and models converge to locally smooth regions. We also show that globally poorly-connected landscapes can arise when models are small or when they are trained to lower quality data; and that, if the loss landscape is globally poorly-connected, then training to zero loss can actually lead to worse test accuracy. Our detailed empirical results shed light on phases of learning (and consequent double descent behavior), fundamental versus incidental determinants of good generalization, the role of load-like and temperature-like parameters in the learning process, different influences on the loss landscape from model and data, and the relationships between local and global metrics, all topics of recent interest.
Yaoqing Yang 0002, Liam Hodgkinson, Ryan Theisen, Joe Zou, Joseph Gonzalez 0001, Kannan Ramchandran, Michael W. Mahoney
NeurIPS3