Alan Jeffares

dblp:304/1985 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Deep learning architectures and training · 39% Learning theory · 28% Trustworthy machine learning · 13%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory › over-parameterization
double descent
1.422024
Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond · NeurIPS 2024
A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning · NeurIPS 2023
Machine learning › Deep learning architectures and training › training dynamics
grokking
0.812024
Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond · NeurIPS 2024
Machine learning › Deep learning architectures and training › loss landscape › mode connectivity
linear mode connectivity
0.812024
Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond · NeurIPS 2024
Machine learning › Trustworthy machine learning › uncertainty estimation
prediction intervals
0.812024
Relaxed Quantile Regression: Prediction Intervals for Asymmetric Noise · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression
quantile regression
0.812024
Relaxed Quantile Regression: Prediction Intervals for Asymmetric Noise · ICML 2024
Machine learning › Deep learning architectures and training
tabular deep learning
0.812024
Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond · NeurIPS 2024
Machine learning › Deep learning architectures and training
training dynamics
0.812024
Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond · NeurIPS 2024
Machine learning › Trustworthy machine learning
uncertainty estimation
0.812024
Relaxed Quantile Regression: Prediction Intervals for Asymmetric Noise · ICML 2024
Machine learning › Kernel, tree and ensemble methods › ensemble learning
deep ensembles
0.712023
Joint Training of Deep Ensembles Fails Due to Learner Collusion · NeurIPS 2023
Machine learning › Learning theory › statistical learning theory › statistical complexity
effective number of parameters
0.712023
A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning · NeurIPS 2023
Machine learning › Kernel, tree and ensemble methods › ensemble learning
ensemble diversity
0.712023
Joint Training of Deep Ensembles Fails Due to Learner Collusion · NeurIPS 2023
Machine learning › Learning theory
generalization
0.712023
A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning · NeurIPS 2023
Machine learning › Learning theory › generalization error
generalization gap
0.712023
Joint Training of Deep Ensembles Fails Due to Learner Collusion · NeurIPS 2023
Machine learning › Deep learning architectures and training
recurrent neural network
0.612022
Spike-inspired rank coding for fast and accurate recurrent neural networks · ICLR 2022
Machine learning › Deep learning architectures and training › regularization
gradient regularization
0.212023
TANGOS: Regularizing Tabular Neural Networks through Gradient Orthogonalization and Specialization · ICLR 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
non-parametric methods
0.212023
A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning · NeurIPS 2023
Natural language and speech › Language models and text generation › language modeling
smoothing
0.212023
A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning · NeurIPS 2023
Machine learning › Deep learning architectures and training
spiking neural network
0.212022
Spike-inspired rank coding for fast and accurate recurrent neural networks · ICLR 2022

Methods — techniques the papers use, named apart from their topics

relaxed quantile regression · 0.8gradient boosting · 0.8first-order approximation · 0.8regularization · 0.7nonparametric statistics · 0.7joint optimization · 0.7independent training · 0.7gradient orthogonalization · 0.7complexity axis analysis · 0.7rank coding · 0.6
YearPublicationVenuePosition
2025 Partition Tree Ensembles for Improving Multi-Class Classification
Miran Özdogan, Alan Jeffares, Sean B. Holden
ICPRAM2
2024 Relaxed Quantile Regression: Prediction Intervals for Asymmetric Noise
abstract
Constructing valid prediction intervals rather than point estimates is a well-established approach for uncertainty quantification in the regression setting. Models equipped with this capacity output an interval of values in which the ground truth target will fall with some prespecified probability. This is an essential requirement in many real-world applications where simple point predictions’ inability to convey the magnitude and frequency of errors renders them insufficient for high-stakes decisions. Quantile regression is a leading approach for obtaining such intervals via the empirical estimation of quantiles in the (non-parametric) distribution of outputs. This method is simple, computationally inexpensive, interpretable, assumption-free, and effective. However, it does require that the specific quantiles being learned are chosen a priori. This results in (a) intervals that are arbitrarily symmetric around the median which is sub-optimal for realistic skewed distributions, or (b) learning an excessive number of intervals. In this work, we propose Relaxed Quantile Regression (RQR), a direct alternative to quantile regression based interval construction that removes this arbitrary constraint whilst maintaining its strengths. We demonstrate that this added flexibility results in intervals with an improvement in desirable qualities (e.g. mean width) whilst retaining the essential coverage guarantees of quantile regression.
Thomas Pouplin, Alan Jeffares, Nabeel Seedat, Mihaela van der Schaar
ICML2
2024 Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond
abstract
Deep learning sometimes appears to work in unexpected ways. In pursuit of a deeper understanding of its surprising behaviors, we investigate the utility of a simple yet accurate model of a trained neural network consisting of a sequence of first-order approximations telescoping out into a single empirically operational tool for practical analysis. Across three case studies, we illustrate how it can be applied to derive new empirical insights on a diverse range of prominent phenomena in the literature -- including double descent, grokking, linear mode connectivity, and the challenges of applying deep learning on tabular data -- highlighting that this model allows us to construct and extract metrics that help predict and understand the a priori unexpected performance of neural networks. We also demonstrate that this model presents a pedagogical formalism allowing us to isolate components of the training process even in complex contemporary settings, providing a lens to reason about the effects of design choices such as architecture & optimization strategy, and reveals surprising parallels between neural network learning and gradient boosting.
Alan Jeffares, Alicia Curth, Mihaela van der Schaar
NeurIPS1
2023 Improving Adaptive Conformal Prediction Using Self-Supervised Learning
abstract
Conformal prediction is a powerful distribution-free tool for uncertainty quantification, establishing valid prediction intervals with finite-sample guarantees. To produce valid intervals which are also adaptive to the difficulty of each instance, a common approach is to compute normalized nonconformity scores on a separate calibration set. Self-supervised learning has been effectively utilized in many domains to learn general representations for downstream predictors. However, the use of self-supervision beyond model pretraining and representation learning has been largely unexplored. In this work, we investigate how self-supervised pretext tasks can improve the quality of the conformal regressors, specifically by improving the adaptability of conformal intervals. We train an auxiliary model with a self-supervised pretext task on top of an existing predictive model and use the self-supervised error as an additional feature to estimate nonconformity scores. We empirically demonstrate the benefit of the additional information using both synthetic and real data on the efficiency (width), deficit, and excess of conformal prediction intervals.
Nabeel Seedat, Alan Jeffares, Fergus Imrie, Mihaela van der Schaar
AISTATS2
2023 TANGOS: Regularizing Tabular Neural Networks through Gradient Orthogonalization and Specialization
Alan Jeffares, Tennison Liu, Jonathan Crabbé, Fergus Imrie, Mihaela van der Schaar
ICLR1
2023 A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning
abstract
Conventional statistical wisdom established a well-understood relationship between model complexity and prediction error, typically presented as a _U-shaped curve_ reflecting a transition between under- and overfitting regimes. However, motivated by the success of overparametrized neural networks, recent influential work has suggested this theory to be generally incomplete, introducing an additional regime that exhibits a second descent in test error as the parameter count $p$ grows past sample size $n$ -- a phenomenon dubbed _double descent_. While most attention has naturally been given to the deep-learning setting, double descent was shown to emerge more generally across non-neural models: known cases include _linear regression, trees, and boosting_. In this work, we take a closer look at the evidence surrounding these more classical statistical machine learning methods and challenge the claim that observed cases of double descent truly extend the limits of a traditional U-shaped complexity-generalization curve therein. We show that once careful consideration is given to _what is being plotted_ on the x-axes of their double descent plots, it becomes apparent that there are implicitly multiple, distinct complexity axes along which the parameter count grows. We demonstrate that the second descent appears exactly (and _only_) when and where the transition between these underlying axes occurs, and that its location is thus _not_ inherently tied to the interpolation threshold $p=n$. We then gain further insight by adopting a classical nonparametric statistics perspective. We interpret the investigated methods as _smoothers_ and propose a generalized measure for the _effective_ number of parameters they use _on unseen examples_, using which we find that their apparent double descent curves do indeed fold back into more traditional convex shapes -- providing a resolution to the ostensible tension between double descent and traditional statistical intuition.
Alicia Curth, Alan Jeffares, Mihaela van der Schaar
NeurIPS2
2023 Joint Training of Deep Ensembles Fails Due to Learner Collusion
abstract
Ensembles of machine learning models have been well established as a powerful method of improving performance over a single model. Traditionally, ensembling algorithms train their base learners independently or sequentially with the goal of optimizing their joint performance. In the case of deep ensembles of neural networks, we are provided with the opportunity to directly optimize the true objective: the joint performance of the ensemble as a whole. Surprisingly, however, directly minimizing the loss of the ensemble appears to rarely be applied in practice. Instead, most previous research trains individual models independently with ensembling performed _post hoc_. In this work, we show that this is for good reason - _joint optimization of ensemble loss results in degenerate behavior_. We approach this problem by decomposing the ensemble objective into the strength of the base learners and the diversity between them. We discover that joint optimization results in a phenomenon in which base learners collude to artificially inflate their apparent diversity. This pseudo-diversity fails to generalize beyond the training data, causing a larger generalization gap. We proceed to comprehensively demonstrate the practical implications of this effect on a range of standard machine learning tasks and architectures by smoothly interpolating between independent training and joint optimization.
Alan Jeffares, Tennison Liu, Jonathan Crabbé, Mihaela van der Schaar
NeurIPS1
2022 Spike-inspired rank coding for fast and accurate recurrent neural networks
Alan Jeffares, Qinghai Guo, Pontus Stenetorp, Timoleon Moraitis
ICLR1