VLDB 2026 Research / reviewers in the wild / expert
Alan Jeffares
dblp:304/1985
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Deep learning architectures and training · 39% Learning theory · 28% Trustworthy machine learning · 13% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory › over-parameterization
double descent |
1.4 | 2 | 2024 | Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond · NeurIPS 2024 A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning · NeurIPS 2023 |
Machine learning › Deep learning architectures and training › training dynamics
grokking |
0.8 | 1 | 2024 | Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › loss landscape › mode connectivity
linear mode connectivity |
0.8 | 1 | 2024 | Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › uncertainty estimation
prediction intervals |
0.8 | 1 | 2024 | Relaxed Quantile Regression: Prediction Intervals for Asymmetric Noise · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression
quantile regression |
0.8 | 1 | 2024 | Relaxed Quantile Regression: Prediction Intervals for Asymmetric Noise · ICML 2024 |
Machine learning › Deep learning architectures and training
tabular deep learning |
0.8 | 1 | 2024 | Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
training dynamics |
0.8 | 1 | 2024 | Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.8 | 1 | 2024 | Relaxed Quantile Regression: Prediction Intervals for Asymmetric Noise · ICML 2024 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning
deep ensembles |
0.7 | 1 | 2023 | Joint Training of Deep Ensembles Fails Due to Learner Collusion · NeurIPS 2023 |
Machine learning › Learning theory › statistical learning theory › statistical complexity
effective number of parameters |
0.7 | 1 | 2023 | A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning · NeurIPS 2023 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning
ensemble diversity |
0.7 | 1 | 2023 | Joint Training of Deep Ensembles Fails Due to Learner Collusion · NeurIPS 2023 |
Machine learning › Learning theory
generalization |
0.7 | 1 | 2023 | A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning · NeurIPS 2023 |
Machine learning › Learning theory › generalization error
generalization gap |
0.7 | 1 | 2023 | Joint Training of Deep Ensembles Fails Due to Learner Collusion · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.6 | 1 | 2022 | Spike-inspired rank coding for fast and accurate recurrent neural networks · ICLR 2022 |
Machine learning › Deep learning architectures and training › regularization
gradient regularization |
0.2 | 1 | 2023 | TANGOS: Regularizing Tabular Neural Networks through Gradient Orthogonalization and Specialization · ICLR 2023 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
non-parametric methods |
0.2 | 1 | 2023 | A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning · NeurIPS 2023 |
Natural language and speech › Language models and text generation › language modeling
smoothing |
0.2 | 1 | 2023 | A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
spiking neural network |
0.2 | 1 | 2022 | Spike-inspired rank coding for fast and accurate recurrent neural networks · ICLR 2022 |
Methods — techniques the papers use, named apart from their topics
relaxed quantile regression · 0.8gradient boosting · 0.8first-order approximation · 0.8regularization · 0.7nonparametric statistics · 0.7joint optimization · 0.7independent training · 0.7gradient orthogonalization · 0.7complexity axis analysis · 0.7rank coding · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Partition Tree Ensembles for Improving Multi-Class Classification
Miran Özdogan, Alan Jeffares, Sean B. Holden |
ICPRAM | 2 |
| 2024 | Relaxed Quantile Regression: Prediction Intervals for Asymmetric NoiseabstractConstructing valid prediction intervals rather than point estimates is a well-established approach for uncertainty quantification in the regression setting. Models equipped with this capacity output an interval of values in which the ground truth target will fall with some prespecified probability. This is an essential requirement in many real-world applications where simple point predictions’ inability to convey the magnitude and frequency of errors renders them insufficient for high-stakes decisions. Quantile regression is a leading approach for obtaining such intervals via the empirical estimation of quantiles in the (non-parametric) distribution of outputs. This method is simple, computationally inexpensive, interpretable, assumption-free, and effective. However, it does require that the specific quantiles being learned are chosen a priori. This results in (a) intervals that are arbitrarily symmetric around the median which is sub-optimal for realistic skewed distributions, or (b) learning an excessive number of intervals. In this work, we propose Relaxed Quantile Regression (RQR), a direct alternative to quantile regression based interval construction that removes this arbitrary constraint whilst maintaining its strengths. We demonstrate that this added flexibility results in intervals with an improvement in desirable qualities (e.g. mean width) whilst retaining the essential coverage guarantees of quantile regression. Thomas Pouplin, Alan Jeffares, Nabeel Seedat, Mihaela van der Schaar |
ICML | 2 |
| 2024 | Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & BeyondabstractDeep learning sometimes appears to work in unexpected ways. In pursuit of a deeper understanding of its surprising behaviors, we investigate the utility of a simple yet accurate model of a trained neural network consisting of a sequence of first-order approximations telescoping out into a single empirically operational tool for practical analysis. Across three case studies, we illustrate how it can be applied to derive new empirical insights on a diverse range of prominent phenomena in the literature -- including double descent, grokking, linear mode connectivity, and the challenges of applying deep learning on tabular data -- highlighting that this model allows us to construct and extract metrics that help predict and understand the a priori unexpected performance of neural networks. We also demonstrate that this model presents a pedagogical formalism allowing us to isolate components of the training process even in complex contemporary settings, providing a lens to reason about the effects of design choices such as architecture & optimization strategy, and reveals surprising parallels between neural network learning and gradient boosting. Alan Jeffares, Alicia Curth, Mihaela van der Schaar |
NeurIPS | 1 |
| 2023 | Improving Adaptive Conformal Prediction Using Self-Supervised LearningabstractConformal prediction is a powerful distribution-free tool for uncertainty quantification, establishing valid prediction intervals with finite-sample guarantees. To produce valid intervals which are also adaptive to the difficulty of each instance, a common approach is to compute normalized nonconformity scores on a separate calibration set. Self-supervised learning has been effectively utilized in many domains to learn general representations for downstream predictors. However, the use of self-supervision beyond model pretraining and representation learning has been largely unexplored. In this work, we investigate how self-supervised pretext tasks can improve the quality of the conformal regressors, specifically by improving the adaptability of conformal intervals. We train an auxiliary model with a self-supervised pretext task on top of an existing predictive model and use the self-supervised error as an additional feature to estimate nonconformity scores. We empirically demonstrate the benefit of the additional information using both synthetic and real data on the efficiency (width), deficit, and excess of conformal prediction intervals. Nabeel Seedat, Alan Jeffares, Fergus Imrie, Mihaela van der Schaar |
AISTATS | 2 |
| 2023 | TANGOS: Regularizing Tabular Neural Networks through Gradient Orthogonalization and Specialization
Alan Jeffares, Tennison Liu, Jonathan Crabbé, Fergus Imrie, Mihaela van der Schaar |
ICLR | 1 |
| 2023 | A U-turn on Double Descent: Rethinking Parameter Counting in Statistical LearningabstractConventional statistical wisdom established a well-understood relationship between model complexity and prediction error, typically presented as a _U-shaped curve_ reflecting a transition between under- and overfitting regimes. However, motivated by the success of overparametrized neural networks, recent influential work has suggested this theory to be generally incomplete, introducing an additional regime that exhibits a second descent in test error as the parameter count $p$ grows past sample size $n$ -- a phenomenon dubbed _double descent_. While most attention has naturally been given to the deep-learning setting, double descent was shown to emerge more generally across non-neural models: known cases include _linear regression, trees, and boosting_. In this work, we take a closer look at the evidence surrounding these more classical statistical machine learning methods and challenge the claim that observed cases of double descent truly extend the limits of a traditional U-shaped complexity-generalization curve therein. We show that once careful consideration is given to _what is being plotted_ on the x-axes of their double descent plots, it becomes apparent that there are implicitly multiple, distinct complexity axes along which the parameter count grows. We demonstrate that the second descent appears exactly (and _only_) when and where the transition between these underlying axes occurs, and that its location is thus _not_ inherently tied to the interpolation threshold $p=n$. We then gain further insight by adopting a classical nonparametric statistics perspective. We interpret the investigated methods as _smoothers_ and propose a generalized measure for the _effective_ number of parameters they use _on unseen examples_, using which we find that their apparent double descent curves do indeed fold back into more traditional convex shapes -- providing a resolution to the ostensible tension between double descent and traditional statistical intuition. Alicia Curth, Alan Jeffares, Mihaela van der Schaar |
NeurIPS | 2 |
| 2023 | Joint Training of Deep Ensembles Fails Due to Learner CollusionabstractEnsembles of machine learning models have been well established as a powerful method of improving performance over a single model. Traditionally, ensembling algorithms train their base learners independently or sequentially with the goal of optimizing their joint performance. In the case of deep ensembles of neural networks, we are provided with the opportunity to directly optimize the true objective: the joint performance of the ensemble as a whole. Surprisingly, however, directly minimizing the loss of the ensemble appears to rarely be applied in practice. Instead, most previous research trains individual models independently with ensembling performed _post hoc_. In this work, we show that this is for good reason - _joint optimization of ensemble loss results in degenerate behavior_. We approach this problem by decomposing the ensemble objective into the strength of the base learners and the diversity between them. We discover that joint optimization results in a phenomenon in which base learners collude to artificially inflate their apparent diversity. This pseudo-diversity fails to generalize beyond the training data, causing a larger generalization gap. We proceed to comprehensively demonstrate the practical implications of this effect on a range of standard machine learning tasks and architectures by smoothly interpolating between independent training and joint optimization. Alan Jeffares, Tennison Liu, Jonathan Crabbé, Mihaela van der Schaar |
NeurIPS | 1 |
| 2022 | Spike-inspired rank coding for fast and accurate recurrent neural networks
Alan Jeffares, Qinghai Guo, Pontus Stenetorp, Timoleon Moraitis |
ICLR | 1 |