EDBT 2026 Demo / reviewers in the wild / expert
Jake Snell
dblp:172/1406 · also Jake C. Snell
· DBLP profile ↗
15ranked-venue papers
7as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
13 papers |
Trustworthy machine learning · 25% Language models and text generation · 14% Transfer learning and domain adaptation · 12% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 30 heaviest of 39, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
uncertainty estimation |
1.5 | 2 | 2025 | Conformal Prediction as Bayesian Quadrature · ICML 2025 Quantile Risk Control: A Flexible Framework for Bounding the Probability of High-Loss Predictions · ICLR 2023 |
Machine learning › Trustworthy machine learning
risk control |
1.4 | 2 | 2024 | Prompt Risk Control: A Rigorous Framework for Responsible Deployment of Large Language Models · ICLR 2024 Quantile Risk Control: A Flexible Framework for Bounding the Probability of High-Loss Predictions · ICLR 2023 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
1.0 | 2 | 2023 | Im-Promptu: In-Context Composition from Image Prompts · NeurIPS 2023 Meta-Learning for Semi-Supervised Few-Shot Classification · ICLR (Poster) 2018 |
Machine learning › Trustworthy machine learning › uncertainty estimation › probabilistic numerics
bayesian quadrature |
0.9 | 1 | 2025 | Conformal Prediction as Bayesian Quadrature · ICML 2025 |
Machine learning › Trustworthy machine learning › uncertainty estimation
conformal prediction |
0.9 | 1 | 2025 | Conformal Prediction as Bayesian Quadrature · ICML 2025 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian filtering |
0.8 | 1 | 2024 | Implicit Maximum a Posteriori Filtering via Adaptive Optimization · ICLR 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model |
0.8 | 1 | 2024 | A Metalearned Neural Circuit for Nonparametric Bayesian Inference · NeurIPS 2024 |
Machine learning › Time series and sequential data › time series analysis › bayesian filtering and smoothing
high-dimensional state estimation |
0.8 | 1 | 2024 | Implicit Maximum a Posteriori Filtering via Adaptive Optimization · ICLR 2024 |
Natural language and speech › Language models and text generation › prompting
prompt engineering |
0.8 | 1 | 2024 | Prompt Risk Control: A Rigorous Framework for Responsible Deployment of Large Language Models · ICLR 2024 |
Natural language and speech › Language models and text generation › prompting › prompt engineering
prompt selection |
0.8 | 1 | 2024 | Prompt Risk Control: A Rigorous Framework for Responsible Deployment of Large Language Models · ICLR 2024 |
Machine learning › Deep learning architectures and training
sequence modeling |
0.8 | 1 | 2024 | A Metalearned Neural Circuit for Nonparametric Bayesian Inference · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning › representation learning
metric learning |
0.7 | 2 | 2019 | Lorentzian Distance Learning for Hyperbolic Representations · ICML 2019 Prototypical Networks for Few-shot Learning · NIPS 2017 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
analogical reasoning |
0.7 | 1 | 2023 | Im-Promptu: In-Context Composition from Image Prompts · NeurIPS 2023 |
Natural language and speech › Language models and text generation
compositional generalization |
0.7 | 1 | 2023 | Im-Promptu: In-Context Composition from Image Prompts · NeurIPS 2023 |
Machine learning › Learning theory › generalization bounds
distribution-free bounds |
0.7 | 1 | 2023 | Quantile Risk Control: A Flexible Framework for Bounding the Probability of High-Loss Predictions · ICLR 2023 |
Machine learning › Trustworthy machine learning
fairness |
0.7 | 1 | 2023 | Distribution-Free Statistical Dispersion Control for Societal Applications · NeurIPS 2023 |
Computer vision › Image recognition and object detection
image composition |
0.7 | 1 | 2023 | Im-Promptu: In-Context Composition from Image Prompts · NeurIPS 2023 |
Machine learning › Generative modeling
image generation |
0.7 | 1 | 2023 | Im-Promptu: In-Context Composition from Image Prompts · NeurIPS 2023 |
Natural language and speech › Language models and text generation
in-context learning |
0.7 | 1 | 2023 | Im-Promptu: In-Context Composition from Image Prompts · NeurIPS 2023 |
Machine learning › Learning theory
statistical learning theory |
0.7 | 1 | 2023 | Quantile Risk Control: A Flexible Framework for Bounding the Probability of High-Loss Predictions · ICLR 2023 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.6 | 2 | 2018 | Meta-Learning for Semi-Supervised Few-Shot Classification · ICLR (Poster) 2018 Prototypical Networks for Few-shot Learning · NIPS 2017 |
Machine learning › Transfer learning and domain adaptation
few-shot classification |
0.5 | 1 | 2021 | Bayesian Few-Shot Classification with One-vs-Each Pólya-Gamma Augmented Gaussian Processes · ICLR 2021 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.5 | 1 | 2021 | Bayesian Few-Shot Classification with One-vs-Each Pólya-Gamma Augmented Gaussian Processes · ICLR 2021 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › geometric representation learning
hyperbolic representation learning |
0.4 | 1 | 2019 | Lorentzian Distance Learning for Hyperbolic Representations · ICML 2019 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.3 | 1 | 2018 | Learning Latent Subspaces in Variational Autoencoders · NeurIPS 2018 |
Machine learning › Transfer learning and domain adaptation › meta-learning
few-shot meta-learning |
0.3 | 1 | 2018 | Meta-Learning for Semi-Supervised Few-Shot Classification · ICLR (Poster) 2018 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › subspace learning
latent subspace learning |
0.3 | 1 | 2018 | Learning Latent Subspaces in Variational Autoencoders · NeurIPS 2018 |
Machine learning › Generative modeling
variational autoencoder |
0.3 | 1 | 2018 | Learning Latent Subspaces in Variational Autoencoders · NeurIPS 2018 |
Machine learning › Representation and self-supervised learning › prototype learning
prototype-based representation |
0.3 | 1 | 2017 | Prototypical Networks for Few-shot Learning · NIPS 2017 |
Machine learning › Learning theory › classification › prototype-based classification
prototypical network |
0.3 | 1 | 2017 | Prototypical Networks for Few-shot Learning · NIPS 2017 |
Methods — techniques the papers use, named apart from their topics
meta-learning · 2.0monte carlo estimation · 1.5kalman filter · 1.5gradient descent · 1.5statistical bounding · 0.8sequence model · 0.8particle filter · 0.8distribution shift · 0.8CVaR · 0.8conformal prediction · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Conformal Prediction as Bayesian QuadratureabstractAs machine learning-based prediction systems are increasingly used in high-stakes situations, it is important to understand how such predictive models will perform upon deployment. Distribution-free uncertainty quantification techniques such as conformal prediction provide guarantees about the loss black-box models will incur even when the details of the models are hidden. However, such methods are based on frequentist probability, which unduly limits their applicability. We revisit the central aspects of conformal prediction from a Bayesian perspective and thereby illuminate the shortcomings of frequentist guarantees. We propose a practical alternative based on Bayesian quadrature that provides interpretable guarantees and offers a richer representation of the likely range of losses to be observed at test time. Jake Snell, Thomas L. Griffiths 0001 |
ICML | 1 |
| 2024 | Implicit Maximum a Posteriori Filtering via Adaptive OptimizationabstractBayesian filtering approximates the true underlying behavior of a time-varying system by inverting an explicit generative model to convert noisy measurements into state estimates. This process typically requires matrix storage, inversion, and multiplication or Monte Carlo estimation, none of which are practical in high-dimensional state spaces such as the weight spaces of artificial neural networks. Here, we consider the standard Bayesian filtering problem as optimization over a time-varying objective. Instead of maintaining matrices for the filtering equations or simulating particles, we specify an optimizer that defines the Bayesian filter implicitly. In the linear-Gaussian setting, we show that every Kalman filter has an equivalent formulation using K steps of gradient descent. In the nonlinear setting, our experiments demonstrate that our framework results in filters that are effective, robust, and scalable to high-dimensional systems, comparing well against the standard toolbox of Bayesian filtering solutions. We suggest that it is easier to fine-tune an optimizer than it is to specify the correct filtering equations, making our framework an attractive option for high-dimensional filtering problems. Gianluca M. Bencomo, Jake Snell, Thomas L. Griffiths 0001 |
ICLR | 2 |
| 2024 | Prompt Risk Control: A Rigorous Framework for Responsible Deployment of Large Language ModelsabstractWith the explosion of the zero-shot capabilities of (and thus interest in) pre-trained large language models, there has come accompanying interest in how best to prompt a language model to perform a given task. While it may be tempting to choose a prompt based on empirical results on a validation set, this can lead to a deployment where an unexpectedly high loss occurs. To mitigate this prospect, we propose a lightweight framework, Prompt Risk Control, for selecting a prompt based on rigorous upper bounds on families of informative risk measures. We provide and compare different methods for producing bounds on a diverse set of risk metrics like mean, CVaR, and the Gini coefficient of the loss distribution. In addition, we extend the underlying statistical bounding techniques to accommodate the possibility of distribution shifts in deployment. Extensive experiments on high-impact applications like chatbots, medical question answering, and news summarization highlight why such a framework is necessary to reduce exposure to the worst outcomes. Thomas P. Zollo, Todd Morrill, Zhun Deng, Jake Snell, Toniann Pitassi, Richard S. Zemel |
ICLR | 4 |
| 2024 | A Metalearned Neural Circuit for Nonparametric Bayesian InferenceabstractMost applications of machine learning to classification assume a closed set of balanced classes. This is at odds with the real world, where class occurrence statistics often follow a long-tailed power-law distribution and it is unlikely that all classes are seen in a single sample. Nonparametric Bayesian models naturally capture this phenomenon, but have significant practical barriers to widespread adoption, namely implementation complexity and computational inefficiency. To address this, we present a method for extracting the inductive bias from a nonparametric Bayesian model and transferring it to an artificial neural network. By simulating data with a nonparametric Bayesian prior, we can metalearn a sequence model that performs inference over an unlimited set of classes. After training, this "neural circuit" has distilled the corresponding inductive bias and can successfully perform sequential inference over an open set of classes. Our experimental results show that the metalearned neural circuit achieves comparable or better performance than particle filter-based methods for inference in these models while being faster and simpler to use than methods that explicitly incorporate Bayesian nonparametric inference. Jake Snell, Gianluca M. Bencomo, Thomas L. Griffiths 0001 |
NeurIPS | 1 |
| 2023 | Quantile Risk Control: A Flexible Framework for Bounding the Probability of High-Loss Predictions
Jake Snell, Thomas P. Zollo, Zhun Deng, Toniann Pitassi, Richard S. Zemel |
ICLR | 1 |
| 2023 | Im-Promptu: In-Context Composition from Image PromptsabstractLarge language models are few-shot learners that can solve diverse tasks from a handful of demonstrations. This implicit understanding of tasks suggests that the attention mechanisms over word tokens may play a role in analogical reasoning. In this work, we investigate whether analogical reasoning can enable in-context composition over composable elements of visual stimuli. First, we introduce a suite of three benchmarks to test the generalization properties of a visual in-context learner. We formalize the notion of an analogy-based in-context learner and use it to design a meta-learning framework called Im-Promptu. Whereas the requisite token granularity for language is well established, the appropriate compositional granularity for enabling in-context generalization in visual stimuli is usually unspecified. To this end, we use Im-Promptu to train multiple agents with different levels of compositionality, including vector representations, patch representations, and object slots. Our experiments reveal tradeoffs between extrapolation abilities and the degree of compositionality, with non-compositional representations extending learned composition rules to unseen domains but performing poorly on combinatorial tasks. Patch-based representations require patches to contain entire objects for robust extrapolation. At the same time, object-centric tokenizers coupled with a cross-attention module generate consistent and high-fidelity solutions, with these inductive biases being particularly crucial for compositional generalization. Lastly, we demonstrate a use case of Im-Promptu as an intuitive programming interface for image generation. Bhishma Dedhia, Michael Chang 0003, Jake Snell, Thomas L. Griffiths 0001, Niraj K. Jha |
NeurIPS | 3 |
| 2023 | Distribution-Free Statistical Dispersion Control for Societal ApplicationsabstractExplicit finite-sample statistical guarantees on model performance are an important ingredient in responsible machine learning. Previous work has focused mainly on bounding either the expected loss of a predictor or the probability that an individual prediction will incur a loss value in a specified range. However, for many high-stakes applications it is crucial to understand and control the \textit{dispersion} of a loss distribution, or the extent to which different members of a population experience unequal effects of algorithmic decisions. We initiate the study of distribution-free control of statistical dispersion measures with societal implications and propose a simple yet flexible framework that allows us to handle a much richer class of statistical functionals beyond previous work. Our methods are verified through experiments in toxic comment detection, medical imaging, and film recommendation. Zhun Deng, Thomas P. Zollo, Jake Snell, Toniann Pitassi, Richard S. Zemel |
NeurIPS | 3 |
| 2021 | Bayesian Few-Shot Classification with One-vs-Each Pólya-Gamma Augmented Gaussian Processes
Jake Snell, Richard S. Zemel |
ICLR | 1 |
| 2019 | Dimensionality Reduction for Representing the Knowledge of Probabilistic Models
Marc T. Law, Jake Snell, Amir-massoud Farahmand, Raquel Urtasun, Richard S. Zemel |
ICLR (Poster) | 2 |
| 2019 | Lorentzian Distance Learning for Hyperbolic RepresentationsabstractWe introduce an approach to learn representations based on the Lorentzian distance in hyperbolic geometry. Hyperbolic geometry is especially suited to hierarchically-structured datasets, which are prevalent in the real world. Current hyperbolic representation learning methods compare examples with the Poincaré distance. They try to minimize the distance of each node in a hierarchy with its descendants while maximizing its distance with other nodes. This formulation produces node representations close to the centroid of their descendants. To obtain efficient and interpretable algorithms, we exploit the fact that the centroid w.r.t the squared Lorentzian distance can be written in closed-form. We show that the Euclidean norm of such a centroid decreases as the curvature of the hyperbolic space decreases. This property makes it appropriate to represent hierarchies where parent nodes minimize the distances to their descendants and have smaller Euclidean norm than their children. Our approach obtains state-of-the-art results in retrieval and classification tasks on different datasets. Marc T. Law, Renjie Liao 0001, Jake Snell, Richard S. Zemel |
ICML | 3 |
| 2018 | Meta-Learning for Semi-Supervised Few-Shot Classification
Mengye Ren, Eleni Triantafillou, Sachin Ravi, Jake Snell, Kevin Swersky, Josh Tenenbaum, Hugo Larochelle, Richard S. Zemel |
ICLR (Poster) | 4 |
| 2018 | Learning Latent Subspaces in Variational AutoencodersabstractVariational autoencoders (VAEs) are widely used deep generative models capable of learning unsupervised latent representations of data. Such representations are often difficult to interpret or control. We consider the problem of unsupervised learning of features correlated to specific labels in a dataset. We propose a VAE-based generative model which we show is capable of extracting features correlated to binary labels in the data and structuring it in a latent subspace which is easy to interpret. Our model, the Conditional Subspace VAE (CSVAE), uses mutual information minimization to learn a low-dimensional latent subspace associated with each label that can easily be inspected and independently manipulated. We demonstrate the utility of the learned representations for attribute manipulation tasks on both the Toronto Face and CelebA datasets. Jack Klys, Jake Snell, Richard S. Zemel |
NeurIPS | 2 |
| 2017 | Learning to generate images with perceptual similarity metricsabstractDeep networks are increasingly being applied to problems involving image synthesis, e.g., generating images from textual descriptions and reconstructing an input image from a compact representation. Supervised training of image-synthesis networks typically uses a pixel-wise loss (PL) to indicate the mismatch between a generated image and its corresponding target image. We propose instead to use a loss function that is better calibrated to human perceptual judgments of image quality: the multiscale structural-similarity score (MS-SSIM) [1]. Because MS-SSIM is differentiable, it is easily incorporated into gradient-descent learning. We compare the consequences of using MS-SSIM versus PL loss on training autoencoders. Human observers reliably prefer images synthesized by MS-SSIM-optimized models over those synthesized by PL-optimized models, for two distinct PL measures (L1and L2distances). We also explore the effect of training objective on image encoding and analyze conditions under which perceptually-optimized representations yield better performance on image classification. Finally, we demonstrate the superiority of perceptually-optimized networks for super-resolution imaging. We argue that significant advances can be made in modeling images through the use of training objectives that are well aligned to characteristics of human perception. Jake Snell, Karl Ridgeway, Renjie Liao 0001, Brett Roads, Michael C. Mozer, Richard S. Zemel |
ICIP | 1 |
| 2017 | Prototypical Networks for Few-shot LearningabstractWe propose Prototypical Networks for the problem of few-shot classification, where a classifier must generalize to new classes not seen in the training set, given only a small number of examples of each new class. Prototypical Networks learn a metric space in which classification can be performed by computing distances to prototype representations of each class. Compared to recent approaches for few-shot learning, they reflect a simpler inductive bias that is beneficial in this limited-data regime, and achieve excellent results. We provide an analysis showing that some simple design decisions can yield substantial improvements over recent approaches involving complicated architectural choices and meta-learning. We further extend Prototypical Networks to zero-shot learning and achieve state-of-the-art results on the CU-Birds dataset. Jake Snell, Kevin Swersky, Richard S. Zemel |
NIPS | 1 |
| 2017 | Stochastic Segmentation Trees for Multiple Ground Truths
Jake Snell, Richard S. Zemel |
UAI | 1 |