Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ghassen Jerfel

dblp:185/0768 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
4since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Trustworthy machine learning · 59% Kernel, tree and ensemble methods · 13% Transfer learning and domain adaptation · 10%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
uncertainty estimation
1.122023
A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness · J. Mach. Learn. Res. 2023
Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors · ICML 2020
Machine learning › Trustworthy machine learning › uncertainty estimation › predictive uncertainty
distance-aware uncertainty
0.712023
A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness · J. Mach. Learn. Res. 2023
Machine learning › Trustworthy machine learning
robustness
0.612022
Underspecification Presents Challenges for Credibility in Modern Machine Learning · J. Mach. Learn. Res. 2022
Machine learning › Trustworthy machine learning › robustness
underspecification
0.612022
Underspecification Presents Challenges for Credibility in Modern Machine Learning · J. Mach. Learn. Res. 2022
Machine learning › Trustworthy machine learning
calibration
0.512021
Combining Ensembles and Data Augmentation Can Harm Your Calibration · ICLR 2021
Machine learning › Kernel, tree and ensemble methods › ensemble learning
ensemble calibration
0.512021
Combining Ensembles and Data Augmentation Can Harm Your Calibration · ICLR 2021
Machine learning › Kernel, tree and ensemble methods
model ensemble
0.512021
Combining Ensembles and Data Augmentation Can Harm Your Calibration · ICLR 2021
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks
0.412020
Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors · ICML 2020
Machine learning › Learning paradigms
continual learning
0.412019
Reconciling meta-learning and continual learning with online mixtures of tasks · NeurIPS 2019
Machine learning › Transfer learning and domain adaptation › meta-learning
gradient-based meta-learning
0.412019
Reconciling meta-learning and continual learning with online mixtures of tasks · NeurIPS 2019
Machine learning › Transfer learning and domain adaptation
meta-learning
0.412019
Reconciling meta-learning and continual learning with online mixtures of tasks · NeurIPS 2019
Machine learning › Trustworthy machine learning › calibration
neural network calibration
0.212023
A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness · J. Mach. Learn. Res. 2023
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.212023
A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness · J. Mach. Learn. Res. 2023
Machine learning › Deep learning architectures and training
data augmentation
0.112021
Combining Ensembles and Data Augmentation Can Harm Your Calibration · ICLR 2021
Machine learning › Efficient and distributed learning
parameter-efficient model
0.112020
Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors · ICML 2020
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
dirichlet process mixture model
0.112019
Reconciling meta-learning and continual learning with online mixtures of tasks · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › hierarchical modeling
hierarchical bayesian model
0.112019
Reconciling meta-learning and continual learning with online mixtures of tasks · NeurIPS 2019

Methods — techniques the papers use, named apart from their topics

spectral normalization · 0.7minimax learning · 0.7gaussian process · 0.7ensembling · 0.5data augmentation · 0.5variational inference · 0.4mixture posterior · 0.4deep ensembles · 0.4hierarchical bayes · 0.4dirichlet process mixture · 0.4
YearPublicationVenuePosition
2023 A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness
abstract
Accurate uncertainty quantification is a major challenge in deep learning, as neural networks can make overconfident errors and assign high confidence predictions to out-of-distribution (OOD) inputs. The most popular approaches to estimate predictive uncertainty in deep learning are methods that combine predictions from multiple neural networks, such as Bayesian neural networks (BNNs) and deep ensembles. However their practicality in real-time, industrial-scale applications are limited due to the high memory and computational cost. Furthermore, ensembles and BNNs do not necessarily fix all the issues with the underlying member networks. In this work, we study principled approaches to improve the uncertainty property of a single network, based on a single, deterministic representation. By formalizing the uncertainty quantification as a minimax learning problem, we first identify distance awareness, i.e., the model's ability to quantify the distance of a testing example from the training data, as a necessary condition for a DNN to achieve high-quality (i.e., minimax optimal) uncertainty estimation. We then propose Spectral-normalized Neural Gaussian Process (SNGP), a simple method that improves the distance-awareness ability of modern DNNs with two simple changes: (1) applying spectral normalization to hidden weights to enforce bi-Lipschitz smoothness in representations and (2) replacing the last output layer with a Gaussian process layer. On a suite of vision and language understanding benchmarks and on modern architectures (Wide-ResNet and BERT), SNGP consistently outperforms other single-model approaches in prediction, calibration and out-of-domain detection. Furthermore, SNGP provides complementary benefits to popular techniques such as deep ensembles and data augmentation, making it a simple and scalable building block for probabilistic deep learning.
Jeremiah Z. Liu, Shreyas Padhy, Jie Ren 0006, Zi Lin, Yeming Wen, Ghassen Jerfel, Zachary Nado, Jasper Snoek, Dustin Tran, Balaji Lakshminarayanan
J. Mach. Learn. Res.6
2022 Underspecification Presents Challenges for Credibility in Modern Machine Learning
abstract
Machine learning (ML) systems often exhibit unexpectedly poor behavior when they are deployed in real-world domains. We identify underspecification in ML pipelines as a key reason for these failures. An ML pipeline is the full procedure followed to train and validate a predictor. Such a pipeline is underspecified when it can return many distinct predictors with equivalently strong test performance. Underspecification is common in modern ML pipelines that primarily validate predictors on held-out data that follow the same distribution as the training data. Predictors returned by underspecified pipelines are often treated as equivalent based on their training domain performance, but we show here that such predictors can behave very differently in deployment domains. This ambiguity can lead to instability and poor model behavior in practice, and is a distinct failure mode from previously identified issues arising from structural mismatch between training and deployment domains. We provide evidence that underspecfication has substantive implications for practical ML pipelines, using examples from computer vision, medical imaging, natural language processing, clinical risk prediction based on electronic health records, and medical genomics. Our results show the need to explicitly account for underspecification in modeling pipelines that are intended for real-world deployment in any domain.
Alexander D'Amour, Katherine A. Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew Hoffman 0001, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yi-An Ma, Cory Y. McLean, Diana Mincu, Akinori Mitani, Andrea Montanari, Zachary Nado, Vivek Natarajan, Christopher Nielson, Thomas F. Osborne, Rajiv Raman 0003, Kim Ramasamy, Rory Sayres, Jessica Schrouff, Martin G. Seneviratne, Shannon Sequeira, Harini Suresh, Victor Veitch, Max Vladymyrov, Xuezhi Wang 0002, Kellie Webster, Steve Yadlowsky, Taedong Yun, Xiaohua Zhai, D. Sculley
J. Mach. Learn. Res.14
2021 Combining Ensembles and Data Augmentation Can Harm Your Calibration
Yeming Wen, Ghassen Jerfel, Rafael Muller, Michael Dusenberry, Jasper Snoek, Balaji Lakshminarayanan, Dustin Tran
ICLR2
2021 Variational refinement for importance sampling using the forward Kullback-Leibler divergence
abstract
Variational Inference (VI) is a popular alternative to asymptotically exact sampling in Bayesian inference. Its main workhorse is optimization over a reverse Kullback-Leibler divergence (RKL), which typically underestimates the tail of the posterior leading to miscalibration and potential degeneracy. Importance sampling (IS), on the other hand, is often used to fine-tune and de-bias the estimates of approximate Bayesian inference procedures. The quality of IS crucially depends on the choice of the proposal distribution. Ideally, the proposal distribution has heavier tails than the target, which is rarely achievable by minimizing the RKL. We thus propose a novel combination of optimization and sampling techniques for approximate Bayesian inference by constructing an IS proposal distribution through the minimization of a forward KL (FKL) divergence. This approach guarantees asymptotic consistency and a fast convergence towards both the optimal IS estimator and the optimal variational approximation. We empirically demonstrate on real data that our method is competitive with variational boosting and MCMC.
Ghassen Jerfel, Serena Lutong Wang, Clara Fannjiang, Katherine A. Heller, Yi-An Ma, Michael I. Jordan
UAI1
2020 Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors
abstract
Bayesian neural networks (BNNs) demonstrate promising success in improving the robustness and uncertainty quantification of modern deep learning. However, they generally struggle with underfitting at scale and parameter efficiency. On the other hand, deep ensembles have emerged as alternatives for uncertainty quantification that, while outperforming BNNs on certain problems, also suffer from efficiency issues. It remains unclear how to combine the strengths of these two approaches and remediate their common issues. To tackle this challenge, we propose a rank-1 parameterization of BNNs, where each weight matrix involves only a distribution on a rank-1 subspace. We also revisit the use of mixture approximate posteriors to capture multiple modes, where unlike typical mixtures, this approach admits a significantly smaller memory increase (e.g., only a 0.4% increase for a ResNet-50 mixture of size 10). We perform a systematic empirical study on the choices of prior, variational posterior, and methods to improve training. For ResNet-50 on ImageNet, Wide ResNet 28-10 on CIFAR-10/100, and an RNN on MIMIC-III, rank-1 BNNs achieve state-of-the-art performance across log-likelihood, accuracy, and calibration on the test sets and out-of-distribution variants.
Michael Dusenberry, Ghassen Jerfel, Yeming Wen, Yi-An Ma, Jasper Snoek, Katherine A. Heller, Balaji Lakshminarayanan, Dustin Tran
ICML2
2019 Reconciling meta-learning and continual learning with online mixtures of tasks
abstract
Learning-to-learn or meta-learning leverages data-driven inductive bias to increase the efficiency of learning on a novel task. This approach encounters difficulty when transfer is not advantageous, for instance, when tasks are considerably dissimilar or change over time. We use the connection between gradient-based meta-learning and hierarchical Bayes to propose a Dirichlet process mixture of hierarchical Bayesian models over the parameters of an arbitrary parametric model such as a neural network. In contrast to consolidating inductive biases into a single set of hyperparameters, our approach of task-dependent hyperparameter selection better handles latent distribution shift, as demonstrated on a set of evolving, image-based, few-shot learning benchmarks.
Ghassen Jerfel, Erin Grant, Thomas L. Griffiths 0001, Katherine A. Heller
NeurIPS1
2017 Dynamic Collaborative Filtering With Compound Poisson Factorization
abstract
Model-based collaborative filtering (CF) analyzes user–item interactions to infer latent factors that represent user preferences and item characteristics in order to predict future interactions. Most CF approaches assume that these latent factors are static; however, in most CF data, user preferences and item perceptions drift over time. Here, we propose a new conjugate and numerically stable dynamic matrix factorization (DCPF) based on hierarchical Poisson factorization that models the smoothly drifting latent factors using gamma-Markov chains. We propose a conjugate gamma chain construction that is numerically stable within our compound-Poisson framework. We then derive a scalable stochastic variational inference approach to estimate the parameters of our model. We apply our model to time-stamped ratings data sets from Netflix, Yelp, and Last.fm. We empirically demonstrate that DCPF achieves a higher predictive accuracy than state-of-the-art static and dynamic factorization algorithms.
Ghassen Jerfel, Mehmet Emin Basbug, Barbara E. Engelhardt
AISTATS1