VLDB 2026 Research / reviewers in the wild / expert
Dustin Tran
dblp:163/2106
· DBLP profile ↗
31ranked-venue papers
8as first author
8since 2021 · last 2024
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 8 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
27 papers |
Trustworthy machine learning · 26% Probabilistic and Bayesian machine learning · 19% Efficient and distributed learning · 12% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 30 heaviest of 69, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
uncertainty estimation |
2.5 | 5 | 2023 | A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness · J. Mach. Learn. Res. 2023 Revisiting the Calibration of Modern Neural Networks · NeurIPS 2021 Hyperparameter Ensembles for Robustness and Uncertainty Quantification · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
1.6 | 6 | 2017 | Automatic Differentiation Variational Inference · J. Mach. Learn. Res. 2017 Hierarchical Implicit Models and Likelihood-Free Variational Inference · NIPS 2017 Variational Inference via \chi Upper Bound Minimization · NIPS 2017 |
Machine learning › Probabilistic and Bayesian machine learning
probabilistic programming |
1.3 | 4 | 2019 | Bayesian Layers: A Module for Neural Network Uncertainty · NeurIPS 2019 Simple, Distributed, and Accelerated Probabilistic Programming · NeurIPS 2018 Automatic Differentiation Variational Inference · J. Mach. Learn. Res. 2017 |
Machine learning › Trustworthy machine learning
calibration |
1.0 | 2 | 2021 | Soft Calibration Objectives for Neural Networks · NeurIPS 2021 Combining Ensembles and Data Augmentation Can Harm Your Calibration · ICLR 2021 |
Machine learning › Kernel, tree and ensemble methods
model ensemble |
0.9 | 2 | 2021 | Combining Ensembles and Data Augmentation Can Harm Your Calibration · ICLR 2021 BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong Learning · ICLR 2020 |
Natural language and speech › Language models and text generation › large language model evaluation
automatic evaluation |
0.8 | 1 | 2024 | Long-form factuality in large language models · NeurIPS 2024 |
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability
factuality |
0.8 | 1 | 2024 | Long-form factuality in large language models · NeurIPS 2024 |
Machine learning › Generative modeling
autoregressive model |
0.7 | 2 | 2019 | Discrete Flows: Invertible Generative Models of Discrete Data · NeurIPS 2019 Image Transformer · ICML 2018 |
Machine learning › Trustworthy machine learning › calibration
neural network calibration |
0.7 | 2 | 2023 | Revisiting the Calibration of Modern Neural Networks · NeurIPS 2021 A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness · J. Mach. Learn. Res. 2023 |
Machine learning › Trustworthy machine learning › uncertainty estimation › predictive uncertainty
distance-aware uncertainty |
0.7 | 1 | 2023 | A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness · J. Mach. Learn. Res. 2023 |
Machine learning › Efficient and distributed learning › large-scale learning
large-scale model training |
0.7 | 1 | 2023 | Scaling Vision Transformers to 22 Billion Parameters · ICML 2023 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.7 | 1 | 2023 | Scaling Vision Transformers to 22 Billion Parameters · ICML 2023 |
Machine learning › Deep learning architectures and training › transformer › vision transformer
vision transformer scaling |
0.7 | 1 | 2023 | Scaling Vision Transformers to 22 Billion Parameters · ICML 2023 |
Machine learning › Transfer learning and domain adaptation › zero-shot learning
zero-shot classification |
0.7 | 1 | 2023 | A Simple Zero-shot Prompt Weighting Technique to Improve Prompt Ensembling in Text-Image Models · ICML 2023 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning
ensemble calibration |
0.5 | 1 | 2021 | Combining Ensembles and Data Augmentation Can Harm Your Calibration · ICLR 2021 |
Computer vision › Image recognition and object detection
image classification |
0.5 | 1 | 2021 | Revisiting the Calibration of Modern Neural Networks · NeurIPS 2021 |
Machine learning › Trustworthy machine learning › calibration
model calibration |
0.5 | 1 | 2021 | Revisiting the Calibration of Modern Neural Networks · NeurIPS 2021 |
Machine learning › Efficient and distributed learning
model compression |
0.5 | 1 | 2021 | Training independent subnetworks for robust prediction · ICLR 2021 |
Machine learning › Trustworthy machine learning
robustness |
0.5 | 1 | 2021 | Training independent subnetworks for robust prediction · ICLR 2021 |
Machine learning › Efficient and distributed learning › efficient training
subnetwork training |
0.5 | 1 | 2021 | Training independent subnetworks for robust prediction · ICLR 2021 |
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks |
0.4 | 1 | 2020 | Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors · ICML 2020 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning
deep ensembles |
0.4 | 1 | 2020 | Hyperparameter Ensembles for Robustness and Uncertainty Quantification · NeurIPS 2020 |
Computer vision › 3D vision › depth perception
distance perception |
0.4 | 1 | 2020 | Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance Awareness · NeurIPS 2020 |
Machine learning › Efficient and distributed learning › distributed inference
distributed bayesian inference |
0.4 | 1 | 2020 | Expectation Propagation as a Way of Life: A Framework for Bayesian Inference on Partitioned Data · J. Mach. Learn. Res. 2020 |
Machine learning › Kernel, tree and ensemble methods
ensemble learning |
0.4 | 1 | 2020 | Hyperparameter Ensembles for Robustness and Uncertainty Quantification · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
expectation propagation |
0.4 | 1 | 2020 | Expectation Propagation as a Way of Life: A Framework for Bayesian Inference on Partitioned Data · J. Mach. Learn. Res. 2020 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.4 | 1 | 2020 | Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance Awareness · NeurIPS 2020 |
Machine learning › Learning paradigms
lifelong learning |
0.4 | 1 | 2020 | BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong Learning · ICLR 2020 |
Machine learning › Deep learning architectures and training › normalization
spectral normalization |
0.4 | 1 | 2020 | Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance Awareness · NeurIPS 2020 |
Machine learning › Trustworthy machine learning › uncertainty estimation › neural network uncertainty
bayesian neural network uncertainty |
0.4 | 1 | 2019 | Bayesian Layers: A Module for Neural Network Uncertainty · NeurIPS 2019 |
Methods — techniques the papers use, named apart from their topics
variational inference · 2.7minimax learning · 1.1batchensemble · 0.9search augmentation · 0.8f1 score · 0.8LLM agents · 0.8expectation propagation · 0.7linear probing on frozen features · 0.7large-scale pretraining · 0.7contrastive learning · 0.7self-attention · 0.3deep learning · 0.3collective communication · 0.3autoregressive modeling · 0.3allreduce · 0.3SPMD programming · 0.3deep neural network · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Long-form factuality in large language modelsabstractLarge language models (LLMs) often generate content that contains factual errors when responding to fact-seeking prompts on open-ended topics. To benchmark a model’s long-form factuality in open domains, we first use GPT-4 to generate LongFact, a prompt set comprising thousands of questions spanning 38 topics. We then propose that LLM agents can be used as automated evaluators for long-form factuality through a method which we call Search-Augmented Factuality Evaluator (SAFE). SAFE utilizes an LLM to break down a long-form response into a set of individual facts and to evaluate the accuracy of each fact using a multi-step reasoning process comprising sending search queries to Google Search and determining whether a fact is supported by the search results. Furthermore, we propose extending F1 score as an aggregated metric for long-form factuality. To do so, we balance the percentage of supported facts in a response (precision) with the percentage of provided facts relative to a hyperparameter representing a user’s preferred response length (recall).
Empirically, we demonstrate that LLM agents can outperform crowdsourced human annotators—on a set of∼16k individual facts, SAFE agrees with crowdsourced human annotators 72% of the time, and on a random subset of 100 disagreement cases, SAFE wins 76% of the time. At the same time, SAFE is more than 20 times cheaper than human annotators. We also benchmark thirteen language models on LongFact across four model families (Gemini, GPT, Claude, and PaLM-2), finding that larger language models generally achieve better long-form factuality. LongFact, SAFE, and all experimental code are available at https://github.com/google-deepmind/long-form-factuality. Jerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu, Nathan Hu, Jie Huang 0009, Dustin Tran, Daiyi Peng, Ruibo Liu, Cosmo Du, Quoc V. Le |
NeurIPS | 7 |
| 2023 | Scaling Vision Transformers to 22 Billion ParametersabstractThe scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Vision Transformers (ViT) have introduced the same architecture to image and video modelling, but these have not yet been successfully scaled to nearly the same degree; the largest dense ViT contains 4B parameters (Chen et al., 2022). We present a recipe for highly efficient and stable training of a 22B-parameter ViT (ViT-22B) and perform a wide variety of experiments on the resulting model. When evaluated on downstream tasks (often with a lightweight linear model on frozen features), ViT-22B demonstrates increasing performance with scale. We further observe other interesting benefits of scale, including an improved tradeoff between fairness and performance, state-of-the-art alignment to human visual perception in terms of shape/texture bias, and improved robustness. ViT-22B demonstrates the potential for "LLM-like" scaling in vision, and provides key steps towards getting there. Mostafa Dehghani 0001, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner 0001, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, Rodolphe Jenatton, Lucas Beyer, Michael Tschannen, Anurag Arnab, Xiao Wang 0038, Carlos Riquelme, Matthias Minderer, Joan Puigcerver, Utku Evci, Sjoerd van Steenkiste, Gamaleldin F. Elsayed, Aravindh Mahendran, Fisher Yu 0001, Avital Oliver, Fantine Huot, Jasmijn Bastings, Mark Collier, Alexey A. Gritsenko, Vighnesh Birodkar, Cristina Nader Vasconcelos, Yi Tay, Thomas Mensink, Alexander Kolesnikov 0003, Filip Pavetic, Dustin Tran, Thomas Kipf, Mario Lucic, Xiaohua Zhai, Daniel Keysers, Jeremiah J. Harmsen, Neil Houlsby |
ICML | 36 |
| 2023 | A Simple Zero-shot Prompt Weighting Technique to Improve Prompt Ensembling in Text-Image ModelsabstractContrastively trained text-image models have the remarkable ability to perform zero-shot classification, that is, classifying previously unseen images into categories that the model has never been explicitly trained to identify. However, these zero-shot classifiers need prompt engineering to achieve high accuracy. Prompt engineering typically requires hand-crafting a set of prompts for individual downstream tasks. In this work, we aim to automate this prompt engineering and improve zero-shot accuracy through prompt ensembling. In particular, we ask “Given a large pool of prompts, can we automatically score the prompts and ensemble those that are most suitable for a particular downstream dataset, without needing access to labeled validation data?". We demonstrate that this is possible. In doing so, we identify several pathologies in a naive prompt scoring method where the score can be easily overconfident due to biases in pre-training and test data, and we propose a novel prompt scoring method that corrects for the biases. Using our proposed scoring method to create a weighted average prompt ensemble, our method overall outperforms equal average ensemble, as well as hand-crafted prompts, on ImageNet, 4 of its variants, and 11 fine-grained classification benchmarks. while being fully automatic, optimization-free, and not requiring access to labeled validation data. James Urquhart Allingham, Jie Ren 0006, Michael Dusenberry, Xiuye Gu, Yin Cui, Dustin Tran, Jeremiah Z. Liu, Balaji Lakshminarayanan |
ICML | 6 |
| 2023 | A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-AwarenessabstractAccurate uncertainty quantification is a major challenge in deep learning, as neural networks can make overconfident errors and assign high confidence predictions to out-of-distribution (OOD) inputs. The most popular approaches to estimate predictive uncertainty in deep learning are methods that combine predictions from multiple neural networks, such as Bayesian neural networks (BNNs) and deep ensembles. However their practicality in real-time, industrial-scale applications are limited due to the high memory and computational cost. Furthermore, ensembles and BNNs do not necessarily fix all the issues with the underlying member networks. In this work, we study principled approaches to improve the uncertainty property of a single network, based on a single, deterministic representation. By formalizing the uncertainty quantification as a minimax learning problem, we first identify distance awareness, i.e., the model's ability to quantify the distance of a testing example from the training data, as a necessary condition for a DNN to achieve high-quality (i.e., minimax optimal) uncertainty estimation. We then propose Spectral-normalized Neural Gaussian Process (SNGP), a simple method that improves the distance-awareness ability of modern DNNs with two simple changes: (1) applying spectral normalization to hidden weights to enforce bi-Lipschitz smoothness in representations and (2) replacing the last output layer with a Gaussian process layer. On a suite of vision and language understanding benchmarks and on modern architectures (Wide-ResNet and BERT), SNGP consistently outperforms other single-model approaches in prediction, calibration and out-of-domain detection. Furthermore, SNGP provides complementary benefits to popular techniques such as deep ensembles and data augmentation, making it a simple and scalable building block for probabilistic deep learning. Jeremiah Z. Liu, Shreyas Padhy, Jie Ren 0006, Zi Lin, Yeming Wen, Ghassen Jerfel, Zachary Nado, Jasper Snoek, Dustin Tran, Balaji Lakshminarayanan |
J. Mach. Learn. Res. | 9 |
| 2021 | Training independent subnetworks for robust prediction
Marton Havasi, Rodolphe Jenatton, Stanislav Fort, Jeremiah Z. Liu, Jasper Snoek, Balaji Lakshminarayanan, Andrew M. Dai, Dustin Tran |
ICLR | 8 |
| 2021 | Combining Ensembles and Data Augmentation Can Harm Your Calibration
Yeming Wen, Ghassen Jerfel, Rafael Muller, Michael Dusenberry, Jasper Snoek, Balaji Lakshminarayanan, Dustin Tran |
ICLR | 7 |
| 2021 | Soft Calibration Objectives for Neural NetworksabstractOptimal decision making requires that classifiers produce uncertainty estimates consistent with their empirical accuracy. However, deep neural networks are often under- or over-confident in their predictions. Consequently, methods have been developed to improve the calibration of their predictive uncertainty both during training and post-hoc. In this work, we propose differentiable losses to improve calibration based on a soft (continuous) version of the binning operation underlying popular calibration-error estimators. When incorporated into training, these soft calibration losses achieve state-of-the-art single-model ECE across multiple datasets with less than 1% decrease in accuracy. For instance, we observe an 82% reduction in ECE (70% relative to the post-hoc rescaled ECE) in exchange for a 0.7% relative decrease in accuracy relative to the cross entropy baseline on CIFAR-100.When incorporated post-training, the soft-binning-based calibration error objective improves upon temperature scaling, a popular recalibration method. Overall, experiments across losses and datasets demonstrate that using calibration-sensitive procedures yield better uncertainty estimates under dataset shift than the standard practice of using a cross entropy loss and post-hoc recalibration methods. Archit Karandikar, Nicholas Cain, Dustin Tran, Balaji Lakshminarayanan, Jonathon Shlens, Michael C. Mozer, Rebecca Roelofs |
NeurIPS | 3 |
| 2021 | Revisiting the Calibration of Modern Neural NetworksabstractAccurate estimation of predictive uncertainty (model calibration) is essential for the safe application of neural networks. Many instances of miscalibration in modern neural networks have been reported, suggesting a trend that newer, more accurate models produce poorly calibrated predictions. Here, we revisit this question for recent state-of-the-art image classification models. We systematically relate model calibration and accuracy, and find that the most recent models, notably those not using convolutions, are among the best calibrated. Trends observed in prior model generations, such as decay of calibration with distribution shift or model size, are less pronounced in recent architectures. We also show that model size and amount of pretraining do not fully explain these differences, suggesting that architecture is a major determinant of calibration properties. Matthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis, Xiaohua Zhai, Neil Houlsby, Dustin Tran, Mario Lucic |
NeurIPS | 7 |
| 2020 | BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong Learning
Yeming Wen, Dustin Tran, Jimmy Ba |
ICLR | 2 |
| 2020 | Efficient and Scalable Bayesian Neural Nets with Rank-1 FactorsabstractBayesian neural networks (BNNs) demonstrate promising success in improving the robustness and uncertainty quantification of modern deep learning. However, they generally struggle with underfitting at scale and parameter efficiency. On the other hand, deep ensembles have emerged as alternatives for uncertainty quantification that, while outperforming BNNs on certain problems, also suffer from efficiency issues. It remains unclear how to combine the strengths of these two approaches and remediate their common issues. To tackle this challenge, we propose a rank-1 parameterization of BNNs, where each weight matrix involves only a distribution on a rank-1 subspace. We also revisit the use of mixture approximate posteriors to capture multiple modes, where unlike typical mixtures, this approach admits a significantly smaller memory increase (e.g., only a 0.4% increase for a ResNet-50 mixture of size 10). We perform a systematic empirical study on the choices of prior, variational posterior, and methods to improve training. For ResNet-50 on ImageNet, Wide ResNet 28-10 on CIFAR-10/100, and an RNN on MIMIC-III, rank-1 BNNs achieve state-of-the-art performance across log-likelihood, accuracy, and calibration on the test sets and out-of-distribution variants. Michael Dusenberry, Ghassen Jerfel, Yeming Wen, Yi-An Ma, Jasper Snoek, Katherine A. Heller, Balaji Lakshminarayanan, Dustin Tran |
ICML | 8 |
| 2020 | Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance AwarenessabstractBayesian neural networks (BNN) and deep ensembles are principled approaches to estimate the predictive uncertainty of a deep learning model. However their practicality in real-time, industrial-scale applications are limited due to their heavy memory and inference cost. This motivates us to study principled approaches to high-quality uncertainty estimation that require only a single deep neural network (DNN). By formalizing the uncertainty quantification as a minimax learning problem, we first identify input distance awareness, i.e., the model’s ability to quantify the distance of a testing example from the training data in the input space, as a necessary condition for a DNN to achieve high-quality (i.e., minimax optimal) uncertainty estimation. We then propose Spectral-normalized Neural Gaussian Process (SNGP), a simple method that improves the distance-awareness ability of modern DNNs, by adding a weight normalization step during training and replacing the output layer. On a suite of vision and language understanding tasks and on modern architectures (Wide-ResNet and BERT), SNGP is competitive with deep ensembles in prediction, calibration and out-of-domain detection, and outperforms the other single-model approaches. Jeremiah Z. Liu, Zi Lin, Shreyas Padhy, Dustin Tran, Tania Bedrax-Weiss, Balaji Lakshminarayanan |
NeurIPS | 4 |
| 2020 | Hyperparameter Ensembles for Robustness and Uncertainty QuantificationabstractEnsembles over neural network weights trained from different random initialization, known as deep ensembles, achieve state-of-the-art accuracy and calibration. The recently introduced batch ensembles provide a drop-in replacement that is more parameter efficient. In this paper, we design ensembles not only over weights, but over hyperparameters to improve the state of the art in both settings. For best performance independent of budget, we propose hyper-deep ensembles, a simple procedure that involves a random search over different hyperparameters, themselves stratified across multiple random initializations. Its strong performance highlights the benefit of combining models with both weight and hyperparameter diversity. We further propose a parameter efficient version, hyper-batch ensembles, which builds on the layer structure of batch ensembles and self-tuning networks. The computational and memory costs of our method are notably lower than typical ensembles. On image classification tasks, with MLP, LeNet, ResNet 20 and Wide ResNet 28-10 architectures, we improve upon both deep and batch ensembles. Florian Wenzel, Jasper Snoek, Dustin Tran, Rodolphe Jenatton |
NeurIPS | 3 |
| 2020 | Demonstrating Principled Uncertainty Modeling for Recommender Ecosystems with RecSim NGabstractWe develop RecSim NG, a probabilistic platform that supports natural, concise specification and learning of models for multi-agent recommender systems simulation. RecSim NG is a scalable, modular, differentiable simulator implemented in Edward2 and TensorFlow. Martin Mladenov, Vihan Jain, Eugene Ie, Christopher Colby, Nicolas Mayoraz, Hubert Pham, Dustin Tran, Ivan Vendrov, Craig Boutilier |
RecSys | 8 |
| 2020 | Expectation Propagation as a Way of Life: A Framework for Bayesian Inference on Partitioned DataabstractA common divide-and-conquer approach for Bayesian computation with big data is to partition the data, perform local inference for each piece separately, and combine the results to obtain a global posterior approximation. While being conceptually and computationally appealing, this method involves the problematic need to also split the prior for the local inferences; these weakened priors may not provide enough regularization for each separate computation, thus eliminating one of the key advantages of Bayesian methods. To resolve this dilemma while still retaining the generalizability of the underlying local inference method, we apply the idea of expectation propagation (EP) as a framework for distributed Bayesian inference. The central idea is to iteratively update approximations to the local likelihoods given the state of the other approximations and the prior. The present paper has two roles: we review the steps that are needed to keep EP algorithms numerically stable, and we suggest a general approach, inspired by EP, for approaching data partitioning problems in a way that achieves the computational benefits of parallelism while allowing each local update to make use of relevant information from the other sites. In addition, we demonstrate how the method can be applied in a hierarchical context to make use of partitioning of both data and parameters. The paper describes a general algorithmic framework, rather than a specific algorithm, and presents an example implementation for it. Aki Vehtari, Andrew Gelman, Tuomas Sivula, Pasi Jylänki, Dustin Tran, Swupnil Sahai, Paul Blomstedt, John P. Cunningham, David Schiminovich, Christian P. Robert |
J. Mach. Learn. Res. | 5 |
| 2019 | Bayesian Layers: A Module for Neural Network UncertaintyabstractWe describe Bayesian Layers, a module designed for fast experimentation with neural network uncertainty. It extends neural network libraries with drop-in replacements for common layers. This enables composition via a unified abstraction over deterministic and stochastic functions and allows for scalability via the underlying system. These layers capture uncertainty over weights (Bayesian neural nets), pre-activation units (dropout), activations (stochastic output layers''), or the function itself (Gaussian processes). They can also be reversible to propagate uncertainty from input to output. We include code examples for common architectures such as Bayesian LSTMs, deep GPs, and flow-based models. As demonstration, we fit a 5-billion parameterBayesian Transformer'' on 512 TPUv2 cores for uncertainty in machine translation and a Bayesian dynamics model for model-based planning. Finally, we show how Bayesian Layers can be used within the Edward2 language for probabilistic programming with stochastic processes. Dustin Tran, Michael Dusenberry, Mark van der Wilk, Danijar Hafner |
NeurIPS | 1 |
| 2019 | Discrete Flows: Invertible Generative Models of Discrete DataabstractWhile normalizing flows have led to significant advances in modeling high-dimensional continuous distributions, their applicability to discrete distributions remains unknown. In this paper, we show that flows can in fact be extended to discrete events---and under a simple change-of-variables formula not requiring log-determinant-Jacobian computations. Discrete flows have numerous applications. We consider two flow architectures: discrete autoregressive flows that enable bidirectionality, allowing, for example, tokens in text to depend on both left-to-right and right-to-left contexts in an exact language model; and discrete bipartite flows that enable efficient non-autoregressive generation as in RealNVP. Empirically, we find that discrete autoregressive flows outperform autoregressive baselines on synthetic discrete distributions, an addition task, and Potts models; and bipartite flows can obtain competitive performance with autoregressive baselines on character-level language modeling for Penn Tree Bank and text8. Dustin Tran, Keyon Vafa, Kumar Krishna Agrawal, Laurent Dinh, Ben Poole |
NeurIPS | 1 |
| 2019 | Noise Contrastive Priors for Functional Uncertainty
Danijar Hafner, Dustin Tran, Timothy P. Lillicrap, Alex Irpan, James Davidson |
UAI | 2 |
| 2018 | Implicit Causal Models for Genome-wide Association Studies
Dustin Tran, David M. Blei |
ICLR (Poster) | 1 |
| 2018 | Flipout: Efficient Pseudo-Independent Weight Perturbations on Mini-Batches
Yeming Wen, Paul Vicol, Jimmy Ba, Dustin Tran, Roger B. Grosse |
ICLR (Poster) | 4 |
| 2018 | Image TransformerabstractImage generation has been successfully cast as an autoregressive sequence generation or transformation problem. Recent work has shown that self-attention is an effective way of modeling textual sequences. In this work, we generalize a recently proposed model architecture based on self-attention, the Transformer, to a sequence modeling formulation of image generation with a tractable likelihood. By restricting the self-attention mechanism to attend to local neighborhoods we significantly increase the size of images the model can process in practice, despite maintaining significantly larger receptive fields per layer than typical convolutional neural networks. While conceptually simple, our generative models significantly outperform the current state of the art in image generation on ImageNet, improving the best published negative log-likelihood on ImageNet from 3.83 to 3.77. We also present results on image super-resolution with a large magnification ratio, applying an encoder-decoder configuration of our architecture. In a human evaluation study, we find that images generated by our super-resolution model fool human observers three times more often than the previous state of the art. Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, Dustin Tran |
ICML | 7 |
| 2018 | Mesh-TensorFlow: Deep Learning for SupercomputersabstractBatch-splitting (data-parallelism) is the dominant distributed Deep Neural Network (DNN) training strategy, due to its universal applicability and its amenability to Single-Program-Multiple-Data (SPMD) programming. However, batch-splitting suffers from problems including the inability to train very large models (due to memory constraints), high latency, and inefficiency at small batch sizes. All of these can be solved by more general distribution strategies (model-parallelism). Unfortunately, efficient model-parallel algorithms tend to be complicated to discover, describe, and to implement, particularly on large clusters. We introduce Mesh-TensorFlow, a language for specifying a general class of distributed tensor computations. Where data-parallelism can be viewed as splitting tensors and operations along the "batch" dimension, in Mesh-TensorFlow, the user can specify any tensor-dimensions to be split across any dimensions of a multi-dimensional mesh of processors. A Mesh-TensorFlow graph compiles into a SPMD program consisting of parallel operations coupled with collective communication primitives such as Allreduce. We use Mesh-TensorFlow to implement an efficient data-parallel, model-parallel version of the Transformer sequence-to-sequence model. Using TPU meshes of up to 512 cores, we train Transformer models with up to 5 billion parameters, surpassing SOTA results on WMT'14 English-to-French translation task and the one-billion-word Language modeling benchmark. Mesh-Tensorflow is available at https://github.com/tensorflow/mesh Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, Ryan Sepassi, Blake A. Hechtman |
NeurIPS | 4 |
| 2018 | Simple, Distributed, and Accelerated Probabilistic ProgrammingabstractWe describe a simple, low-level approach for embedding probabilistic programming in a deep learning ecosystem. In particular, we distill probabilistic programming down to a single abstraction—the random variable. Our lightweight implementation in TensorFlow enables numerous applications: a model-parallel variational auto-encoder (VAE) with 2nd-generation tensor processing units (TPUv2s); a data-parallel autoregressive model (Image Transformer) with TPUv2s; and multi-GPU No-U-Turn Sampler (NUTS). For both a state-of-the-art VAE on 64x64 ImageNet and Image Transformer on 256x256 CelebA-HQ, our approach achieves an optimal linear speedup from 1 to 256 TPUv2 chips. With NUTS, we see a 100x speedup on GPUs over Stan and 37x over PyMC3. Dustin Tran, Matthew Hoffman 0001, Dave Moore, Christopher Suter, Srinivas Vasudevan, Alexey Radul |
NeurIPS | 1 |
| 2017 | Deep Probabilistic Programming
Dustin Tran, Matthew Hoffman 0001, Rif A. Saurous, Eugene Brevdo, Kevin Murphy 0002, David M. Blei |
ICLR (Poster) | 1 |
| 2017 | Variational Inference via \chi Upper Bound MinimizationabstractVariational inference (VI) is widely used as an efficient alternative to Markov chain Monte Carlo. It posits a family of approximating distributions $q$ and finds the closest member to the exact posterior $p$. Closeness is usually measured via a divergence $D(q || p)$ from $q$ to $p$. While successful, this approach also has problems. Notably, it typically leads to underestimation of the posterior variance. In this paper we propose CHIVI, a black-box variational inference algorithm that minimizes $D_{\chi}(p || q)$, the $\chi$-divergence from $p$ to $q$. CHIVI minimizes an upper bound of the model evidence, which we term the $\chi$ upper bound (CUBO). Minimizing the CUBO leads to improved posterior uncertainty, and it can also be used with the classical VI lower bound (ELBO) to provide a sandwich estimate of the model evidence. We study CHIVI on three models: probit regression, Gaussian process classification, and a Cox process model of basketball plays. When compared to expectation propagation and classical VI, CHIVI produces better error rates and more accurate estimates of posterior variance. Adji B. Dieng, Dustin Tran, Rajesh Ranganath, John W. Paisley, David M. Blei |
NIPS | 2 |
| 2017 | Hierarchical Implicit Models and Likelihood-Free Variational InferenceabstractImplicit probabilistic models are a flexible class of models defined by a simulation process for data. They form the basis for models which encompass our understanding of the physical word. Despite this fundamental nature, the use of implicit models remains limited due to challenge in positing complex latent structure in them, and the ability to inference in such models with large data sets. In this paper, we first introduce the hierarchical implicit models (HIMs). HIMs combine the idea of implicit densities with hierarchical Bayesian modeling thereby defining models via simulators of data with rich hidden structure. Next, we develop likelihood-free variational inference (LFVI), a scalable variational inference algorithm for HIMs. Key to LFVI is specifying a variational family that is also implicit. This matches the model's flexibility and allows for accurate approximation of the posterior. We demonstrate diverse applications: a large-scale physical simulator for predator-prey populations in ecology; a Bayesian generative adversarial network for discrete data; and a deep implicit model for symbol generation. Dustin Tran, Rajesh Ranganath, David M. Blei |
NIPS | 1 |
| 2017 | Automatic Differentiation Variational InferenceabstractProbabilistic modeling is iterative. A scientist posits a simple model, fits it to her data, refines it according to her analysis, and repeats. However, fitting complex models to large data is a bottleneck in this process. Deriving algorithms for new models can be both mathematically and computationally challenging, which makes it difficult to efficiently cycle through the steps. To this end, we develop ADVI. Using our method, the scientist only provides a probabilistic model and a dataset, nothing else. ADVI automatically derives an efficient variational inference algorithm, freeing the scientist to refine and explore many models. ADVI supports a broad class of models ---no conjugacy assumptions are required. We study ADVI across ten modern probabilistic models and apply it to a dataset with millions of observations. We deploy ADVI as part of Stan, a probabilistic programming system. Alp Kucukelbir, Dustin Tran, Rajesh Ranganath, Andrew Gelman, David M. Blei |
J. Mach. Learn. Res. | 2 |
| 2016 | Towards Stability and Optimality in Stochastic Gradient DescentabstractIterative procedures for parameter estimation based on stochastic gradient descent (SGD) allow the estimation to scale to massive data sets. However, they typically suffer from numerical instability, while estimators based on SGD are statistically inefficient as they do not use all the information in the data set. To address these two issues we propose an iterative estimation procedure termed averaged implicit SGD (AI-SGD). For statistical efficiency AI-SGD employs averaging of the iterates, which achieves the Cramer-Rao bound under strong convexity, i.e., it is asymptotically an optimal unbiased estimator of the true parameter value. For numerical stability AI-SGD employs an implicit update at each iteration, which is similar to updates performed by proximal operators in optimization. In practice, AI-SGD achieves competitive performance with state-of-the-art procedures. Furthermore, it is more stable than averaging procedures that do not employ proximal updates, and is simple to implement as it requires fewer tunable hyperparameters than procedures that do employ proximal updates. Panos Toulis, Dustin Tran, Edoardo M. Airoldi |
AISTATS | 2 |
| 2016 | Spectral M-estimation with Applications to Hidden Markov ModelsabstractMethod of moment estimators exhibit appealing statistical properties, such as asymptotic unbiasedness, for nonconvex problems. However, they typically require a large number of samples and are extremely sensitive to model misspecification. In this paper, we apply the framework of M-estimation to develop both a generalized method of moments procedure and a principled method for regularization. Our proposed M-estimator obtains optimal sample efficiency rates (in the class of moment-based estimators) and the same well-known rates on prediction accuracy as other spectral estimators. It also makes it straightforward to incorporate regularization into the sample moment conditions. We demonstrate empirically the gains in sample efficiency from our approach on hidden Markov models. Dustin Tran, Finale Doshi-Velez |
AISTATS | 1 |
| 2016 | Hierarchical Variational ModelsabstractBlack box variational inference allows researchers to easily prototype and evaluate an array of models. Recent advances allow such algorithms to scale to high dimensions. However, a central question remains: How to specify an expressive variational distribution that maintains efficient computation? To address this, we develop hierarchical variational models (HVMs). HVMs augment a variational approximation with a prior on its parameters, which allows it to capture complex structure for both discrete and continuous latent variables. The algorithm we develop is black box, can be used for any HVM, and has the same computational efficiency as the original approximation. We study HVMs on a variety of deep discrete latent variable models. HVMs generalize other expressive variational distributions and maintains higher fidelity to the posterior. Rajesh Ranganath, Dustin Tran, David M. Blei |
ICML | 2 |
| 2016 | Operator Variational InferenceabstractVariational inference is an umbrella term for algorithms which cast Bayesian inference as optimization. Classically, variational inference uses the Kullback-Leibler divergence to define the optimization. Though this divergence has been widely used, the resultant posterior approximation can suffer from undesirable statistical properties. To address this, we reexamine variational inference from its roots as an optimization problem. We use operators, or functions of functions, to design variational objectives. As one example, we design a variational objective with a Langevin-Stein operator. We develop a black box algorithm, operator variational inference (OPVI), for optimizing any operator objective. Importantly, operators enable us to make explicit the statistical and computational tradeoffs for variational inference. We can characterize different properties of variational objectives, such as objectives that admit data subsampling---allowing inference to scale to massive data---as well as objectives that admit variational programs---a rich class of posterior approximations that does not require a tractable density. We illustrate the benefits of OPVI on a mixture model and a generative model of images. Rajesh Ranganath, Dustin Tran, Jaan Altosaar, David M. Blei |
NIPS | 2 |
| 2015 | Copula variational inferenceabstractWe develop a general variational inference method that preserves dependency among the latent variables. Our method uses copulas to augment the families of distributions used in mean-field and structured approximations. Copulas model the dependency that is not captured by the original variational distribution, and thus the augmented variational family guarantees better approximations to the posterior. With stochastic optimization, inference on the augmented distribution is scalable. Furthermore, our strategy is generic: it can be applied to any inference procedure that currently uses the mean-field or structured approach. Copula variational inference has many advantages: it reduces bias; it is less sensitive to local optima; it is less sensitive to hyperparameters; and it helps characterize and interpret the dependency among the latent variables. Dustin Tran, David M. Blei, Edoardo M. Airoldi |
NIPS | 1 |