Jiri Hron

dblp:209/4975 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
3since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Probabilistic and Bayesian machine learning · 58% Learning theory · 16% Deep learning architectures and training · 15%
Databases, data mining, and information retrieval
2 papers
Recommender systems · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%
Software engineering, system software, and programming languages
1 paper
Programming languages and type systems · 100%

Topics — the 28 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
neural network gaussian process
1.432022
Wide Bayesian neural networks have a simple weight posterior: theory and accelerated sampling · ICML 2022
Infinite attention: NNGP and NTK for deep attention networks · ICML 2020
Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes · ICLR (Poster) 2019
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks
1.332022
Wide Bayesian neural networks have a simple weight posterior: theory and accelerated sampling · ICML 2022
Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes · ICLR (Poster) 2019
Variational Bayesian dropout: pitfalls and fixes · ICML 2018
Machine learning › Learning theory
neural network theory
1.022022
Wide Bayesian neural networks have a simple weight posterior: theory and accelerated sampling · ICML 2022
Infinite attention: NNGP and NTK for deep attention networks · ICML 2020
Machine learning › Probabilistic and Bayesian machine learning › sampling
posterior sampling
1.022022
Wide Bayesian neural networks have a simple weight posterior: theory and accelerated sampling · ICML 2022
Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning · NeurIPS 2019
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel
0.922020
Infinite attention: NNGP and NTK for deep attention networks · ICML 2020
Neural Tangents: Fast and Easy Infinite Neural Networks in Python · ICLR 2020
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.822020
Infinite attention: NNGP and NTK for deep attention networks · ICML 2020
Gaussian Process Behaviour in Wide Deep Neural Networks · ICLR (Poster) 2018
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo
0.612022
Wide Bayesian neural networks have a simple weight posterior: theory and accelerated sampling · ICML 2022
Recommender systems › large-scale recommendation › multi-stage recommender systems
candidate generation
0.512021
On Component Interactions in Two-Stage Recommender Systems · NeurIPS 2021
Recommender systems › large-scale recommendation › multi-stage recommender systems
two-stage recommender systems
0.512021
On Component Interactions in Two-Stage Recommender Systems · NeurIPS 2021
Machine learning › Deep learning architectures and training › attention mechanism
attention network
0.412020
Infinite attention: NNGP and NTK for deep attention networks · ICML 2020
Programming languages and type systems
library design
0.412020
Neural Tangents: Fast and Easy Infinite Neural Networks in Python · ICLR 2020
Machine learning › Deep learning architectures and training › convolutional neural network
bayesian convolutional neural networks
0.412019
Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes · ICLR (Poster) 2019
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff
0.412019
Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning · NeurIPS 2019
Machine learning › Reinforcement learning › exploration
randomized value functions
0.412019
Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning · NeurIPS 2019
Machine learning › Reinforcement learning
value function
0.412019
Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference
0.312018
Variational Bayesian dropout: pitfalls and fixes · ICML 2018
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
variational dropout
0.312018
Variational Bayesian dropout: pitfalls and fixes · ICML 2018
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.312018
Variational Bayesian dropout: pitfalls and fixes · ICML 2018
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
variational objective
0.312018
Variational Bayesian dropout: pitfalls and fixes · ICML 2018
Machine learning › Deep learning architectures and training › overparameterized neural network
wide neural networks
0.312018
Gaussian Process Behaviour in Wide Deep Neural Networks · ICLR (Poster) 2018
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models
bayesian deep learning
0.312017
Concrete Dropout · NIPS 2017
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.212022
Wide Bayesian neural networks have a simple weight posterior: theory and accelerated sampling · ICML 2022
Machine learning › Deep learning architectures and training
reparameterization
0.212022
Wide Bayesian neural networks have a simple weight posterior: theory and accelerated sampling · ICML 2022
Machine learning › Deep learning architectures and training
attention mechanism
0.112020
Infinite attention: NNGP and NTK for deep attention networks · ICML 2020
Machine learning › Deep learning architectures and training › attention mechanism
multi-head attention
0.112020
Infinite attention: NNGP and NTK for deep attention networks · ICML 2020
Machine learning › Deep learning architectures and training › regularization
dropout
0.112018
Variational Bayesian dropout: pitfalls and fixes · ICML 2018
Machine learning › Deep learning architectures and training
regularization
0.112018
Variational Bayesian dropout: pitfalls and fixes · ICML 2018
Machine learning › Trustworthy machine learning
uncertainty estimation
0.112017
Concrete Dropout · NIPS 2017

Methods — techniques the papers use, named apart from their topics

game theory · 1.3neural tangent kernel · 1.3gaussian process · 0.8repriorisation · 0.6reparameterization · 0.6markov chain monte carlo · 0.6mixture of experts · 0.5generalization bounds · 0.5bayesian neural network · 0.4temporal difference learning · 0.4neural network function approximation · 0.4bayesian inference · 0.4kullback-leibler divergence · 0.3
YearPublicationVenuePosition
2023 Modeling content creator incentives on algorithm-curated platforms
Jiri Hron, Karl Krauth, Michael I. Jordan, Niki Kilbertus, Sarah Dean
ICLR1
2022 Wide Bayesian neural networks have a simple weight posterior: theory and accelerated sampling
abstract
We introduce repriorisation, a data-dependent reparameterisation which transforms a Bayesian neural network (BNN) posterior to a distribution whose KL divergence to the BNN prior vanishes as layer widths grow. The repriorisation map acts directly on parameters, and its analytic simplicity complements the known neural network Gaussian process (NNGP) behaviour of wide BNNs in function space. Exploiting the repriorisation, we develop a Markov chain Monte Carlo (MCMC) posterior sampling algorithm which mixes faster the wider the BNN. This contrasts with the typically poor performance of MCMC in high dimensions. We observe up to 50x higher effective sample size relative to no reparametrisation for both fully-connected and residual networks. Improvements are achieved at all widths, with the margin between reparametrised and standard BNNs growing with layer width.
Jiri Hron, Roman Novak, Jeffrey Pennington, Jascha Sohl-Dickstein
ICML1
2021 On Component Interactions in Two-Stage Recommender Systems
abstract
Thanks to their scalability, two-stage recommenders are used by many of today's largest online platforms, including YouTube, LinkedIn, and Pinterest. These systems produce recommendations in two steps: (i) multiple nominators—tuned for low prediction latency—preselect a small subset of candidates from the whole item pool; (ii) a slower but more accurate ranker further narrows down the nominated items, and serves to the user. Despite their popularity, the literature on two-stage recommenders is relatively scarce, and the algorithms are often treated as mere sums of their parts. Such treatment presupposes that the two-stage performance is explained by the behavior of the individual components in isolation. This is not the case: using synthetic and real-world data, we demonstrate that interactions between the ranker and the nominators substantially affect the overall performance. Motivated by these findings, we derive a generalization lower bound which shows that independent nominator training can lead to performance on par with uniformly random recommendations. We find that careful design of item pools, each assigned to a different nominator, alleviates these issues. As manual search for a good pool allocation is difficult, we propose to learn one instead using a Mixture-of-Experts based approach. This significantly improves both precision and recall at $K$.
Jiri Hron, Karl Krauth, Michael I. Jordan, Niki Kilbertus
NeurIPS1
2020 Neural Tangents: Fast and Easy Infinite Neural Networks in Python
Roman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee 0001, Alexander A. Alemi, Jascha Sohl-Dickstein, Samuel S. Schoenholz
ICLR3
2020 Infinite attention: NNGP and NTK for deep attention networks
abstract
There is a growing amount of literature on the relationship between wide neural networks (NNs) and Gaussian processes (GPs), identifying an equivalence between the two for a variety of NN architectures. This equivalence enables, for instance, accurate approximation of the behaviour of wide Bayesian NNs without MCMC or variational approximations, or characterisation of the distribution of randomly initialised wide NNs optimised by gradient descent without ever running an optimiser. We provide a rigorous extension of these results to NNs involving attention layers, showing that unlike single-head attention, which induces non-Gaussian behaviour, multi-head attention architectures behave as GPs as the number of heads tends to infinity. We further discuss the effects of positional encodings and layer normalisation, and propose modifications of the attention mechanism which lead to improved results for both finite and infinitely wide NNs. We evaluate attention kernels empirically, leading to a moderate improvement upon the previous state-of-the-art on CIFAR-10 for GPs without trainable kernels and advanced data preprocessing. Finally, we introduce new features to the Neural Tangents library (Novak et al.,2020) allowing applications of NNGP/NTK models, with and without attention, to variable-length sequences, with an example on the IMDb reviews dataset.
Jiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, Roman Novak
ICML1
2019 Orthogonal Estimation of Wasserstein Distances
abstract
Wasserstein distances are increasingly used in a wide variety of applications in machine learning. Sliced Wasserstein distances form an important subclass which may be estimated efficiently through one-dimensional sorting operations. In this paper, we propose a new variant of sliced Wasserstein distance, study the use of orthogonal coupling in Monte Carlo estimation of Wasserstein distances and draw connections with stratified sampling, and evaluate our approaches experimentally in a range of large-scale experiments in generative modelling and reinforcement learning.
Mark Rowland 0001, Jiri Hron, Yunhao Tang, Krzysztof Choromanski, Tamás Sarlós, Adrian Weller
AISTATS2
2019 Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes
Roman Novak, Lechao Xiao, Yasaman Bahri, Jaehoon Lee 0001, Greg Yang, Jiri Hron, Daniel A. Abolafia, Jeffrey Pennington, Jascha Sohl-Dickstein
ICLR (Poster)6
2019 Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning
abstract
Posterior sampling for reinforcement learning (PSRL) is an effective method for balancing exploration and exploitation in reinforcement learning. Randomised value functions (RVF) can be viewed as a promising approach to scaling PSRL. However, we show that most contemporary algorithms combining RVF with neural network function approximation do not possess the properties which make PSRL effective, and provably fail in sparse reward problems. Moreover, we find that propagation of uncertainty, a property of PSRL previously thought important for exploration, does not preclude this failure. We use these insights to design Successor Uncertainties (SU), a cheap and easy to implement RVF algorithm that retains key properties of PSRL. SU is highly effective on hard tabular exploration benchmarks. Furthermore, on the Atari 2600 domain, it surpasses human performance on 38 of 49 games tested (achieving a median human normalised score of 2.09), and outperforms its closest RVF competitor, Bootstrapped DQN, on 36 of those.
David Janz, Jiri Hron, Przemyslaw Mazur, Katja Hofmann, José Miguel Hernández-Lobato, Sebastian Tschiatschek
NeurIPS2
2018 Gaussian Process Behaviour in Wide Deep Neural Networks
Alexander G. de G. Matthews, Jiri Hron, Mark Rowland 0001, Richard E. Turner, Zoubin Ghahramani
ICLR (Poster)2
2018 Variational Bayesian dropout: pitfalls and fixes
abstract
Dropout, a stochastic regularisation technique for training of neural networks, has recently been reinterpreted as a specific type of approximate inference algorithm for Bayesian neural networks. The main contribution of the reinterpretation is in providing a theoretical framework useful for analysing and extending the algorithm. We show that the proposed framework suffers from several issues; from undefined or pathological behaviour of the true posterior related to use of improper priors, to an ill-defined variational objective due to singularity of the approximating distribution relative to the true posterior. Our analysis of the improper log uniform prior used in variational Gaussian dropout suggests the pathologies are generally irredeemable, and that the algorithm still works only because the variational formulation annuls some of the pathologies. To address the singularity issue, we proffer Quasi-KL (QKL) divergence, a new approximate inference objective for approximation of high-dimensional distributions. We show that motivations for variational Bernoulli dropout based on discretisation and noise have QKL as a limit. Properties of QKL are studied both theoretically and on a simple practical example which shows that the QKL-optimal approximation of a full rank Gaussian with a degenerate one naturally leads to the Principal Component Analysis solution.
Jiri Hron, Alexander G. de G. Matthews, Zoubin Ghahramani
ICML1
2017 Concrete Dropout
abstract
Dropout is used as a practical tool to obtain uncertainty estimates in large vision models and reinforcement learning (RL) tasks. But to obtain well-calibrated uncertainty estimates, a grid-search over the dropout probabilities is necessary—a prohibitive operation with large models, and an impossible one with RL. We propose a new dropout variant which gives improved performance and better calibrated uncertainties. Relying on recent developments in Bayesian deep learning, we use a continuous relaxation of dropout’s discrete masks. Together with a principled optimisation objective, this allows for automatic tuning of the dropout probability in large models, and as a result faster experimentation cycles. In RL this allows the agent to adapt its uncertainty dynamically as more data is observed. We analyse the proposed variant extensively on a range of tasks, and give insights into common practice in the field where larger dropout probabilities are often used in deeper model layers.
Yarin Gal, Jiri Hron, Alex Kendall
NIPS2