Tom Blau

dblp:233/0118 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
2since 2021 · last 2022
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 48% Probabilistic and Bayesian machine learning · 24% Robot manipulation · 21%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › experimental design
bayesian experimental design
0.612022
Optimizing Sequential Experimental Design with Deep Reinforcement Learning · ICML 2022
Machine learning › Reinforcement learning
deep reinforcement learning
0.612022
Optimizing Sequential Experimental Design with Deep Reinforcement Learning · ICML 2022
Machine learning › Reinforcement learning
sequential experimental design
0.612022
Optimizing Sequential Experimental Design with Deep Reinforcement Learning · ICML 2022
Robotics › Robot manipulation
learning from demonstration
0.512021
Learning from Demonstration without Demonstrations · ICRA 2021
Robotics › Motion planning and robot control
robot learning
0.112021
Learning from Demonstration without Demonstrations · ICRA 2021

Methods — techniques the papers use, named apart from their topics

markov decision process · 0.6deep reinforcement learning · 0.6reinforcement learning · 0.5imitation learning · 0.5
YearPublicationVenuePosition
2022 Optimizing Sequential Experimental Design with Deep Reinforcement Learning
abstract
Bayesian approaches developed to solve the optimal design of sequential experiments are mathematically elegant but computationally challenging. Recently, techniques using amortization have been proposed to make these Bayesian approaches practical, by training a parameterized policy that proposes designs efficiently at deployment time. However, these methods may not sufficiently explore the design space, require access to a differentiable probabilistic model and can only optimize over continuous design spaces. Here, we address these limitations by showing that the problem of optimizing policies can be reduced to solving a Markov decision process (MDP). We solve the equivalent MDP with modern deep reinforcement learning techniques. Our experiments show that our approach is also computationally efficient at deployment time and exhibits state-of-the-art performance on both continuous and discrete design spaces, even when the probabilistic model is a black box.
Tom Blau, Edwin V. Bonilla, Iadine Chades, Amir Dezfouli
ICML1
2021 Learning from Demonstration without Demonstrations
Tom Blau, Philippe Morere, Gilad Francis
ICRA1
2018 Improving Reinforcement Learning Pre-Training with Variational Dropout
abstract
Reinforcement learning has been very successful at learning control policies for robotic agents in order to perform various tasks, such as driving around a track, navigating a maze, and bipedal locomotion. One significant drawback of reinforcement learning methods is that they require a large number of data points in order to learn good policies, a trait known as poor data efficiency or poor sample efficiency. One approach for improving sample efficiency is supervised pre-training of policies to directly clone the behavior of an expert, but this suffers from poor generalization far from the training data. We propose to improve this by using Gaussian dropout networks with a regularization term based on variational inference in the pre-training step. We show that this initializes policy parameters to significantly better values than standard supervised learning or random initialization, thus greatly reducing sample complexity compared with state-of-the-art methods, and enabling an RL algorithm to learn optimal policies for high-dimensional continuous control problems in a practical time frame.
Tom Blau, Lionel Ott, Fabio Ramos 0001
IROS1