VLDB 2026 Research / reviewers in the wild / expert
Tom Blau
dblp:233/0118
· DBLP profile ↗
3ranked-venue papers
3as first author
2since 2021 · last 2022
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 48% Probabilistic and Bayesian machine learning · 24% Robot manipulation · 21% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › experimental design
bayesian experimental design |
0.6 | 1 | 2022 | Optimizing Sequential Experimental Design with Deep Reinforcement Learning · ICML 2022 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.6 | 1 | 2022 | Optimizing Sequential Experimental Design with Deep Reinforcement Learning · ICML 2022 |
Machine learning › Reinforcement learning
sequential experimental design |
0.6 | 1 | 2022 | Optimizing Sequential Experimental Design with Deep Reinforcement Learning · ICML 2022 |
Robotics › Robot manipulation
learning from demonstration |
0.5 | 1 | 2021 | Learning from Demonstration without Demonstrations · ICRA 2021 |
Robotics › Motion planning and robot control
robot learning |
0.1 | 1 | 2021 | Learning from Demonstration without Demonstrations · ICRA 2021 |
Methods — techniques the papers use, named apart from their topics
markov decision process · 0.6deep reinforcement learning · 0.6reinforcement learning · 0.5imitation learning · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Optimizing Sequential Experimental Design with Deep Reinforcement LearningabstractBayesian approaches developed to solve the optimal design of sequential experiments are mathematically elegant but computationally challenging. Recently, techniques using amortization have been proposed to make these Bayesian approaches practical, by training a parameterized policy that proposes designs efficiently at deployment time. However, these methods may not sufficiently explore the design space, require access to a differentiable probabilistic model and can only optimize over continuous design spaces. Here, we address these limitations by showing that the problem of optimizing policies can be reduced to solving a Markov decision process (MDP). We solve the equivalent MDP with modern deep reinforcement learning techniques. Our experiments show that our approach is also computationally efficient at deployment time and exhibits state-of-the-art performance on both continuous and discrete design spaces, even when the probabilistic model is a black box. Tom Blau, Edwin V. Bonilla, Iadine Chades, Amir Dezfouli |
ICML | 1 |
| 2021 | Learning from Demonstration without Demonstrations
Tom Blau, Philippe Morere, Gilad Francis |
ICRA | 1 |
| 2018 | Improving Reinforcement Learning Pre-Training with Variational DropoutabstractReinforcement learning has been very successful at learning control policies for robotic agents in order to perform various tasks, such as driving around a track, navigating a maze, and bipedal locomotion. One significant drawback of reinforcement learning methods is that they require a large number of data points in order to learn good policies, a trait known as poor data efficiency or poor sample efficiency. One approach for improving sample efficiency is supervised pre-training of policies to directly clone the behavior of an expert, but this suffers from poor generalization far from the training data. We propose to improve this by using Gaussian dropout networks with a regularization term based on variational inference in the pre-training step. We show that this initializes policy parameters to significantly better values than standard supervised learning or random initialization, thus greatly reducing sample complexity compared with state-of-the-art methods, and enabling an RL algorithm to learn optimal policies for high-dimensional continuous control problems in a practical time frame. Tom Blau, Lionel Ott, Fabio Ramos 0001 |
IROS | 1 |