EDBT 2026 Demo / reviewers in the wild / expert
John W. Roberts
dblp:76/6779
· DBLP profile ↗
5ranked-venue papers
3as first author
0since 2021 · last 2013
0000-0002-0870-952XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-authorSystems, architecture and hardware · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 68% Optimization for machine learning · 18% Probabilistic and Bayesian machine learning · 14% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 62% Parallel and multicore computing · 38% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.2 | 1 | 2013 | Reinforcement learning with misspecified model classes · ICRA 2013 |
Machine learning › Reinforcement learning
policy learning |
0.2 | 1 | 2013 | Reinforcement learning with misspecified model classes · ICRA 2013 |
Machine learning › Reinforcement learning › policy optimization
policy gradient |
0.1 | 1 | 2008 | Signal-to-Noise Ratio Analysis of Policy Gradient Algorithms · NIPS 2008 |
Machine learning › Probabilistic and Bayesian machine learning
signal-to-noise ratio analysis |
0.1 | 1 | 2008 | Signal-to-Noise Ratio Analysis of Policy Gradient Algorithms · NIPS 2008 |
Machine learning › Optimization for machine learning
variance reduction |
0.1 | 1 | 2008 | Signal-to-Noise Ratio Analysis of Policy Gradient Algorithms · NIPS 2008 |
Machine learning › Optimization for machine learning › gradient-based optimization › gradient descent
natural gradient descent |
0.0 | 1 | 2008 | Signal-to-Noise Ratio Analysis of Policy Gradient Algorithms · NIPS 2008 |
Performance modeling and evaluation
performance instrumentation |
0.0 | 1 | 1989 | Hybrid Performance Measurement Instrumentation for Loosely-Cpupled MIMD Architectures · SIGMETRICS 1989 |
Parallel and multicore computing › multiprocessor system
loosely coupled multiprocessor |
0.0 | 1 | 1989 | Hybrid Performance Measurement Instrumentation for Loosely-Cpupled MIMD Architectures · SIGMETRICS 1989 |
Parallel and multicore computing › parallel architecture
MIMD architecture |
0.0 | 1 | 1989 | Hybrid Performance Measurement Instrumentation for Loosely-Cpupled MIMD Architectures · SIGMETRICS 1989 |
Methods — techniques the papers use, named apart from their topics
maximum likelihood estimation · 0.2batch reinforcement learning · 0.2weight perturbation · 0.1non-gaussian distributions · 0.1anisotropic sampling · 0.1software instrumentation · 0.0hardware performance counters · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | Reinforcement learning with misspecified model classesabstractReal-world robots commonly have to act in complex, poorly understood environments where the true world dynamics are unknown. To compensate for the unknown world dynamics, we often provide a class of models to a learner so it may select a model, typically using a minimum prediction error metric over a set of training data. Often in real-world domains the model class is unable to capture the true dynamics, due to either limited domain knowledge or a desire to use a small model. In these cases we call the model class misspecified, and an unfortunate consequence of misspecification is that even with unlimited data and computation there is no guarantee the model with minimum prediction error leads to the best performing policy. In this work, our approach improves upon the standard maximum likelihood model selection metric by explicitly selecting the model which achieves the highest expected reward, rather than the most likely model. We present an algorithm for which the highest performing model from the model class is guaranteed to be found given unlimited data and computation. Empirically, we demonstrate that our algorithm is often superior to the maximum likelihood learner in a batch learning setting for two common RL benchmark problems and a third real-world system, the hydrodynamic cart-pole, a domain whose complex dynamics cannot be known exactly. Joshua Mason Joseph, Alborz Geramifard, John W. Roberts, Jonathan P. How, Nicholas Roy |
ICRA | 3 |
| 2011 | Feedback controller parameterizations for Reinforcement LearningabstractReinforcement Learning offers a very general framework for learning controllers, but its effectiveness is closely tied to the controller parameterization used. Especially when learning feedback controllers for weakly stable systems, ineffective parameterizations can result in unstable controllers and poor performance both in terms of learning convergence and in the cost of the resulting policy. In this paper we explore four linear controller parameterizations in the context of REINFORCE, applying them to the control of a reaching task with a linearized flexible manipulator. We find that some natural but naive parameterizations perform very poorly, while the Youla Parameterization (a popular parameterization from the controls literature) offers a number of robustness and performance advantages. John W. Roberts, Ian R. Manchester, Russ Tedrake |
ADPRL | 1 |
| 2008 | Signal-to-Noise Ratio Analysis of Policy Gradient AlgorithmsabstractPolicy gradient (PG) reinforcement learning algorithms have strong (local) convergence guarantees, but their learning performance is typically limited by a large variance in the estimate of the gradient. In this paper, we formulate the variance reduction problem by describing a signal-to-noise ratio (SNR) for policy gradient algorithms, and evaluate this SNR carefully for the popular Weight Perturbation (WP) algorithm. We confirm that SNR is a good predictor of long-term learning performance, and that in our episodic formulation, the cost-to-go function is indeed the optimal baseline. We then propose two modifications to traditional model-free policy gradient algorithms in order to optimize the SNR. First, we examine WP using anisotropic sampling distributions, which introduces a bias into the update but increases the SNR; this bias can be interpretted as following the natural gradient of the cost function. Second, we show that non-Gaussian distributions can also increase the SNR, and argue that the optimal isotropic distribution is a âshellâ distribution with a constant magnitude and uniform distribution in direction. We demonstrate that both modifications produce substantial improvements in learning performance in challenging policy gradient experiments. John W. Roberts, Russ Tedrake |
NIPS | 1 |
| 1989 | Hybrid Performance Measurement Instrumentation for Loosely-Cpupled MIMD Architectures
John W. Roberts, John Antonishek, Alan Mink |
SIGMETRICS | 1 |
| 1987 | Hardware-Assisted Multiprocessor Performance Measurement
Alan Mink, Jesse M. Draper, John W. Roberts, Robert J. Carpenter |
Performance | 3 |