Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

John W. Roberts

dblp:76/6779 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
0since 2021 · last 2013
0000-0002-0870-952XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-authorSystems, architecture and hardware · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 68% Optimization for machine learning · 18% Probabilistic and Bayesian machine learning · 14%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 62% Parallel and multicore computing · 38%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
model-based reinforcement learning
0.212013
Reinforcement learning with misspecified model classes · ICRA 2013
Machine learning › Reinforcement learning
policy learning
0.212013
Reinforcement learning with misspecified model classes · ICRA 2013
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.112008
Signal-to-Noise Ratio Analysis of Policy Gradient Algorithms · NIPS 2008
Machine learning › Probabilistic and Bayesian machine learning
signal-to-noise ratio analysis
0.112008
Signal-to-Noise Ratio Analysis of Policy Gradient Algorithms · NIPS 2008
Machine learning › Optimization for machine learning
variance reduction
0.112008
Signal-to-Noise Ratio Analysis of Policy Gradient Algorithms · NIPS 2008
Machine learning › Optimization for machine learning › gradient-based optimization › gradient descent
natural gradient descent
0.012008
Signal-to-Noise Ratio Analysis of Policy Gradient Algorithms · NIPS 2008
Performance modeling and evaluation
performance instrumentation
0.011989
Hybrid Performance Measurement Instrumentation for Loosely-Cpupled MIMD Architectures · SIGMETRICS 1989
Parallel and multicore computing › multiprocessor system
loosely coupled multiprocessor
0.011989
Hybrid Performance Measurement Instrumentation for Loosely-Cpupled MIMD Architectures · SIGMETRICS 1989
Parallel and multicore computing › parallel architecture
MIMD architecture
0.011989
Hybrid Performance Measurement Instrumentation for Loosely-Cpupled MIMD Architectures · SIGMETRICS 1989

Methods — techniques the papers use, named apart from their topics

maximum likelihood estimation · 0.2batch reinforcement learning · 0.2weight perturbation · 0.1non-gaussian distributions · 0.1anisotropic sampling · 0.1software instrumentation · 0.0hardware performance counters · 0.0
YearPublicationVenuePosition
2013 Reinforcement learning with misspecified model classes
abstract
Real-world robots commonly have to act in complex, poorly understood environments where the true world dynamics are unknown. To compensate for the unknown world dynamics, we often provide a class of models to a learner so it may select a model, typically using a minimum prediction error metric over a set of training data. Often in real-world domains the model class is unable to capture the true dynamics, due to either limited domain knowledge or a desire to use a small model. In these cases we call the model class misspecified, and an unfortunate consequence of misspecification is that even with unlimited data and computation there is no guarantee the model with minimum prediction error leads to the best performing policy. In this work, our approach improves upon the standard maximum likelihood model selection metric by explicitly selecting the model which achieves the highest expected reward, rather than the most likely model. We present an algorithm for which the highest performing model from the model class is guaranteed to be found given unlimited data and computation. Empirically, we demonstrate that our algorithm is often superior to the maximum likelihood learner in a batch learning setting for two common RL benchmark problems and a third real-world system, the hydrodynamic cart-pole, a domain whose complex dynamics cannot be known exactly.
Joshua Mason Joseph, Alborz Geramifard, John W. Roberts, Jonathan P. How, Nicholas Roy
ICRA3
2011 Feedback controller parameterizations for Reinforcement Learning
abstract
Reinforcement Learning offers a very general framework for learning controllers, but its effectiveness is closely tied to the controller parameterization used. Especially when learning feedback controllers for weakly stable systems, ineffective parameterizations can result in unstable controllers and poor performance both in terms of learning convergence and in the cost of the resulting policy. In this paper we explore four linear controller parameterizations in the context of REINFORCE, applying them to the control of a reaching task with a linearized flexible manipulator. We find that some natural but naive parameterizations perform very poorly, while the Youla Parameterization (a popular parameterization from the controls literature) offers a number of robustness and performance advantages.
John W. Roberts, Ian R. Manchester, Russ Tedrake
ADPRL1
2008 Signal-to-Noise Ratio Analysis of Policy Gradient Algorithms
abstract
Policy gradient (PG) reinforcement learning algorithms have strong (local) convergence guarantees, but their learning performance is typically limited by a large variance in the estimate of the gradient. In this paper, we formulate the variance reduction problem by describing a signal-to-noise ratio (SNR) for policy gradient algorithms, and evaluate this SNR carefully for the popular Weight Perturbation (WP) algorithm. We confirm that SNR is a good predictor of long-term learning performance, and that in our episodic formulation, the cost-to-go function is indeed the optimal baseline. We then propose two modifications to traditional model-free policy gradient algorithms in order to optimize the SNR. First, we examine WP using anisotropic sampling distributions, which introduces a bias into the update but increases the SNR; this bias can be interpretted as following the natural gradient of the cost function. Second, we show that non-Gaussian distributions can also increase the SNR, and argue that the optimal isotropic distribution is a ‘shell’ distribution with a constant magnitude and uniform distribution in direction. We demonstrate that both modifications produce substantial improvements in learning performance in challenging policy gradient experiments.
John W. Roberts, Russ Tedrake
NIPS1
1989 Hybrid Performance Measurement Instrumentation for Loosely-Cpupled MIMD Architectures
John W. Roberts, John Antonishek, Alan Mink
SIGMETRICS1
1987 Hardware-Assisted Multiprocessor Performance Measurement
Alan Mink, Jesse M. Draper, John W. Roberts, Robert J. Carpenter
Performance3