Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

William Montgomery

dblp:183/6382 · also William H. Montgomery · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
0since 2021 · last 2017
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 72% Motion planning and robot control · 15% Optimization for machine learning · 13%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › policy search
guided policy search
0.522017
Reset-free guided policy search: Efficient deep reinforcement learning with stochastic initial states · ICRA 2017
Guided Policy Search via Approximate Mirror Descent · NIPS 2016
Machine learning › Reinforcement learning
policy search
0.522017
Reset-free guided policy search: Efficient deep reinforcement learning with stochastic initial states · ICRA 2017
Guided Policy Search via Approximate Mirror Descent · NIPS 2016
Machine learning › Reinforcement learning › deep reinforcement learning
deep reinforcement learning for manipulation
0.312017
Reset-free guided policy search: Efficient deep reinforcement learning with stochastic initial states · ICRA 2017
Robotics › Motion planning and robot control
robot learning
0.312017
Reset-free guided policy search: Efficient deep reinforcement learning with stochastic initial states · ICRA 2017
Machine learning › Optimization for machine learning
mirror descent
0.212016
Guided Policy Search via Approximate Mirror Descent · NIPS 2016

Methods — techniques the papers use, named apart from their topics

neural network policy · 0.3guided policy search · 0.3trajectory optimization · 0.2supervised learning · 0.2mirror descent · 0.2
YearPublicationVenuePosition
2017 Reset-free guided policy search: Efficient deep reinforcement learning with stochastic initial states
abstract
Autonomous learning of robotic skills can allow general-purpose robots to learn wide behavioral repertoires without extensive manual engineering. However, robotic skill learning must typically make trade-offs to enable practical real-world learning, such as requiring manually designed policy or value function representations, initialization from human demonstrations, instrumentation of the training environment, or extremely long training times. We propose a new reinforcement learning algorithm that can train general-purpose neural network policies with minimal human engineering, while still allowing for fast, efficient learning in stochastic environments. We build on the guided policy search (GPS) algorithm, which transforms the reinforcement learning problem into supervised learning from a computational teacher (without human demonstrations). In contrast to prior GPS methods, which require a consistent set of initial states to which the system must be reset after each episode, our approach can handle random initial states, allowing it to be used even when deterministic resets are impossible. We compare our method to existing policy search algorithms in simulation, showing that it can train high-dimensional neural network policies with the same sample efficiency as prior GPS methods, and can learn policies directly from image pixels. We also present real-world robot results that show that our method can learn manipulation policies with visual features and random initial states.
William Montgomery, Anurag Ajay, Chelsea Finn, Pieter Abbeel, Sergey Levine
ICRA1
2016 Guided Policy Search via Approximate Mirror Descent
abstract
Guided policy search algorithms can be used to optimize complex nonlinear policies, such as deep neural networks, without directly computing policy gradients in the high-dimensional parameter space. Instead, these methods use supervised learning to train the policy to mimic a “teacher” algorithm, such as a trajectory optimizer or a trajectory-centric reinforcement learning method. Guided policy search methods provide asymptotic local convergence guarantees by construction, but it is not clear how much the policy improves within a small, finite number of iterations. We show that guided policy search algorithms can be interpreted as an approximate variant of mirror descent, where the projection onto the constraint manifold is not exact. We derive a new guided policy search algorithm that is simpler and provides appealing improvement and convergence guarantees in simplified convex and linear settings, and show that in the more general nonlinear setting, the error in the projection step can be bounded. We provide empirical results on several simulated robotic manipulation tasks that show that our method is stable and achieves similar or better performance when compared to prior guided policy search methods, with a simpler formulation and fewer hyperparameters.
William Montgomery, Sergey Levine
NIPS1