Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Carlos Florensa

dblp:199/1800 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 1 since 2021Systems, architecture and hardware · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Reinforcement learning · 61% Representation and self-supervised learning · 15% Motion planning and robot control · 6%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.722020
Sub-policy Adaptation for Hierarchical Reinforcement Learning · ICLR 2020
Stochastic Neural Networks for Hierarchical Reinforcement Learning · ICLR (Poster) 2017
Machine learning › Representation and self-supervised learning
mutual information maximization
0.512021
Which Mutual-Information Representation Learning Objectives are Sufficient for Control? · NeurIPS 2021
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning
state representation learning
0.512021
Which Mutual-Information Representation Learning Objectives are Sufficient for Control? · NeurIPS 2021
Machine learning › Reinforcement learning
model-based reinforcement learning
0.412020
Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning · ICRA 2020
Machine learning › Reinforcement learning
policy learning
0.412020
Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning · ICRA 2020
Robotics › Motion planning and robot control
robot learning
0.412020
Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning · ICRA 2020
Machine learning › Reinforcement learning › policy optimization
sample-efficient policy optimization
0.412020
Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning · ICRA 2020
Machine learning › Reinforcement learning › policy optimization
uncertainty-aware policy optimization
0.412020
Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning · ICRA 2020
Machine learning › Reinforcement learning › imitation learning
goal-conditioned imitation learning
0.412019
Goal-conditioned Imitation Learning · NeurIPS 2019
Machine learning › Reinforcement learning
goal-conditioned reinforcement learning
0.412019
Goal-conditioned Imitation Learning · NeurIPS 2019
Machine learning › Reinforcement learning
imitation learning
0.412019
Goal-conditioned Imitation Learning · NeurIPS 2019
Machine learning › Learning paradigms
curriculum learning
0.312018
Automatic Goal Generation for Reinforcement Learning Agents · ICML 2018
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › goal reasoning
goal generation
0.312018
Automatic Goal Generation for Reinforcement Learning Agents · ICML 2018
Machine learning › Deep learning architectures and training
stochastic neural network
0.312017
Stochastic Neural Networks for Hierarchical Reinforcement Learning · ICLR (Poster) 2017
Robotics › Robot manipulation › assembly
peg insertion
0.112020
Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning · ICRA 2020
Robotics › Robot manipulation
learning from demonstration
0.112019
Goal-conditioned Imitation Learning · NeurIPS 2019
Machine learning › Reinforcement learning
policy optimization
0.112019
Adaptive Variance for Changing Sparse-Reward Environments · ICRA 2019

Methods — techniques the papers use, named apart from their topics

mutual information estimation · 0.5uncertainty estimation · 0.4model-based policy · 0.4locally learned policy · 0.4value function analysis · 0.4hindsight experience replay · 0.4generative adversarial imitation learning · 0.4generative modeling · 0.3adversarial training · 0.3hierarchical reinforcement learning · 0.3
YearPublicationVenuePosition
2021 Which Mutual-Information Representation Learning Objectives are Sufficient for Control?
abstract
Mutual information (MI) maximization provides an appealing formalism for learning representations of data. In the context of reinforcement learning (RL), such representations can accelerate learning by discarding irrelevant and redundant information, while retaining the information necessary for control. Much prior work on these methods has addressed the practical difficulties of estimating MI from samples of high-dimensional observations, while comparatively less is understood about which MI objectives yield representations that are sufficient for RL from a theoretical perspective. In this paper, we formalize the sufficiency of a state representation for learning and representing the optimal policy, and study several popular MI based objectives through this lens. Surprisingly, we find that two of these objectives can yield insufficient representations given mild and common assumptions on the structure of the MDP. We corroborate our theoretical results with empirical experiments on a simulated game environment with visual observations.
Kate Rakelly, Abhishek Gupta 0004, Carlos Florensa, Sergey Levine
NeurIPS3
2020 Sub-policy Adaptation for Hierarchical Reinforcement Learning
Alexander C. Li, Carlos Florensa, Ignasi Clavera, Pieter Abbeel
ICLR2
2020 Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning
abstract
Traditional robotic approaches rely on an accurate model of the environment, a detailed description of how to perform the task, and a robust perception system to keep track of the current state. On the other hand, reinforcement learning approaches can operate directly from raw sensory inputs with only a reward signal to describe the task, but are extremely sampleinefficient and brittle. In this work, we combine the strengths of model-based methods with the flexibility of learning-based methods to obtain a general method that is able to overcome inaccuracies in the robotics perception/actuation pipeline, while requiring minimal interactions with the environment. This is achieved by leveraging uncertainty estimates to divide the space in regions where the given model-based policy is reliable, and regions where it may have flaws or not be well defined. In these uncertain regions, we show that a locally learned-policy can be used directly with raw sensory inputs. We test our algorithm, Guided Uncertainty-Aware Policy Optimization (GUAPO), on a real-world robot performing peg insertion. Videos are available at: https://sites.google.com/view/guapo-rl.
Michelle A. Lee, Carlos Florensa, Jonathan Tremblay, Nathan D. Ratliff, Animesh Garg, Fabio Ramos 0001, Dieter Fox
ICRA2
2019 Adaptive Variance for Changing Sparse-Reward Environments
abstract
Robots that are trained to perform a task in a fixed environment often fail when facing unexpected changes to the environment due to a lack of exploration. We propose a principled way to adapt the policy for better exploration in changing sparse-reward environments. Unlike previous works which explicitly model environmental changes, we analyze the relationship between the value function and the optimal exploration for a Gaussian-parameterized policy and show that our theory leads to an effective strategy for adjusting the variance of the policy, enabling fast adapt to changes in a variety of sparse-reward environments.
Pengsheng Guo, Carlos Florensa, David Held
ICRA3
2019 Goal-conditioned Imitation Learning
abstract
Designing rewards for Reinforcement Learning (RL) is challenging because it needs to convey the desired task, be efficient to optimize, and be easy to compute. The latter is particularly problematic when applying RL to robotics, where detecting whether the desired configuration is reached might require considerable supervision and instrumentation. Furthermore, we are often interested in being able to reach a wide range of configurations, hence setting up a different reward every time might be unpractical. Methods like Hindsight Experience Replay (HER) have recently shown promise to learn policies able to reach many goals, without the need of a reward. Unfortunately, without tricks like resetting to points along the trajectory, HER might require many samples to discover how to reach certain areas of the state-space. In this work we propose a novel algorithm goalGAIL, which incorporates demonstrations to drastically speed up the convergence to a policy able to reach any goal, surpassing the performance of an agent trained with other Imitation Learning algorithms. Furthermore, we show our method can also be used when the available expert trajectories do not contain the actions or when the expert is suboptimal, which makes it applicable when only kinesthetic, third person or noisy demonstration is available.
Carlos Florensa, Pieter Abbeel, Mariano Phielipp
NeurIPS2
2018 Automatic Goal Generation for Reinforcement Learning Agents
abstract
Reinforcement learning (RL) is a powerful technique to train an agent to perform a task; however, an agent that is trained using RL is only capable of achieving the single task that is specified via its reward function. Such an approach does not scale well to settings in which an agent needs to perform a diverse set of tasks, such as navigating to varying positions in a room or moving objects to varying locations. Instead, we propose a method that allows an agent to automatically discover the range of tasks that it is capable of performing in its environment. We use a generator network to propose tasks for the agent to try to accomplish, each task being specified as reaching a certain parametrized subset of the state-space. The generator network is optimized using adversarial training to produce tasks that are always at the appropriate level of difficulty for the agent, thus automatically producing a curriculum. We show that, by using this framework, an agent can efficiently and automatically learn to perform a wide set of tasks without requiring any prior knowledge of its environment, even when only sparse rewards are available. Videos and code available at https://sites.google.com/view/goalgeneration4rl.
Carlos Florensa, David Held, Xinyang Geng, Pieter Abbeel
ICML1
2017 Stochastic Neural Networks for Hierarchical Reinforcement Learning
Carlos Florensa, Yan Duan, Pieter Abbeel
ICLR (Poster)1