EDBT 2026 Demo / reviewers in the wild / expert
Carlos Florensa
dblp:199/1800
· DBLP profile ↗
7ranked-venue papers
2as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 1 since 2021Systems, architecture and hardware · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Reinforcement learning · 61% Representation and self-supervised learning · 15% Motion planning and robot control · 6% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.7 | 2 | 2020 | Sub-policy Adaptation for Hierarchical Reinforcement Learning · ICLR 2020 Stochastic Neural Networks for Hierarchical Reinforcement Learning · ICLR (Poster) 2017 |
Machine learning › Representation and self-supervised learning
mutual information maximization |
0.5 | 1 | 2021 | Which Mutual-Information Representation Learning Objectives are Sufficient for Control? · NeurIPS 2021 |
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning
state representation learning |
0.5 | 1 | 2021 | Which Mutual-Information Representation Learning Objectives are Sufficient for Control? · NeurIPS 2021 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.4 | 1 | 2020 | Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning · ICRA 2020 |
Machine learning › Reinforcement learning
policy learning |
0.4 | 1 | 2020 | Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning · ICRA 2020 |
Robotics › Motion planning and robot control
robot learning |
0.4 | 1 | 2020 | Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning · ICRA 2020 |
Machine learning › Reinforcement learning › policy optimization
sample-efficient policy optimization |
0.4 | 1 | 2020 | Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning · ICRA 2020 |
Machine learning › Reinforcement learning › policy optimization
uncertainty-aware policy optimization |
0.4 | 1 | 2020 | Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning · ICRA 2020 |
Machine learning › Reinforcement learning › imitation learning
goal-conditioned imitation learning |
0.4 | 1 | 2019 | Goal-conditioned Imitation Learning · NeurIPS 2019 |
Machine learning › Reinforcement learning
goal-conditioned reinforcement learning |
0.4 | 1 | 2019 | Goal-conditioned Imitation Learning · NeurIPS 2019 |
Machine learning › Reinforcement learning
imitation learning |
0.4 | 1 | 2019 | Goal-conditioned Imitation Learning · NeurIPS 2019 |
Machine learning › Learning paradigms
curriculum learning |
0.3 | 1 | 2018 | Automatic Goal Generation for Reinforcement Learning Agents · ICML 2018 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › goal reasoning
goal generation |
0.3 | 1 | 2018 | Automatic Goal Generation for Reinforcement Learning Agents · ICML 2018 |
Machine learning › Deep learning architectures and training
stochastic neural network |
0.3 | 1 | 2017 | Stochastic Neural Networks for Hierarchical Reinforcement Learning · ICLR (Poster) 2017 |
Robotics › Robot manipulation › assembly
peg insertion |
0.1 | 1 | 2020 | Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning · ICRA 2020 |
Robotics › Robot manipulation
learning from demonstration |
0.1 | 1 | 2019 | Goal-conditioned Imitation Learning · NeurIPS 2019 |
Machine learning › Reinforcement learning
policy optimization |
0.1 | 1 | 2019 | Adaptive Variance for Changing Sparse-Reward Environments · ICRA 2019 |
Methods — techniques the papers use, named apart from their topics
mutual information estimation · 0.5uncertainty estimation · 0.4model-based policy · 0.4locally learned policy · 0.4value function analysis · 0.4hindsight experience replay · 0.4generative adversarial imitation learning · 0.4generative modeling · 0.3adversarial training · 0.3hierarchical reinforcement learning · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Which Mutual-Information Representation Learning Objectives are Sufficient for Control?abstractMutual information (MI) maximization provides an appealing formalism for learning representations of data. In the context of reinforcement learning (RL), such representations can accelerate learning by discarding irrelevant and redundant information, while retaining the information necessary for control. Much prior work on these methods has addressed the practical difficulties of estimating MI from samples of high-dimensional observations, while comparatively less is understood about which MI objectives yield representations that are sufficient for RL from a theoretical perspective. In this paper, we formalize the sufficiency of a state representation for learning and representing the optimal policy, and study several popular MI based objectives through this lens. Surprisingly, we find that two of these objectives can yield insufficient representations given mild and common assumptions on the structure of the MDP. We corroborate our theoretical results with empirical experiments on a simulated game environment with visual observations. Kate Rakelly, Abhishek Gupta 0004, Carlos Florensa, Sergey Levine |
NeurIPS | 3 |
| 2020 | Sub-policy Adaptation for Hierarchical Reinforcement Learning
Alexander C. Li, Carlos Florensa, Ignasi Clavera, Pieter Abbeel |
ICLR | 2 |
| 2020 | Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy LearningabstractTraditional robotic approaches rely on an accurate model of the environment, a detailed description of how to perform the task, and a robust perception system to keep track of the current state. On the other hand, reinforcement learning approaches can operate directly from raw sensory inputs with only a reward signal to describe the task, but are extremely sampleinefficient and brittle. In this work, we combine the strengths of model-based methods with the flexibility of learning-based methods to obtain a general method that is able to overcome inaccuracies in the robotics perception/actuation pipeline, while requiring minimal interactions with the environment. This is achieved by leveraging uncertainty estimates to divide the space in regions where the given model-based policy is reliable, and regions where it may have flaws or not be well defined. In these uncertain regions, we show that a locally learned-policy can be used directly with raw sensory inputs. We test our algorithm, Guided Uncertainty-Aware Policy Optimization (GUAPO), on a real-world robot performing peg insertion. Videos are available at: https://sites.google.com/view/guapo-rl. Michelle A. Lee, Carlos Florensa, Jonathan Tremblay, Nathan D. Ratliff, Animesh Garg, Fabio Ramos 0001, Dieter Fox |
ICRA | 2 |
| 2019 | Adaptive Variance for Changing Sparse-Reward EnvironmentsabstractRobots that are trained to perform a task in a fixed environment often fail when facing unexpected changes to the environment due to a lack of exploration. We propose a principled way to adapt the policy for better exploration in changing sparse-reward environments. Unlike previous works which explicitly model environmental changes, we analyze the relationship between the value function and the optimal exploration for a Gaussian-parameterized policy and show that our theory leads to an effective strategy for adjusting the variance of the policy, enabling fast adapt to changes in a variety of sparse-reward environments. Pengsheng Guo, Carlos Florensa, David Held |
ICRA | 3 |
| 2019 | Goal-conditioned Imitation LearningabstractDesigning rewards for Reinforcement Learning (RL) is challenging because it needs to convey the desired task, be efficient to optimize, and be easy to compute. The latter is particularly problematic when applying RL to robotics, where detecting whether the desired configuration is reached might require considerable supervision and instrumentation. Furthermore, we are often interested in being able to reach a wide range of configurations, hence setting up a different reward every time might be unpractical. Methods like Hindsight Experience Replay (HER) have recently shown promise to learn policies able to reach many goals, without the need of a reward. Unfortunately, without tricks like resetting to points along the trajectory, HER might require many samples to discover how to reach certain areas of the state-space. In this work we propose a novel algorithm goalGAIL, which incorporates demonstrations to drastically speed up the convergence to a policy able to reach any goal, surpassing the performance of an agent trained with other Imitation Learning algorithms. Furthermore, we show our method can also be used when the available expert trajectories do not contain the actions or when the expert is suboptimal, which makes it applicable when only kinesthetic, third person or noisy demonstration is available. Carlos Florensa, Pieter Abbeel, Mariano Phielipp |
NeurIPS | 2 |
| 2018 | Automatic Goal Generation for Reinforcement Learning AgentsabstractReinforcement learning (RL) is a powerful technique to train an agent to perform a task; however, an agent that is trained using RL is only capable of achieving the single task that is specified via its reward function. Such an approach does not scale well to settings in which an agent needs to perform a diverse set of tasks, such as navigating to varying positions in a room or moving objects to varying locations. Instead, we propose a method that allows an agent to automatically discover the range of tasks that it is capable of performing in its environment. We use a generator network to propose tasks for the agent to try to accomplish, each task being specified as reaching a certain parametrized subset of the state-space. The generator network is optimized using adversarial training to produce tasks that are always at the appropriate level of difficulty for the agent, thus automatically producing a curriculum. We show that, by using this framework, an agent can efficiently and automatically learn to perform a wide set of tasks without requiring any prior knowledge of its environment, even when only sparse rewards are available. Videos and code available at https://sites.google.com/view/goalgeneration4rl. Carlos Florensa, David Held, Xinyang Geng, Pieter Abbeel |
ICML | 1 |
| 2017 | Stochastic Neural Networks for Hierarchical Reinforcement Learning
Carlos Florensa, Yan Duan, Pieter Abbeel |
ICLR (Poster) | 1 |