Fabio Pardo

dblp:210/1005 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 68% Legged, aerial and field robots · 10% Language models and text generation · 10%

Topics — the 19 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
imitation learning
0.912025
LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations · ICML 2025
Machine learning › Reinforcement learning › imitation learning › few-shot imitation learning
in-context imitation learning
0.912025
LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations · ICML 2025
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context language model
0.912025
LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations · ICML 2025
Machine learning › Reinforcement learning
value-based reinforcement learning
0.822020
Scaling All-Goals Updates in Reinforcement Learning Using Convolutional Neural Networks · AAAI 2020
Action Branching Architectures for Deep Reinforcement Learning · AAAI 2018
Machine learning › Reinforcement learning
exploration
0.412020
Scaling All-Goals Updates in Reinforcement Learning Using Convolutional Neural Networks · AAAI 2020
Machine learning › Reinforcement learning
goal-conditioned reinforcement learning
0.412020
Scaling All-Goals Updates in Reinforcement Learning Using Convolutional Neural Networks · AAAI 2020
Machine learning › Reinforcement learning › exploration › directed exploration
goal-directed exploration
0.412020
Scaling All-Goals Updates in Reinforcement Learning Using Convolutional Neural Networks · AAAI 2020
Robotics › Legged, aerial and field robots › legged robots
humanoid locomotion
0.412020
CoMic: Complementary Task Learning & Mimicry for Reusable Skills · ICML 2020
Robotics › Legged, aerial and field robots
locomotion
0.412020
CoMic: Complementary Task Learning & Mimicry for Reusable Skills · ICML 2020
Computer vision › 3D vision
motion capture
0.412020
CoMic: Complementary Task Learning & Mimicry for Reusable Skills · ICML 2020
Robotics › Robot manipulation › learning from demonstration
motion imitation
0.412020
CoMic: Complementary Task Learning & Mimicry for Reusable Skills · ICML 2020
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning
0.412020
Scaling All-Goals Updates in Reinforcement Learning Using Convolutional Neural Networks · AAAI 2020
Machine learning › Reinforcement learning › hierarchical reinforcement learning › skill learning
reusable skill learning
0.412020
CoMic: Complementary Task Learning & Mimicry for Reusable Skills · ICML 2020
Machine learning › Reinforcement learning › hierarchical reinforcement learning
skill learning
0.412020
CoMic: Complementary Task Learning & Mimicry for Reusable Skills · ICML 2020
Natural language and speech › Information extraction and text analysis
bootstrapping
0.312018
Time Limits in Reinforcement Learning · ICML 2018
Machine learning › Reinforcement learning
deep reinforcement learning
0.312018
Action Branching Architectures for Deep Reinforcement Learning · AAAI 2018
Machine learning › Reinforcement learning
markov decision process
0.312018
Time Limits in Reinforcement Learning · ICML 2018
Machine learning › Reinforcement learning
value function estimation
0.312018
Time Limits in Reinforcement Learning · ICML 2018
Machine learning › Reinforcement learning › off-policy reinforcement learning
experience replay
0.112018
Time Limits in Reinforcement Learning · ICML 2018

Methods — techniques the papers use, named apart from their topics

chain-of-thought prompting · 0.9reinforcement learning · 0.4motion capture imitation · 0.4convolutional neural network · 0.4temporal difference learning · 0.3
YearPublicationVenuePosition
2025 LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations
abstract
In this paper, we present a benchmark to pressure-test today’s frontier models’ multimodal decision-making capabilities in the very long-context regime (up to one million tokens) and investigate whether these models can learn from large numbers of expert demonstrations in their context. We evaluate the performance of Claude 3.5 Sonnet, Gemini 1.5 Flash, Gemini 1.5 Pro, Gemini 2.0 Flash Experimental, GPT-4o, o1-mini, o1-preview, and o1 as policies across a battery of simple interactive decision-making tasks: playing tic-tac-toe, chess, and Atari, navigating grid worlds, solving crosswords, and controlling a simulated cheetah. We study increasing amounts of expert demonstrations in the context — from no demonstrations to 512 full episodes. Across our tasks, models rarely manage to fully reach expert performance, and often, presenting more demonstrations has little effect. Some models steadily improve with more demonstrations on a few tasks. We investigate the effect of encoding observations as text or images and the impact of chain-of-thought prompting. To help quantify the impact of other approaches and future innovations, we open source our benchmark that covers the zero-, few-, and many-shot regimes in a unified evaluation.
Anian Ruoss, Fabio Pardo, Harris Chan, Bonnie Li, Volodymyr Mnih, Tim Genewein
ICML2
2020 Scaling All-Goals Updates in Reinforcement Learning Using Convolutional Neural Networks
abstract
Being able to reach any desired location in the environment can be a valuable asset for an agent. Learning a policy to navigate between all pairs of states individually is often not feasible. An all-goals updating algorithm uses each transition to learn Q-values towards all goals simultaneously and off-policy. However the expensive numerous updates in parallel limited the approach to small tabular cases so far. To tackle this problem we propose to use convolutional network architectures to generate Q-values and updates for a large number of goals at once. We demonstrate the accuracy and generalization qualities of the proposed method on randomly generated mazes and Sokoban puzzles. In the case of on-screen goal coordinates the resulting mapping from frames to distance-maps directly informs the agent about which places are reachable and in how many steps. As an example of application we show that replacing the random actions in ε-greedy exploration by several actions towards feasible goals generates better exploratory trajectories on Montezuma's Revenge and Super Mario All-Stars games.
Fabio Pardo, Vitaly Levdik, Petar Kormushev
AAAI1
2020 CoMic: Complementary Task Learning & Mimicry for Reusable Skills
abstract
Learning to control complex bodies and reuse learned behaviors is a longstanding challenge in continuous control. We study the problem of learning reusable humanoid skills by imitating motion capture data and joint training with complementary tasks. We show that it is possible to learn reusable skills through reinforcement learning on 50 times more motion capture data than prior work. We systematically compare a variety of different network architectures across different data regimes both in terms of imitation performance as well as transfer to challenging locomotion tasks. Finally we show that it is possible to interleave the motion capture tracking with training on complementary tasks, enriching the resulting skill space, and enabling the reuse of skills not well covered by the motion capture data such as getting up from the ground or catching a ball.
Leonard Hasenclever, Fabio Pardo, Raia Hadsell, Nicolas Heess, Josh Merel
ICML2
2018 Action Branching Architectures for Deep Reinforcement Learning
abstract
Discrete-action algorithms have been central to numerous recent successes of deep reinforcement learning. However, applying these algorithms to high-dimensional action tasks requires tackling the combinatorial increase of the number of possible actions with the number of action dimensions. This problem is further exacerbated for continuous-action tasks that require fine control of actions via discretization. In this paper, we propose a novel neural architecture featuring a shared decision module followed by several network branches, one for each action dimension. This approach achieves a linear increase of the number of network outputs with the number of degrees of freedom by allowing a level of independence for each individual action dimension. To illustrate the approach, we present a novel agent, called Branching Dueling Q-Network (BDQ), as a branching variant of the Dueling Double Deep Q-Network (Dueling DDQN). We evaluate the performance of our agent on a set of challenging continuous control tasks. The empirical results show that the proposed agent scales gracefully to environments with increasing action dimensionality and indicate the significance of the shared decision module in coordination of the distributed action branches. Furthermore, we show that the proposed agent performs competitively against a state-of-the-art continuous control algorithm, Deep Deterministic Policy Gradient (DDPG).
Arash Tavakoli, Fabio Pardo, Petar Kormushev
AAAI2
2018 Time Limits in Reinforcement Learning
abstract
In reinforcement learning, it is common to let an agent interact for a fixed amount of time with its environment before resetting it and repeating the process in a series of episodes. The task that the agent has to learn can either be to maximize its performance over (i) that fixed period, or (ii) an indefinite period where time limits are only used during training to diversify experience. In this paper, we provide a formal account for how time limits could effectively be handled in each of the two cases and explain why not doing so can cause state-aliasing and invalidation of experience replay, leading to suboptimal policies and training instability. In case (i), we argue that the terminations due to time limits are in fact part of the environment, and thus a notion of the remaining time should be included as part of the agent’s input to avoid violation of the Markov property. In case (ii), the time limits are not part of the environment and are only used to facilitate learning. We argue that this insight should be incorporated by bootstrapping from the value of the state at the end of each partial episode. For both cases, we illustrate empirically the significance of our considerations in improving the performance and stability of existing reinforcement learning algorithms, showing state-of-the-art results on several control tasks.
Fabio Pardo, Arash Tavakoli, Vitaly Levdik, Petar Kormushev
ICML1