Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Risto Vuorio

dblp:222/2614 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 57% Motion planning and robot control · 14% Transfer learning and domain adaptation · 14%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
meta-reinforcement learning
1.432023
Discovering General Reinforcement Learning Algorithms with Adversarial Environment Design · NeurIPS 2023
Recurrent Hypernetworks are Surprisingly Strong in Meta-RL · NeurIPS 2023
Multimodal Model-Agnostic Meta-Learning via Task-Aware Modulation · NeurIPS 2019
Robotics › Motion planning and robot control
robot control
1.122025
Action-Constrained Imitation Learning · ICML 2025
Distilling Morphology-Conditioned Hypernetworks for Efficient Universal Morphology Control · ICML 2024
Machine learning › Reinforcement learning › safe reinforcement learning
constrained policy learning
0.912025
Action-Constrained Imitation Learning · ICML 2025
Machine learning › Reinforcement learning
imitation learning
0.912025
Action-Constrained Imitation Learning · ICML 2025
Machine learning › Efficient and distributed learning
model compression
0.812024
Distilling Morphology-Conditioned Hypernetworks for Efficient Universal Morphology Control · ICML 2024
Machine learning › Reinforcement learning › transfer learning in reinforcement learning › policy transfer
policy distillation
0.812024
Distilling Morphology-Conditioned Hypernetworks for Efficient Universal Morphology Control · ICML 2024
Robotics › Motion planning and robot control
robot learning
0.812024
Distilling Morphology-Conditioned Hypernetworks for Efficient Universal Morphology Control · ICML 2024
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot transfer
0.712023
Recurrent Hypernetworks are Surprisingly Strong in Meta-RL · NeurIPS 2023
Machine learning › Deep learning architectures and training
hypernetwork
0.712023
Recurrent Hypernetworks are Surprisingly Strong in Meta-RL · NeurIPS 2023
Machine learning › Reinforcement learning › reinforcement learning environment › environment design
unsupervised environment design
0.712023
Discovering General Reinforcement Learning Algorithms with Adversarial Environment Design · NeurIPS 2023
Machine learning › Reinforcement learning › multi-agent reinforcement learning › credit assignment
temporal credit assignment
0.612022
Adaptive Pairwise Weights for Temporal Credit Assignment · AAAI 2022
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
auxiliary tasks
0.512021
Learning State Representations from Random Deep Action-conditional Predictions · NeurIPS 2021
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning
0.512021
Learning State Representations from Random Deep Action-conditional Predictions · NeurIPS 2021
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning
state representation learning
0.512021
Learning State Representations from Random Deep Action-conditional Predictions · NeurIPS 2021
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.412019
Multimodal Model-Agnostic Meta-Learning via Task-Aware Modulation · NeurIPS 2019
Machine learning › Reinforcement learning
fleet management
0.412019
Deep Reinforcement Learning for Multi-driver Vehicle Dispatching and Repositioning Problem · ICDM 2019
Machine learning › Transfer learning and domain adaptation
meta-learning
0.412019
Multimodal Model-Agnostic Meta-Learning via Task-Aware Modulation · NeurIPS 2019
Machine learning › Transfer learning and domain adaptation › meta-learning › gradient-based meta-learning
model-agnostic meta-learning
0.412019
Multimodal Model-Agnostic Meta-Learning via Task-Aware Modulation · NeurIPS 2019
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.412019
Deep Reinforcement Learning for Multi-driver Vehicle Dispatching and Repositioning Problem · ICDM 2019
Machine learning › Reinforcement learning
actor-critic methods
0.112021
Learning State Representations from Random Deep Action-conditional Predictions · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

hypernetwork · 1.4trajectory alignment · 0.9model predictive control · 0.9dynamic time warping · 0.9transformer · 0.8knowledge distillation · 0.8regret approximation · 0.7recurrent network · 0.7end-to-end learning · 0.7adversarial environment design · 0.7
YearPublicationVenuePosition
2025 Action-Constrained Imitation Learning
abstract
Policy learning under action constraints plays a central role in ensuring safe behaviors in various robot control and resource allocation applications. In this paper, we study a new problem setting termed Action-Constrained Imitation Learning (ACIL), where an action-constrained imitator aims to learn from a demonstrative expert with larger action space. The fundamental challenge of ACIL lies in the unavoidable mismatch of occupancy measure between the expert and the imitator caused by the action constraints. We tackle this mismatch through trajectory alignment and propose DTWIL, which replaces the original expert demonstrations with a surrogate dataset that follows similar state trajectories while adhering to the action constraints. Specifically, we recast trajectory alignment as a planning problem and solve it via Model Predictive Control, which aligns the surrogate trajectories with the expert trajectories based on the Dynamic Time Warping (DTW) distance. Through extensive experiments, we demonstrate that learning from the dataset generated by DTWIL significantly enhances performance across multiple robot control tasks and outperforms various benchmark imitation learning algorithms in terms of sample efficiency.
Chia-Han Yeh, Tse-Sheng Nan, Risto Vuorio, Wei Hung, Hung-Yen Wu, Shao-Hua Sun, Ping-Chun Hsieh
ICML3
2025 IGDrivSim: A Benchmark for the Imitation Gap in Autonomous Driving
abstract
Developing autonomous vehicles that can navigate complex environments with human-level safety and efficiency is a central goal in self-driving research. A common approach to achieving this is imitation learning, where agents are trained to mimic human expert demonstrations collected from real- world driving scenarios. However, discrepancies between human perception and the self-driving car's sensors can introduce an imitation gap, leading to imitation learning failures. In this work, we introduce IGDrivSim, a benchmark built on top of the Waymax simulator, designed to investigate the effects of the imitation gap in learning autonomous driving policy from human expert demonstrations. Our experiments show that this perception gap between human experts and selfdriving agents can hinder the learning of safe and effective driving behaviors. We further show that combining imitation with reinforcement learning, using a simple penalty reward for prohibited behaviors, effectively mitigates these failures. All code developed for this work is released as open source1.
Clémence Grislain, Risto Vuorio, Cong Lu, Shimon Whiteson
IROS2
2024 Distilling Morphology-Conditioned Hypernetworks for Efficient Universal Morphology Control
abstract
Learning a universal policy across different robot morphologies can significantly improve learning efficiency and enable zero-shot generalization to unseen morphologies. However, learning a highly performant universal policy requires sophisticated architectures like transformers (TF) that have larger memory and computational cost than simpler multi-layer perceptrons (MLP). To achieve both good performance like TF and high efficiency like MLP at inference time, we propose HyperDistill, which consists of: (1) A morphology-conditioned hypernetwork (HN) that generates robot-wise MLP policies, and (2) A policy distillation approach that is essential for successful training. We show that on UNIMAL, a benchmark with hundreds of diverse morphologies, HyperDistill performs as well as a universal TF teacher policy on both training and unseen test robots, but reduces model size by 6-14 times, and computational cost by 67-160 times in different environments. Our analysis attributes the efficiency advantage of HyperDistill at inference time to knowledge decoupling, i.e., the ability to decouple inter-task and intra-task knowledge, a general principle that could also be applied to improve inference efficiency in other domains. The code is publicly available at https://github.com/MasterXiong/Universal-Morphology-Control.
Zheng Xiong, Risto Vuorio, Jacob Beck, Matthieu Zimmer, Kun Shao, Shimon Whiteson
ICML2
2023 Recurrent Hypernetworks are Surprisingly Strong in Meta-RL
abstract
Deep reinforcement learning (RL) is notoriously impractical to deploy due to sample inefficiency. Meta-RL directly addresses this sample inefficiency by learning to perform few-shot learning when a distribution of related tasks is available for meta-training. While many specialized meta-RL methods have been proposed, recent work suggests that end-to-end learning in conjunction with an off-the-shelf sequential model, such as a recurrent network, is a surprisingly strong baseline. However, such claims have been controversial due to limited supporting evidence, particularly in the face of prior work establishing precisely the opposite. In this paper, we conduct an empirical investigation. While we likewise find that a recurrent network can achieve strong performance, we demonstrate that the use of hypernetworks is crucial to maximizing their potential. Surprisingly, when combined with hypernetworks, the recurrent baselines that are far simpler than existing specialized methods actually achieve the strongest performance of all methods evaluated. We provide code at https://github.com/jacooba/hyper.
Jacob Beck, Risto Vuorio, Zheng Xiong, Shimon Whiteson
NeurIPS2
2023 Discovering General Reinforcement Learning Algorithms with Adversarial Environment Design
abstract
The past decade has seen vast progress in deep reinforcement learning (RL) on the back of algorithms manually designed by human researchers. Recently, it has been shown that it is possible to meta-learn update rules, with the hope of discovering algorithms that can perform well on a wide range of RL tasks. Despite impressive initial results from algorithms such as Learned Policy Gradient (LPG), there remains a generalization gap when these algorithms are applied to unseen environments. In this work, we examine how characteristics of the meta-training distribution impact the generalization performance of these algorithms. Motivated by this analysis and building on ideas from Unsupervised Environment Design (UED), we propose a novel approach for automatically generating curricula to maximize the regret of a meta-learned optimizer, in addition to a novel approximation of regret, which we name algorithmic regret (AR). The result is our method, General RL Optimizers Obtained Via Environment Design (GROOVE). In a series of experiments, we show that GROOVE achieves superior generalization to LPG, and evaluate AR against baseline metrics from UED, identifying it as a critical component of environment design in this setting. We believe this approach is a step towards the discovery of truly general RL algorithms, capable of solving a wide range of real-world environments.
Matthew Thomas Jackson, Minqi Jiang, Jack Parker-Holder, Risto Vuorio, Chris Lu 0001, Gregory Farquhar, Shimon Whiteson, Jakob N. Foerster
NeurIPS4
2022 Adaptive Pairwise Weights for Temporal Credit Assignment
abstract
How much credit (or blame) should an action taken in a state get for a future reward? This is the fundamental temporal credit assignment problem in Reinforcement Learning (RL). One of the earliest and still most widely used heuristics is to assign this credit based on a scalar coefficient, lambda (treated as a hyperparameter), raised to the power of the time interval between the state-action and the reward. In this empirical paper, we explore heuristics based on more general pairwise weightings that are functions of the state in which the action was taken, the state at the time of the reward, as well as the time interval between the two. Of course it isn't clear what these pairwise weight functions should be, and because they are too complex to be treated as hyperparameters we develop a metagradient procedure for learning these weight functions during the usual RL training of a policy. Our empirical work shows that it is often possible to learn these pairwise weight functions during learning of the policy to achieve better performance than competing approaches.
Risto Vuorio, Richard L. Lewis, Satinder Singh 0001
AAAI2
2021 Learning State Representations from Random Deep Action-conditional Predictions
abstract
Our main contribution in this work is an empirical finding that random General Value Functions (GVFs), i.e., deep action-conditional predictions---random both in what feature of observations they predict as well as in the sequence of actions the predictions are conditioned upon---form good auxiliary tasks for reinforcement learning (RL) problems. In particular, we show that random deep action-conditional predictions when used as auxiliary tasks yield state representations that produce control performance competitive with state-of-the-art hand-crafted auxiliary tasks like value prediction, pixel control, and CURL in both Atari and DeepMind Lab tasks. In another set of experiments we stop the gradients from the RL part of the network to the state representation learning part of the network and show, perhaps surprisingly, that the auxiliary tasks alone are sufficient to learn state representations good enough to outperform an end-to-end trained actor-critic baseline. We opensourced our code at https://github.com/Hwhitetooth/random_gvfs.
Vivek Veeriah, Risto Vuorio, Richard L. Lewis, Satinder Singh 0001
NeurIPS3
2019 Deep Reinforcement Learning for Multi-driver Vehicle Dispatching and Repositioning Problem
abstract
Order dispatching and driver repositioning (also known as fleet management) in the face of spatially and temporally varying supply and demand are central to a ride-sharing platform marketplace. Hand-crafting heuristic solutions that account for the dynamics in these resource allocation problems is difficult, and may be better handled by an end-to-end machine learning method. Previous works have explored machine learning methods to the problem from a high-level perspective, where the learning method is responsible for either repositioning the drivers or dispatching orders, and as a further simplification, the drivers are considered independent agents maximizing their own reward functions. In this paper we present a deep reinforcement learning approach for tackling the full fleet management and dispatching problems. In addition to treating the drivers as individual agents, we consider the problem from a system-centric perspective, where a central fleet management agent is responsible for decision-making for all drivers.
John Holler, Risto Vuorio, Zhiwei (Tony) Qin, Xiaocheng Tang, Yan Jiao, Tiancheng Jin, Satinder Singh 0001, Jieping Ye
ICDM2
2019 Multimodal Model-Agnostic Meta-Learning via Task-Aware Modulation
abstract
Model-agnostic meta-learners aim to acquire meta-learned parameters from similar tasks to adapt to novel tasks from the same distribution with few gradient updates. With the flexibility in the choice of models, those frameworks demonstrate appealing performance on a variety of domains such as few-shot image classification and reinforcement learning. However, one important limitation of such frameworks is that they seek a common initialization shared across the entire task distribution, substantially limiting the diversity of the task distributions that they are able to learn from. In this paper, we augment MAML with the capability to identify the mode of tasks sampled from a multimodal task distribution and adapt quickly through gradient updates. Specifically, we propose a multimodal MAML (MMAML) framework, which is able to modulate its meta-learned prior parameters according to the identified mode, allowing more efficient fast adaptation. We evaluate the proposed model on a diverse set of few-shot learning tasks, including regression, image classification, and reinforcement learning. The results not only demonstrate the effectiveness of our model in modulating the meta-learned prior in response to the characteristics of tasks but also show that training on a multimodal distribution can produce an improvement over unimodal training. The code for this project is publicly available at https://vuoristo.github.io/MMAML.
Risto Vuorio, Shao-Hua Sun, Hexiang Hu, Joseph J. Lim
NeurIPS1