Kevin S. Luck

dblp:153/7680 · also Kevin Sebastian Luck · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0003-2228-203XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 6 since 2021Systems, architecture and hardware · 4 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 40% Representation and self-supervised learning · 24% Probabilistic and Bayesian machine learning · 12%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
1.322023
Conditional Mutual Information for Disentangled Representations in Reinforcement Learning · NeurIPS 2023
Temporal Disentanglement of Representations for Improved Generalisation in Reinforcement Learning · ICLR 2023
Machine learning › Reinforcement learning › model-based reinforcement learning › model-based planning
decision-time planning
0.912025
Discrete Codebook World Models for Continuous Control · ICLR 2025
Machine learning › Representation and self-supervised learning › representation learning › discrete representation learning
discrete latent representation
0.912025
Discrete Codebook World Models for Continuous Control · ICLR 2025
Machine learning › Reinforcement learning
model-based reinforcement learning
0.912025
Discrete Codebook World Models for Continuous Control · ICLR 2025
Robotics › Motion planning and robot control › robot control
model predictive control
0.912025
Discrete Codebook World Models for Continuous Control · ICLR 2025
Machine learning › Reinforcement learning › model-based reinforcement learning
world model
0.912025
Discrete Codebook World Models for Continuous Control · ICLR 2025
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › neural processes
conditional neural process
0.712023
Practical Equivariances via Relational Conditional Neural Processes · NeurIPS 2023
Machine learning › Representation and self-supervised learning
equivariance
0.712023
Practical Equivariances via Relational Conditional Neural Processes · NeurIPS 2023
Machine learning › Deep learning architectures and training
equivariant neural network
0.712023
Practical Equivariances via Relational Conditional Neural Processes · NeurIPS 2023
Machine learning › Learning theory
generalization
0.712023
Conditional Mutual Information for Disentangled Representations in Reinforcement Learning · NeurIPS 2023
Machine learning › Reinforcement learning
generalization in reinforcement learning
0.712023
Temporal Disentanglement of Representations for Improved Generalisation in Reinforcement Learning · ICLR 2023
Machine learning › Reinforcement learning
imitation learning
0.712023
Co-imitation: Learning Design and Behaviour by Imitation · AAAI 2023
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
neural processes
0.712023
Practical Equivariances via Relational Conditional Neural Processes · NeurIPS 2023
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning
0.712023
Conditional Mutual Information for Disentangled Representations in Reinforcement Learning · NeurIPS 2023
Machine learning › Reinforcement learning
policy search
0.212016
Sparse Latent Space Policy Search · AAAI 2016
Machine learning › Reinforcement learning › sample efficiency
sample-efficient policy learning
0.212016
Sparse Latent Space Policy Search · AAAI 2016
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
amortized inference
0.212023
Practical Equivariances via Relational Conditional Neural Processes · NeurIPS 2023
Machine learning › Representation and self-supervised learning
mutual information minimization
0.212023
Conditional Mutual Information for Disentangled Representations in Reinforcement Learning · NeurIPS 2023
Robotics › Motion planning and robot control
robot learning
0.212023
Co-imitation: Learning Design and Behaviour by Imitation · AAAI 2023
Machine learning › Trustworthy machine learning
uncertainty estimation
0.212023
Practical Equivariances via Relational Conditional Neural Processes · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

model predictive control · 0.9discrete codebook · 0.9state-distribution matching · 0.7relational conditional neural process · 0.7reinforcement learning · 0.7policy learning · 0.7meta-learning · 0.7conditional mutual information · 0.7auxiliary task learning · 0.7sparsity constraint · 0.2
YearPublicationVenuePosition
2025 Discrete Codebook World Models for Continuous Control
abstract
In reinforcement learning (RL), world models serve as internal simulators, enabling agents to predict environment dynamics and future outcomes in order to make informed decisions. While previous approaches leveraging discrete latent spaces, such as DreamerV3, have demonstrated strong performance in discrete action settings and visual control tasks, their comparative performance in state-based continuous control remains underexplored. In contrast, methods with continuous latent spaces, such as TD-MPC2, have shown notable success in state-based continuous control benchmarks. In this paper, we demonstrate that modeling discrete latent states has benefits over continuous latent states and that discrete codebook encodings are more effective representations for continuous control, compared to alternative encodings, such as one-hot and label-based encodings. Based on these insights, we introduce DCWM: Discrete Codebook World Model, a self-supervised world model with a discrete and stochastic latent space, where latent states are codes from a codebook. We combine DCWM with decision-time planning to get our model-based RL algorithm, named DC-MPC: Discrete Codebook Model Predictive Control, which performs competitively against recent state-of-the-art algorithms, including TD-MPC2 and DreamerV3, on continuous control benchmarks.
Aidan Scannell, Mohammadreza Nakhaei, Kalle Kujanpää, Yi Zhao 0014, Kevin S. Luck, Arno Solin, Joni Pajarinen
ICLR5
2025 Co-Adaptation of Embodiment and Control with Self-Imitation Learning
abstract
The task of co-optimizing the body and behaviour of agents has been a long-standing problem in the fields of evolutionary robotics and embodied AI. Previous work has largely focused on the development of learning methods exploiting massive parallelization of agent evaluations with large population sizes, a paradigm which is applicable to simulated agents but cannot be transferred to the real world due to the assoicated costs with the production of embodiments and robots. Furthermore, recent data-efficient approaches utilizing reinforcement learning can suffer from distributional shifts in transition dynamics as well as in state and action spaces when experiencing new body morphologies. In this work, we propose a new co-adaptation method combining reinforcement learning and State-Aligned Self-Imitation Learning to co-design embodiment and behavioural policies withing a handful of design iterations. We show that the integration of a self-imitation signal improves the data-efficiency of the co-adaptation process as well as the behavioural recovery when adapting morphological parameters.
Sergio Hernández-Gutiérrez, Ville Kyrki, Kevin S. Luck
IROS3
2023 Co-imitation: Learning Design and Behaviour by Imitation
abstract
The co-adaptation of robots has been a long-standing research endeavour with the goal of adapting both body and behaviour of a robot for a given task, inspired by the natural evolution of animals. Co-adaptation has the potential to eliminate costly manual hardware engineering as well as improve the performance of systems. The standard approach to co-adaptation is to use a reward function for optimizing behaviour and morphology. However, defining and constructing such reward functions is notoriously difficult and often a significant engineering effort. This paper introduces a new viewpoint on the co-adaptation problem, which we call co-imitation: finding a morphology and a policy that allow an imitator to closely match the behaviour of a demonstrator. To this end we propose a co-imitation methodology for adapting behaviour and morphology by matching state-distributions of the demonstrator. Specifically, we focus on the challenging scenario with mismatched state- and action-spaces between both agents. We find that co-imitation increases behaviour similarity across a variety of tasks and settings, and demonstrate co-imitation by transferring human walking, jogging and kicking skills onto a simulated humanoid.
Chang Rajani, Karol Arndt, David Blanco Mulero, Kevin S. Luck, Ville Kyrki
AAAI4
2023 Temporal Disentanglement of Representations for Improved Generalisation in Reinforcement Learning
Mhairi Dunion, Trevor McInroe, Kevin S. Luck, Josiah Hanna, Stefano V. Albrecht
ICLR3
2023 Conditional Mutual Information for Disentangled Representations in Reinforcement Learning
abstract
Reinforcement Learning (RL) environments can produce training data with spurious correlations between features due to the amount of training data or its limited feature coverage. This can lead to RL agents encoding these misleading correlations in their latent representation, preventing the agent from generalising if the correlation changes within the environment or when deployed in the real world. Disentangled representations can improve robustness, but existing disentanglement techniques that minimise mutual information between features require independent features, thus they cannot disentangle correlated features. We propose an auxiliary task for RL algorithms that learns a disentangled representation of high-dimensional observations with correlated features by minimising the conditional mutual information between features in the representation. We demonstrate experimentally, using continuous control tasks, that our approach improves generalisation under correlation shifts, as well as improving the training performance of RL algorithms in the presence of correlated features.
Mhairi Dunion, Trevor McInroe, Kevin S. Luck, Josiah Hanna, Stefano V. Albrecht
NeurIPS3
2023 Practical Equivariances via Relational Conditional Neural Processes
abstract
Conditional Neural Processes (CNPs) are a class of metalearning models popular for combining the runtime efficiency of amortized inference with reliable uncertainty quantification. Many relevant machine learning tasks, such as in spatio-temporal modeling, Bayesian Optimization and continuous control, inherently contain equivariances – for example to translation – which the model can exploit for maximal performance. However, prior attempts to include equivariances in CNPs do not scale effectively beyond two input dimensions. In this work, we propose Relational Conditional Neural Processes (RCNPs), an effective approach to incorporate equivariances into any neural process model. Our proposed method extends the applicability and impact of equivariant neural processes to higher dimensions. We empirically demonstrate the competitive performance of RCNPs on a large array of tasks naturally containing equivariances.
Daolang Huang, Manuel Haußmann, Ulpu Remes, St John, Gregoire Clarte, Kevin S. Luck, Samuel Kaski, Luigi Acerbi
NeurIPS6
2019 Improved Exploration through Latent Trajectory Optimization in Deep Deterministic Policy Gradient
abstract
Model-free reinforcement learning algorithms such as Deep Deterministic Policy Gradient (DDPG) often require additional exploration strategies, especially if the actor is of deterministic nature. This work evaluates the use of model-based trajectory optimization methods used for exploration in Deep Deterministic Policy Gradient when trained on a latent image embedding. In addition, an extension of DDPG is derived using a value function as critic, making use of a learned deep dynamics model to compute the policy gradient. This approach leads to a symbiotic relationship between the deep reinforcement learning algorithm and the latent trajectory optimizer. The trajectory optimizer benefits from the critic learned by the RL algorithm and the latter from the enhanced exploration generated by the planner. The developed methods are evaluated on two continuous control tasks, one in simulation and one in the real world. In particular, a Baxter robot is trained to perform an insertion task, while only receiving sparse rewards and images as observations from the environment.
Kevin S. Luck, Mel Vecerík, Simon Stepputtis, Heni Ben Amor, Jonathan Scholz
IROS1
2017 Extracting bimanual synergies with reinforcement learning
abstract
Motor synergies are an important concept in human motor control. Through the co-activation of multiple muscles, complex motion involving many degrees-of-freedom can be generated. However, leveraging this concept in robotics typically entails using human data that may be incompatible for the kinematics of the robot. In this paper, our goal is to enable a robot to identify synergies for low-dimensional control using trial-and-error only. We discuss how synergies can be learned through latent space policy search and introduce an extension of the algorithm for the re-use of previously learned synergies for exploration. The application of the algorithm on a bimanual manipulation task for the Baxter robot shows that performance can be increased by reusing learned synergies intra-task when learning to lift objects. But the reuse of synergies between two tasks with different objects did not lead to a significant improvement.
Kevin S. Luck, Heni Ben Amor
IROS1
2016 Sparse Latent Space Policy Search
abstract
Computational agents often need to learn policies that involve many control variables, e.g., a robot needs to control several joints simultaneously. Learning a policy with a high number of parameters, however, usually requires a large number of training samples. We introduce a reinforcement learning method for sample-efficient policy search that exploits correlations between control variables. Such correlations are particularly frequent in motor skill learning tasks. The introduced method uses Variational Inference to estimate policy parameters, while at the same time uncovering a low-dimensional latent space of controls. Prior knowledge about the task and the structure of the learning agent can be provided by specifying groups of potentially correlated parameters. This information is then used to impose sparsity constraints on the mapping between the high-dimensional space of controls and a lower-dimensional latent space. In experiments with a simulated bi-manual manipulator, the new approach effectively identifies synergies between joints, performs efficient low-dimensional policy search, and outperforms state-of-the-art policy search methods.
Kevin S. Luck, Joni Pajarinen, Erik Berger, Ville Kyrki, Heni Ben Amor
AAAI1
2014 Latent space policy search for robotics
abstract
Learning motor skills for robots is a hard task. In particular, a high number of degrees-of-freedom in the robot can pose serious challenges to existing reinforcement learning methods, since it leads to a high-dimensional search space. However, complex robots are often intrinsically redundant systems and, therefore, can be controlled using a latent manifold of much smaller dimensionality. In this paper, we present a novel policy search method that performs efficient reinforcement learning by uncovering the low-dimensional latent space of actuator redundancies. In contrast to previous attempts at combining reinforcement learning and dimensionality reduction, our approach does not perform dimensionality reduction as a preprocessing step but naturally combines it with policy search. Our evaluations show that the new approach outperforms existing algorithms for learning motor skills with high-dimensional robots.
Kevin S. Luck, Gerhard Neumann, Erik Berger, Jan Peters 0001, Heni Ben Amor
IROS1