Daniel Palenicek

dblp:267/9480 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-8292-1318ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 65% Probabilistic and Bayesian machine learning · 16% Robot manipulation · 7%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
sample efficiency
1.622025
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization · NeurIPS 2025
CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity · ICLR 2024
Machine learning › Reinforcement learning
continuous control
0.912025
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.912025
DIME: Diffusion-Based Maximum Entropy Reinforcement Learning · ICML 2025
Robotics › Robot manipulation
diffusion policy
0.912025
DIME: Diffusion-Based Maximum Entropy Reinforcement Learning · ICML 2025
Machine learning › Reinforcement learning
maximum entropy reinforcement learning
0.912025
DIME: Diffusion-Based Maximum Entropy Reinforcement Learning · ICML 2025
Machine learning › Reinforcement learning
model-free reinforcement learning
0.912025
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization · NeurIPS 2025
Machine learning › Reinforcement learning
off-policy reinforcement learning
0.912025
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization · NeurIPS 2025
Machine learning › Reinforcement learning › policy learning
policy parameterization
0.912025
DIME: Diffusion-Based Maximum Entropy Reinforcement Learning · ICML 2025
Machine learning › Deep learning architectures and training › normalization
batch normalization
0.812024
CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity · ICLR 2024
Machine learning › Reinforcement learning
deep reinforcement learning
0.812024
CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity · ICLR 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference
approximate bayesian computation
0.712023
Pseudo-Likelihood Inference · NeurIPS 2023
Machine learning › Reinforcement learning
model-based reinforcement learning
0.712023
Diminishing Return of Value Expansion Methods in Model-Based Reinforcement Learning · ICLR 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference › simulation-based inference
neural posterior estimation
0.712023
Pseudo-Likelihood Inference · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference
simulation-based inference
0.712023
Pseudo-Likelihood Inference · NeurIPS 2023
Machine learning › Reinforcement learning › model-based reinforcement learning
value expansion
0.712023
Diminishing Return of Value Expansion Methods in Model-Based Reinforcement Learning · ICLR 2023
Machine learning › Reinforcement learning › dynamic programming
policy iteration
0.312025
DIME: Diffusion-Based Maximum Entropy Reinforcement Learning · ICML 2025

Methods — techniques the papers use, named apart from their topics

batch normalization · 1.6weight normalization · 0.9policy gradient · 0.9diffusion model · 0.9crossq · 0.9approximate inference · 0.9target networks · 0.8value expansion · 0.7integral probability metrics · 0.7gradient descent · 0.7
YearPublicationVenuePosition
2025 DIME: Diffusion-Based Maximum Entropy Reinforcement Learning
abstract
Maximum entropy reinforcement learning (MaxEnt-RL) has become the standard approach to RL due to its beneficial exploration properties. Traditionally, policies are parameterized using Gaussian distributions, which significantly limits their representational capacity. Diffusion-based policies offer a more expressive alternative, yet integrating them into MaxEnt-RL poses challenges—primarily due to the intractability of computing their marginal entropy. To overcome this, we propose Diffusion-Based Maximum Entropy RL (DIME). DIME leverages recent advances in approximate inference with diffusion models to derive a lower bound on the maximum entropy objective. Additionally, we propose a policy iteration scheme that provably converges to the optimal diffusion policy. Our method enables the use of expressive diffusion-based policies while retaining the principled exploration benefits of MaxEnt-RL, significantly outperforming other diffusion-based methods on challenging high-dimensional control benchmarks. It is also competitive with state-of-the-art non-diffusion based RL methods while requiring fewer algorithmic design choices and smaller update-to-data ratios, reducing computational complexity.
Onur Celik, Zechu Li, Denis Blessing, Daniel Palenicek, Jan Peters 0001, Georgia Chalvatzaki, Gerhard Neumann
ICML5
2025 Gait in Eight: Efficient On-Robot Learning for Omnidirectional Quadruped Locomotion
abstract
On-robot Reinforcement Learning is a promising approach to train embodiment-aware policies for legged robots. However, the computational constraints of real-time learning on robots pose a significant challenge. We present a framework for efficiently learning quadruped locomotion in just 8 minutes of raw real-time training utilizing the sample efficiency and minimal computational overhead of the new off-policy algorithm CrossQ. We investigate two control architectures: Predicting joint target positions for agile, high-speed locomotion and Central Pattern Generators for stable, natural gaits. While prior work focused on learning simple forward gaits, our framework extends on-robot learning to omnidirectional locomotion. We demonstrate the robustness of our approach in different indoor and outdoor environments and provide the videos and code for our experiments at: https://nico-bohlinger.github.io/gait_in_eight_website
Nico Bohlinger, Jonathan Kinzel, Daniel Palenicek, Lukasz Antczak, Jan Peters 0001
IROS3
2025 Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
abstract
Reinforcement learning has achieved significant milestones, but sample efficiency remains a bottleneck for real-world applications. Recently, CrossQ has demonstrated state-of-the-art sample efficiency with a low update-to-data (UTD) ratio of 1. In this work, we explore CrossQ's scaling behavior with higher UTD ratios. We identify challenges in the training dynamics, which are emphasized by higher UTD ratios. To address these, we integrate weight normalization into the CrossQ framework, a solution that stabilizes training, has been shown to prevent potential loss of plasticity, and keeps the effective learning rate constant. Our proposed approach reliably scales with increasing UTD ratios, achieving competitive performance across 25 challenging continuous control tasks on the DeepMind Control Suite and Myosuite benchmarks, notably the complex dog and humanoid environments. This work eliminates the need for drastic interventions, such as network resets, and offers a simple yet robust pathway for improving sample efficiency and scalability in model-free reinforcement learning.
Daniel Palenicek, Florian Vogt, Joe Watson, Jan Peters 0001
NeurIPS1
2024 CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity
abstract
Sample efficiency is a crucial problem in deep reinforcement learning. Recent algorithms, such as REDQ and DroQ, found a way to improve the sample efficiency by increasing the update-to-data (UTD) ratio to 20 gradient update steps on the critic per environment sample. However, this comes at the expense of a greatly increased computational cost. To reduce this computational burden, we introduce CrossQ: A lightweight algorithm for continuous control tasks that makes careful use of Batch Normalization and removes target networks to surpass the current state-of-the-art in sample efficiency while maintaining a low UTD ratio of 1. Notably, CrossQ does not rely on advanced bias-reduction schemes used in current methods. CrossQ's contributions are threefold: (1) it matches or surpasses current state-of-the-art methods in terms of sample efficiency, (2) it substantially reduces the computational cost compared to REDQ and DroQ, (3) it is easy to implement, requiring just a few lines of code on top of SAC.
Aditya Bhatt 0001, Daniel Palenicek, Boris Belousov, Max Argus, Artemij Amiranashvili, Thomas Brox, Jan Peters 0001
ICLR2
2023 Diminishing Return of Value Expansion Methods in Model-Based Reinforcement Learning
Daniel Palenicek, Michael Lutter, Jan Peters 0001
ICLR1
2023 Pseudo-Likelihood Inference
abstract
Simulation-Based Inference (SBI) is a common name for an emerging family of approaches that infer the model parameters when the likelihood is intractable. Existing SBI methods either approximate the likelihood, such as Approximate Bayesian Computation (ABC) or directly model the posterior, such as Sequential Neural Posterior Estimation (SNPE). While ABC is efficient on low-dimensional problems, on higher-dimensional tasks, it is generally outperformed by SNPE, which leverages function approximation. In this paper, we propose Pseudo-Likelihood Inference (PLI), a new method that brings neural approximation into ABC, making it competitive on challenging Bayesian system identification tasks. By utilizing integral probability metrics, we introduce a smooth likelihood kernel with an adaptive bandwidth that is updated based on information-theoretic trust regions. Thanks to this formulation, our method (i) allows for optimizing neural posteriors via gradient descent, (ii) does not rely on summary statistics, and (iii) enables multiple observations as input. In comparison to SNPE, it leads to improved performance when more data is available. The effectiveness of PLI is evaluated on four classical SBI benchmark tasks and on a highly dynamic physical system, showing particular advantages on stochastic simulations and multi-modal posterior landscapes.
Theo Gruner, Boris Belousov, Fabio Muratore, Daniel Palenicek, Jan Peters 0001
NeurIPS4
2022 SAMBA: safe model-based & active reinforcement learning
Alexander I. Cowen-Rivers, Daniel Palenicek, Vincent Moens, Mohammed Amin Abdullah 0001, Aivar Sootla, Jun Wang 0012, Haitham Bou-Ammar
Mach. Learn.2