VLDB 2026 Research / reviewers in the wild / expert
Daniel Palenicek
dblp:267/9480
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-8292-1318ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 65% Probabilistic and Bayesian machine learning · 16% Robot manipulation · 7% |
Topics — the 16 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
sample efficiency |
1.6 | 2 | 2025 | Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization · NeurIPS 2025 CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity · ICLR 2024 |
Machine learning › Reinforcement learning
continuous control |
0.9 | 1 | 2025 | Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization · NeurIPS 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | DIME: Diffusion-Based Maximum Entropy Reinforcement Learning · ICML 2025 |
Robotics › Robot manipulation
diffusion policy |
0.9 | 1 | 2025 | DIME: Diffusion-Based Maximum Entropy Reinforcement Learning · ICML 2025 |
Machine learning › Reinforcement learning
maximum entropy reinforcement learning |
0.9 | 1 | 2025 | DIME: Diffusion-Based Maximum Entropy Reinforcement Learning · ICML 2025 |
Machine learning › Reinforcement learning
model-free reinforcement learning |
0.9 | 1 | 2025 | Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization · NeurIPS 2025 |
Machine learning › Reinforcement learning
off-policy reinforcement learning |
0.9 | 1 | 2025 | Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization · NeurIPS 2025 |
Machine learning › Reinforcement learning › policy learning
policy parameterization |
0.9 | 1 | 2025 | DIME: Diffusion-Based Maximum Entropy Reinforcement Learning · ICML 2025 |
Machine learning › Deep learning architectures and training › normalization
batch normalization |
0.8 | 1 | 2024 | CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity · ICLR 2024 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.8 | 1 | 2024 | CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity · ICLR 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference
approximate bayesian computation |
0.7 | 1 | 2023 | Pseudo-Likelihood Inference · NeurIPS 2023 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.7 | 1 | 2023 | Diminishing Return of Value Expansion Methods in Model-Based Reinforcement Learning · ICLR 2023 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference › simulation-based inference
neural posterior estimation |
0.7 | 1 | 2023 | Pseudo-Likelihood Inference · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference
simulation-based inference |
0.7 | 1 | 2023 | Pseudo-Likelihood Inference · NeurIPS 2023 |
Machine learning › Reinforcement learning › model-based reinforcement learning
value expansion |
0.7 | 1 | 2023 | Diminishing Return of Value Expansion Methods in Model-Based Reinforcement Learning · ICLR 2023 |
Machine learning › Reinforcement learning › dynamic programming
policy iteration |
0.3 | 1 | 2025 | DIME: Diffusion-Based Maximum Entropy Reinforcement Learning · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
batch normalization · 1.6weight normalization · 0.9policy gradient · 0.9diffusion model · 0.9crossq · 0.9approximate inference · 0.9target networks · 0.8value expansion · 0.7integral probability metrics · 0.7gradient descent · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DIME: Diffusion-Based Maximum Entropy Reinforcement LearningabstractMaximum entropy reinforcement learning (MaxEnt-RL) has become the standard approach to RL due to its beneficial exploration properties. Traditionally, policies are parameterized using Gaussian distributions, which significantly limits their representational capacity. Diffusion-based policies offer a more expressive alternative, yet integrating them into MaxEnt-RL poses challenges—primarily due to the intractability of computing their marginal entropy. To overcome this, we propose Diffusion-Based Maximum Entropy RL (DIME). DIME leverages recent advances in approximate inference with diffusion models to derive a lower bound on the maximum entropy objective. Additionally, we propose a policy iteration scheme that provably converges to the optimal diffusion policy. Our method enables the use of expressive diffusion-based policies while retaining the principled exploration benefits of MaxEnt-RL, significantly outperforming other diffusion-based methods on challenging high-dimensional control benchmarks. It is also competitive with state-of-the-art non-diffusion based RL methods while requiring fewer algorithmic design choices and smaller update-to-data ratios, reducing computational complexity. Onur Celik, Zechu Li, Denis Blessing, Daniel Palenicek, Jan Peters 0001, Georgia Chalvatzaki, Gerhard Neumann |
ICML | 5 |
| 2025 | Gait in Eight: Efficient On-Robot Learning for Omnidirectional Quadruped LocomotionabstractOn-robot Reinforcement Learning is a promising approach to train embodiment-aware policies for legged robots. However, the computational constraints of real-time learning on robots pose a significant challenge. We present a framework for efficiently learning quadruped locomotion in just 8 minutes of raw real-time training utilizing the sample efficiency and minimal computational overhead of the new off-policy algorithm CrossQ. We investigate two control architectures: Predicting joint target positions for agile, high-speed locomotion and Central Pattern Generators for stable, natural gaits. While prior work focused on learning simple forward gaits, our framework extends on-robot learning to omnidirectional locomotion. We demonstrate the robustness of our approach in different indoor and outdoor environments and provide the videos and code for our experiments at: https://nico-bohlinger.github.io/gait_in_eight_website Nico Bohlinger, Jonathan Kinzel, Daniel Palenicek, Lukasz Antczak, Jan Peters 0001 |
IROS | 3 |
| 2025 | Scaling Off-Policy Reinforcement Learning with Batch and Weight NormalizationabstractReinforcement learning has achieved significant milestones, but sample efficiency remains a bottleneck for real-world applications.
Recently, CrossQ has demonstrated state-of-the-art sample efficiency with a low update-to-data (UTD) ratio of 1.
In this work, we explore CrossQ's scaling behavior with higher UTD ratios.
We identify challenges in the training dynamics, which are emphasized by higher UTD ratios.
To address these, we integrate weight normalization into the CrossQ framework, a solution that stabilizes training, has been shown to prevent potential loss of plasticity, and keeps the effective learning rate constant.
Our proposed approach reliably scales with increasing UTD ratios, achieving competitive performance across 25 challenging continuous control tasks on the DeepMind Control Suite and Myosuite benchmarks, notably the complex dog and humanoid environments.
This work eliminates the need for drastic interventions, such as network resets, and offers a simple yet robust pathway for improving sample efficiency and scalability in model-free reinforcement learning. Daniel Palenicek, Florian Vogt, Joe Watson, Jan Peters 0001 |
NeurIPS | 1 |
| 2024 | CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and SimplicityabstractSample efficiency is a crucial problem in deep reinforcement learning. Recent algorithms, such as REDQ and DroQ, found a way to improve the sample efficiency by increasing the update-to-data (UTD) ratio to 20 gradient update steps on the critic per environment sample.
However, this comes at the expense of a greatly increased computational cost. To reduce this computational burden, we introduce CrossQ:
A lightweight algorithm for continuous control tasks that makes careful use of Batch Normalization and removes target networks to surpass the current state-of-the-art in sample efficiency while maintaining a low UTD ratio of 1. Notably, CrossQ does not rely on advanced bias-reduction schemes used in current methods. CrossQ's contributions are threefold: (1) it matches or surpasses current state-of-the-art methods in terms of sample efficiency, (2) it substantially reduces the computational cost compared to REDQ and DroQ, (3) it is easy to implement, requiring just a few lines of code on top of SAC. Aditya Bhatt 0001, Daniel Palenicek, Boris Belousov, Max Argus, Artemij Amiranashvili, Thomas Brox, Jan Peters 0001 |
ICLR | 2 |
| 2023 | Diminishing Return of Value Expansion Methods in Model-Based Reinforcement Learning
Daniel Palenicek, Michael Lutter, Jan Peters 0001 |
ICLR | 1 |
| 2023 | Pseudo-Likelihood InferenceabstractSimulation-Based Inference (SBI) is a common name for an emerging family of approaches that infer the model parameters when the likelihood is intractable. Existing SBI methods either approximate the likelihood, such as Approximate Bayesian Computation (ABC) or directly model the posterior, such as Sequential Neural Posterior Estimation (SNPE). While ABC is efficient on low-dimensional problems, on higher-dimensional tasks, it is generally outperformed by SNPE, which leverages function approximation. In this paper, we propose Pseudo-Likelihood Inference (PLI), a new method that brings neural approximation into ABC, making it competitive on challenging Bayesian system identification tasks. By utilizing integral probability metrics, we introduce a smooth likelihood kernel with an adaptive bandwidth that is updated based on information-theoretic trust regions. Thanks to this formulation, our method (i) allows for optimizing neural posteriors via gradient descent, (ii) does not rely on summary statistics, and (iii) enables multiple observations as input. In comparison to SNPE, it leads to improved performance when more data is available. The effectiveness of PLI is evaluated on four classical SBI benchmark tasks and on a highly dynamic physical system, showing particular advantages on stochastic simulations and multi-modal posterior landscapes. Theo Gruner, Boris Belousov, Fabio Muratore, Daniel Palenicek, Jan Peters 0001 |
NeurIPS | 4 |
| 2022 | SAMBA: safe model-based & active reinforcement learning
Alexander I. Cowen-Rivers, Daniel Palenicek, Vincent Moens, Mohammed Amin Abdullah 0001, Aivar Sootla, Jun Wang 0012, Haitham Bou-Ammar |
Mach. Learn. | 2 |