EDBT 2026 Demo / reviewers in the wild / expert
Michal Nauman
dblp:277/6250
· DBLP profile ↗
8ranked-venue papers
6as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Reinforcement learning · 80% Deep learning architectures and training · 9% Planning, search and constraint satisfaction · 4% |
Topics — the 20 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
actor-critic methods |
2.5 | 3 | 2025 | A Case for Validation Buffer in Pessimistic Actor-Critic · IJCAI 2025 Decoupled Policy Actor-Critic: Bridging Pessimism and Risk Awareness in Reinforcement Learning · AAAI 2025 Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning · ICML 2024 |
Machine learning › Reinforcement learning
value function approximation |
2.5 | 3 | 2025 | Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners · NeurIPS 2025 A Case for Validation Buffer in Pessimistic Actor-Critic · IJCAI 2025 Bigger, Regularized, Optimistic: scaling for compute and sample efficient continuous control · NeurIPS 2024 |
Machine learning › Reinforcement learning › actor-critic methods
pessimistic actor-critic |
1.7 | 2 | 2025 | A Case for Validation Buffer in Pessimistic Actor-Critic · IJCAI 2025 Decoupled Policy Actor-Critic: Bridging Pessimism and Risk Awareness in Reinforcement Learning · AAAI 2025 |
Machine learning › Reinforcement learning
continuous control |
1.0 | 2 | 2025 | Bigger, Regularized, Optimistic: scaling for compute and sample efficient continuous control · NeurIPS 2024 Decoupled Policy Actor-Critic: Bridging Pessimism and Risk Awareness in Reinforcement Learning · AAAI 2025 |
Machine learning › Deep learning architectures and training › scaling laws
compute-optimal scaling |
0.9 | 1 | 2025 | Compute-Optimal Scaling for Value-Based Deep RL · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › efficient training
compute-optimal training |
0.9 | 1 | 2025 | Value-Based Deep RL Scales Predictably · ICML 2025 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.9 | 1 | 2025 | Compute-Optimal Scaling for Value-Based Deep RL · NeurIPS 2025 |
Machine learning › Reinforcement learning
multi-task reinforcement learning |
0.9 | 1 | 2025 | Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners · NeurIPS 2025 |
Machine learning › Reinforcement learning
off-policy reinforcement learning |
0.9 | 1 | 2025 | Value-Based Deep RL Scales Predictably · ICML 2025 |
Machine learning › Reinforcement learning › safe reinforcement learning
risk-sensitive reinforcement learning |
0.9 | 1 | 2025 | Decoupled Policy Actor-Critic: Bridging Pessimism and Risk Awareness in Reinforcement Learning · AAAI 2025 |
Machine learning › Deep learning architectures and training
scaling laws |
0.9 | 1 | 2025 | Value-Based Deep RL Scales Predictably · ICML 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › decision making under uncertainty
utility maximization |
0.9 | 1 | 2025 | Decoupled Policy Actor-Critic: Bridging Pessimism and Risk Awareness in Reinforcement Learning · AAAI 2025 |
Machine learning › Reinforcement learning
value-based reinforcement learning |
0.9 | 1 | 2025 | Value-Based Deep RL Scales Predictably · ICML 2025 |
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning |
0.8 | 1 | 2024 | Bigger, Regularized, Optimistic: scaling for compute and sample efficient continuous control · NeurIPS 2024 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.7 | 1 | 2023 | On Many-Actions Policy Gradient · ICML 2023 |
Machine learning › Reinforcement learning › policy optimization
policy gradient |
0.7 | 1 | 2023 | On Many-Actions Policy Gradient · ICML 2023 |
Machine learning › Reinforcement learning › policy optimization › policy gradient
stochastic policy gradient |
0.7 | 1 | 2023 | On Many-Actions Policy Gradient · ICML 2023 |
Machine learning › Learning theory
sample complexity |
0.3 | 1 | 2025 | Compute-Optimal Scaling for Value-Based Deep RL · NeurIPS 2025 |
Machine learning › Transfer learning and domain adaptation
task embedding |
0.3 | 1 | 2025 | Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners · NeurIPS 2025 |
Machine learning › Reinforcement learning › exploration
optimistic exploration |
0.2 | 1 | 2024 | Bigger, Regularized, Optimistic: scaling for compute and sample efficient continuous control · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
temporal difference learning · 1.7regularization · 1.5validation buffer · 0.9task embedding · 0.9scaling laws · 0.9pareto frontier estimation · 0.9hyperparameter tuning · 0.9exponential utility function · 0.9expected utility hypothesis · 0.9cross-entropy temporal difference learning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Decoupled Policy Actor-Critic: Bridging Pessimism and Risk Awareness in Reinforcement LearningabstractActor-Critic (AC) algorithms like SAC and TD3 were shown to perform well in a variety of continuous-action tasks. However, the theoretical basis for the pessimistic objectives these algorithms employ remains unestablished, raising questions about the specific class of policies they are implementing. In this work, we apply the expected utility hypothesis, a fundamental concept in economics, to illustrate that both pessimistic and non-pessimistic RL objectives can be interpreted through expected utility maximization using an exponential utility function. This approach reveals that pessimistic policies effectively maximize value certainty equivalent, aligning them with the optimization of risk-aware objectives. Furthermore, we propose Decoupled Policy Actor-Critic (DAC). DAC is a model-free algorithm that features two distinct actor networks: a pessimistic actor for temporal-difference learning and an optimistic actor for exploration. Our evaluations of DAC across various locomotion and manipulation tasks demonstrate improvements in sample efficiency and final performance. Remarkably, DAC, while requiring significantly fewer computational resources, matches the performance of leading model-based methods in the complex dog and humanoid domains. Michal Nauman, Marek Cygan |
AAAI | 1 |
| 2025 | Value-Based Deep RL Scales PredictablyabstractScaling data and compute is critical in modern machine learning. However, scaling also demands _predictability_: we want methods to not only perform well with more compute or data, but also have their performance be predictable from low compute or low data runs, without ever running the large-scale experiment. In this paper, we show predictability of value-based off-policy deep RL. First, we show that data and compute requirements to reach a given performance level lie on a _Pareto frontier_, controlled by the updates-to-data (UTD) ratio. By estimating this frontier, we can extrapolate data requirements into a higher compute regime, and compute requirements into a higher data regime. Second, we determine the optimal allocation of total _budget_ across data and compute to obtain given performance and use it to determine hyperparameters that maximize performance for a given budget. Third, this scaling behavior is enabled by first estimating predictable relationships between different _hyperparameters_, which is used to counteract effects of overfitting and plasticity loss unique to RL. We validate our approach using three algorithms: SAC, BRO, and PQL on DeepMind Control, OpenAI gym, and IsaacGym, when extrapolating to higher levels of data, compute, budget, or performance. Oleh Rybkin, Michal Nauman, Preston Fu, Charlie Snell, Pieter Abbeel, Sergey Levine, Aviral Kumar |
ICML | 2 |
| 2025 | A Case for Validation Buffer in Pessimistic Actor-CriticabstractIn this paper, we investigate the issue of error accumulation in critic networks updated via pessimistic temporal difference objectives. We show that the critic approximation error can be approximated via a recursive fixed-point model similar to that of the Bellman value. We use such recursive definition to retrieve the conditions under which the pessimistic critic is unbiased. Building on these insights, we propose Validation Pessimism Learning (VPL) algorithm. VPL uses a small validation buffer to adjust the levels of pessimism throughout the agent training, with the pessimism set such that the approximation error of the critic targets is minimized. We investigate the proposed approach on a variety of locomotion and manipulation tasks and report improvements in sample efficiency and performance. Michal Nauman, Mateusz Ostaszewski, Marek Cygan |
IJCAI | 1 |
| 2025 | Compute-Optimal Scaling for Value-Based Deep RLabstractAs models grow larger and training them becomes expensive, it becomes increasingly important to scale training recipes not just to larger models and more data, but to do so in a compute-optimal manner that extracts maximal performance per unit of compute. While such scaling has been well studied for language modeling, reinforcement learning (RL) has received less attention in this regard. In this paper, we investigate compute scaling for online, value-based deep RL. These methods present two primary axes for compute allocation: model capacity and the update-to-data (UTD) ratio. Given a fixed compute budget, we ask: how should resources be partitioned across these axes to maximize data efficiency? Our analysis reveals a nuanced interplay between model size, batch size, and UTD. In particular, we identify a phenomenon we call TD-overfitting: increasing the batch quickly harms Q-function accuracy for small models, but this effect is absent in large models, enabling effective use of large batch size at scale. We provide a mental model for understanding this phenomenon and build guidelines for choosing batch size and UTD to optimize compute usage. Our findings provide a grounded starting point for compute-optimal scaling in deep RL, mirroring studies in supervised learning but adapted to TD learning. Project page: https://value-scaling.github.io/. Preston Fu, Oleh Rybkin, Michal Nauman, Pieter Abbeel, Sergey Levine, Aviral Kumar |
NeurIPS | 4 |
| 2025 | Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task LearnersabstractRecent advances in language modeling and vision stem from training large models on diverse, multi‑task data. This paradigm has had limited impact in value-based reinforcement learning (RL), where improvements are often driven by small models trained in a single-task context. This is because in multi-task RL sparse rewards and gradient conflicts make optimization of temporal difference brittle. Practical workflows for generalist policies therefore avoid online training, instead cloning expert trajectories or distilling collections of single‑task policies into one agent. In this work, we show that the use of high-capacity value models trained via cross-entropy and conditioned on learnable task embeddings addresses the problem of task interference in online RL, allowing for robust and scalable multi‑task training. We test our approach on 7 multi-task benchmarks with over 280 unique tasks, spanning high degree-of-freedom humanoid control and discrete vision-based RL. We find that, despite its simplicity, the proposed approach leads to state-of-the-art single and multi-task performance, as well as sample-efficient transfer to new tasks. Michal Nauman, Marek Cygan, Carmelo Sferrazza, Aviral Kumar, Pieter Abbeel |
NeurIPS | 1 |
| 2024 | Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement LearningabstractRecent advancements in off-policy Reinforcement Learning (RL) have significantly improved sample efficiency, primarily due to the incorporation of various forms of regularization that enable more gradient update steps than traditional agents. However, many of these techniques have been tested in limited settings, often on tasks from single simulation benchmarks and against well-known algorithms rather than a range of regularization approaches. This limits our understanding of the specific mechanisms driving RL improvements. To address this, we implemented over 60 different off-policy agents, each integrating established regularization techniques from recent state-of-the-art algorithms. We tested these agents across 14 diverse tasks from 2 simulation benchmarks, measuring training metrics related to overestimation, overfitting, and plasticity loss — issues that motivate the examined regularization techniques. Our findings reveal that while the effectiveness of a specific regularization setup varies with the task, certain combinations consistently demonstrate robust and superior performance. Notably, a simple Soft Actor-Critic agent, appropriately regularized, reliably finds a better-performing policy within the training regime, which previously was achieved mainly through model-based approaches. Michal Nauman, Michal Bortkiewicz, Piotr Milos, Tomasz Trzcinski, Mateusz Ostaszewski, Marek Cygan |
ICML | 1 |
| 2024 | Bigger, Regularized, Optimistic: scaling for compute and sample efficient continuous controlabstractSample efficiency in Reinforcement Learning (RL) has traditionally been driven by algorithmic enhancements. In this work, we demonstrate that scaling can also lead to substantial improvements. We conduct a thorough investigation into the interplay of scaling model capacity and domain-specific RL enhancements. These empirical findings inform the design choices underlying our proposed BRO (Bigger, Regularized, Optimistic) algorithm. The key innovation behind BRO is that strong regularization allows for effective scaling of the critic networks, which, paired with optimistic exploration, leads to superior performance. BRO achieves state-of-the-art results, significantly outperforming the leading model-based and model-free algorithms across 40 complex tasks from the DeepMind Control, MetaWorld, and MyoSuite benchmarks. BRO is the first model-free algorithm to achieve near-optimal policies in the notoriously challenging Dog and Humanoid tasks. Michal Nauman, Mateusz Ostaszewski, Krzysztof Jankowski, Piotr Milos, Marek Cygan |
NeurIPS | 1 |
| 2023 | On Many-Actions Policy GradientabstractWe study the variance of stochastic policy gradients (SPGs) with many action samples per state. We derive a many-actions optimality condition, which determines when many-actions SPG yields lower variance as compared to a single-action agent with proportionally extended trajectory. We propose Model-Based Many-Actions (MBMA), an approach leveraging dynamics models for many-actions sampling in the context of SPG. MBMA addresses issues associated with existing implementations of many-actions SPG and yields lower bias and comparable variance to SPG estimated from states in model-simulated rollouts. We find that MBMA bias and variance structure matches that predicted by theory. As a result, MBMA achieves improved sample efficiency and higher returns on a range of continuous action environments as compared to model-free, many-actions, and model-based on-policy SPG baselines. Michal Nauman, Marek Cygan |
ICML | 1 |