Michal Nauman

dblp:277/6250 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 80% Deep learning architectures and training · 9% Planning, search and constraint satisfaction · 4%

Topics — the 20 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
actor-critic methods
2.532025
A Case for Validation Buffer in Pessimistic Actor-Critic · IJCAI 2025
Decoupled Policy Actor-Critic: Bridging Pessimism and Risk Awareness in Reinforcement Learning · AAAI 2025
Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning · ICML 2024
Machine learning › Reinforcement learning
value function approximation
2.532025
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners · NeurIPS 2025
A Case for Validation Buffer in Pessimistic Actor-Critic · IJCAI 2025
Bigger, Regularized, Optimistic: scaling for compute and sample efficient continuous control · NeurIPS 2024
Machine learning › Reinforcement learning › actor-critic methods
pessimistic actor-critic
1.722025
A Case for Validation Buffer in Pessimistic Actor-Critic · IJCAI 2025
Decoupled Policy Actor-Critic: Bridging Pessimism and Risk Awareness in Reinforcement Learning · AAAI 2025
Machine learning › Reinforcement learning
continuous control
1.022025
Bigger, Regularized, Optimistic: scaling for compute and sample efficient continuous control · NeurIPS 2024
Decoupled Policy Actor-Critic: Bridging Pessimism and Risk Awareness in Reinforcement Learning · AAAI 2025
Machine learning › Deep learning architectures and training › scaling laws
compute-optimal scaling
0.912025
Compute-Optimal Scaling for Value-Based Deep RL · NeurIPS 2025
Machine learning › Efficient and distributed learning › efficient training
compute-optimal training
0.912025
Value-Based Deep RL Scales Predictably · ICML 2025
Machine learning › Reinforcement learning
deep reinforcement learning
0.912025
Compute-Optimal Scaling for Value-Based Deep RL · NeurIPS 2025
Machine learning › Reinforcement learning
multi-task reinforcement learning
0.912025
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners · NeurIPS 2025
Machine learning › Reinforcement learning
off-policy reinforcement learning
0.912025
Value-Based Deep RL Scales Predictably · ICML 2025
Machine learning › Reinforcement learning › safe reinforcement learning
risk-sensitive reinforcement learning
0.912025
Decoupled Policy Actor-Critic: Bridging Pessimism and Risk Awareness in Reinforcement Learning · AAAI 2025
Machine learning › Deep learning architectures and training
scaling laws
0.912025
Value-Based Deep RL Scales Predictably · ICML 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › decision making under uncertainty
utility maximization
0.912025
Decoupled Policy Actor-Critic: Bridging Pessimism and Risk Awareness in Reinforcement Learning · AAAI 2025
Machine learning › Reinforcement learning
value-based reinforcement learning
0.912025
Value-Based Deep RL Scales Predictably · ICML 2025
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning
0.812024
Bigger, Regularized, Optimistic: scaling for compute and sample efficient continuous control · NeurIPS 2024
Machine learning › Reinforcement learning
model-based reinforcement learning
0.712023
On Many-Actions Policy Gradient · ICML 2023
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.712023
On Many-Actions Policy Gradient · ICML 2023
Machine learning › Reinforcement learning › policy optimization › policy gradient
stochastic policy gradient
0.712023
On Many-Actions Policy Gradient · ICML 2023
Machine learning › Learning theory
sample complexity
0.312025
Compute-Optimal Scaling for Value-Based Deep RL · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation
task embedding
0.312025
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners · NeurIPS 2025
Machine learning › Reinforcement learning › exploration
optimistic exploration
0.212024
Bigger, Regularized, Optimistic: scaling for compute and sample efficient continuous control · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

temporal difference learning · 1.7regularization · 1.5validation buffer · 0.9task embedding · 0.9scaling laws · 0.9pareto frontier estimation · 0.9hyperparameter tuning · 0.9exponential utility function · 0.9expected utility hypothesis · 0.9cross-entropy temporal difference learning · 0.9
YearPublicationVenuePosition
2025 Decoupled Policy Actor-Critic: Bridging Pessimism and Risk Awareness in Reinforcement Learning
abstract
Actor-Critic (AC) algorithms like SAC and TD3 were shown to perform well in a variety of continuous-action tasks. However, the theoretical basis for the pessimistic objectives these algorithms employ remains unestablished, raising questions about the specific class of policies they are implementing. In this work, we apply the expected utility hypothesis, a fundamental concept in economics, to illustrate that both pessimistic and non-pessimistic RL objectives can be interpreted through expected utility maximization using an exponential utility function. This approach reveals that pessimistic policies effectively maximize value certainty equivalent, aligning them with the optimization of risk-aware objectives. Furthermore, we propose Decoupled Policy Actor-Critic (DAC). DAC is a model-free algorithm that features two distinct actor networks: a pessimistic actor for temporal-difference learning and an optimistic actor for exploration. Our evaluations of DAC across various locomotion and manipulation tasks demonstrate improvements in sample efficiency and final performance. Remarkably, DAC, while requiring significantly fewer computational resources, matches the performance of leading model-based methods in the complex dog and humanoid domains.
Michal Nauman, Marek Cygan
AAAI1
2025 Value-Based Deep RL Scales Predictably
abstract
Scaling data and compute is critical in modern machine learning. However, scaling also demands _predictability_: we want methods to not only perform well with more compute or data, but also have their performance be predictable from low compute or low data runs, without ever running the large-scale experiment. In this paper, we show predictability of value-based off-policy deep RL. First, we show that data and compute requirements to reach a given performance level lie on a _Pareto frontier_, controlled by the updates-to-data (UTD) ratio. By estimating this frontier, we can extrapolate data requirements into a higher compute regime, and compute requirements into a higher data regime. Second, we determine the optimal allocation of total _budget_ across data and compute to obtain given performance and use it to determine hyperparameters that maximize performance for a given budget. Third, this scaling behavior is enabled by first estimating predictable relationships between different _hyperparameters_, which is used to counteract effects of overfitting and plasticity loss unique to RL. We validate our approach using three algorithms: SAC, BRO, and PQL on DeepMind Control, OpenAI gym, and IsaacGym, when extrapolating to higher levels of data, compute, budget, or performance.
Oleh Rybkin, Michal Nauman, Preston Fu, Charlie Snell, Pieter Abbeel, Sergey Levine, Aviral Kumar
ICML2
2025 A Case for Validation Buffer in Pessimistic Actor-Critic
abstract
In this paper, we investigate the issue of error accumulation in critic networks updated via pessimistic temporal difference objectives. We show that the critic approximation error can be approximated via a recursive fixed-point model similar to that of the Bellman value. We use such recursive definition to retrieve the conditions under which the pessimistic critic is unbiased. Building on these insights, we propose Validation Pessimism Learning (VPL) algorithm. VPL uses a small validation buffer to adjust the levels of pessimism throughout the agent training, with the pessimism set such that the approximation error of the critic targets is minimized. We investigate the proposed approach on a variety of locomotion and manipulation tasks and report improvements in sample efficiency and performance.
Michal Nauman, Mateusz Ostaszewski, Marek Cygan
IJCAI1
2025 Compute-Optimal Scaling for Value-Based Deep RL
abstract
As models grow larger and training them becomes expensive, it becomes increasingly important to scale training recipes not just to larger models and more data, but to do so in a compute-optimal manner that extracts maximal performance per unit of compute. While such scaling has been well studied for language modeling, reinforcement learning (RL) has received less attention in this regard. In this paper, we investigate compute scaling for online, value-based deep RL. These methods present two primary axes for compute allocation: model capacity and the update-to-data (UTD) ratio. Given a fixed compute budget, we ask: how should resources be partitioned across these axes to maximize data efficiency? Our analysis reveals a nuanced interplay between model size, batch size, and UTD. In particular, we identify a phenomenon we call TD-overfitting: increasing the batch quickly harms Q-function accuracy for small models, but this effect is absent in large models, enabling effective use of large batch size at scale. We provide a mental model for understanding this phenomenon and build guidelines for choosing batch size and UTD to optimize compute usage. Our findings provide a grounded starting point for compute-optimal scaling in deep RL, mirroring studies in supervised learning but adapted to TD learning. Project page: https://value-scaling.github.io/.
Preston Fu, Oleh Rybkin, Michal Nauman, Pieter Abbeel, Sergey Levine, Aviral Kumar
NeurIPS4
2025 Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
abstract
Recent advances in language modeling and vision stem from training large models on diverse, multi‑task data. This paradigm has had limited impact in value-based reinforcement learning (RL), where improvements are often driven by small models trained in a single-task context. This is because in multi-task RL sparse rewards and gradient conflicts make optimization of temporal difference brittle. Practical workflows for generalist policies therefore avoid online training, instead cloning expert trajectories or distilling collections of single‑task policies into one agent. In this work, we show that the use of high-capacity value models trained via cross-entropy and conditioned on learnable task embeddings addresses the problem of task interference in online RL, allowing for robust and scalable multi‑task training. We test our approach on 7 multi-task benchmarks with over 280 unique tasks, spanning high degree-of-freedom humanoid control and discrete vision-based RL. We find that, despite its simplicity, the proposed approach leads to state-of-the-art single and multi-task performance, as well as sample-efficient transfer to new tasks.
Michal Nauman, Marek Cygan, Carmelo Sferrazza, Aviral Kumar, Pieter Abbeel
NeurIPS1
2024 Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning
abstract
Recent advancements in off-policy Reinforcement Learning (RL) have significantly improved sample efficiency, primarily due to the incorporation of various forms of regularization that enable more gradient update steps than traditional agents. However, many of these techniques have been tested in limited settings, often on tasks from single simulation benchmarks and against well-known algorithms rather than a range of regularization approaches. This limits our understanding of the specific mechanisms driving RL improvements. To address this, we implemented over 60 different off-policy agents, each integrating established regularization techniques from recent state-of-the-art algorithms. We tested these agents across 14 diverse tasks from 2 simulation benchmarks, measuring training metrics related to overestimation, overfitting, and plasticity loss — issues that motivate the examined regularization techniques. Our findings reveal that while the effectiveness of a specific regularization setup varies with the task, certain combinations consistently demonstrate robust and superior performance. Notably, a simple Soft Actor-Critic agent, appropriately regularized, reliably finds a better-performing policy within the training regime, which previously was achieved mainly through model-based approaches.
Michal Nauman, Michal Bortkiewicz, Piotr Milos, Tomasz Trzcinski, Mateusz Ostaszewski, Marek Cygan
ICML1
2024 Bigger, Regularized, Optimistic: scaling for compute and sample efficient continuous control
abstract
Sample efficiency in Reinforcement Learning (RL) has traditionally been driven by algorithmic enhancements. In this work, we demonstrate that scaling can also lead to substantial improvements. We conduct a thorough investigation into the interplay of scaling model capacity and domain-specific RL enhancements. These empirical findings inform the design choices underlying our proposed BRO (Bigger, Regularized, Optimistic) algorithm. The key innovation behind BRO is that strong regularization allows for effective scaling of the critic networks, which, paired with optimistic exploration, leads to superior performance. BRO achieves state-of-the-art results, significantly outperforming the leading model-based and model-free algorithms across 40 complex tasks from the DeepMind Control, MetaWorld, and MyoSuite benchmarks. BRO is the first model-free algorithm to achieve near-optimal policies in the notoriously challenging Dog and Humanoid tasks.
Michal Nauman, Mateusz Ostaszewski, Krzysztof Jankowski, Piotr Milos, Marek Cygan
NeurIPS1
2023 On Many-Actions Policy Gradient
abstract
We study the variance of stochastic policy gradients (SPGs) with many action samples per state. We derive a many-actions optimality condition, which determines when many-actions SPG yields lower variance as compared to a single-action agent with proportionally extended trajectory. We propose Model-Based Many-Actions (MBMA), an approach leveraging dynamics models for many-actions sampling in the context of SPG. MBMA addresses issues associated with existing implementations of many-actions SPG and yields lower bias and comparable variance to SPG estimated from states in model-simulated rollouts. We find that MBMA bias and variance structure matches that predicted by theory. As a result, MBMA achieves improved sample efficiency and higher returns on a range of continuous action environments as compared to model-free, many-actions, and model-based on-policy SPG baselines.
Michal Nauman, Marek Cygan
ICML1