Tomasz Arczewski

dblp:385/8264 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › non-stationary reinforcement learning
continual reinforcement learning
0.912025
On-Policy Algorithms for Continual Reinforcement Learning (Student Abstract) · AAAI 2025
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
contrastive reinforcement learning
0.912025
Accelerating Goal-Conditioned Reinforcement Learning Algorithms and Research · ICLR 2025
Machine learning › Reinforcement learning
goal-conditioned reinforcement learning
0.912025
Accelerating Goal-Conditioned Reinforcement Learning Algorithms and Research · ICLR 2025
Machine learning › Reinforcement learning
on-policy reinforcement learning
0.912025
On-Policy Algorithms for Continual Reinforcement Learning (Student Abstract) · AAAI 2025
Machine learning › Reinforcement learning
policy optimization
0.312025
On-Policy Algorithms for Continual Reinforcement Learning (Student Abstract) · AAAI 2025
Machine learning › Reinforcement learning › policy optimization
proximal policy optimization
0.312025
On-Policy Algorithms for Continual Reinforcement Learning (Student Abstract) · AAAI 2025

Methods — techniques the papers use, named apart from their topics

proximal policy optimization · 0.9contrastive learning · 0.9GPU-accelerated replay buffer · 0.9
YearPublicationVenuePosition
2025 On-Policy Algorithms for Continual Reinforcement Learning (Student Abstract)
abstract
Continual reinforcement learning (CRL) is the study of optimal strategies for maximizing rewards in sequential environments that change over time. This is particularly crucial in domains such as robotics, where the operational environment is inherently dynamic and subject to continual change. Nevertheless, research in this area has thus far concentrated on off-policy algorithms with replay buffers that are capable of amortizing the impact of distribution shifts. Such an approach is not feasible with on-policy reinforcement learning algorithms that learn solely from the data obtained from the current policy. In this paper, we examine the performance of proximal policy optimization (PPO), a prevalent on-policy reinforcement learning (RL) algorithm, in a classical CRL benchmark. Our findings suggest that the current methods are suboptimal in terms of average performance. Nevertheless, they demonstrate encouraging competitive outcomes with respect to forward transfer and forgetting metrics. This highlights the need for further research into continual on-policy reinforcement learning. The source code is available at https://github.com/Teddy298/continualworld-ppo.
Tadeusz Dziarmaga, Tomasz Arczewski, Marcin Mazur, Maciej Wolczyk
AAAI2
2025 Accelerating Goal-Conditioned Reinforcement Learning Algorithms and Research
abstract
Self-supervision has the potential to transform reinforcement learning (RL), paralleling the breakthroughs it has enabled in other areas of machine learning. While self-supervised learning in other domains aims to find patterns in a fixed dataset, self-supervised goal-conditioned reinforcement learning (GCRL) agents discover *new* behaviors by learning from the goals achieved during unstructured interaction with the environment. However, these methods have failed to see similar success, both due to a lack of data from slow environment simulations as well as a lack of stable algorithms. We take a step toward addressing both of these issues by releasing a high-performance codebase and benchmark (`JaxGCRL`) for self-supervised GCRL, enabling researchers to train agents for millions of environment steps in minutes on a single GPU. By utilizing GPU-accelerated replay buffers, environments, and a stable contrastive RL algorithm, we reduce training time by up to $22\times$. Additionally, we assess key design choices in contrastive RL, identifying those that most effectively stabilize and enhance training performance. With this approach, we provide a foundation for future research in self-supervised GCRL, enabling researchers to quickly iterate on new ideas and evaluate them in diverse and challenging environments. Code: [https://anonymous.4open.science/r/JaxGCRL-2316/README.md](https://anonymous.4open.science/r/JaxGCRL-2316/README.md)
Michal Bortkiewicz, Wladyslaw Palucki, Vivek Myers, Tadeusz Dziarmaga, Tomasz Arczewski, Lukasz Kucinski, Benjamin Eysenbach
ICLR5