VLDB 2026 Research / reviewers in the wild / expert
Pawel Wawrzynski
dblp:88/4802
· DBLP profile ↗
23ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-1154-0470ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 7 first-author · 11 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SACn: Soft Actor-Critic with n-step Returns
Jakub Lyskawa, Jakub Lewandowski, Pawel Wawrzynski |
ICAART (3) | 3 |
| 2025 | EnEnv 1.0: Energy Grid Environment for Multi-Agent Reinforcement Learning Benchmarking
Dominik Jacek Bogucki, Lukasz Lepak, Sonam Parashar, Bartlomiej Blachowski, Pawel Wawrzynski |
AAMAS | 5 |
| 2025 | HINT: Hypernetwork approach to training weight interval regions in continual learning
Patryk Krukowski, Anna Bielawska, Kamil Ksiazek, Pawel Wawrzynski, Pawel Batorski, Przemyslaw Spurek |
Inf. Sci. | 4 |
| 2024 | Subgoal Reachability in Goal Conditioned Hierarchical Reinforcement Learning
Michal Bortkiewicz, Jakub Lyskawa, Pawel Wawrzynski, Mateusz Ostaszewski, Artur Grudkowski, Bartlomiej Sobieski, Tomasz Trzcinski |
ICAART (1) | 3 |
| 2024 | Reinforcement Learning Meets Microeconomics: Learning to Designate Price-Dependent Supply and Demand for Automated Trading
Lukasz Lepak, Pawel Wawrzynski |
ECML/PKDD (9) | 2 |
| 2024 | ACERAC: Efficient Reinforcement Learning in Fine Time DiscretizationabstractOne of the main goals of reinforcement learning (RL) is to provide a way for physical machines to learn optimal behavior instead of being programmed. However, effective control of the machines usually requires fine time discretization. The most common RL methods apply independent random elements to each action, which is not suitable in that setting. It is not feasible because it causes the controlled system to jerk and does not ensure sufficient exploration since a single action is not long enough to create a significant experience that could be translated into policy improvement. In our view, these are the main obstacles that prevent the application of RL in contemporary control systems. To address these pitfalls, in this article, we introduce an RL framework and adequate analytical tools for actions that may be stochastically dependent in subsequent time instances. We also introduce an RL algorithm that approximately optimizes a policy that produces such actions. It applies experience replay (ER) to adjust the likelihood of sequences of previous actions to optimize expected n -step returns that the policy yields. The efficiency of this algorithm is verified against four other RL methods [continuous deep advantage updating (CDAU), proximal policy optimization (PPO), soft actor-critic (SAC), and actor-critic with ER (ACER)] in four simulated learning control problems (Ant, HalfCheetah, Hopper, and Walker2D) in diverse time discretization. The algorithm introduced here outperforms the competitors in most cases considered. Jakub Lyskawa, Pawel Wawrzynski |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Actor-Critic with Variable Time Discretization via Sustained Actions
Jakub Lyskawa, Pawel Wawrzynski |
ICONIP (1) | 2 |
| 2023 | Least Redundant Gated Recurrent Neural NetworkabstractRecurrent neural networks are important tools for sequential data processing. However, they are notorious for problems regarding their training. Challenges include capturing complex relations between consecutive states and stability and efficiency of training. In this paper, we introduce a recurrent neural architecture called Deep Memory Update (DMU). It is based on updating the previous memory state with a deep transformation of the lagged state and the network input. The architecture is able to learn to transform its internal state using any nonlinear function. Its training is stable and fast due to relating its learning rate to the size of the module. Even though DMU is based on standard components, experimental results presented here confirm that it can compete with and often outperform state-of-the-art architectures such as Long Short-Term Memory, Gated Recurrent Units, and Recurrent Highway Networks. Lukasz Neumann, Lukasz Lepak, Pawel Wawrzynski |
IJCNN | 3 |
| 2022 | Reinforcement Learning for on-line Sequence TransformationabstractIn simultaneous machine translation (SMT), an output sequence should be produced as soon as possible, without reading the whole input sequence.This requirement creates a trade-off between translation delay and quality because less context may be known during translation.In most SMT methods, this trade-off is controlled with parameters whose values need to be tuned.In this paper, we introduce an SMT system that learns with reinforcement and is able to find the optimal delay in training.We conduct experiments on Tatoeba and IWSLT2014 datasets against state-of-the-art translation architectures.Our method achieves comparable results on the former dataset, with better results on long sentences and worse but comparable results on the latter dataset. Grzegorz Rypesc, Lukasz Lepak, Pawel Wawrzynski |
FedCSIS | 3 |
| 2022 | ReGAE: Graph Autoencoder Based on Recursive Neural Networks
Adam Malkowski, Jakub Grzechocinski, Pawel Wawrzynski |
ICONIP (4) | 3 |
| 2022 | Multiband VAE: Latent Space Alignment for Knowledge Consolidation in Continual LearningabstractWe propose a new method for unsupervised generative continual learning through realignment of Variational Autoencoder's latent space. Deep generative models suffer from catastrophic forgetting in the same way as other neural structures. Recent generative continual learning works approach this problem and try to learn from new data without forgetting previous knowledge. However, those methods usually focus on artificial scenarios where examples share almost no similarity between subsequent portions of data - an assumption not realistic in the real-life applications of continual learning. In this work, we identify this limitation and posit the goal of generative continual learning as a knowledge accumulation task. We solve it by continuously aligning latent representations of new data that we call bands in additional latent space where examples are encoded independently of their source task. In addition, we introduce a method for controlled forgetting of past data that simplifies this process. On top of the standard continual learning benchmarks, we propose a novel challenging knowledge consolidation scenario and show that the proposed approach outperforms state-of-the-art by up to twofold across all experiments and additional real-life evaluation. To our knowledge, Multiband VAE is the first method to show forward and backward knowledge transfer in generative continual learning. Kamil Deja, Pawel Wawrzynski, Wojciech Masarczyk, Daniel Marczak, Tomasz Trzcinski |
IJCAI | 2 |
| 2021 | BinPlay: A Binary Latent Autoencoder for Generative Replay Continual LearningabstractWe introduce a novel binary latent space autoen-coder architecture to rehearse training samples for the continual learning of neural networks. The ability to extend the knowledge of a model with new data without forgetting previously learned samples is a fundamental requirement in continual learning. Existing solutions address it by regularizing network weights, adjusting its architecture, or retraining with past data samples, regenerated from memory or reconstructed with generative models. Unfortunately, recreating past data from memory requires an infinite buffer, while the reconstructions of generative models tend to miss details of individual samples when generalizing beyond the training set. In this paper, we aim to overcome these limitations and introduce a novel generative rehearsal approach called BinPlay. Its main objective is to find a quality-preserving encoding of past samples into precomputed binary codes living in the autoencoder's binary latent space. Since we parametrize the formula for precomputing the codes only on the training samples' chronological indices, the autoencoder is able to compute the binary codes of rehearsed samples on the fly without the need to keep them in memory. Evaluation on three benchmark datasets shows up to a twofold accuracy improvement of BinPlay versus competing generative replay methods. Kamil Deja, Pawel Wawrzynski, Daniel Marczak, Wojciech Masarczyk, Tomasz Trzcinski |
IJCNN | 2 |
| 2020 | A Framework for Reinforcement Learning with Autocorrelated Actions
Marcin Szulc, Jakub Lyskawa, Pawel Wawrzynski |
ICONIP (2) | 3 |
| 2020 | DCT-Conv: Coding filters in convolutional networks with Discrete Cosine TransformabstractConvolutional neural networks are based on a huge number of trained weights. Consequently, they are often data-greedy, sensitive to overtraining, and learn slowly. We follow the line of research in which filters of convolutional neural layers are determined on the basis of a smaller number of trained parameters. In this paper, the trained parameters define a frequency spectrum which is transformed into convolutional filters with Inverse Discrete Cosine Transform (IDCT, the same is applied in decompression from JPEG). We analyze how switching off selected components of the spectra, thereby reducing the number of trained weights of the network, affects its performance. Our experiments show that coding the filters with trained DCT parameters leads to improvement over traditional convolution. Also, the performance of the networks modified this way decreases very slowly with the increasing extent of switching off these parameters. In some experiments, a good performance is observed when even 99.9% of these parameters are switched off. Karol Cheinski, Pawel Wawrzynski |
IJCNN | 2 |
| 2020 | Automatic hyperparameter tuning in on-line learning: Classic Momentum and ADAMabstractWe propose a method that adapts hyperparameters, namely step-sizes and momentum decay factors, in on-line learning with classic momentum and ADAM. The approach is based on the estimation of the short- and long-term influence of these hyperparameters on the loss value. In the experimental study, our approach is applied to on-line learning in small neural networks and deep autoencoders. Automatically tuned coefficients surpass or roughly match the best ones selected manually in terms of learning speed. As a result, on-line learning can be a fully automatic process, producing results from the first run, without preliminary experiments aimed at manual hyperparameter tuning. Pawel Wawrzynski, Pawel Zawistowski, Lukasz Lepak |
IJCNN | 1 |
| 2019 | Efficient on-line learning with diagonal approximation of loss function HessianabstractThe subject of this paper is stochastic optimization as a tool for on-line learning. New ingredients are introduced to Nesterov's Accelerated Gradient that increase efficiency of this algorithm and determine its parameters that are otherwise tuned manually: step-size and momentum decay factor. In this order a diagonal approximation of the Hessian of the loss function is estimated. In the experimental study the approach is applied to various types of neural networks, deep ones among others. Pawel Wawrzynski |
IJCNN | 1 |
| 2017 | ASD+M: Automatic parameter tuning in stochastic optimization and on-line learning
Pawel Wawrzynski |
Neural Networks | 1 |
| 2013 | dotRL: A platform for rapid Reinforcement Learning methods development and validation
Bartosz Papis, Pawel Wawrzynski |
FedCSIS | 2 |
| 2013 | Learning population of spiking neural networks with perturbation of conductancesabstractIn this paper a method is presented for learning of spiking neural networks. It is based on perturbation of synaptic conductances. While this approach is known to be model-free, it is also known to be slow, because it applies improvement direction estimates with large variance. Two ideas are analysed to alleviate this problem: First, learning of many networks at the same time instead of one. Second, autocorrelation of perturbations in time. In the experimental study the method is validated on three learning tasks in which information is conveyed with frequency and spike timing. Piotr Suszynski, Pawel Wawrzynski |
IJCNN | 2 |
| 2011 | Fixed point method for autonomous on-line neural network training
Pawel Wawrzynski, Bartosz Papis |
Neurocomputing | 1 |
| 2010 | Fixed point method of step-size estimation for on-line neural network trainingabstractThis paper considers on-line training of feadforward neural networks. Training examples are only available sampled randomly from a given generator. What emerges in this setting is the problem of step-sizes, or learning rates, adaptation. A scheme of determining step-sizes is introduced here that satisfies the following requirements: (i) it does not need any auxiliary problem-dependent parameters, (ii) it does not assume any particular loss function that the training process is intended to minimize, (iii) it makes the learning process stable and efficient. An experimental study with the 2D Gabor function approximation is presented. Pawel Wawrzynski |
IJCNN | 1 |
| 2009 | Real-time reinforcement learning by sequential Actor-Critics and experience replay
Pawel Wawrzynski |
Neural Networks | 1 |
| 2004 | Model-free off-policy reinforcement learning in continuous environmentabstractWe introduce an algorithm of reinforcement learning in continuous state and action spaces. In order to construct a control policy, the algorithm utilizes the entire history of agent-environment interaction. The policy is a result of an estimation process based on all available information rather than the result of stochastic convergence as in classical reinforcement learning approaches. The policy is derived from the history directly, not through any kind of a model of the environment. We test our algorithm in the cart-pole swing-up simulated environment. The algorithm learns to control this plant in about 100 trials, which corresponds to 15 minutes of plant's real time. This is several times shorter than the one required by other algorithms. Pawel Wawrzynski, Andrzej Pacut |
IJCNN | 1 |