Marek Grzes

dblp:81/6759 · DBLP profile ↗
← Back
26ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 The effect of attention in cooperative MARL environments with shared rewards
abstract
Scalability and coordination remain major challenges in training Multi-Agent Reinforcement Learning (MARL) algorithms. One approach postulates the Centralized Training and Decentralized Execution, which assumes full access to observations from the environment during training but limits agents' reliance on the joint observations within the execution phase. However, this often leads to a rapid increase in input dimensions of the centralized component (critic). Previous studies have suggested using attention mechanisms to enhance scalability and coordination in domains like Treasure Collection and Rover-Tower. This paper aims to complement these findings and offer new insights into the role of attention in MARL, focusing on these two domains on which attention was shown to be beneficial. We show that the impact of attention is very specific and different in the two domains studied. We use manually designed policies to inform our analysis, and explore the challenges concealed in the domains. We argue that the role of attention in the first domain is mainly to provide convenient inductive bias because the local observations of the agents surprisingly contain the same information as the joint observations. In the second domain, the local observations make the exploration challenging due to partial observability in one type of agents and the 'lazy agent' phenomenon. In this case, the role of centralized critic with attention is to mitigate the lazy agent phenomena and partial observability, and the attention itself acts as a simple averaging mechanism.
Aidin Kazempour, Marek Grzes
Neural Networks2
2025 Can Large Language Models Outperform Non-Experts in Poetry Evaluation? A Comparative Study Using the Consensual Assessment Technique
abstract
This study adapts the Consensual Assessment Technique (CAT) for Large Language Models (LLMs), introducing a novel methodology for poetry evaluation.Using a 90-poem dataset with a ground truth based on publication venue, we demonstrate that this approach allows LLMs to significantly surpass the performance of non-expert human judges.Our method, which leverages forced-choice ranking within small, randomized batches, enabled Claude-3-Opus to achieve a Spearman's Rank Correlation of 0.87 with the ground truth, dramatically outperforming the best human nonexpert evaluation (SRC = 0.38).The LLM assessments also exhibited high inter-rater reliability, underscoring the methodology's robustness.These findings establish that LLMs, when guided by a comparative framework, can be effective and reliable tools for assessing poetry, paving the way for their broader application in other creative domains.
Piotr Sawicki 0001, Marek Grzes, Dan Brown 0001, Fabrício Góes
EMNLP2
2024 Deconstructing Deep Active Inference: A Contrarian Information Gatherer
abstract
Active inference is a theory of perception, learning, and decision making that can be applied to neuroscience, robotics, psychology, and machine learning. Recently, intensive research has been taking place to scale up this framework using Monte Carlo tree search and deep learning. The goal of this activity is to solve more complicated tasks using deep active inference. First, we review the existing literature and then progressively build a deep active inference agent as follows: we (1) implement a variational autoencoder (VAE), (2) implement a deep hidden Markov model (HMM), and (3) implement a deep critical hidden Markov model (CHMM). For the CHMM, we implemented two versions, one minimizing expected free energy, CHMM[EFE] and one maximizing rewards, CHMM[reward]. Then we experimented with three different action selection strategies: the ε-greedy algorithm as well as softmax and best action selection. According to our experiments, the models able to solve the dSprites environment are the ones that maximize rewards. On further inspection, we found that the CHMM minimizing expected free energy almost always picks the same action, which makes it unable to solve the dSprites environment. In contrast, the CHMM maximizing reward keeps on selecting all the actions, enabling it to successfully solve the task. The only difference between those two CHMMs is the epistemic value, which aims to make the outputs of the transition and encoder networks as close as possible. Thus, the CHMM minimizing expected free energy repeatedly picks a single action and becomes an expert at predicting the future when selecting this action. This effectively makes the KL divergence between the output of the transition and encoder networks small. Additionally, when selecting the action down the average reward is zero, while for all the other actions, the expected reward will be negative. Therefore, if the CHMM has to stick to a single action to keep the KL divergence small, then the action down is the most rewarding. We also show in simulation that the epistemic value used in deep active inference can behave degenerately and in certain circumstances effectively lose, rather than gain, information. As the agent minimizing EFE is not able to explore its environment, the appropriate formulation of the epistemic value in deep active inference remains an open question.
Théophile Champion, Marek Grzes, Lisa Bonheme, Howard Bowman
Neural Comput.2
2023 Is GPT-4 Good Enough to Evaluate Jokes?
Fabrício Góes, Piotr Sawicki 0001, Marek Grzes, Marco Volpe 0001, Dan Brown 0001
ICCC3
2023 Pushing GPT's Creativity to Its Limits: Alternative Uses and Torrance Tests
Fabrício Góes, Piotr Sawicki 0001, Marek Grzes, Marco Volpe 0001, Jacob Watson
ICCC3
2023 Bits of Grass: Does GPT already know how to write like Whitman?
Piotr Sawicki 0001, Marek Grzes, Fabrício Góes, Dan Brown 0001, Max Peeperkorn, Aisha Khatun
ICCC2
2023 On the power of special-purpose GPT models to create and evaluate new poetry in old styles
Piotr Sawicki 0001, Marek Grzes, Fabrício Góes, Anna Jordanous, Dan Brown 0001, Simona Paraskevopoulou, Max Peeperkorn, Aisha Khatun
ICCC2
2023 Posterior Collapse in Variational Gradient Origin Networks
abstract
Posterior collapse is a phenomenon that occurs when the posterior distribution degenerates to the prior, leading to a decline in the quality of latent encodings and generative models. While it is known to occur in Variational Autoencoders (VAEs), it is unknown whether it occurs in Variational Gradient Origin Networks (VGONs). The goal of this paper is to compare the posterior collapse of Variational Gradient Origin Networks and Variational Autoencoders. By checking the latent encodings of VGONs against the key posterior collapse metrics, our experiments reveal that VGONs do exhibit posterior collapse both in the decline of the Kullback-Leibler divergence (KLD) and the collapse of individual variables. Furthermore, the results show that VGONs and VAEs have a similar polarized regime, suggesting that the cause of posterior collapse is not specific to the architecture of the model used to find an encoding. These findings support the claim made in previous research that posterior collapse is a general issue that affects a wide range of latent variable models.
Peter Clapham, Marek Grzes
ICMLA2
2023 Be More Active! Understanding the Differences Between Mean and Sampled Representations of Variational Autoencoders
abstract
The ability of Variational Autoencoders to learn disentangled representations has made them appealing for practical applications. However, their mean representations, which are generally used for downstream tasks, have recently been shown to be more correlated than their sampled counterpart, on which disentanglement is usually measured. In this paper, we refine this observation through the lens of selective posterior collapse, which states that only a subset of the learned representations, the active variables, is encoding useful information while the rest (the passive variables) is discarded. We first extend the existing definition to multiple data examples and show that active variables are equally disentangled in mean and sampled representations. Based on this extension and the pre-trained models from disentanglement_lib}, we then isolate the passive variables and show that they are responsible for the discrepancies between mean and sampled representations. Specifically, passive variables exhibit high correlation scores with other variables in mean representations while being fully uncorrelated in sampled ones. We thus conclude that despite what their higher correlation might suggest, mean representations are still good candidates for downstream tasks applications. However, it may be beneficial to remove their passive variables, especially when used with models sensitive to correlated features.
Lisa Bonheme, Marek Grzes
J. Mach. Learn. Res.2
2022 Training GPT-2 to represent two Romantic-era authors: challenges, evaluations and pitfalls
Piotr Sawicki 0001, Marek Grzes, Anna Jordanous, Dan Brown 0001, Max Peeperkorn
ICCC2
2022 Branching Time Active Inference with Bayesian Filtering
abstract
Branching time active inference is a framework proposing to look at planning as a form of Bayesian model expansion. Its root can be found in active inference, a neuroscientific framework widely used for brain modeling, as well as in Monte Carlo tree search, a method broadly applied in the reinforcement learning literature. Up to now, the inference of the latent variables was carried out by taking advantage of the flexibility offered by variational message passing, an iterative process that can be understood as sending messages along the edges of a factor graph. In this letter, we harness the efficiency of an alternative method for inference, Bayesian filtering, which does not require the iteration of the update equations until convergence of the variational free energy. Instead, this scheme alternates between two phases: integration of evidence and prediction of future states. Both phases can be performed efficiently, and this provides a forty times speedup over the state of the art.
Théophile Champion, Marek Grzes, Howard Bowman
Neural Comput.2
2022 Branching time active inference: Empirical study and complexity class analysis
abstract
Active inference is a state-of-the-art framework for modelling the brain that explains a wide range of mechanisms such as habit formation, dopaminergic discharge and curiosity. However, recent implementations suffer from an exponential (space and time) complexity class when computing the prior over all the possible policies up to the time horizon. Fountas et al. (2020) used Monte Carlo tree search to address this problem, leading to very good results in two different tasks. Additionally, Champion et al. (2021a) proposed a tree search approach based on (temporal) structure learning. This was enabled by the development of a variational message passing approach to active inference (Champion, Bowman, Grześ, 2021), which enables compositional construction of Bayesian networks for active inference. However, this message passing tree search approach, which we call branching-time active inference (BTAI), has never been tested empirically. In this paper, we present an experimental study of the approach (Champion, Grześ, Bowman, 2021) in the context of a maze solving agent. In this context, we show that both improved prior preferences and deeper search help mitigate the vulnerability to local minima. Then, we compare BTAI to standard active inference (AcI) on a graph navigation task. We show that for small graphs, both BTAI and AcI successfully solve the task. For larger graphs, AcI exhibits an exponential (space) complexity class, making the approach intractable. However, BTAI explores the space of policies more efficiently, successfully scaling to larger graphs. Then, BTAI was compared to the POMCP algorithm (Silver and Veness, 2010) on the frozen lake environment. The experiments suggest that BTAI and the POMCP algorithm accumulate a similar amount of reward. Also, we describe when BTAI receives more rewards than the POMCP agent, and when the opposite is true. Finally, we compared BTAI to the approach of Fountas et al. (2020) on the dSprites dataset, and we discussed the pros and cons of each approach.
Théophile Champion, Howard Bowman, Marek Grzes
Neural Networks3
2022 Branching Time Active Inference: The theory and its generality
abstract
Over the last 10 to 15 years, active inference has helped to explain various brain mechanisms from habit formation to dopaminergic discharge and even modelling curiosity. However, the current implementations suffer from an exponential (space and time) complexity class when computing the prior over all the possible policies up to the time-horizon. Fountas et al. (2020) used Monte Carlo tree search to address this problem, leading to impressive results in two different tasks. In this paper, we present an alternative framework that aims to unify tree search and active inference by casting planning as a structure learning problem. Two tree search algorithms are then presented. The first propagates the expected free energy forward in time (i.e., towards the leaves), while the second propagates it backward (i.e., towards the root). Then, we demonstrate that forward and backward propagations are related to active inference and sophisticated inference, respectively, thereby clarifying the differences between those two planning strategies.
Théophile Champion, Lancelot Da Costa, Howard Bowman, Marek Grzes
Neural Networks4
2021 Realizing Active Inference in Variational Message Passing: The Outcome-Blind Certainty Seeker
abstract
Active inference is a state-of-the-art framework in neuroscience that offers a unified theory of brain function. It is also proposed as a framework for planning in AI. Unfortunately, the complex mathematics required to create new models can impede application of active inference in neuroscience and AI research. This letter addresses this problem by providing a complete mathematical treatment of the active inference framework in discrete time and state spaces and the derivation of the update equations for any new model. We leverage the theoretical connection between active inference and variational message passing as described by John Winn and Christopher M. Bishop in 2005. Since variational message passing is a well-defined methodology for deriving Bayesian belief update equations, this letter opens the door to advanced generative models for active inference. We show that using a fully factorized variational distribution simplifies the expected free energy, which furnishes priors over policies so that agents seek unambiguous states. Finally, we consider future extensions that support deep tree searches for sequential policy optimization based on structure learning and belief propagation.
Théophile Champion, Marek Grzes, Howard Bowman
Neural Comput.2
2019 Comparing Explanations between Random Forests and Artificial Neural Networks
abstract
The decisions made by machines are increasingly comparable in predictive performance to those made by humans, but these decision making processes are often concealed as black boxes. Additional techniques are required to extract understanding, and one such category are explanation methods. This research compares the explanations of two popular forms of artificial intelligence; neural networks and random forests. Researchers in either field often have divided opinions on transparency, and comparing explanations may discover similar ground truths between models. Similarity can help to encourage trust in predictive accuracy alongside transparent structure and unite the respective research fields. This research explores a variety of simulated and real-world datasets that ensure fair applicability to both learning algorithms. A new heuristic explanation method that extends an existing technique is introduced, and our results show that this is somewhat similar to the other methods examined whilst also offering an alternative perspective towards least-important features.
Lee Harris, Marek Grzes
SMC2
2018 Improving Language Modelling with Noise Contrastive Estimation
abstract
Neural language models do not scale well when the vocabulary is large. Noise contrastive estimation (NCE) is a sampling-based method that allows for fast learning with large vocabularies. Although NCE has shown promising performance in neural machine translation, its full potential has not been demonstrated in the language modelling literature. A sufficient investigation of the hyperparameters in the NCE-based neural language models was clearly missing. In this paper, we showed that NCE can be a very successful approach in neural language modelling when the hyperparameters of a neural network are tuned appropriately. We introduced the `search-then-converge' learning rate schedule for NCE and designed a heuristic that specifies how to use this schedule. The impact of the other important hyperparameters, such as the dropout rate and the weight initialisation range, was also demonstrated. Using a popular benchmark, we showed that appropriate tuning of NCE in neural language models outperforms the state-of-the-art single-model methods based on standard dropout and the standard LSTM recurrent neural networks.
Farhana Ferdousi Liza, Marek Grzes
AAAI2
2015 Energy Efficient Execution of POMDP Policies
abstract
Recent advances in planning techniques for partially observable Markov decision processes (POMDPs) have focused on online search techniques and offline point-based value iteration. While these techniques allow practitioners to obtain policies for fairly large problems, they assume that a nonnegligible amount of computation can be done between each decision point. In contrast, the recent proliferation of mobile and embedded devices has lead to a surge of applications that could benefit from state-of-the-art planning techniques if they can operate under severe constraints on computational resources. To that effect, we describe two techniques to compile policies into controllers that can be executed by a mere table lookup at each decision point. The first approach compiles policies induced by a set of alpha vectors (such as those obtained by point-based techniques) into approximately equivalent controllers, while the second approach performs a simulation to compile arbitrary policies into approximately equivalent controllers. We also describe an approach to compress controllers by removing redundant and dominated nodes, often yielding smaller and yet better controllers. Further compression and higher value can sometimes be obtained by considering stochastic controllers. The compilation and compression techniques are demonstrated on benchmark problems as well as a mobile application to help persons with Alzheimer's to way-find. The battery consumption of several POMDP policies is compared against finite-state controllers learned using methods introduced in this paper. Experiments performed on the Nexus 4 phone show that finite-state controllers are the least battery consuming POMDP policies.
Marek Grzes, Pascal Poupart, Jesse Hoey
IEEE Trans. Cybern.1
2014 Multi-test decision tree and its application to microarray data classification
Marcin Czajkowski, Marek Grzes, Marek Kretowski
Artif. Intell. Medicine2
2014 Relational approach to knowledge engineering for POMDP-based assistance systems as a translation of a psychological model
Marek Grzes, Jesse Hoey, Shehroz S. Khan, Alex Mihailidis, Stephen Czarnuch, Daniel Jackson 0002, Andrew F. Monk
Int. J. Approx. Reason.1
2013 Isomorph-Free Branch and Bound Search for Finite State Controllers
Marek Grzes, Pascal Poupart, Jesse Hoey
IJCAI1
2013 On the convergence of techniques that improve value iteration
abstract
Prioritisation of Bellman backups or updating only a small subset of actions represent important techniques for speeding up planning in MDPs. The recent literature showed new efficient approaches which exploit these directions. Backward value iteration and backing up only the best actions were shown to lead to a significant reduction of the planning time. This paper conducts a theoretical and empirical analysis of these techniques and shows new important proofs. In particular, (1) it identifies weaker requirements for the convergence of backups based on best actions only, (2) a new method for evaluation of the Bellman error is shown for the update that updates one best action once, (3) it presents the theoretical proof of backward value iteration and establishes required initialisation, (4) and shows that the default state ordering of backups in standard value iteration can significantly influence its performance. Additionally, (5) the existing literature did not compare these methods, either empirically or analytically, against policy iteration. The rigorous empirical and novel theoretical parts of the paper reveal important associations and allow drawing guidelines on which type of value or policy iteration is suitable for a given domain. Finally, our chief message is that standard value iteration can be made far more efficient by simple modifications shown in the paper.
Marek Grzes, Jesse Hoey
IJCNN1
2010 Online learning of shaping rewards in reinforcement learning
Marek Grzes, Daniel Kudenko
Neural Networks1
2009 Theoretical and Empirical Analysis of Reward Shaping in Reinforcement Learning
abstract
Reinforcement learning suffers scalability problems due to the state space explosion and the temporal credit assignment problem. Knowledge-based approaches have received a significant attention in the area. Reward shaping is a particular approach to incorporate domain knowledge into reinforcement learning. Theoretical and empirical analysis of this paper reveals important properties of this principle, especially the influence of the reward type, MDP discount factor, and the way of evaluating the potential function on the performance.
Marek Grzes, Daniel Kudenko
ICMLA1
2008 Multigrid Reinforcement Learning with Reward Shaping
Marek Grzes, Daniel Kudenko
ICANN (1)1
2006 Mixed Decision Trees: An Evolutionary Approach
Marek Kretowski, Marek Grzes
DaWaK2
2006 Evolutionary Induction of Cost-Sensitive Decision Trees
Marek Kretowski, Marek Grzes
ISMIS2