Ryan Kortvelesy

dblp:289/0863 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 58% Knowledge representation and reasoning · 17% Graph learning · 14%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Knowledge representation and reasoning › cognitive modeling
memory models
1.422024
Recurrent Reinforcement Learning with Memoroids · NeurIPS 2024
Reinforcement Learning with Fast and Forgetful Memory · NeurIPS 2023
Machine learning › Reinforcement learning › partially observable reinforcement learning › memory-based reinforcement learning
recurrent reinforcement learning
1.422024
Recurrent Reinforcement Learning with Memoroids · NeurIPS 2024
Reinforcement Learning with Fast and Forgetful Memory · NeurIPS 2023
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.322024
Controlling Behavioral Diversity in Multi-Agent Reinforcement Learning · ICML 2024
ModGNN: Expert Policy Approximation in Multi-Agent Systems with a Modular Graph Neural Network Architecture · ICRA 2021
Machine learning › Graph learning
graph neural network
1.222023
Generalised f-Mean Aggregation for Graph Neural Networks · NeurIPS 2023
ModGNN: Expert Policy Approximation in Multi-Agent Systems with a Modular Graph Neural Network Architecture · ICRA 2021
Machine learning › Reinforcement learning
actor-critic methods
0.812024
Controlling Behavioral Diversity in Multi-Agent Reinforcement Learning · ICML 2024
Machine learning › Reinforcement learning › multi-agent reinforcement learning
behavioral diversity
0.812024
Controlling Behavioral Diversity in Multi-Agent Reinforcement Learning · ICML 2024
Machine learning › Reinforcement learning
partially observable reinforcement learning
0.712023
POPGym: Benchmarking Partially Observable Reinforcement Learning · ICLR 2023
Knowledge, reasoning and agents › Multi-agent systems
multi-agent coordination
0.512021
ModGNN: Expert Policy Approximation in Multi-Agent Systems with a Modular Graph Neural Network Architecture · ICRA 2021
Machine learning › Deep learning architectures and training
recurrent neural network
0.422024
Recurrent Reinforcement Learning with Memoroids · NeurIPS 2024
Reinforcement Learning with Fast and Forgetful Memory · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

policy architecture constraints · 0.8monoid-based framework · 0.8memoroids · 0.8linear recurrent models · 0.8intrinsic reward · 0.8parametrized aggregation functions · 0.7fast and forgetful memory · 0.7benchmarking · 0.7graph convolutional network · 0.5ablation study · 0.5
YearPublicationVenuePosition
2024 Controlling Behavioral Diversity in Multi-Agent Reinforcement Learning
abstract
The study of behavioral diversity in Multi-Agent Reinforcement Learning (MARL) is a nascent yet promising field. In this context, the present work deals with the question of how to control the diversity of a multi-agent system. With no existing approaches to control diversity to a set value, current solutions focus on blindly promoting it via intrinsic rewards or additional loss functions, effectively changing the learning objective and lacking a principled measure for it. To address this, we introduce Diversity Control (DiCo), a method able to control diversity to an exact value of a given metric by representing policies as the sum of a parameter-shared component and dynamically scaled per-agent components. By applying constraints directly to the policy architecture, DiCo leaves the learning objective unchanged, enabling its applicability to any actor-critic MARL algorithm. We theoretically prove that DiCo achieves the desired diversity, and we provide several experiments, both in cooperative and competitive tasks, that show how DiCo can be employed as a novel paradigm to increase performance and sample efficiency in MARL. Multimedia results are available on the paper’s website: https://sites.google.com/view/dico-marl
Matteo Bettini, Ryan Kortvelesy, Amanda Prorok
ICML2
2024 Recurrent Reinforcement Learning with Memoroids
abstract
Memory models such as Recurrent Neural Networks (RNNs) and Transformers address Partially Observable Markov Decision Processes (POMDPs) by mapping trajectories to latent Markov states. Neither model scales particularly well to long sequences, especially compared to an emerging class of memory models called Linear Recurrent Models. We discover that the recurrent update of these models resembles a monoid, leading us to reformulate existing models using a novel monoid-based framework that we call memoroids. We revisit the traditional approach to batching in recurrent reinforcement learning, highlighting theoretical and empirical deficiencies. We leverage memoroids to propose a batching method that improves sample efficiency, increases the return, and simplifies the implementation of recurrent loss functions in reinforcement learning.
Steven D. Morad, Chris Lu 0001, Ryan Kortvelesy, Stephan Liwicki, Jakob N. Foerster, Amanda Prorok
NeurIPS3
2023 POPGym: Benchmarking Partially Observable Reinforcement Learning
Steven D. Morad, Ryan Kortvelesy, Matteo Bettini, Stephan Liwicki, Amanda Prorok
ICLR2
2023 Generalised f-Mean Aggregation for Graph Neural Networks
abstract
Graph Neural Network (GNN) architectures are defined by their implementations of update and aggregation modules. While many works focus on new ways to parametrise the update modules, the aggregation modules receive comparatively little attention. Because it is difficult to parametrise aggregation functions, currently most methods select a ``standard aggregator'' such as mean, sum, or max. While this selection is often made without any reasoning, it has been shown that the choice in aggregator has a significant impact on performance, and the best choice in aggregator is problem-dependent. Since aggregation is a lossy operation, it is crucial to select the most appropriate aggregator in order to minimise information loss. In this paper, we present GenAgg, a generalised aggregation operator, which parametrises a function space that includes all standard aggregators. In our experiments, we show that GenAgg is able to represent the standard aggregators with much higher accuracy than baseline methods. We also show that using GenAgg as a drop-in replacement for an existing aggregator in a GNN often leads to a significant boost in performance across various tasks.
Ryan Kortvelesy, Steven D. Morad, Amanda Prorok
NeurIPS1
2023 Reinforcement Learning with Fast and Forgetful Memory
abstract
Nearly all real world tasks are inherently partially observable, necessitating the use of memory in Reinforcement Learning (RL). Most model-free approaches summarize the trajectory into a latent Markov state using memory models borrowed from Supervised Learning (SL), even though RL tends to exhibit different training and efficiency characteristics. Addressing this discrepancy, we introduce Fast and Forgetful Memory, an algorithm-agnostic memory model designed specifically for RL. Our approach constrains the model search space via strong structural priors inspired by computational psychology. It is a drop-in replacement for recurrent neural networks (RNNs) in recurrent RL algorithms, achieving greater reward than RNNs across various recurrent benchmarks and algorithms _without changing any hyperparameters_. Moreover, Fast and Forgetful Memory exhibits training speeds two orders of magnitude faster than RNNs, attributed to its logarithmic time and linear space complexity. Our implementation is available at https://github.com/proroklab/ffm.
Steven D. Morad, Ryan Kortvelesy, Stephan Liwicki, Amanda Prorok
NeurIPS2
2021 ModGNN: Expert Policy Approximation in Multi-Agent Systems with a Modular Graph Neural Network Architecture
abstract
Recent work in the multi-agent domain has shown the promise of Graph Neural Networks (GNNs) to learn complex coordination strategies. However, most current approaches use minor variants of a Graph Convolutional Network (GCN), which applies a convolution to the communication graph formed by the multi-agent system. In this paper, we investigate whether the performance and generalization of GCNs can be improved upon. We introduce ModGNN, a decentralized framework which serves as a generalization of GCNs, providing more flexibility. To test our hypothesis, we evaluate an implementation of ModGNN against several baselines in the multi-agent flocking problem. We perform an ablation analysis to show that the most important component of our framework is one that does not exist in a GCN. By varying the number of agents, we also demonstrate that an application-agnostic implementation of ModGNN possesses an improved ability to generalize to new environments.
Ryan Kortvelesy, Amanda Prorok
ICRA1