VLDB 2026 Research / reviewers in the wild / expert
Marco Pleines
dblp:249/7666
· DBLP profile ↗
7ranked-venue papers
6as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pokémon Red via Reinforcement LearningabstractWe present a Deep Reinforcement Learning (DRL) agent that successfully completes the first several hours of Pokémon Red, a classic Game Boy JRPG that exposes significant challenges as a testbed for agents, including multitasking, long horizons of tens of thousands of steps, hard exploration, and a vast array of potential policies. Our agent completes an initial segment of the game, up to Cerulean City, a location requiring progression through two cities, battle-filled passages, a maze-like cave, and defeating the first gym leader. Our experiments include various ablations that reveal vulnerabilities in reward shaping. We argue that long-form games like Pokémon hold strong potential for future research, presenting long-term coherent reasoning challenges absent from simpler arcade games. Our environment wrapper, training algorithm, human replay data, and pretrained agent are available at (REDACTED FOR REVIEW). Marco Pleines, Daniel Addis, David Rubinstein, Frank Zimmer, Mike Preuss, Peter Whidden |
CoG | 1 |
| 2025 | Memory Gym: Towards Endless Tasks to Benchmark Memory Capabilities of AgentsabstractMemory Gym presents a suite of 2D partially observable environments, namely Mortar Mayhem, Mystery Path, and Searing Spotlights, designed to benchmark memory capabilities in decision-making agents. These environments, originally with finite tasks, are expanded into innovative, endless formats, mirroring the escalating challenges of cumulative memory games such as “I packed my bag”. This progression in task design shifts the focus from merely assessing sample efficiency to also probing the levels of memory effectiveness in dynamic, prolonged scenarios. To address the gap in available memory-based Deep Reinforcement Learning baselines, we introduce an implementation within the open-source CleanRL library that integrates Transformer-XL (TrXL) with Proximal Policy Optimization. This approach utilizes TrXL as a form of episodic memory, employing a sliding window technique. Our comparative study between the Gated Recurrent Unit (GRU) and TrXL reveals varied performances across our finite and endless tasks. TrXL, on the finite environments, demonstrates superior effectiveness over GRU, but only when utilizing an auxiliary loss to reconstruct observations. Notably, GRU makes a remarkable resurgence in all endless tasks, consistently outperforming TrXL by significant margins. Website and Source Code: https://marcometer.github.io/jmlr_2024.github.io/ Marco Pleines, Matthias Pallasch, Frank Zimmer, Mike Preuss |
J. Mach. Learn. Res. | 1 |
| 2023 | Memory Gym: Partially Observable Challenges to Memory-Based Agents
Marco Pleines, Matthias Pallasch, Frank Zimmer, Mike Preuss |
ICLR | 1 |
| 2022 | On the Verge of Solving Rocket League using Deep Reinforcement Learning and Sim-to-sim TransferabstractAutonomously trained agents that are supposed to play video games reasonably well rely either on fast simulation speeds or heavy parallelization across thousands of machines running concurrently. This work explores a third way that is established in robotics, namely sim-to-real transfer, or if the game is considered a simulation itself, sim-to-sim transfer. In the case of Rocket League, we demonstrate that single behaviors of goalies and strikers can be successfully learned using Deep Reinforcement Learning in the simulation environment and transferred back to the original game. Although the implemented training simulation is to some extent inaccurate, the goalkeeping agent saves nearly 100% of its faced shots once transferred, while the striking agent scores in about 75% of cases. Therefore, the trained agent is robust enough and able to generalize to the target domain of Rocket League. Marco Pleines, Konstantin Ramthun, Yannik Wegener, Hendrik Meyer, Matthias Pallasch, Sebastian Prior, Jannik Drögemüller, Leon Büttinghaus, Thilo Röthemeyer, Alexander Kaschwig, Oliver Chmurzynski, Frederik Rohkrähmer, Roman Kalkreuth, Frank Zimmer, Mike Preuss |
CoG | 1 |
| 2022 | Improving Bidding and Playing Strategies in the Trick-Taking game Wizard using Deep Q-NetworksabstractIn this work, the trick-taking game Wizard with a separate bidding and playing phase is modeled by two interleaved partially observable Markov decision processes (POMDP). Deep Q-Networks (DQN) are used to empower self-improving agents, which are capable of tackling the challenges of a highly non-stationary environment. To compare algorithms between each other, the accuracy between bid and trick count is monitored, which strongly correlates with the actual rewards and provides a well-defined upper and lower performance bound. The trained DQN agents achieve accuracies between 66% and 87% in self-play, leaving behind both a random baseline and a rule-based heuristic. The conducted analysis also reveals a strong information asymmetry concerning player positions during bidding. To overcome the missing Markov property of imperfect-information games, a long short-term memory (LSTM) network is implemented to integrate historic information into the decision-making process. Additionally, a forward-directed tree search is conducted by sampling a state of the environment and thereby turning the game into a perfect information setting. To our surprise, both approaches do not surpass the performance of the basic DQN agent. Jonas Schumacher, Marco Pleines |
CoG | 2 |
| 2020 | Obstacle Tower Without Human Demonstrations: How Far a Deep Feed-Forward Network Goes with Reinforcement LearningabstractThe Obstacle Tower Challenge is the task to master a procedurally generated chain of levels that subsequently get harder to complete. Whereas the most top performing entries of last year's competition used human demonstrations or reward shaping to learn how to cope with the challenge, we present an approach that performed competitively (placed 7th) but starts completely from scratch by means of Deep Reinforcement Learning with a relatively simple feed-forward deep network structure. We especially look at the generalization performance of the taken approach concerning different seeds and various visual themes that have become available after the competition, and investigate where the agent fails and why. Note that our approach does not possess a short-term memory like employing recurrent hidden states. With this work, we hope to contribute to a better understanding of what is possible with a relatively simple, flexible solution that can be applied to learning in environments featuring complex 3D visual input where the abstract task structure itself is still fairly simple. Marco Pleines, Jenia Jitsev, Mike Preuss, Frank Zimmer |
CoG | 1 |
| 2019 | Action Spaces in Deep Reinforcement Learning to Mimic Human Input DevicesabstractEnabling agents to generally play video games requires to implement a common action space that mimics human input devices like a gamepad. Such action spaces have to support concurrent discrete and continuous actions. To solve this problem, this work investigates three approaches to examine the application of concurrent discrete and continuous actions in Deep Reinforcement Learning (DRL). One approach implements a threshold to discretize a continuous action, while another one divides a continuous action into multiple discrete actions. The third approach creates a multiagent to combine both action kinds. These approaches are benchmarked by two novel environments. In the first environment (Shooting Birds) the goal of the agent is to accurately shoot birds by controlling a cross-hair. The second environment is a simplification of the game Beastly Rivals On-slaught, where the agent is in charge of its controlled character's survival. Throughout multiple experiments, the bucket approach is recommended, because it is trained faster than the multiagent and is more stable than the threshold approach. Due to the contributions of this paper, consecutive work can start training agents using visual observations. Marco Pleines, Frank Zimmer, Vincent-Pierre Berges |
CoG | 1 |