Markus Hofmarcher

dblp:224/9960 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
3since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 44% Deep learning architectures and training · 19% Trustworthy machine learning · 12%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning paradigms
continual learning
0.712023
Learning to Modulate pre-trained Models in RL · NeurIPS 2023
Machine learning › Trustworthy machine learning
interpretability
0.712023
Semantic HELM: A Human-Readable Memory for Reinforcement Learning · NeurIPS 2023
Machine learning › Deep learning architectures and training
memory mechanism
0.712023
Semantic HELM: A Human-Readable Memory for Reinforcement Learning · NeurIPS 2023
Machine learning › Reinforcement learning
partially observable reinforcement learning
0.712023
Semantic HELM: A Human-Readable Memory for Reinforcement Learning · NeurIPS 2023
Machine learning › Reinforcement learning
transfer learning in reinforcement learning
0.712023
Learning to Modulate pre-trained Models in RL · NeurIPS 2023
Robotics › Robot manipulation
learning from demonstration
0.612022
Align-RUDDER: Learning From Few Demonstrations by Reward Redistribution · ICML 2022
Machine learning › Reinforcement learning › reward design › reward shaping
reward redistribution
0.612022
Align-RUDDER: Learning From Few Demonstrations by Reward Redistribution · ICML 2022
Machine learning › Reinforcement learning › reward design
reward shaping
0.612022
Align-RUDDER: Learning From Few Demonstrations by Reward Redistribution · ICML 2022
Machine learning › Deep learning architectures and training
convolutional neural network
0.412019
Human-level Protein Localization with Convolutional Neural Networks · ICLR (Poster) 2019
Bioinformatics and computational biology › protein function prediction
protein localization
0.412019
Human-level Protein Localization with Convolutional Neural Networks · ICLR (Poster) 2019
Bioinformatics and computational biology › protein function prediction
protein subcellular localization prediction
0.412019
Human-level Protein Localization with Convolutional Neural Networks · ICLR (Poster) 2019
Natural language and speech › Language models and text generation › LLM agents
agent memory
0.212023
Semantic HELM: A Human-Readable Memory for Reinforcement Learning · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

convolutional neural network · 0.8pre-training · 0.7pre-trained language model · 0.7modulation · 0.7CLIP · 0.7profile model · 0.6multiple sequence alignment · 0.6
YearPublicationVenuePosition
2023 Semantic HELM: A Human-Readable Memory for Reinforcement Learning
abstract
Reinforcement learning agents deployed in the real world often have to cope with partially observable environments. Therefore, most agents employ memory mechanisms to approximate the state of the environment. Recently, there have been impressive success stories in mastering partially observable environments, mostly in the realm of computer games like Dota 2, StarCraft II, or MineCraft. However, existing methods lack interpretability in the sense that it is not comprehensible for humans what the agent stores in its memory. In this regard, we propose a novel memory mechanism that represents past events in human language. Our method uses CLIP to associate visual inputs with language tokens. Then we feed these tokens to a pretrained language model that serves the agent as memory and provides it with a coherent and human-readable representation of the past. We train our memory mechanism on a set of partially observable environments and find that it excels on tasks that require a memory component, while mostly attaining performance on-par with strong baselines on tasks that do not. On a challenging continuous recognition task, where memorizing the past is crucial, our memory mechanism converges two orders of magnitude faster than prior methods. Since our memory mechanism is human-readable, we can peek at an agent's memory and check whether crucial pieces of information have been stored. This significantly enhances troubleshooting and paves the way toward more interpretable agents.
Fabian Paischer, Thomas Adler, Markus Hofmarcher, Sepp Hochreiter
NeurIPS3
2023 Learning to Modulate pre-trained Models in RL
abstract
Reinforcement Learning (RL) has been successful in various domains like robotics, game playing, and simulation. While RL agents have shown impressive capabilities in their specific tasks, they insufficiently adapt to new tasks. In supervised learning, this adaptation problem is addressed by large-scale pre-training followed by fine-tuning to new down-stream tasks. Recently, pre-training on multiple tasks has been gaining traction in RL. However, fine-tuning a pre-trained model often suffers from catastrophic forgetting. That is, the performance on the pre-training tasks deteriorates when fine-tuning on new tasks. To investigate the catastrophic forgetting phenomenon, we first jointly pre-train a model on datasets from two benchmark suites, namely Meta-World and DMControl. Then, we evaluate and compare a variety of fine-tuning methods prevalent in natural language processing, both in terms of performance on new tasks, and how well performance on pre-training tasks is retained. Our study shows that with most fine-tuning approaches, the performance on pre-training tasks deteriorates significantly. Therefore, we propose a novel method, Learning-to-Modulate (L2M), that avoids the degradation of learned skills by modulating the information flow of the frozen pre-trained model via a learnable modulation pool. Our method achieves state-of-the-art performance on the Continual-World benchmark, while retaining performance on the pre-training tasks. Finally, to aid future research in this area, we release a dataset encompassing 50 Meta-World and 16 DMControl tasks.
Thomas Schmied, Markus Hofmarcher, Fabian Paischer, Razvan Pascanu, Sepp Hochreiter
NeurIPS2
2022 Align-RUDDER: Learning From Few Demonstrations by Reward Redistribution
abstract
Reinforcement learning algorithms require many samples when solving complex hierarchical tasks with sparse and delayed rewards. For such complex tasks, the recently proposed RUDDER uses reward redistribution to leverage steps in the Q-function that are associated with accomplishing sub-tasks. However, often only few episodes with high rewards are available as demonstrations since current exploration strategies cannot discover them in reasonable time. In this work, we introduce Align-RUDDER, which utilizes a profile model for reward redistribution that is obtained from multiple sequence alignment of demonstrations. Consequently, Align-RUDDER employs reward redistribution effectively and, thereby, drastically improves learning on few demonstrations. Align-RUDDER outperforms competitors on complex artificial tasks with delayed rewards and few demonstrations. On the Minecraft ObtainDiamond task, Align-RUDDER is able to mine a diamond, though not frequently. Code is available at github.com/ml-jku/align-rudder.
Vihang Patil, Markus Hofmarcher, Marius-Constantin Dinu, Matthias Dorfer, Patrick M. Blies, Johannes Brandstetter, Jose A. Arjona-Medina, Sepp Hochreiter
ICML2
2019 Human-level Protein Localization with Convolutional Neural Networks
Elisabeth Rumetshofer, Markus Hofmarcher, Clemens Röhrl, Sepp Hochreiter, Günter Klambauer
ICLR (Poster)2
2018 Estimating Collective Attention toward a Public Display
abstract
Enticing groups of passers-by to focused interaction with a public display requires the display system to take appropriate action that depends on how much attention the group is already paying to the display. In the design of such a system, we might want to present the content so that it indicates that a part of the group that is looking head-on at the display has already been registered and is addressed individually, whereas it simultaneously emits a strong audio signal that makes the inattentive rest of the group turn toward it. The challenge here is to define and delimit adequate mixed attention states for groups of people, allowing for classifying collective attention based on inhomogeneous variants of individual attention, i.e., where some group members might be highly attentive, others even interacting with the public display, and some unperceptive. In this article, we present a model for estimating collective human attention toward a public display and investigate technical methods for practical implementation that employs measurement of physical expressive features of people appearing within the display's field of view (i.e., the basis for deriving a person's attention). We delineate strengths and weaknesses and prove the potentials of our model by experimentally exerting influence on the attention of groups of passers-by in a public gaming scenario.
Wolfgang Narzt, Otto Weichselbaum, Gustav Pomberger, Markus Hofmarcher, Michael Strauss, Peter Holzkorn, Roland Haring, Monika Sturm 0001
ACM Trans. Interact. Intell. Syst.4