Fabian Paischer

dblp:309/5971 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 35% Deep learning architectures and training · 24% Efficient and distributed learning · 20%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
partially observable reinforcement learning
1.222023
Semantic HELM: A Human-Readable Memory for Reinforcement Learning · NeurIPS 2023
History Compression via Language Models in Reinforcement Learning · ICML 2022
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation
0.912025
Parameter Efficient Fine-tuning via Explained Variance Adaptation · NeurIPS 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.912025
Parameter Efficient Fine-tuning via Explained Variance Adaptation · NeurIPS 2025
Machine learning › Deep learning architectures and training
weight initialization
0.912025
Parameter Efficient Fine-tuning via Explained Variance Adaptation · NeurIPS 2025
Computational science and engineering › computational physics
plasma physics simulation
0.912025
GyroSwin: 5D Surrogates for Gyrokinetic Plasma Turbulence Simulations · NeurIPS 2025
Machine learning › Learning paradigms
continual learning
0.712023
Learning to Modulate pre-trained Models in RL · NeurIPS 2023
Machine learning › Trustworthy machine learning
interpretability
0.712023
Semantic HELM: A Human-Readable Memory for Reinforcement Learning · NeurIPS 2023
Machine learning › Deep learning architectures and training
memory mechanism
0.712023
Semantic HELM: A Human-Readable Memory for Reinforcement Learning · NeurIPS 2023
Machine learning › Reinforcement learning
transfer learning in reinforcement learning
0.712023
Learning to Modulate pre-trained Models in RL · NeurIPS 2023
Machine learning › Reinforcement learning › partially observable reinforcement learning
history representation
0.612022
History Compression via Language Models in Reinforcement Learning · ICML 2022
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning
0.612022
History Compression via Language Models in Reinforcement Learning · ICML 2022
Machine learning › Deep learning architectures and training
foundation model
0.312025
Parameter Efficient Fine-tuning via Explained Variance Adaptation · NeurIPS 2025
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.312025
GyroSwin: 5D Surrogates for Gyrokinetic Plasma Turbulence Simulations · NeurIPS 2025
Natural language and speech › Language models and text generation › LLM agents
agent memory
0.212023
Semantic HELM: A Human-Readable Memory for Reinforcement Learning · NeurIPS 2023
Natural language and speech › Language models and text generation › pre-trained language model › efficient pre-trained language model
frozen language model
0.212022
History Compression via Language Models in Reinforcement Learning · ICML 2022
Natural language and speech › Language models and text generation
pre-trained language model
0.212022
History Compression via Language Models in Reinforcement Learning · ICML 2022

Methods — techniques the papers use, named apart from their topics

vision transformer · 1.7mode separation · 1.7cross-attention · 1.7singular value decomposition · 0.9incremental SVD · 0.9explained variance · 0.9pre-training · 0.7pre-trained language model · 0.7modulation · 0.7CLIP · 0.7
YearPublicationVenuePosition
2025 GyroSwin: 5D Surrogates for Gyrokinetic Plasma Turbulence Simulations
abstract
Nuclear fusion plays a pivotal role in the quest for reliable and sustainable energy production. A major roadblock to viable fusion power is understanding plasma turbulence, which significantly impairs plasma confinement, and is vital for next generation reactor design. Plasma turbulence is governed by the nonlinear gyrokinetic equation, which evolves a 5D distribution function over time. Due to its high computational cost, reduced-order models are often employed in practice to approximate turbulent transport of energy. However, they omit nonlinear effects unique to the full 5D dynamics. To tackle this, we introduce GyroSwin, the first scalable 5D neural surrogate that can model 5D nonlinear gyrokinetic simulations, thereby capturing the physical phenomena neglected by reduced models, while providing accurate estimates of turbulent heat transport. GyroSwin (i) extends hierarchical Vision Transformers to 5D, (ii) introduces cross-attention and integration modules for latent 3D$\leftrightarrow$5D interactions between electrostatic potential fields and the distribution function, and (iii) performs channelwise mode separation inspired by nonlinear physics. We demonstrate that GyroSwin outperforms widely used reduced numerics on heat flux prediction, captures the turbulent energy cascade, and reduces the cost of fully resolved nonlinear gyrokinetics by three orders of magnitude while remaining physically verifiable. GyroSwin shows promising scaling laws, tested up to one billion parameters, paving the way for scalable neural surrogates for gyrokinetic simulations of plasma turbulence.
Fabian Paischer, Gianluca Galletti, William Hornsby, Paul Setinek, Lorenzo Zanisi, Naomi Carey, Stanislas Pamela, Johannes Brandstetter
NeurIPS1
2025 Parameter Efficient Fine-tuning via Explained Variance Adaptation
abstract
Foundation models (FMs) are pre-trained on large-scale datasets and then fine-tuned for a specific downstream task. The most common fine-tuning method is to update pretrained weights via low-rank adaptation (LoRA). Existing initialization strategies for LoRA often rely on singular value decompositions (SVD) of gradients or weight matrices. However, they do not provably maximize the expected gradient signal, which is critical for fast adaptation. To this end, we introduce **E**xplained **V**ariance **A**daptation (EVA), an initialization scheme that uses the directions capturing the most activation variance, provably maximizing the expected gradient signal and accelerating fine-tuning. EVA performs incremental SVD on minibatches of activation vectors and selects the right-singular vectors for initialization once they converged. Further, by selecting the directions that capture the most activation-variance for a given rank budget, EVA accommodates adaptive ranks that reduce the number of trainable parameters. We apply EVA to a variety of fine-tuning tasks as language generation and understanding, image classification, and reinforcement learning. EVA exhibits faster convergence than competitors and achieves the highest average score across a multitude of tasks per domain while reducing the number of trainable parameters through rank redistribution. In summary, EVA establishes a new Pareto frontier compared to existing LoRA initialization schemes in both accuracy and efficiency.
Fabian Paischer, Lukas Hauzenberger, Thomas Schmied, Benedikt Alkin, Marc Peter Deisenroth, Sepp Hochreiter
NeurIPS1
2023 Semantic HELM: A Human-Readable Memory for Reinforcement Learning
abstract
Reinforcement learning agents deployed in the real world often have to cope with partially observable environments. Therefore, most agents employ memory mechanisms to approximate the state of the environment. Recently, there have been impressive success stories in mastering partially observable environments, mostly in the realm of computer games like Dota 2, StarCraft II, or MineCraft. However, existing methods lack interpretability in the sense that it is not comprehensible for humans what the agent stores in its memory. In this regard, we propose a novel memory mechanism that represents past events in human language. Our method uses CLIP to associate visual inputs with language tokens. Then we feed these tokens to a pretrained language model that serves the agent as memory and provides it with a coherent and human-readable representation of the past. We train our memory mechanism on a set of partially observable environments and find that it excels on tasks that require a memory component, while mostly attaining performance on-par with strong baselines on tasks that do not. On a challenging continuous recognition task, where memorizing the past is crucial, our memory mechanism converges two orders of magnitude faster than prior methods. Since our memory mechanism is human-readable, we can peek at an agent's memory and check whether crucial pieces of information have been stored. This significantly enhances troubleshooting and paves the way toward more interpretable agents.
Fabian Paischer, Thomas Adler, Markus Hofmarcher, Sepp Hochreiter
NeurIPS1
2023 Learning to Modulate pre-trained Models in RL
abstract
Reinforcement Learning (RL) has been successful in various domains like robotics, game playing, and simulation. While RL agents have shown impressive capabilities in their specific tasks, they insufficiently adapt to new tasks. In supervised learning, this adaptation problem is addressed by large-scale pre-training followed by fine-tuning to new down-stream tasks. Recently, pre-training on multiple tasks has been gaining traction in RL. However, fine-tuning a pre-trained model often suffers from catastrophic forgetting. That is, the performance on the pre-training tasks deteriorates when fine-tuning on new tasks. To investigate the catastrophic forgetting phenomenon, we first jointly pre-train a model on datasets from two benchmark suites, namely Meta-World and DMControl. Then, we evaluate and compare a variety of fine-tuning methods prevalent in natural language processing, both in terms of performance on new tasks, and how well performance on pre-training tasks is retained. Our study shows that with most fine-tuning approaches, the performance on pre-training tasks deteriorates significantly. Therefore, we propose a novel method, Learning-to-Modulate (L2M), that avoids the degradation of learned skills by modulating the information flow of the frozen pre-trained model via a learnable modulation pool. Our method achieves state-of-the-art performance on the Continual-World benchmark, while retaining performance on the pre-training tasks. Finally, to aid future research in this area, we release a dataset encompassing 50 Meta-World and 16 DMControl tasks.
Thomas Schmied, Markus Hofmarcher, Fabian Paischer, Razvan Pascanu, Sepp Hochreiter
NeurIPS3
2022 History Compression via Language Models in Reinforcement Learning
abstract
In a partially observable Markov decision process (POMDP), an agent typically uses a representation of the past to approximate the underlying MDP. We propose to utilize a frozen Pretrained Language Transformer (PLT) for history representation and compression to improve sample efficiency. To avoid training of the Transformer, we introduce FrozenHopfield, which automatically associates observations with pretrained token embeddings. To form these associations, a modern Hopfield network stores these token embeddings, which are retrieved by queries that are obtained by a random but fixed projection of observations. Our new method, HELM, enables actor-critic network architectures that contain a pretrained language Transformer for history representation as a memory module. Since a representation of the past need not be learned, HELM is much more sample efficient than competitors. On Minigrid and Procgen environments HELM achieves new state-of-the-art results. Our code is available at https://github.com/ml-jku/helm.
Fabian Paischer, Thomas Adler, Vihang Patil, Angela Bitto-Nemling, Markus Holzleitner, Sebastian Lehner, Hamid Eghbalzadeh, Sepp Hochreiter
ICML1
2022 WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models
abstract
Large pretrained language models (LMs) have become the central building block of many NLP applications. Training these models requires ever more computational resources and most of the existing models are trained on English text only. It is exceedingly expensive to train these models in other languages. To alleviate this problem, we introduce a novel method -- called WECHSEL -- to efficiently and effectively transfer pretrained LMs to new languages. WECHSEL can be applied to any model which uses subword-based tokenization and learns an embedding for each subword. The tokenizer of the source model (in English) is replaced with a tokenizer in the target language and token embeddings are initialized such that they are semantically similar to the English tokens by utilizing multilingual static word embeddings covering English and the target language. We use WECHSEL to transfer the English RoBERTa and GPT-2 models to four languages (French, German, Chinese and Swahili). We also study the benefits of our method on very low-resource languages. WECHSEL improves over proposed methods for cross-lingual parameter transfer and outperforms models of comparable size trained from scratch with up to 64x less training effort. Our method makes training large language models for new languages more accessible and less damaging to the environment. We make our code and models publicly available.
Benjamin Minixhofer, Fabian Paischer, Navid Rekabsaz
NAACL-HLT2