Jonas Pfeiffer

dblp:222/9866 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
14since 2021 · last 2025
0000-0002-8634-6170ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Deliberation in Latent Space via Differentiable Cache Augmentation
abstract
Techniques enabling large language models (LLMs) to "think more" by generating and attending to intermediate reasoning steps have shown promise in solving complex problems. However, the standard approaches generate sequences of discrete tokens immediately before responding, and so they can incur significant latency costs and be challenging to optimize. In this work, we demonstrate that a frozen LLM can be augmented with an offline coprocessor that operates on the model's key-value (kv) cache. This coprocessor augments the cache with a set of latent embeddings designed to improve the fidelity of subsequent decoding. We train this coprocessor using the language modeling loss from the decoder on standard pretraining data, while keeping the decoder itself frozen. This approach enables the model to learn, in an end-to-end differentiable fashion, how to distill additional computation into its kv-cache. Because the decoder remains unchanged, the coprocessor can operate offline and asynchronously, and the language model can function normally if the coprocessor is unavailable or if a given cache is deemed not to require extra computation. We show experimentally that when a cache is augmented, the decoder achieves lower perplexity on numerous subsequent tokens. Furthermore, even without any task-specific training, our experiments demonstrate that cache augmentation consistently improves performance across a range of reasoning-intensive tasks.
Jonas Pfeiffer, Jiaxing Wu, Arthur Szlam
ICML2
2025 DARE: Diverse Visual Question Answering with Robustness Evaluation
abstract
Abstract Vision Language Models (VLMs) extend remarkable capabilities of text-only large language models and vision-only models, being able to learn from and process multi-modal vision-text input. While modern VLMs perform well on a number of standard image classification and image-text matching tasks, they still struggle with a number of crucial vision-language (VL) reasoning abilities such as counting and spatial reasoning. Moreover, while they might be very brittle to small variations in instructions and/or evaluation protocols, existing benchmarks fail to evaluate their robustness (or rather the lack of it). In order to couple challenging VL scenarios with comprehensive robustness evaluation, we introduce DARE, Diverse Visual Question Answering with Robustness Evaluation, a carefully created and curated multiple-choice VQA benchmark. DARE evaluates VLM performance on five diverse categories and includes four robustness-oriented evaluations based on the variations of prompts, the subsets of answer options, the output format, and the number of correct answers. Among a spectrum of other findings, we report that state-of-the-art VLMs still struggle with questions in most categories and are unable to consistently deliver their peak performance across the tested robustness evaluations. Consequently, our work calls for the systematic addition of robustness evaluations in future VLM research.
Hannah Sterz, Jonas Pfeiffer, Ivan Vulic
Trans. Assoc. Comput. Linguistics2
2024 FUN with Fisher: Improving Generalization of Adapter-Based Cross-lingual Transfer with Scheduled Unfreezing
abstract
Chen Cecilia Liu, Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jonas Pfeiffer, Ivan Vulic, Iryna Gurevych
NAACL-HLT2
2023 Where's the Point? Self-Supervised Multilingual Punctuation-Agnostic Sentence Segmentation
abstract
Many NLP pipelines split text into sentences as one of the crucial preprocessing steps.Prior sentence segmentation tools either rely on punctuation or require a considerable amount of sentence-segmented training data: both central assumptions might fail when porting sentence segmenters to diverse languages on a massive scale.In this work, we thus introduce a multilingual punctuation-agnostic sentence segmentation method, currently covering 85 languages, trained in a self-supervised fashion on unsegmented text, by making use of newline characters which implicitly perform segmentation into paragraphs.We further propose an approach that adapts our method to the segmentation in a given corpus by using only a small number (64-256) of sentence-segmented examples.The main results indicate that our method outperforms all the prior best sentence-segmentation tools by an average of 6.1% F1 points.Furthermore, we demonstrate that proper sentence segmentation has a point: the use of a (powerful) sentence segmenter makes a considerable difference for a downstream application such as machine translation (MT).By using our method to match sentence segmentation to the segmentation used during training of MT models, we achieve an average improvement of 2.3 BLEU points over the best prior segmentation tool, as well as massive gains over a trivial segmenter that splits text into equally sized blocks.
Benjamin Minixhofer, Jonas Pfeiffer, Ivan Vulic
ACL (1)2
2023 CompoundPiece: Evaluating and Improving Decompounding Performance of Language Models
abstract
While many languages possess processes of joining two or more words to create compound words, previous studies have been typically limited only to languages with excessively productive compound formation (e.g., German, Dutch) and there is no public dataset containing compound and non-compound words across a large number of languages.In this work, we systematically study decompounding, the task of splitting compound words into their constituents, at a wide scale.We first address the data gap by introducing a dataset of 255k compound and noncompound words across 56 diverse languages obtained from Wiktionary.We then use this dataset to evaluate an array of Large Language Models (LLMs) on the decompounding task.We find that LLMs perform poorly, especially on words which are tokenized unfavorably by subword tokenization.We thus introduce a novel methodology to train dedicated models for decompounding.The proposed two-stage procedure relies on a fully self-supervised objective in the first stage, while the second, supervised learning stage optionally fine-tunes the model on the annotated Wiktionary data.Our self-supervised models outperform the prior best unsupervised decompounding models by 13.9% accuracy on average.Our fine-tuned models outperform all prior (language-specific) decompounding tools.Furthermore, we use our models to leverage decompounding during the creation of a subword tokenizer, which we refer to as CompoundPiece.CompoundPiece tokenizes compound words more favorably on average, leading to improved performance on decompounding over an otherwise equivalent model using SentencePiece tokenization.
Benjamin Minixhofer, Jonas Pfeiffer, Ivan Vulic
EMNLP2
2022 IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and Languages
abstract
Reliable evaluation benchmarks designed for replicability and comprehensiveness have driven progress in machine learning. Due to the lack of a multilingual benchmark, however, vision-and-language research has mostly focused on English language tasks. To fill this gap, we introduce the Image-Grounded Language Understanding Evaluation benchmark. IGLUE brings together{—}by both aggregating pre-existing datasets and creating new ones{—}visual question answering, cross-modal retrieval, grounded reasoning, and grounded entailment tasks across 20 diverse languages. Our benchmark enables the evaluation of multilingual multimodal models for transfer learning, not only in a zero-shot setting, but also in newly defined few-shot learning setups. Based on the evaluation of the available state-of-the-art models, we find that translate-test transfer is superior to zero-shot transfer and that few-shot learning is hard to harness for many tasks. Moreover, downstream performance is partially explained by the amount of available unlabelled textual data for pretraining, and only weakly by the typological distance of target{–}source languages. We hope to encourage future research efforts in this area by releasing the benchmark to the community.
Emanuele Bugliarello, Fangyu Liu 0001, Jonas Pfeiffer, Siva Reddy, Desmond Elliott, Edoardo Maria Ponti, Ivan Vulic
ICML3
2022 Lifting the Curse of Multilinguality by Pre-training Modular Transformers
abstract
Jonas Pfeiffer, Naman Goyal, Xi Lin, Xian Li, James Cross, Sebastian Riedel, Mikel Artetxe. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Jonas Pfeiffer, Naman Goyal 0001, Xi Victoria Lin, Xian Li 0003, James Cross 0003, Sebastian Riedel 0001, Mikel Artetxe
NAACL-HLT1
2022 Retrieve Fast, Rerank Smart: Cooperative and Joint Approaches for Improved Cross-Modal Retrieval
abstract
Abstract Current state-of-the-art approaches to cross- modal retrieval process text and visual input jointly, relying on Transformer-based architectures with cross-attention mechanisms that attend over all words and objects in an image. While offering unmatched retrieval performance, such models: 1) are typically pretrained from scratch and thus less scalable, 2) suffer from huge retrieval latency and inefficiency issues, which makes them impractical in realistic applications. To address these crucial gaps towards both improved and efficient cross- modal retrieval, we propose a novel fine-tuning framework that turns any pretrained text-image multi-modal model into an efficient retrieval model. The framework is based on a cooperative retrieve-and-rerank approach that combines: 1) twin networks (i.e., a bi-encoder) to separately encode all items of a corpus, enabling efficient initial retrieval, and 2) a cross-encoder component for a more nuanced (i.e., smarter) ranking of the retrieved small set of items. We also propose to jointly fine- tune the two components with shared weights, yielding a more parameter-efficient model. Our experiments on a series of standard cross-modal retrieval benchmarks in monolingual, multilingual, and zero-shot setups, demonstrate improved accuracy and huge efficiency benefits over the state-of-the-art cross- encoders.1
Gregor Geigle, Jonas Pfeiffer, Nils Reimers 0001, Ivan Vulic, Iryna Gurevych
Trans. Assoc. Comput. Linguistics2
2021 How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models
abstract
Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, Iryna Gurevych. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Phillip Rust, Jonas Pfeiffer, Ivan Vulic, Sebastian Ruder, Iryna Gurevych
ACL/IJCNLP (1)2
2021 AdapterFusion: Non-Destructive Task Composition for Transfer Learning
abstract
Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, Iryna Gurevych. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, Iryna Gurevych
EACL1
2021 UNKs Everywhere: Adapting Multilingual Language Models to New Scripts
abstract
Massively multilingual language models such as multilingual BERT offer state-of-the-art cross-lingual transfer performance on a range of NLP tasks.However, due to limited capacity and large differences in pretraining data sizes, there is a profound performance gap between resource-rich and resource-poor target languages.The ultimate challenge is dealing with under-resourced languages not covered at all by the models and written in scripts unseen during pretraining.In this work, we propose a series of novel data-efficient methods that enable quick and effective adaptation of pretrained multilingual models to such lowresource languages and unseen scripts.Relying on matrix factorization, our methods capitalize on the existing latent knowledge about multiple languages already available in the pretrained model's embedding matrix.Furthermore, we show that learning of the new dedicated embedding matrix in the target language can be improved by leveraging a small number of vocabulary items (i.e., the so-called lexically overlapping tokens) shared between mBERT's and target language vocabulary.Our adaptation techniques offer substantial performance gains for languages with unseen scripts.We also demonstrate that they can yield improvements for low-resource languages written in scripts covered by the pretrained model.
Jonas Pfeiffer, Ivan Vulic, Iryna Gurevych, Sebastian Ruder
EMNLP (1)1
2021 What to Pre-Train on? Efficient Intermediate Task Selection
abstract
Intermediate task fine-tuning has been shown to culminate in large transfer gains across many NLP tasks.With an abundance of candidate datasets as well as pre-trained language models, it has become infeasible to experiment with all combinations to find the best transfer setting.In this work, we provide a comprehensive comparison of different methods for efficiently identifying beneficial tasks for intermediate transfer learning.We focus on parameter and computationally efficient adapter settings, highlight different data-availability scenarios, and provide expense estimates for each method.We experiment with a diverse set of 42 intermediate and 11 target English classification, multiple choice, question answering, and sequence tagging tasks.Our results demonstrate that efficient embedding based methods, which rely solely on the respective datasets, outperform computational expensive few-shot fine-tuning approaches.Our best methods achieve an average Regret@3 of 1% across all target tasks, demonstrating that we are able to efficiently identify the best datasets for intermediate training.1
Clifton Poth, Jonas Pfeiffer, Andreas Rücklé, Iryna Gurevych
EMNLP (1)2
2021 Smelting Gold and Silver for Improved Multilingual AMR-to-Text Generation
abstract
Recent work on multilingual AMR-to-text generation has exclusively focused on data augmentation strategies that utilize silver AMR.However, this assumes a high quality of generated AMRs, potentially limiting the transferability to the target task.In this paper, we investigate different techniques for automatically generating AMR annotations, where we aim to study which source of information yields better multilingual results.Our models trained on gold AMR with silver (machine translated) sentences outperform approaches which leverage generated silver AMR.We find that combining both complementary sources of information further improves multilingual AMR-to-text generation.Our models surpass the previous state of the art for German, Italian, Spanish, and Chinese by a large margin. 1
Leonardo F. R. Ribeiro, Jonas Pfeiffer, Iryna Gurevych
EMNLP (1)2
2021 AdapterDrop: On the Efficiency of Adapters in Transformers
abstract
Transformer models are expensive to fine-tune, slow for inference, and have large storage requirements.Recent approaches tackle these shortcomings by training smaller models, dynamically reducing the model size, and by training light-weight adapters.In this paper, we propose AdapterDrop, removing adapters from lower transformer layers during training and inference, which incorporates concepts from all three directions.We show that Adap-terDrop can dynamically reduce the computational overhead when performing inference over multiple tasks simultaneously, with minimal decrease in task performances.We further prune adapters from AdapterFusion, which improves the inference efficiency while maintaining the task performances entirely.
Andreas Rücklé, Gregor Geigle, Max Glockner, Tilman Beck, Jonas Pfeiffer, Nils Reimers 0001, Iryna Gurevych
EMNLP (1)5
2020 Low Resource Sequence Tagging with Weak Labels
abstract
Current methods for sequence tagging depend on large quantities of domain-specific training data, limiting their use in new, user-defined tasks with few or no annotations. While crowdsourcing can be a cheap source of labels, it often introduces errors that degrade the performance of models trained on such crowdsourced data. Another solution is to use transfer learning to tackle low resource sequence labelling, but current approaches rely heavily on similar high resource datasets in different languages. In this paper, we propose a domain adaptation method using Bayesian sequence combination to exploit pre-trained models and unreliable crowdsourced data that does not require high resource data in a different language. Our method boosts performance by learning the relationship between each labeller and the target task and trains a sequence labeller on the target domain with little or no gold-standard data. We apply our approach to labelling diagnostic classes in medical and educational case studies, showing that the model achieves strong performance though zero-shot transfer learning and is more effective than alternative ensemble methods. Using NER and information extraction tasks, we show how our approach can train a model directly from crowdsourced labels, outperforming pipeline approaches that first aggregate the crowdsourced data, then train on the aggregated labels.
Edwin Simpson, Jonas Pfeiffer, Iryna Gurevych
AAAI2
2020 MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer
abstract
The main goal behind state-of-the-art pretrained multilingual models such as multilingual BERT and XLM-R is enabling and bootstrapping NLP applications in low-resource languages through zero-shot or few-shot crosslingual transfer.However, due to limited model capacity, their transfer performance is the weakest exactly on such low-resource languages and languages unseen during pretraining.We propose MAD-X, an adapter-based framework that enables high portability and parameter-efficient transfer to arbitrary tasks and languages by learning modular language and task representations.In addition, we introduce a novel invertible adapter architecture and a strong baseline method for adapting a pretrained multilingual model to a new language.MAD-X outperforms the state of the art in cross-lingual transfer across a representative set of typologically diverse languages on named entity recognition and causal commonsense reasoning, and achieves competitive results on question answering.Our code and adapters are available at AdapterHub.ml.
Jonas Pfeiffer, Ivan Vulic, Iryna Gurevych, Sebastian Ruder
EMNLP (1)1
2020 MultiCQA: Zero-Shot Transfer of Self-Supervised Text Matching Models on a Massive Scale
abstract
We study the zero-shot transfer capabilities of text matching models on a massive scale, by self-supervised training on 140 source domains from community question answering forums in English.We investigate the model performances on nine benchmarks of answer selection and question similarity tasks, and show that all 140 models transfer surprisingly well, where the large majority of models substantially outperforms common IR baselines.We also demonstrate that considering a broad selection of source domains is crucial for obtaining the best zero-shot transfer performances, which contrasts the standard procedure that merely relies on the largest and most similar domains.In addition, we extensively study how to best combine multiple source domains.We propose to incorporate self-supervised with supervised multi-task learning on all available source domains.Our best zero-shot transfer model considerably outperforms in-domain BERT and the previous state of the art on six benchmarks.Fine-tuning of our model with in-domain data results in additional large gains and achieves the new state of the art on all nine benchmarks.
Andreas Rücklé, Jonas Pfeiffer, Iryna Gurevych
EMNLP (1)2