Philipp Dufter

dblp:213/8070 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Multimodal Autoregressive Pre-training of Large Vision Encoders
abstract
We introduce a novel method for pre-training of large-scale vision encoders. Building on recent advancements in autoregressive pre-training of vision models, we extend this framework to a multimodal setting, i.e., images and text. In this paper, we present AIMV2, a family of generalist vision encoders characterized by a straightforward pre-training process, scalability, and remarkable performance across a range of downstream tasks. This is achieved by pairing the vision encoder with a multimodal decoder that autoregressively generates raw image patches and text tokens. Our encoders excel not only in multimodal evaluations but also in vision benchmarks such as localization, grounding, and classification. Notably, our AIMV2-3B encoder achieves 89.5% accuracy on ImageNet-1k with a frozen trunk. Furthermore, AIMV2 consistently outperforms state-of-the-art contrastive models (e.g., CLIP, SigLIP) in multimodal image understanding across diverse settings.
Enrico Fini, Mustafa Shukor, Xiujun Li, Philipp Dufter, Michal Klein, David Haldimann, Sai Aitharaju, Victor G. T. da Costa, Louis Béthune, Zhe Gan, Alexander Toshev, Marcin Eichner, Moin Nabi, Yinfei Yang, Joshua M. Susskind, Alaaeldin El-Nouby
CVPR4
2025 MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
abstract
We present MM1.5, a new family of multimodal large language models (MLLMs) designed to enhance capabilities in text-rich image understanding, visual referring and grounding, and multi-image reasoning. Building upon the MM1 architecture, MM1.5 adopts a data-centric approach to model training, systematically exploring the impact of diverse data mixtures across the entire model training lifecycle. This includes high-quality OCR data and synthetic captions for continual pre-training, as well as an optimized visual instruction-tuning data mixture for supervised fine-tuning. Our models range from 1B to 30B parameters, encompassing both dense and mixture-of-experts (MoE) variants, and demonstrate that careful data curation and training strategies can yield strong performance even at small scales (1B and 3B). Additionally, we introduce two specialized variants: MM1.5-Video, designed for video understanding, and MM1.5-UI, tailored for mobile UI understanding. Through extensive empirical studies and ablations, we provide detailed insights into the training processes and decisions that inform our final designs, offering valuable guidance for future research in MLLM development.
Haotian Zhang 0005, Mingfei Gao, Zhe Gan, Philipp Dufter, Nina Wenzel, Forrest Huang, Dhruti Shah, Xianzhi Du, Bowen Zhang 0002, Yanghao Li, Sam Dodge, Keen You, Aleksei Timofeev, Hong-You Chen, Jean-Philippe Fauconnier, Zhengfeng Lai, Haoxuan You
ICLR4
2024 MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang 0002, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Anton Belyi, Haotian Zhang 0005, Karanjeet Singh 0003, Doug Kang, Hongyu Hè, Max Schwarzer, Tom Gunter, Xiang Kong, Aonan Zhang, Nan Du 0002, Tao Lei 0001, Sam Wiseman, Mark Lee 0003, Ruoming Pang, Peter Grasch, Alexander Toshev, Yinfei Yang
ECCV (29)6
2022 Towards a Broad Coverage Named Entity Resource: A Data-Efficient Approach for Many Diverse Languages
abstract
Parallel corpora are ideal for extracting a multilingual named entity (MNE) resource, i.e., a dataset of names translated into multiple languages. Prior work on extracting MNE datasets from parallel corpora required resources such as large monolingual corpora or word aligners that are unavailable or perform poorly for underresourced languages. We present CLC-BN, a new method for creating an MNE resource, and apply it to the Parallel Bible Corpus, a corpus of more than 1000 languages. CLC-BN learns a neural transliteration model from parallel-corpus statistics, without requiring any other bilingual resources, word aligners, or seed data. Experimental results show that CLC-BN clearly outperforms prior work. We release an MNE resource for 1340 languages and demonstrate its effectiveness in two downstream tasks: knowledge graph augmentation and bilingual lexicon induction.
Silvia Severini, Ayyoob Imani, Philipp Dufter, Hinrich Schütze
LREC3
2022 Position Information in Transformers: An Overview
abstract
Abstract Transformers are arguably the main workhorse in recent natural language processing research. By definition, a Transformer is invariant with respect to reordering of the input. However, language is inherently sequential and word order is essential to the semantics and syntax of an utterance. In this article, we provide an overview and theoretical comparison of existing methods to incorporate position information into Transformer models. The objectives of this survey are to (1) showcase that position information in Transformer is a vibrant and extensive research area; (2) enable the reader to compare existing methods by providing a unified notation and systematization of different approaches along important model dimensions; (3) indicate what characteristics of an application should be taken into account when selecting a position encoding; and (4) provide stimuli for future research.
Philipp Dufter, Martin Schmitt, Hinrich Schütze
Comput. Linguistics1
2021 Multilingual LAMA: Investigating Knowledge in Multilingual Pretrained Language Models
abstract
Recently, it has been found that monolingual English language models can be used as knowledge bases.Instead of structural knowledge base queries, masked sentences such as "Paris is the capital of [MASK]" are used as probes.We translate the established benchmarks TREx and GoogleRE into 53 languages.Working with mBERT, we investigate three questions.(i) Can mBERT be used as a multilingual knowledge base?Most prior work only considers English.Extending research to multiple languages is important for diversity and accessibility.(ii) Is mBERT's performance as knowledge base language-independent or does it vary from language to language?(iii) A multilingual model is trained on more text, e.g., mBERT is trained on 104 Wikipedias.Can mBERT leverage this for better performance?We find that using mBERT as a knowledge base yields varying performance across languages and pooling predictions across languages improves performance.Conversely, mBERT exhibits a language bias; e.g., when queried in Italian, it tends to predict Italy as the country of origin.
Nora Kassner, Philipp Dufter, Hinrich Schütze
EACL2
2021 Graph Algorithms for Multiparallel Word Alignment
abstract
Ayyoob ImaniGooghari, Masoud Jalili Sabet, Lutfi Kerem Senel, Philipp Dufter, François Yvon, Hinrich Schütze. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Ayyoob Imani, Masoud Jalili Sabet, Lutfi Kerem Senel, Philipp Dufter, François Yvon, Hinrich Schütze
EMNLP (1)4
2021 Static Embeddings as Efficient Knowledge Bases?
abstract
Recent research investigates factual knowledge stored in large pretrained language models (PLMs).Instead of structural knowledge base (KB) queries, masked sentences such as "Paris is the capital of [MASK]" are used as probes.The good performance on this analysis task has been interpreted as PLMs becoming potential repositories of factual knowledge.In experiments across ten linguistically diverse languages, we study knowledge contained in static embeddings.We show that, when restricting the output space to a candidate set, simple nearest neighbor matching using static embeddings performs better than PLMs.E.g., static embeddings perform 1.6% points better than BERT while just using 0.3% of energy for training.One important factor in their good comparative performance is that static embeddings are standardly learned for a large vocabulary.In contrast, BERT exploits its more sophisticated, but expensive ability to compose meaningful representations from a much smaller subword vocabulary.
Philipp Dufter, Nora Kassner, Hinrich Schütze
NAACL-HLT1
2021 Semantic Text Segment Classification of Structured Technical Content
Julian Höllig, Philipp Dufter, Michaela Geierhos, Wolfgang Ziegler, Hinrich Schütze
NLDB2
2020 Increasing Learning Efficiency of Self-Attention Networks through Direct Position Interactions, Learnable Temperature, and Convoluted Attention
abstract
Self-Attention Networks (SANs) are an integral part of successful neural architectures such as Transformer (Vaswani et al., 2017), and thus of pretrained language models such as BERT (Devlin et al., 2019) or GPT-3 (Brown et al., 2020).Training SANs on a task or pretraining them on language modeling requires large amounts of data and compute resources.We are searching for modifications to SANs that enable faster learning, i.e., higher accuracies after fewer update steps.We investigate three modifications to SANs: direct position interactions, learnable temperature, and convoluted attention.When evaluating them on part-of-speech tagging, we find that direct position interactions are an alternative to position embeddings, and convoluted attention has the potential to speed up the learning process.
Philipp Dufter, Martin Schmitt, Hinrich Schütze
COLING1
2020 Monolingual and Multilingual Reduction of Gender Bias in Contextualized Representations
abstract
Pretrained language models (PLMs) learn stereotypes held by humans and reflected in text from their training corpora, including gender bias.When PLMs are used for downstream tasks such as picking candidates for a job, people's lives can be negatively affected by these learned stereotypes.Prior work usually identifies a linear gender subspace and removes gender information by eliminating the subspace.Following this line of work, we propose to use DensRay, an analytical method for obtaining interpretable dense subspaces.We show that DensRay performs on-par with prior approaches, but provide arguments that it is more robust and provide indications that it preserves language model performance better.By applying DensRay to attention heads and layers of BERT we show that gender information is spread across all attention heads and most of the layers.Also we show that DensRay can obtain gender bias scores on both token and sentence levels.Finally, we demonstrate that we can remove bias multilingually, e.g.,
Sheng Liang, Philipp Dufter, Hinrich Schütze
COLING2
2020 Identifying Elements Essential for BERT's Multilinguality
abstract
It has been shown that multilingual BERT (mBERT) yields high quality multilingual representations and enables effective zero-shot transfer.This is surprising given that mBERT does not use any crosslingual signal during training.While recent literature has studied this phenomenon, the reasons for the multilinguality are still somewhat obscure.We aim to identify architectural properties of BERT and linguistic properties of languages that are necessary for BERT to become multilingual.To allow for fast experimentation we propose an efficient setup with small BERT models trained on a mix of synthetic and natural data.Overall, we identify four architectural and two linguistic elements that influence multilinguality.Based on our insights, we experiment with a multilingual pretraining setup that modifies the masking strategy using VecMap, i.e., unsupervised embedding alignment.Experiments on XNLI with three languages indicate that our findings transfer from our small setup to larger scale settings.
Philipp Dufter, Hinrich Schütze
EMNLP (1)1
2019 Analytical Methods for Interpretable Ultradense Word Embeddings
abstract
Philipp Dufter, Hinrich Schütze. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Philipp Dufter, Hinrich Schütze
EMNLP/IJCNLP (1)1
2018 Embedding Learning Through Multilingual Concept Induction
abstract
We present a new method for estimating vector space representations of words: embedding learning by concept induction.We test this method on a highly parallel corpus and learn semantic representations of words in 1259 different languages in a single common space.An extensive experimental evaluation on crosslingual word similarity and sentiment analysis indicates that concept-based multilingual embedding learning performs better than previous approaches.
Philipp Dufter, Martin Schmitt, Alexander Fraser 0001, Hinrich Schütze
ACL (1)1