Phillip Rust

dblp:263/9843 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0001-5123-821XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation
abstract
Mahi Luthra, Jiayi Shen, Maxime Poli, Angelo Ortiz Tandazo, Yosuke Higuchi, Youssef Benchekroun, Martin Gleize, Charles-Éric Saint-James, Dongyan Lin, Phillip Rust, Angel Villar-Corrales, Surya, Vanessa Stark, Rashel Moritz, Juan Pino, Yann LeCun, Emmanuel Dupoux. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Mahi Luthra, Maxime Poli, Angelo Ortiz Tandazo, Yosuke Higuchi, Youssef Benchekroun, Martin Gleize, Charles-Éric Saint-James, Dongyan Lin, Phillip Rust, Angel Villar-Corrales, Surya Parimi, Vanessa Stark, Rashel Moritz, Juan Pino 0001, Yann LeCun, Emmanuel Dupoux
ACL (1)10
2025 Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users
abstract
Antonia Karamolegkou, Malvina Nikandrou, Georgios Pantazopoulos, Danae Sanchez Villegas, Phillip Rust, Ruchira Dhar, Daniel Hershcovich, Anders Søgaard. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Antonia Karamolegkou, Malvina Nikandrou, Georgios Pantazopoulos, Danae Sanchez Villegas, Phillip Rust, Ruchira Dhar, Daniel Hershcovich, Anders Søgaard
ACL (1)5
2025 Multilingual Pretraining for Pixel Language Models
abstract
Pixel language models operate directly on images of rendered text, eliminating the need for a fixed vocabulary.While these models have demonstrated strong capabilities for downstream cross-lingual transfer, multilingual pretraining remains underexplored.We introduce PIXEL-M4, a model pretrained on four visually and linguistically diverse languages: English, Hindi, Ukrainian, and Simplified Chinese.Multilingual evaluations on semantic and syntactic tasks show that PIXEL-M4 outperforms an English-only counterpart on non-Latin scripts.Word-level probing analyses confirm that PIXEL-M4 captures rich linguistic features, even in languages not seen during pretraining.Furthermore, an analysis of its hidden representations shows that multilingual pretraining yields a semantic embedding space closely aligned across the languages used for pretraining.This work demonstrates that multilingual pretraining substantially enhances the capability of pixel language models to effectively support a diverse set of languages.
Ilker Kesen, Jonas F. Lotz, Ingo Ziegler, Phillip Rust, Desmond Elliott
EMNLP4
2024 Towards Privacy-Aware Sign Language Translation at Scale
abstract
A major impediment to the advancement of sign language translation (SLT) is data scarcity.Much of the sign language data currently available on the web cannot be used for training supervised models due to the lack of aligned captions.Furthermore, scaling SLT using large-scale web-scraped datasets bears privacy risks due to the presence of biometric information, which the responsible development of SLT technologies should account for.In this work, we propose a twostage framework for privacy-aware SLT at scale that addresses both of these issues.We introduce SSVP-SLT, which leverages selfsupervised video pretraining on anonymized and unannotated videos, followed by supervised SLT finetuning on a curated parallel dataset.SSVP-SLT achieves state-of-the-art finetuned and zero-shot gloss-free SLT performance on the How2Sign dataset, outperforming the strongest respective baselines by over 3 BLEU-4.Based on controlled experiments, we further discuss the advantages and limitations of self-supervised pretraining and anonymization via facial obfuscation for SLT. 1
Phillip Rust, Skyler Wang, Necati Cihan Camgöz, Jean Maillard
ACL (1)1
2023 PHD: Pixel-Based Language Modeling of Historical Documents
abstract
The digitisation of historical documents has provided historians with unprecedented research opportunities. Yet, the conventional approach to analysing historical documents involves converting them from images to text using OCR, a process that overlooks the potential benefits of treating them as images and introduces high levels of noise. To bridge this gap, we take advantage of recent advancements in pixel-based language models trained to reconstruct masked patches of pixels instead of predicting token distributions. Due to the scarcity of real historical scans, we propose a novel method for generating synthetic scans to resemble real historical documents. We then pre-train our model, PHD, on a combination of synthetic scans and real historical newspapers from the 1700-1900 period. Through our experiments, we demonstrate that PHD exhibits high proficiency in reconstructing masked image patches and provide evidence of our model's noteworthy language understanding capabilities. Notably, we successfully apply our model to a historical QA task, highlighting its usefulness in this domain.
Nadav Borenstein, Phillip Rust, Desmond Elliott, Isabelle Augenstein
EMNLP2
2023 Text Rendering Strategies for Pixel Language Models
abstract
Pixel-based language models process text rendered as images, which allows them to handle any script, making them a promising approach to open vocabulary language modelling.However, recent approaches use text renderers that produce a large set of almost-equivalent input patches, which may prove sub-optimal for downstream tasks, due to redundancy in the input representations.In this paper, we investigate four approaches to rendering text in the PIXEL model (Rust et al., 2023), and find that simple character bigram rendering brings improved performance on sentence-level tasks without compromising performance on tokenlevel or multilingual tasks.This new rendering strategy also makes it possible to train a more compact model with only 22M parameters that performs on par with the original 86M parameter model.Our analyses show that character bigram rendering leads to a consistently better model but with an anisotropic patch embedding space, driven by a patch frequency bias, highlighting the connections between image patchand tokenization-based language models.
Jonas F. Lotz, Elizabeth Salesky, Phillip Rust, Desmond Elliott
EMNLP3
2023 Language Modelling with Pixels
Phillip Rust, Jonas F. Lotz, Emanuele Bugliarello, Elizabeth Salesky, Miryam de Lhoneux, Desmond Elliott
ICLR1
2023 Differential Privacy, Linguistic Fairness, and Training Data Influence: Impossibility and Possibility Theorems for Multilingual Language Models
abstract
Language models such as mBERT, XLM-R, and BLOOM aim to achieve multilingual generalization or compression to facilitate transfer to a large number of (potentially unseen) languages. However, these models should ideally also be private, linguistically fair, and transparent, by relating their predictions to training data. Can these requirements be simultaneously satisfied? We show that multilingual compression and linguistic fairness are compatible with differential privacy, but that differential privacy is at odds with training data influence sparsity, an objective for transparency. We further present a series of experiments on two common NLP tasks and evaluate multilingual compression and training data influence sparsity under different privacy guarantees, exploring these trade-offs in more detail. Our results suggest that we need to develop ways to jointly optimize for these objectives in order to find practical trade-offs.
Phillip Rust, Anders Søgaard
ICML1
2022 Challenges and Strategies in Cross-Cultural NLP
abstract
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, Anders Søgaard. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Daniel Hershcovich, Stella Frank, Heather C. Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Aikaterini Margatina, Phillip Rust, Anders Søgaard
ACL (1)13
2022 Date Recognition in Historical Parish Records
Laura Cabello Piqueras, Constanza Fierro, Jonas F. Lotz, Phillip Rust, Joen Rommedahl, Jeppe Klok Due, Christian Igel, Desmond Elliott, Carsten B. Pedersen, Israfel Salazar, Anders Søgaard
ICFHR4
2021 How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models
abstract
Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, Iryna Gurevych. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Phillip Rust, Jonas Pfeiffer, Ivan Vulic, Sebastian Ruder, Iryna Gurevych
ACL/IJCNLP (1)1
2020 PuzzLing Machines: A Challenge on Learning From Small Data
abstract
Deep neural models have repeatedly proved excellent at memorizing surface patterns from large datasets for various ML and NLP benchmarks.They struggle to achieve human-like thinking, however, because they lack the skill of iterative reasoning upon knowledge.To expose this problem in a new light, we introduce a challenge on learning from small data, PuzzLing Machines, which consists of Rosetta Stone puzzles from Linguistic Olympiads for high school students.These puzzles are carefully designed to contain only the minimal amount of parallel text necessary to deduce the form of unseen expressions.Solving them does not require external information (e.g., knowledge bases, visual signals) or linguistic expertise, but meta-linguistic awareness and deductive skills.Our challenge contains around 100 puzzles covering a wide range of linguistic phenomena from 81 languages.We show that both simple statistical algorithms and state-of-the-art deep neural models perform inadequately on this challenge, as expected.We hope that this benchmark, available at https://ukplab.github.io/ PuzzLing-Machines/, inspires further efforts towards a new paradigm in NLP-one that is grounded in human-like reasoning and understanding.
Gözde Gül Sahin, Yova Kementchedjhieva, Phillip Rust, Iryna Gurevych
ACL3