VLDB 2026 Research / reviewers in the wild / expert
Katharina von der Wense
dblp:359/0314
· DBLP profile ↗
22ranked-venue papers
0as first author
22since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 22 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Meenz bleibt Meenz, but Large Language Models Do Not Speak Its Dialect
Minh Duc Bui, Manuel Mager, Peter Herbert Kann, Katharina von der Wense |
LREC | 4 |
| 2025 | On Generalization across Measurement Systems: LLMs Entail More Test-Time Compute for Underrepresented CulturesabstractMeasurement systems (e.g., currencies) differ across cultures, but the conversions between them are well defined so that humans can state using any measurement system of their choice. Being available to users from diverse cultural backgrounds, Large Language Models (LLMs) should also be able to provide accurate information irrespective of the measurement system at hand. Using newly compiled datasets we test if this is truly the case for seven open-source LLMs, addressing three key research questions: (RQ1) What is the default system used by LLMs for each type of measurement? (RQ2) Do LLMs’ answers and their accuracy vary across different measurement systems? (RQ3) Can LLMs mitigate potential challenges w.r.t. underrepresented systems via reasoning? Our findings show that LLMs default to the measurement system predominantly used in the data. Additionally, we observe considerable instability and variance in performance across different measurement systems. While this instability can in part be mitigated by employing reasoning methods such as chain-of-thought (CoT), this implies longer responses and thereby significantly increases test-time compute (and inference costs), marginalizing users from cultural backgrounds that use underrepresented measurement systems. Minh Duc Bui, Kyung Eun Park, Goran Glavas, Fabian David Schmidt, Katharina von der Wense |
ACL (1) | 5 |
| 2025 | Improving Low-Resource Morphological Inflection via Self-Supervised ObjectivesabstractSelf-supervised objectives have driven major advances in NLP by leveraging large-scale unlabeled data, but such resources are scarce for many of the world's languages.Surprisingly, they have not been explored much for characterlevel tasks, where smaller amounts of data have the potential to be beneficial.We investigate the effectiveness of self-supervised auxiliary tasks for morphological inflection -a character-level task highly relevant for language documentation -in extremely low-resource settings, training encoder-decoder transformers for 19 languages and 13 auxiliary objectives.Autoencoding yields the best performance when unlabeled data is very limited, while character masked language modeling (CMLM) becomes more effective as data availability increases.Though objectives with stronger inductive biases influence model predictions intuitively, they rarely outperform standard CMLM.However, sampling masks based on known morpheme boundaries consistently improves performance, highlighting a promising direction for low-resource morphological modeling. Adam Wiemerslage, Katharina von der Wense |
ACL (1) | 2 |
| 2025 | From Priest to Doctor: Domain Adaptation for Low-Resource Neural Machine TranslationabstractMany of the world’s languages have insufficient data to train high-performing general neural machine translation (NMT) models, let alone domain-specific models, and often the only available parallel data are small amounts of religious texts. Hence, domain adaptation (DA) is a crucial issue faced by contemporary NMT and has, so far, been underexplored for low-resource languages. In this paper, we evaluate a set of methods from both low-resource NMT and DA in a realistic setting, in which we aim to translate between a high-resource and a low-resource language with access to only: a) parallel Bible data, b) a bilingual dictionary, and c) a monolingual target-domain corpus in the high-resource language. Our results show that the effectiveness of the tested methods varies, with the simplest one, DALI, being most effective. We follow up with a small human evaluation of DALI, which shows that there is still a need for more careful investigation of how to accomplish DA for low-resource NMT. Ali Marashian, Enora Rice, Luke Gessler, Alexis Palmer, Katharina von der Wense |
COLING | 5 |
| 2025 | Measuring Contextual Informativeness in Child-Directed TextabstractTo address an important gap in creating children’s stories for vocabulary enrichment, we investigate the automatic evaluation of how well stories convey the semantics of target vocabulary words, a task with substantial implications for generating educational content. We motivate this task, which we call measuring contextual informativeness in children’s stories, and provide a formal task definition as well as a dataset for the task. We further propose a method for automating the task using a large language model (LLM). Our experiments show that our approach reaches a Spearman correlation of 0.4983 with human judgments of informativeness, while the strongest baseline only obtains a correlation of 0.3534. An additional analysis shows that the LLM-based approach is able to generalize to measuring contextual informativeness in adult-directed text, on which it also outperforms all baselines. Maria R. Valentini, Téa Wright, Ali Marashian, Jennifer Weber, Eliana Colunga, Katharina von der Wense |
COLING | 6 |
| 2025 | ReSeeding Latent States for Sequential Language UnderstandingabstractWe introduce Refeeding State Embeddings aligned using Environmental Data (RESEED), a novel method for grounding language in environmental data.While large language models (LLMs) excel at many tasks, they continue to struggle with multi-step sequential reasoning.RESEED addresses this by producing latent embeddings aligned with the true state of the environment and refeeding these embeddings into the model before generating its output.To evaluate its effectiveness, we develop three new sequential reasoning benchmarks, each with a training set of paired state-text trajectories and several text-only evaluation sets that test generalization to longer trajectories.Across all benchmarks, RESEED significantly improves generalization and scalability over a text-only baseline.We further show that RE-SEED outperforms commercial LLMs on our benchmarks, highlighting the value of grounding language in the environment.1 Stephane Aroca-Ouellette, Katharina von der Wense, Alessandro Roncone |
EMNLP | 2 |
| 2025 | Molecular String Representation Preferences in Pretrained LLMs: A Comparative Study in Zero- & Few-Shot Molecular Property PredictionabstractLarge Language Models (LLMs) have demonstrated capabilities for natural language formulations of molecular property prediction tasks, but little is known about how performance depends on the representation of input molecules to the model; the status quo approach is to use SMILES strings, although alternative chemical notations convey molecular information differently, each with their own strengths and weaknesses.To learn more about molecular string representation preferences in LLMs, we compare the performance of four recent models-GPT-4o, Gemini 1.5 Pro, Llama 3.1 405b, and Mistral Large 2-on molecular property prediction tasks from the MoleculeNet benchmark across five different molecular string representations: SMILES, DeepSMILES, SELF-IES, InChI, and IUPAC names.We find statistically significant zero-and few-shot preferences for InChI and IUPAC names, potentially due to representation granularity, favorable tokenization, and prevalence in pretraining corpora.This contradicts previous assumptions that molecules should be presented to LLMs as SMILES strings.When these preferences are taken advantage of, few-shot performance rivals or surpasses many previous conventional approaches to property prediction, with the advantage of explainable predictions through chain-of-thought reasoning not held by taskspecific models. George Arthur Baker, Mario Sanz-Guerrero, Katharina von der Wense |
EMNLP | 3 |
| 2025 | Large Language Models Discriminate Against Speakers of German DialectsabstractDialects represent a significant component of human culture and are found across all regions of the world.In Germany, more than 40% of the population speaks a regional dialect (Adler and Hansen, 2022).However, despite cultural importance, individuals speaking dialects often face negative societal stereotypes.We examine whether such stereotypes are mirrored by large language models (LLMs).We draw on the sociolinguistic literature on dialect perception to analyze traits commonly associated with dialect speakers.Based on these traits, we assess the dialect naming bias and dialect usage bias expressed by LLMs in two tasks: an association task and a decision task.To assess a model's dialect usage bias, we construct a novel evaluation corpus that pairs sentences from seven regional German dialects (e.g., Alemannic and Bavarian) with their standard German counterparts.We find that: (1) in the association task, all evaluated LLMs exhibit significant dialect naming and dialect usage bias against German dialect speakers, reflected in negative adjective associations; (2) all models reproduce these dialect naming and dialect usage biases in their decision making; and (3) contrary to prior work showing minimal bias with explicit demographic mentions, we find that explicitly labeling linguistic demographics-German dialect speakers-amplifies bias more than implicit cues like dialect usage.* Equal contribution. 1 The literature disagrees on an exact definition; we give more information in Appendix A.1. Minh Duc Bui, Carolin Holtermann, Valentin Hofmann, Anne Lauscher, Katharina von der Wense |
EMNLP | 5 |
| 2025 | Model-Based Ranking of Source Languages for Zero-Shot Cross-Lingual TransferabstractWe present NN-RANK, an algorithm for ranking source languages for cross-lingual transfer, which leverages hidden representations from multilingual models and unlabeled targetlanguage data.We experiment with two pretrained multilingual models and two tasks: partof-speech tagging (POS) and named entity recognition (NER).We consider 51 source languages and evaluate on 56 and 72 target languages for POS and NER, respectively.When using in-domain data, NN-RANK beats stateof-the-art baselines that leverage lexical and linguistic features, with average improvements of up to 35.56 NDCG for POS and 18.14 NDCG for NER.As prior approaches can fall back to language-level features if target language data is not available, we show that NN-RANK remains competitive using only the Bible, an out-of-domain corpus available for a large number of languages.Ablations on the amount of unlabeled target data show that, for subsets consisting of as few as 25 examples, NN-RANK produces high-quality rankings which achieve 92.8% of the NDCG achieved using all available target data for ranking.POS test-all NER test-all Abteen Ebrahimi, Adam Wiemerslage, Katharina von der Wense |
EMNLP | 3 |
| 2025 | Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language DocumentationabstractComputational morphology has the potential to support language documentation through tasks like morphological segmentation and the generation of Interlinear Glossed Text (IGT).However, our research outputs have seen limited use in real-world language documentation settings.This position paper situates the disconnect between computational morphology and language documentation within a broader misalignment between research and practice in NLP and argues that the field risks becoming decontextualized and ineffectual without systematic integration of User-Centered Design (UCD).To demonstrate how principles from UCD can reshape the research agenda, we present a case study of GlossLM, a stateof-the-art multilingual IGT generation model.Through a small-scale user study with three documentary linguists, we find that, despite strong metric-based performance, the system fails to meet core usability needs in real documentation contexts.These insights raise new research questions around model constraints, label standardization, segmentation, and personalization.We argue that centering users not only produces more effective tools, but surfaces richer, more relevant research directions. Enora Rice, Katharina von der Wense, Alexis Palmer |
EMNLP | 2 |
| 2025 | Mind the Gap: A Closer Look at Tokenization for Multiple-Choice Question Answering with LLMsabstractWhen evaluating large language models (LLMs) with multiple-choice question answering (MCQA), it is common to end the prompt with the string "Answer:" to facilitate automated answer extraction via next-token probabilities.However, there is no consensus on how to tokenize the space following the colon, often overlooked as a trivial choice.In this paper, we uncover accuracy differences of up to 11% due to this (seemingly irrelevant) tokenization variation as well as reshuffled model rankings, raising concerns about the reliability of LLM comparisons in prior work.Surprisingly, we are able to recommend one specific strategy -tokenizing the space together with the answer letter -as we observe consistent and statistically significant performance improvements.Additionally, it improves model calibration, enhancing the reliability of the model's confidence estimates.Our findings underscore the importance of careful evaluation design and highlight the need for standardized, transparent evaluation protocols to ensure reliable and comparable results. Mario Sanz-Guerrero, Minh Duc Bui, Katharina von der Wense |
EMNLP | 3 |
| 2025 | More Experts Than Galaxies: Conditionally-Overlapping Experts with Biologically-Inspired Fixed RoutingabstractThe evolution of biological neural systems has led to both modularity and sparse coding, which enables energy efficiency and robustness across the diversity of tasks in the lifespan. In contrast, standard neural networks rely on dense, non-specialized architectures, where all model parameters are simultaneously updated to learn multiple tasks, leading to interference. Current sparse neural network approaches aim to alleviate this issue but are hindered by limitations such as 1) trainable gating functions that cause representation collapse, 2) disjoint experts that result in redundant computation and slow learning, and 3) reliance on explicit input or task IDs that limit flexibility and scalability.
In this paper we propose Conditionally Overlapping Mixture of ExperTs (COMET), a general deep learning method that addresses these challenges by inducing a modular, sparse architecture with an exponential number of overlapping experts. COMET replaces the trainable gating function used in Sparse Mixture of Experts with a fixed, biologically inspired random projection applied to individual input representations. This design causes the degree of expert overlap to depend on input similarity, so that similar inputs tend to share more parameters. This results in faster learning per update step and improved out-of-sample generalization.
We demonstrate the effectiveness of COMET on a range of tasks, including image classification, language modeling, and regression, using several popular deep learning architectures. Sagi Shaier, Francisco Pereira 0001, Katharina von der Wense, Lawrence Hunter, Matt Jones 0002 |
ICLR | 3 |
| 2025 | Implicitly Aligning Humans and Autonomous Agents through Shared Task AbstractionsabstractIn collaborative tasks, autonomous agents fall short of humans in their capability to quickly adapt to new and unfamiliar teammates. We posit that a limiting factor for zero-shot coordination is the lack of shared task abstractions, a mechanism humans rely on to implicitly align with teammates. To address this gap, we introduce HA^2: Hierarchical Ad Hoc Agents, a framework leveraging hierarchical reinforcement learning to mimic the structured approach humans use in collaboration. We evaluate HA^2 in the Overcooked environment, demonstrating statistically significant improvement over existing baselines when paired with both unseen agents and humans, providing better resilience to environmental shifts, and outperforming all state-of-the-art methods. Stephane Aroca-Ouellette, Miguel Aroca-Ouellette, Katharina von der Wense, Alessandro Roncone |
IJCAI | 3 |
| 2025 | Multi³Hate: Multimodal, Multilingual, and Multicultural Hate Speech Detection with Vision-Language ModelsabstractMinh Duc Bui, Katharina Von Der Wense, Anne Lauscher. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Minh Duc Bui, Katharina von der Wense, Anne Lauscher |
NAACL (Long Papers) | 2 |
| 2024 | TAMS: Translation-Assisted Morphological SegmentationabstractCanonical morphological segmentation is the process of analyzing words into the standard (aka underlying) forms of their constituent morphemes.This is a core task in endangered language documentation, and NLP systems have the potential to dramatically speed up this process.In typical language documentation settings, training data for canonical morpheme segmentation is scarce, making it difficult to train high quality models.However, translation data is often much more abundant, and, in this work, we present a method that attempts to leverage translation data in the canonical segmentation task.We propose a character-level sequence-to-sequence model that incorporates representations of translations obtained from pretrained high-resource monolingual language models as an additional signal.Our model outperforms the baseline in a super-low resource setting but yields mixed results on training splits with more data.Additionally, we find that we can achieve strong performance even without needing difficult-to-obtain word level alignments.While further work is needed to make translations useful in higher-resource settings, our model shows promise in severely resource-constrained settings. Enora Rice, Ali Marashian, Luke Gessler, Alexis Palmer, Katharina von der Wense |
ACL (1) | 5 |
| 2024 | Evaluating LLMs as Tools to Support Early Vocabulary Learning
Jennifer Weber, Maria R. Valentini, Téa Wright, Katharina von der Wense, Eliana Colunga |
CogSci | 4 |
| 2024 | Comparing Template-based and Template-free Language Model ProbingabstractSagi Shaier, Kevin Bennett, Lawrence Hunter, Katharina von der Wense. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Sagi Shaier, Kevin Bennett, Lawrence Hunter, Katharina von der Wense |
EACL (1) | 4 |
| 2024 | Desiderata For The Context Use Of Question Answering SystemsabstractPrior work has uncovered a set of common problems in state-of-the-art context-based question answering (QA) systems: a lack of attention to the context when the latter conflicts with a model's parametric knowledge, little robustness to noise, and a lack of consistency with their answers.However, most prior work focus on one or two of those problems in isolation, which makes it difficult to see trends across them.We aim to close this gap, by first outlining a set of -previously discussed as well as novel -desiderata for QA models.We then survey relevant analysis and methods papers to provide an overview of the state of the field.The second part of our work presents experiments where we evaluate 15 QA systems on 5 datasets according to all desiderata at once.We find many novel trends, including (1) systems that are less susceptible to noise are not necessarily more consistent with their answers when given irrelevant context; (2) most systems that are more susceptible to noise are more likely to correctly answer according to a context that conflicts with their parametric knowledge; and (3) the combination of conflicting knowledge and noise can reduce system performance by up to 96%.As such, our desiderata help increase our understanding of how these models work and reveal potential avenues for improvements.Code and data can be found here: https://github.com/Shaier/ context_usage_desiderata.git. Sagi Shaier, Lawrence Hunter, Katharina von der Wense |
EACL (1) | 3 |
| 2024 | Quantifying the Hyperparameter Sensitivity of Neural Networks for Character-level Sequence-to-Sequence TasksabstractHyperparameter tuning, the process of searching for suitable hyperparameters, becomes more difficult as the computing resources required to train neural networks continue to grow.This topic continues to receive little attention and discussion-much of it hearsaydespite its obvious importance.We attempt to formalize hyperparameter sensitivity using two metrics: similarity-based sensitivity and performance-based sensitivity.We then use these metrics to quantify two such claims: (1) transformers are more sensitive to hyperparameter choices than LSTMs and (2) transformers are particularly sensitive to batch size.We conduct experiments on two different characterlevel sequence-to-sequence tasks and find that, indeed, the transformer is slightly more sensitive to hyperparameters according to both of our metrics.However, we do not find that it is more sensitive to batch size in particular. Adam Wiemerslage, Kyle Gorman, Katharina von der Wense |
EACL (1) | 3 |
| 2024 | Getting The Most Out of Your Training Data: Exploring Unsupervised Tasks for Morphological InflectionabstractPretrained transformers such as BERT (Devlin et al., 2019) have been shown to be effective in many natural language tasks.However, they are under-explored for character-level sequence-tosequence tasks.In this work, we investigate pretraining transformers for the character-level task of morphological inflection in several languages.We compare various training setups and secondary tasks where unsupervised data taken directly from the target task is used.We show that training on secondary unsupervised tasks increases inflection performance even without any external data, suggesting that models learn from additional unsupervised tasks themselves-not just from additional data.We also find that this does not hold true for specific combinations of secondary task and training setup, which has interesting implications for unsupervised training and denoising objectives in character-level tasks. Abhishek Purushothama, Adam Wiemerslage, Katharina von der Wense |
EMNLP | 3 |
| 2024 | Eyes on the Game: Deciphering Implicit Human Signals to Infer Human Proficiency, Trust, and IntentabstractEffective collaboration between humans and AIs hinges on transparent communication and alignment of mental models. However, explicit, verbal communication is not always feasible. Under such circumstances, human-human teams often depend on implicit, nonverbal cues to glean important information about their teammates such as intent and expertise, thereby bolstering team alignment and adaptability. Among these implicit cues, two of the most salient and fundamental are a human’s actions in the environment and their visual attention. In this paper, we present a novel method to combine eye gaze data and behavioral data, and evaluate their respective predictive power for human proficiency, trust, and intent. We first collect a dataset of paired eye gaze and gameplay data in the fast-paced collaborative "Overcooked" environment. We then train models on this dataset to compare how the predictive powers differ between gaze data, gameplay data, and their combination. We additionally compare our method to prior works that aggregate eye gaze data and demonstrate how these aggregation methods can substantially reduce the predictive ability of eye gaze. Our results indicate that, while eye gaze data and gameplay data excel in different situations, a model that integrates both types consistently outperforms all baselines. This work paves the way for developing intuitive and responsive agents that can efficiently adapt to new teammates. Nikhil Hulle, Stephane Aroca-Ouellette, Anthony J. Ries, Jake Brawer, Katharina von der Wense, Alessandro Roncone |
RO-MAN | 5 |
| 2023 | On the Automatic Generation and Simplification of Children's StoriesabstractWith recent advances in large language models (LLMs), the concept of automatically generating children's educational materials has become increasingly realistic.Working toward the goal of age-appropriate simplicity in generated educational texts, we first examine the ability of several popular LLMs to generate stories with properly adjusted lexical and readability levels.We find that, in spite of the growing capabilities of LLMs, they do not yet possess the ability to limit their vocabulary to levels appropriate for younger age groups.As a second experiment, we explore the ability of state-ofthe-art lexical simplification models to generalize to the domain of children's stories and, thus, create an efficient pipeline for their automatic generation.In order to test these models, we develop a dataset of child-directed lexical simplification instances, with examples taken from the LLM-generated stories in our first experiment.We find that, while the strongest-performing lexical simplification models do not perform as well on material designed for children due to their reliance on LLMs, a model that performs well on general data strongly improves its performance on children-directed data with proper finetuning, which we conduct using our newly created child-directed simplification dataset. Maria R. Valentini, Jennifer Weber, Jesus Salcido, Téa Wright, Eliana Colunga, Katharina von der Wense |
EMNLP | 6 |