EDBT 2026 Demo / reviewers in the wild / expert
Niklas Stoehr
dblp:234/2996
· DBLP profile ↗
15ranked-venue papers
6as first author
14since 2021 · last 2025
0000-0003-2867-0236ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 6 first-author · 14 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Taxonomy-Aware Evaluation of Vision-Language ModelsabstractWhen a vision–language model (VLM) is prompted to identify an entity depicted in an image, it may answer "I see a conifer," rather than the specific label NORWAY SPRUCE. This raises two issues for evaluation: Firstly, the unconstrained generated text needs to be mapped to the evaluation label space (i.e., CONIFER). Secondly, a useful classification measure should give partial credit to lessspecific, but not incorrect, answers (NORWAY SPRUCE being a type of CONIFER). To meet these requirements, we propose a framework for evaluating unconstrained text predictions such as those generated from a vision–language model against a taxonomy. Specifically, we propose the use of hierarchical precision and recall measures to assess the level of correctness and specificity of predictions with regard to a taxonomy. Experimentally, we first show that existing text similarity measures do not capture taxonomic similarity well. We then develop and compare different methods to map textual VLM predictions onto a taxonomy. This allows us to compute hierarchical similarity measures between the generated text and the ground truth labels. Finally, we analyze modern VLMs on fine-grained visual classification tasks based on our proposed taxonomic evaluation scheme. Data and code are made available at https://github.com/vesteinn/vlm-eval. Vésteinn Snæbjarnarson, Kevin Du, Niklas Stoehr, Serge J. Belongie, Ryan Cotterell, Nico Lang, Stella Frank |
CVPR | 3 |
| 2025 | Measuring scalar constructs in social science with LLMsabstractHauke Licht, Rupak Sarkar, Patrick Y. Wu, Pranav Goel, Niklas Stoehr, Elliott Ash, Alexander Miserlis Hoyle. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Hauke Licht, Rupak Sarkar, Patrick Y. Wu, Pranav Goel 0001, Niklas Stoehr, Elliott Ash, Alexander Miserlis Hoyle |
EMNLP | 5 |
| 2025 | Controllable Context Sensitivity and the Knob Behind ItabstractWhen making predictions, a language model must trade off how much it relies on its context vs. its prior knowledge.
Choosing how sensitive the model is to its context is a fundamental functionality, as it enables the model to excel at tasks like retrieval-augmented generation and question-answering.
In this paper, we search for a knob which controls this sensitivity, determining whether language models answer from the context or their prior knowledge.
To guide this search, we design a task for controllable context sensitivity.
In this task, we first feed the model a context ("Paris is in England") and a question ("Where is Paris?"); we then instruct the model to either use its prior or contextual knowledge and evaluate whether it generates the correct answer for both intents (either "France" or "England").
When fine-tuned on this task, instruct versions of Llama-3.1, Mistral-v0.3, and Gemma-2 can solve it with high accuracy (85-95%).
Analyzing these high-performing models, we narrow down which layers may be important to context sensitivity using a novel linear time algorithm.
Then, in each model, we identify a 1-D subspace in a single layer that encodes whether the model follows context or prior knowledge.
Interestingly, while we identify this subspace in a fine-tuned model, we find that the exact same subspace serves as an effective knob in not only that model but also non-fine-tuned instruct and base models of that model family.
Finally, we show a strong correlation between a model's performance and how distinctly it separates context-agreeing from context-ignoring answers in this subspace.
These results suggest a single fundamental subspace facilitates how the model chooses between context and prior knowledge. Julian Minder, Kevin Du, Niklas Stoehr, Giovanni Monea, Chris Wendler, Robert West 0001, Ryan Cotterell |
ICLR | 3 |
| 2024 | Context versus Prior Knowledge in Language ModelsabstractKevin Du, Vésteinn Snæbjarnarson, Niklas Stoehr, Jennifer White, Aaron Schein, Ryan Cotterell. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Kevin Du, Vésteinn Snæbjarnarson, Niklas Stoehr, Jennifer C. White, Aaron Schein, Ryan Cotterell |
ACL (1) | 3 |
| 2024 | Unsupervised Contrast-Consistent Ranking with Language ModelsabstractNiklas Stoehr, Pengxiang Cheng, Jing Wang, Daniel Preotiuc-Pietro, Rajarshi Bhowmik. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Niklas Stoehr, Pengxiang Cheng 0001, Jing Wang 0069, Daniel Preotiuc-Pietro, Rajarshi Bhowmik |
EACL (1) | 1 |
| 2023 | Generalizing Backpropagation for Gradient-Based InterpretabilityabstractMany popular feature-attribution methods for interpreting deep neural networks rely on computing the gradients of a model's output with respect to its inputs.While these methods can indicate which input features may be important for the model's prediction, they reveal little about the inner workings of the model itself.In this paper, we observe that the gradient computation of a model is a special case of a more general formulation using semirings.This observation allows us to generalize the backpropagation algorithm to efficiently compute other interpretable statistics about the gradient graph of a neural network, such as the highest-weighted path and entropy.We implement this generalized algorithm, evaluate it on synthetic datasets to better understand the statistics it computes, and apply it to study BERT's behavior on the subject-verb number agreement task (SVA).With this method, we (a) validate that the amount of gradient flow through a component of a model reflects its importance to a prediction and (b) for SVA, identify which pathways of the self-attention mechanism are most important. Kevin Du, Lucas Torroba Hennigen, Niklas Stoehr, Alex Warstadt, Ryan Cotterell |
ACL (1) | 3 |
| 2023 | An Ordinal Latent Variable Model of Conflict IntensityabstractNiklas Stoehr, Lucas Torroba Hennigen, Josef Valvoda, Robert West, Ryan Cotterell, Aaron Schein. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Niklas Stoehr, Lucas Torroba Hennigen, Josef Valvoda, Robert West 0001, Ryan Cotterell, Aaron Schein |
ACL (1) | 1 |
| 2023 | The Ordered Matrix Dirichlet for State-Space ModelsabstractMany dynamical systems in the real world are naturally described by latent states with intrinsic ordering, such as “ally”, “neutral”, and “enemy” relationships in international relations. These latent states manifest through countries’ cooperative versus conflictual interactions over time. State-space models (SSMs) explicitly relate the dynamics of observed measurements to transitions in latent states. For discrete data, SSMs commonly do so through a state-to-action emission matrix and a state-to-state transition matrix. This paper introduces the Ordered Matrix Dirichlet (OMD) as a prior distribution over ordered stochastic matrices wherein the discrete distribution in the kth row is stochastically dominated by the (k+1)th, such that probability mass is shifted to the right when moving down rows. We illustrate the OMD prior within two SSMs: a hidden Markov model, and a novel dynamic Poisson Tucker decomposition model tailored to international relations data. We find that models built on the OMD recover interpretable ordered latent structure without forfeiting predictive performance. We suggest future applications to other domains where models with stochastic matrices are popular (e.g., topic modeling), and publish user-friendly code. Niklas Stoehr, Benjamin J. Radford, Ryan Cotterell, Aaron Schein |
AISTATS | 1 |
| 2023 | Sentiment as an Ordinal Latent VariableabstractSentiment analysis has become a central tool in various disciplines outside of natural language processing.In particular in applied and domain-specific settings with strong requirements for interpretable methods, dictionary-based approaches are still a popular choice.However, existing dictionaries are often limited in coverage, static once annotation is completed and sentiment scales differ widely; some are discrete others continuous.We propose a Bayesian generative model that learns a composite sentiment dictionary as an interpolation between six existing dictionaries with different scales.We argue that sentiment is a latent concept with intrinsically ranking-based characteristics -the word "excellent" may be ranked more positive than "great" and "okay", but it is hard to express how much more exactly.This prompts us to enforce an ordinal scale of ordered discrete sentiment values in our dictionary.We achieve this through an ordering transformation in the priors of our model.We evaluate the model intrinsically by imputing missing values in existing dictionaries.Moreover, we conduct extrinsic evaluations through sentiment classification tasks.Finally, we present two extension: first, we present a method to augment dictionary-based approaches with word embeddings to construct sentiment scales along new semantic axes.Second, we demonstrate a Latent Dirichlet Allocation-inspired variant of our model that learns document topics that are ordered by sentiment. Niklas Stoehr, Ryan Cotterell, Aaron Schein |
EACL | 1 |
| 2023 | Extracting Victim Counts from TextabstractDecision-makers in the humanitarian sector rely on timely and exact information during crisis events. Knowing how many civilians were injured during an earthquake is vital to allocate aids properly. Information about such victim counts are however often only available within full-text event descriptions from newspapers and other reports. Extracting numbers from text is challenging: numbers have different formats and may require numeric reasoning. This renders purely tagging approaches insufficient. As a consequence, fine-grained counts of injured, displaced, or abused victims beyond fatalities are often not extracted and remain unseen. We cast victim count extraction as a question answering (QA) task with a regression or classification objective. We compare tagging approaches: regex, dependency parsing, semantic role labeling, and advanced text-to-text models. Beyond model accuracy, we analyze extraction reliability and robustness which are key for this sensitive task. In particular, we discuss model calibration and investigate out-of-distribution and few-shot performance. Ultimately, we make a comprehensive recommendation on which model to select for different desiderata and data domains. Our work is among the first to apply numeracy-focused large language models in a real-world use case with a positive impact. Mian Zhong, Shehzaad Dhuliawala, Niklas Stoehr |
EACL | 3 |
| 2022 | Attentional Probe: Estimating a Module's Functional PotentialabstractIn this paper, we seek to measure how much information a component in a neural network could extract from the representations fed into it.Our work stands in contrast to prior probing work, most of which investigates how much information a model's representations contain.This shift in perspective leads us to propose a new principle for probing, the architectural bottleneck principle: In order to estimate how much information a given component could extract, a probe should look exactly like the component.Relying on this principle, we estimate how much syntactic information is available to transformers through our attentional probe, a probe that exactly resembles a transformer's self-attention head.Experimentally, we find that, in three models (BERT, ALBERT, and RoBERTa), a sentence's syntax tree is mostly extractable by our probe, suggesting these models have access to syntactic information while composing their contextual representations.Whether this information is actually used by these models, however, remains an open question. Tiago Pimentel, Josef Valvoda, Niklas Stoehr, Ryan Cotterell |
EMNLP | 3 |
| 2022 | UniMorph 4.0: Universal MorphologyabstractThe Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation, and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements on several fronts that were made in the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 66 new languages, including 24 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g., missing gender and macrons information. We have amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet. Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieras, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina J. Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Lane 0002, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóga, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer C. White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo Maria Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar 0002, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Tucker Prud'hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova |
LREC | 68 |
| 2021 | Classifying Dyads for Militarized Conflict AnalysisabstractUnderstanding the origins of militarized conflict is a complex, yet important undertaking.Existing research seeks to build this understanding by considering bi-lateral relationships between entity pairs (dyadic causes) and multi-lateral relationships among multiple entities (systemic causes).The aim of this work is to compare these two causes in terms of how they correlate with conflict between two entities.We do this by devising a set of textual and graph-based features which represent each of the causes.The features are extracted from Wikipedia and modeled as a large graph.Nodes in this graph represent entities connected by labeled edges representing ally or enemy-relationships.This allows casting the problem as an edge classification task, which we term dyad classification.We propose and evaluate classifiers to determine if a particular pair of entities are allies or enemies.Our results suggest that our systemic features might be slightly better correlates of conflict.Further, we find that Wikipedia articles of allies are semantically more similar than enemies. 1 Niklas Stoehr, Lucas Torroba Hennigen, Samin Ahbab, Robert West 0001, Ryan Cotterell |
EMNLP (1) | 1 |
| 2021 | What About the Precedent: An Information-Theoretic Analysis of Common LawabstractJosef Valvoda, Tiago Pimentel, Niklas Stoehr, Ryan Cotterell, Simone Teufel. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Josef Valvoda, Tiago Pimentel, Niklas Stoehr, Ryan Cotterell, Simone Teufel |
NAACL-HLT | 3 |
| 2018 | Heatflip: Temporal-Spatial Sampling for Progressive Heat Maps on Social Media DataabstractKeyword-based heat maps are a natural way to explore and analyze the spatial properties of social media data. Dealing with large datasets, there may be many different keywords, making offline pre-computations very hard. Interactive frameworks that exploit database sampling can address this challenge. We present a novel middleware technique called Heatflip, which issues diametrically opposed samples into the temporal and spatial dimensions of the data stored in an external database. Spatial samples provide insights into the temporal distribution and vice versa. The progressive exploration approach benefits from adaptive indexing and combines the retrieval and visualization of the data in a middleware layer. Without any a priori knowledge of the underlying data, the middleware can generate accurate heat maps in 85% shorter processing times than conventional systems. In this paper, we discuss the analytical background of Heatflip, showcase its scalability, and validate its performance when visualizing large amounts of social media data. Niklas Stoehr, Johannes Jakob Meyer, Volker Markl, Qiushi Bai, Taewoo Kim 0001, De-Yu Chen, Chen Li 0001 |
IEEE BigData | 1 |