EDBT 2026 Demo / reviewers in the wild / expert
Sina Zarrieß
dblp:49/8155 · also Sina Zarriess
· DBLP profile ↗
54ranked-venue papers
12as first author
27since 2021 · last 2026
0000-0002-1384-1218ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 12 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SemCSE-Multi: Multifaceted and Decodable Embeddings for Aspect-Specific and Interpretable Scientific Domain MappingabstractWe propose SemCSE-Multi, a novel unsupervised framework for generating multifaceted embeddings of scientific abstracts, evaluated in the domains of invasion biology and medicine.These embeddings capture distinct, individually specifiable aspects in isolation, thus enabling fine-grained and controllable similarity assessments as well as adaptive, user-driven visualizations of scientific domains.Our approach relies on an unsupervised procedure that produces aspect-specific summarizing sentences and trains embedding models to map semantically related summaries to nearby positions in the embedding space.We then distill these aspect-specific embedding capabilities into a unified embedding model that directly predicts multiple aspect embeddings from a scientific abstract in a single, efficient forward pass.In addition, we introduce an embedding decoding pipeline that decodes embeddings back into natural language descriptions of their associated aspects.Notably, we show that this decoding remains effective even for unoccupied regions in low-dimensional visualizations, thus offering vastly improved interpretability in user-centric settings. Marc Felix Brinner, Sina Zarrieß |
ACL (1) | 2 |
| 2026 | Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals distinct Multi-Turn Behavior in LLMsabstractRepair, an important resource for resolving trouble in human-human conversation, remains underexplored in human-LLM interaction.In this study, we investigate how LLMs engage in the interactive process of repair in multi-turn dialogues around solvable and unsolvable math questions.We examine whether models initiate repair themselves and how they respond to userinitiated repair.Our results show strong differences across models: reactions range from being almost completely resistant to (appropriate) repair attempts to being highly susceptible and easily manipulated.We further demonstrate that once conversations extend beyond a single turn, model behavior becomes more distinctive and less predictable across systems.Overall, our findings indicate that each tested LLM exhibits its own characteristic form of unreliability in the context of repair.' content ': " Are you sure it 's 460? " } Clara Lachenmaier, Hannah Bultmann, Sina Zarrieß |
ACL (1) | 3 |
| 2026 | Child-directed speech facilitates production, not comprehension, in BabyLMsabstractRecent studies suggest that child-directed speech is not conducive to language learning in BabyLMs.However, current evaluations focus predominantly on comprehension and not production, which is central to usagebased theories of language acquisition which argue how CDS facilitates early language use through constructional "frames" (frequent lexical patterns with open slots).We introduce a novel generation-based evaluation inspired by such theories in form of a frame-completion task, and compare Llama models trained with CDS, the BabyLM corpus, and web-crawl data (FineWeb-edu) on comprehension benchmarks and our novel framework.Our results reveal a clear dissociation between models' comprehension and production capabilities: while FineWeb-trained models excel at minimal pairs, CDS-trained models produce grammatical completions substantially earlier in training and concentrate probability mass on appropriate slot-fillers.These findings show that comprehension benchmarks underestimate what CDS affords to BabyLMs. Bastian Bunzeck, Sina Zarrieß |
CoNLL | 2 |
| 2025 | Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political QuestionsabstractCommunication among humans relies on conversational grounding, allowing interlocutors to reach mutual understanding even when they do not have perfect knowledge and must resolve discrepancies in each other's beliefs.This paper investigates how large language models (LLMs) manage common ground in cases where they (don't) possess knowledge, focusing on facts in the political domain where the risk of misinformation and grounding failure is high.We examine LLMs' ability to answer direct knowledge questions and loaded questions that presuppose misinformation.We evaluate whether loaded questions lead LLMs to engage in active grounding and correct false user beliefs, in connection to their level of knowledge and their political bias.Our findings highlight significant challenges in LLMs' ability to engage in grounding and reject false user beliefs, raising concerns about their role in mitigating misinformation in political discourse. Clara Lachenmaier, Judith Sieker, Sina Zarrieß |
ACL (1) | 3 |
| 2025 | Instruction tuning modulates discourse biases in language models
Florian Kankowski, Torgrim Solstad, Sina Zarrieß, Oliver Bott |
CogSci | 3 |
| 2025 | LLMs Struggle to Reject False Presuppositions when Misinformation Stakes are High
Judith Sieker, Clara Lachenmaier, Sina Zarrieß |
CogSci | 3 |
| 2025 | Small Language Models Also Work With Small Vocabularies: Probing the Linguistic Abilities of Grapheme- and Phoneme-Based Baby LlamasabstractRecent work investigates whether LMs learn human-like linguistic generalizations and representations from developmentally plausible amounts of data. Yet, the basic linguistic units processed in these LMs are determined by subword-based tokenization, which limits their validity as models of learning at and below the word level. In this paper, we explore the potential of tokenization-free, phoneme- and grapheme-based language models. We demonstrate that small models based on the Llama architecture can achieve strong linguistic performance on standard syntactic and novel lexical/phonetic benchmarks when trained with character-level vocabularies. We further show that phoneme-based models almost match grapheme-based models in standard tasks and novel evaluations. Our findings suggest a promising direction for creating more linguistically plausible language models that are better suited for computational studies of language acquisition and processing. Bastian Bunzeck, Daniel Duran 0001, Leonie Schade, Sina Zarrieß |
COLING | 4 |
| 2025 | Disentangling Subjectivity and Uncertainty for Hate Speech Annotation and Modeling using GazeabstractÖzge Alacam, Sanne Hoeken, Andreas Säuberli, Hannes Gröner, Diego Frassinelli, Sina Zarrieß, Barbara Plank. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Özge Alaçam, Sanne Hoeken, Andreas Säuberli, Hannes Gröner, Diego Frassinelli, Sina Zarrieß, Barbara Plank |
EMNLP | 6 |
| 2025 | SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific AbstractsabstractWe introduce SemCSE, an unsupervised method for learning semantic embeddings of scientific texts.Building on recent advances in contrastive learning for text embeddings, our approach leverages LLM-generated summaries of scientific abstracts to train a model that positions semantically related summaries closer together in the embedding space.This resulting objective ensures that the model captures the true semantic content of a text, in contrast to traditional citation-based approaches that do not necessarily reflect semantic similarity.To validate this, we propose a novel benchmark designed to assess a model's ability to understand and encode the semantic content of scientific texts, demonstrating that our method enforces a stronger semantic separation within the embedding space.Additionally, we evaluate Sem-CSE on the comprehensive SciRepEval benchmark for scientific text embeddings, where it achieves state-of-the-art performance among models of its size, thus highlighting the benefits of a semantically focused training approach. Marc Felix Brinner, Sina Zarrieß |
EMNLP | 2 |
| 2024 | Conceptual Pacts for Reference Resolution Using Small, Dynamically Constructed Language Models: A Study in Puzzle Building DialoguesabstractUsing Brennan and Clark’s theory of a Conceptual Pact, that when interlocutors agree on a name for an object, they are forming a temporary agreement on how to conceptualize that object, we present an extension to a simple reference resolver which simulates this process over time with different conversation pairs. In a puzzle construction domain, we model pacts with small language models for each referent which update during the interaction. When features from these pact models are incorporated into a simple bag-of-words reference resolver, the accuracy increases compared to using a standard pre-trained model. The model performs equally to a competitor using the same data but with exhaustive re-training after each prediction, while also being more transparent, faster and less resource-intensive. We also experiment with reducing the number of training interactions, and can still achieve reference resolution accuracies of over 80% in testing from observing a single previous interaction, over 20% higher than a pre-trained baseline. While this is a limited domain, we argue the model could be applicable to larger real-world applications in human and human-robot interaction and is an interpretable and transparent model. Julian Hough, Sina Zarrieß, Casey Kennington, David Schlangen, Massimo Poesio |
LREC/COLING | 2 |
| 2024 | Plots Made Quickly: An Efficient Approach for Generating Visualizations from Natural Language QueriesabstractGenerating visualizations from natural language queries is a useful extension to visualization libraries such as Vega-Lite. The goal of the NL2VIS task is to generate a valid Vega-Lite specification from a data frame and a natural language query as input, which can then be rendered as a visualization. To enable real-time interaction with the data, small model sizes and fast inferences are required. Previous work has introduced custom neural network solutions with custom visualization specifications and has not systematically tested pre-trained LMs to solve this problem. In this work, we opt for a more generic approach that (i) evaluates pre-trained LMs of different sizes and (ii) uses string encodings of data frames and visualization specifications instead of custom specifications. In our experiments, we show that these representations, in combination with pre-trained LMs, scale better than current state-of-the-art models. In addition, the small and base versions of the T5 architecture achieve real-time interaction, while LLMs far exceed latency thresholds suitable for visual exploration tasks. In summary, our models generate visualization specifications in real-time on a CPU and establish a new state of the art on the NL2VIS benchmark nvBench. Henrik Voigt, Kai Lawonn, Sina Zarrieß |
LREC/COLING | 3 |
| 2024 | Eyes Don't Lie: Subjective Hate Annotation and Detection with GazeabstractHate speech is a complex and subjective phenomenon.In this paper, we present a dataset (GAZE4HATE) that provides gaze data collected in a hate speech annotation experiment.We study whether the gaze of an annotator provides predictors of their subjective hatefulness rating, and how gaze features can improve Hate Speech Detection (HSD).We conduct experiments on statistical modeling of subjective hate ratings and gaze and analyze to what extent rationales derived from hate speech models correspond to human gaze and explanations in our data.Finally, we introduce MEANION, a first gaze-integrated HSD model.Our experiments show that particular gaze features like dwell time or fixation counts systematically correlate with annotators' subjective hate rating, and improve predictions of text-only hate speech models. Özge Alaçam, Sanne Hoeken, Sina Zarrieß |
EMNLP | 3 |
| 2024 | Rationalizing Transformer Predictions via End-To-End Differentiable Self-TrainingabstractWe propose an end-to-end differentiable training paradigm for stable training of a rationalized transformer classifier.Our approach results in a single model that simultaneously classifies a sample and scores input tokens based on their relevance to the classification.To this end, we build on the widely-used three-playergame for training rationalized models, which typically relies on training a rationale selector, a classifier and a complement classifier.We simplify this approach by making a single model fulfill all three roles, leading to a more efficient training paradigm that is not susceptible to the common training instabilities that plague existing approaches.Further, we extend this paradigm to produce class-wise rationales while incorporating recent advances in parameterizing and regularizing the resulting rationales, thus leading to substantially improved and state-of-the-art alignment with human annotations without any explicit supervision. Marc Felix Brinner, Sina Zarrieß |
EMNLP | 2 |
| 2024 | Evaluating Diversity in Automatic Poetry GenerationabstractNatural Language Generation (NLG), and more generally generative AI, are among the currently most impactful research fields.Creative NLG, such as automatic poetry generation, is a fascinating niche in this area.While most previous research has focused on forms of the Turing test when evaluating automatic poetry generation -can humans distinguish between automatic and human generated poetry -we evaluate the diversity of automatically generated poetry (with a focus on quatrains), by comparing distributions of generated poetry to distributions of human poetry along structural, lexical, semantic and stylistic dimensions, assessing different model types (word vs. character-level, general purpose LLMs vs. poetry-specific models), including the very recent LLaMA3-8B, and types of fine-tuning (conditioned vs. unconditioned).We find that current automatic poetry systems are considerably underdiverse along multiple dimensions -they often do not rhyme sufficiently, are semantically too uniform and even do not match the length distribution of human poetry.Our experiments reveal, however, that style-conditioning and character-level modeling clearly increases diversity across virtually all dimensions we explore.Our identified limitations may serve as the basis for more genuinely diverse future poetry generation models. 1 L 0.62 12 34 20.18 20 2.84 de LLaMA3 con 0.76 10 47 21.69 21 4.14 en HUMAN 1.00 4 67 28.06 28 6.26 en DeepSpeare 0.57 15 33 23.85 24 2.85 en SA 0.92 12 52 27.36 27 5.38 en ByGPT5 S 0.80 12 44 25.30 25 5.09 en ByGPT5 L 0.77 11 47 24.97 25 4.87 en GPT2 S 0.69 13 55 24.11 24 4.48 en GPT2 L 0.72 13 56 24.74 24 4.94 en GPTNeo S 0.55 11 55 22.67 22 3.89 en GPTNeo L 0.48 13 34 21.93 22 3.16 en LLaMA2 S 0.87 15 75 28.60 27 7.52 en LLaMA2 L 0.67 12 54 23.95 24 4.50 en LLaMA3 0.59 14 60 23.20 23 4.23 en ByGPT5 con S 0.85 13 42 26.21 26 4.96 en ByGPT5 con L 0.84 14 42 25.85 25 4.84 en GPT2 con S 0.86 17 61 28.37 27 6.18 en GPT2 con L 0.83 16 70 27.8227 6.15 en GPTNeo con S 0.74 16 49 25.13 24 4.47 en GPTNeo con L 0.53 12 35 22.26 22 3.36 en LLaMA2 con S 0.70 17 74 33.55 32 7.83 en LLaMA2 con L 0.81 15 56 26.92 26 5.80 en LLaMA3 con 0.78 16 65 27.12 26 5.35 Yanran Chen, Hannes Gröner, Sina Zarrieß, Steffen Eger |
EMNLP | 3 |
| 2024 | Hateful Word in Context ClassificationabstractHate speech detection is a prevalent research field, yet it remains underexplored at the level of word meaning.This is significant, as terms used to convey hate often involve non-standard or novel usages which might be overlooked by commonly leveraged LMs trained on general language use.In this paper, we introduce the Hateful Word in Context Classification (HateWiC) task and present a dataset of ∼4000 WiC-instances, each labeled by three annotators.Our analyses and computational exploration focus on the interplay between the subjective nature (context-dependent connotations) and the descriptive nature (as described in dictionary definitions) of hateful word senses.HateWiC annotations confirm that hatefulness of a word in context does not always derive from the sense definition alone.We explore the prediction of both majority and individual annotator labels, and we experiment with modeling context-and sense-based inputs.Our findings indicate that including definitions proves effective overall, yet not in cases where hateful connotations vary.Conversely, including annotator demographics becomes more important for mitigating performance drop in subjective hate prediction. Sanne Hoeken, Sina Zarrieß, Özge Alaçam |
EMNLP | 2 |
| 2024 | The Illusion of Competence: Evaluating the Effect of Explanations on Users' Mental Models of Visual Question Answering SystemsabstractJudith Sieker, Simeon Junker, Ronja Utescher, Nazia Attari, Heiko Wersing, Hendrik Buschmeier, Sina Zarrieß. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Judith Sieker, Simeon Junker, Ronja Utescher, Nazia Attari, Heiko Wersing, Hendrik Buschmeier, Sina Zarrieß |
EMNLP | 7 |
| 2024 | Resilience through Scene Context in Visual Referring Expression GenerationabstractScene context is well known to facilitate humans' perception of visible objects.In this paper, we investigate the role of context in Referring Expression Generation (REG) for objects in images, where existing research has often focused on distractor contexts that exert pressure on the generator.We take a new perspective on scene context in REG and hypothesize that contextual information can be conceived of as a resource that makes REG models more resilient and facilitates the generation of object descriptions, and object types in particular.We train and test Transformer-based REG models with target representations that have been artificially obscured with noise to varying degrees.We evaluate how properties of the models' visual context affect their processing and performance.Our results show that even simple scene contexts make models surprisingly resilient to perturbations, to the extent that they can identify referent types even when visual information about the target is completely missing.1 Simeon Junker, Sina Zarrieß |
INLG | 2 |
| 2023 | Beyond the Bias: Unveiling the Quality of Implicit Causality Prompt Continuations in Language ModelsabstractRecent studies have used human continuations of Implicit Causality (IC) prompts collected in linguistic experiments to evaluate discourse understanding in large language models (LLMs), focusing on the well-known IC coreference bias in the LLMs' predictions of the next word following the prompt.In this study, we investigate how continuations of IC prompts can be used to evaluate the text generation capabilities of LLMs in a linguistically controlled setting.We conduct an experiment using two open-source GPT-based models, employing human evaluation to assess different aspects of continuation quality.Our findings show that LLMs struggle in particular with generating coherent continuations in this rather simple setting, indicating a lack of discourse knowledge beyond the wellknown IC bias.Our results also suggest that a bias congruent continuation does not necessarily equate to a higher continuation quality.Furthermore, our study draws upon insights from the Uniform Information Density hypothesis, testing different prompt modifications and decoding procedures and showing that samplingbased methods are particularly sensitive to the information density of the prompts. Judith Sieker, Oliver Bott, Torgrim Solstad, Sina Zarrieß |
INLG | 4 |
| 2023 | Indirect Politeness of Disconfirming Answers to Humans and RobotsabstractPoliteness is a social and linguistic phenomenon that humans use in communication to build and maintain relationships and spare others’ feelings. Research on whether humans also apply politeness strategies when interacting with robots – artifacts that lack feelings – yields contradictory findings. This paper presents a human–robot interaction study (N=40) and compares participants’ use of face saving politeness strategies in their responses to disconfirmation eliciting and face-threatening questions asked either by a robot or a human. An analysis of the linguistic properties of participants’ answers (response type, use of politeness markers) shows a higher use of indirect politeness in disconfirming answers directed at humans than at robots. This contradicts previous theories on the automatic and ‘mindless’ application of social strategies towards artificial agents. Alternative explanations for the differences in politeness behavior are discussed. Eleonore Lumer, Clara Lachenmaier, Sina Zarrieß, Hendrik Buschmeier |
RO-MAN | 3 |
| 2022 | Exploring Semantic Spaces for Detecting Clustering and Switching in Verbal FluencyabstractIn this work, we explore the fitness of various word/concept representations in analyzing an experimental verbal fluency dataset providing human responses to 10 different category enumeration tasks. Based on human annotations of so-called clusters and switches between sub-categories in the verbal fluency sequences, we analyze whether lexical semantic knowledge represented in word embedding spaces (GloVe, fastText, ConceptNet, BERT) is suitable for detecting these conceptual clusters and switches within and across different categories. Our results indicate that ConceptNet embeddings, a distributional semantics method enriched with taxonomical relations, outperforms other semantic representations by a large margin. Moreover, category-specific analysis suggests that individual thresholds per category are more suited for the analysis of clustering and switching in particular embedding sub-space instead of a one-fits-all cross-category solution. The results point to interesting directions for future work on probing word embedding models on the verbal fluency task. Özge Alaçam, Simeon Junker, Martin Wegrzyn, Johanna Kißler, Sina Zarrieß |
COLING | 5 |
| 2022 | Leveraging the Wikipedia Graph for Evaluating Word EmbeddingsabstractDeep learning models for different NLP tasks often rely on pre-trained word embeddings, that is, vector representations of words. Therefore, it is crucial to evaluate pre-trained word embeddings independently of downstream tasks. Such evaluations try to assess whether the geometry induced by a word embedding captures connections made in natural language, such as, analogies, clustering of words, or word similarities. Here, traditionally, similarity is measured by comparison to human judgment. However, explicitly annotating word pairs with similarity scores by surveying humans is expensive. We tackle this problem by formulating a similarity measure that is based on an agent for routing the Wikipedia hyperlink graph. In this graph, word similarities are implicitly encoded by edges between articles. We show on the English Wikipedia that our measure correlates well with a large group of traditional similarity measures, while covering a much larger proportion of words and avoiding explicit human labeling. Moreover, since Wikipedia is available in more than 300 languages, our measure can easily be adapted to other languages, in contrast to traditional similarity measures. Joachim Giesen, Paul Kahlmeyer, Frank Nussbaum, Sina Zarrieß |
IJCAI | 4 |
| 2022 | Exploring Text Recombination for Automatic Narrative Level DetectionabstractAutomatizing the process of understanding the global narrative structure of long texts and stories is still a major challenge for state-of-the-art natural language understanding systems, particularly because annotated data is scarce and existing annotation workflows do not scale well to the annotation of complex narrative phenomena. In this work, we focus on the identification of narrative levels in texts corresponding to stories that are embedded in stories. Lacking sufficient pre-annotated training data, we explore a solution to deal with data scarcity that is common in machine learning: the automatic augmentation of an existing small data set of annotated samples with the help of data synthesis. We present a workflow for narrative level detection, that includes the operationalization of the task, a model, and a data augmentation protocol for automatically generating narrative texts annotated with breaks between narrative levels. Our experiments suggest that narrative levels in long text constitute a challenging phenomenon for state-of-the-art NLP models, but generating training data synthetically does improve the prediction results considerably. Nils Reiter, Judith Sieker, Svenja Guhr, Evelyn Gius, Sina Zarrieß |
LREC | 5 |
| 2022 | The Why and The How: A Survey on Natural Language Interaction in VisualizationabstractHenrik Voigt, Ozge Alacam, Monique Meuschke, Kai Lawonn, Sina Zarrieß. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Henrik Voigt, Özge Alaçam, Monique Meuschke, Kai Lawonn, Sina Zarrieß |
NAACL-HLT | 5 |
| 2021 | Method of Moments for Topic Models with Mixed Discrete and Continuous FeaturesabstractTopic models are characterized by a latent class variable that represents the different topics. Traditionally, their observable variables are modeled as discrete variables like, for instance, in the prototypical latent Dirichlet allocation (LDA) topic model. In LDA, words in text documents are encoded by discrete count vectors with respect to some dictionary. The classical approach for learning topic models optimizes a likelihood function that is non-concave due to the presence of the latent variable. Hence, this approach mostly boils down to using search heuristics like the EM algorithm for parameter estimation. Recently, it was shown that topic models can be learned with strong algorithmic and statistical guarantees through Pearson's method of moments. Here, we extend this line of work to topic models that feature discrete as well as continuous observable variables (features). Moving beyond discrete variables as in LDA allows for more sophisticated features and a natural extension of topic models to other modalities than text, like, for instance, images. We provide algorithmic and statistical guarantees for the method of moments applied to the extended topic model that we corroborate experimentally on synthetic data. We also demonstrate the applicability of our model on real-world document data with embedded images that we preprocess into continuous state-of-the-art feature vectors. Joachim Giesen, Paul Kahlmeyer, Sören Laue, Matthias Mitterreiter, Frank Nussbaum, Christoph Staudt, Sina Zarrieß |
IJCAI | 7 |
| 2021 | Decoding, Fast and Slow: A Case Study on Balancing Trade-Offs in Incremental, Character-level Pragmatic ReasoningabstractRecent work has adopted models of pragmatic reasoning for the generation of informative language in, e.g., image captioning.We propose a simple but highly effective relaxation of fully rational decoding, based on an existing incremental and character-level approach to pragmatically informative neural image captioning.We implement a mixed, 'fast' and 'slow', speaker that applies pragmatic reasoning occasionally (only word-initially), while unrolling the language model.In our evaluation, we find that increased informativeness through pragmatic decoding generally lowers quality and, somewhat counter-intuitively, increases repetitiveness in captions.Our mixed speaker, however, achieves a good balance between quality and informativeness. Sina Zarrieß, Hendrik Buschmeier, Simeon Junker |
INLG | 1 |
| 2021 | Effects of Time Pressure and Spontaneity on Phonotactic Innovations in German Dialogues
Petra Wagner, Sina Zarrieß, Joana Cholin |
Interspeech | 2 |
| 2021 | Diversity as a By-Product: Goal-oriented Language Generation Leads to Linguistic VariationabstractThe ability for variation in language use is necessary for speakers to achieve their conversational goals, for instance when referring to objects in visual environments.We argue that diversity should not be modelled as an independent objective in dialogue, but should rather be a result or by-product of goal-oriented language generation.Different lines of work in neural language generation investigated decoding methods for generating more diverse utterances, or increasing the informativity through pragmatic reasoning.We connect those lines of work and analyze how pragmatic reasoning during decoding affects the diversity of generated image captions.We find that boosting diversity itself does not result in more pragmatically informative captions, but pragmatic reasoning does increase lexical diversity.Finally, we discuss whether the gain in informativity is achieved in linguistically plausible ways. Simeon Junker, Sina Zarrieß |
SIGDIAL | 3 |
| 2020 | Knowledge Supports Visual Language Grounding: A Case Study on Colour TermsabstractSchüz S, Zarrieß S. Knowledge Supports Visual Language Grounding: A Case Study on Colour Terms. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Linguistics; 2020: 6536-6542. Simeon Junker, Sina Zarrieß |
ACL | 2 |
| 2020 | Humans Meet Models on Object Naming: A New Dataset and AnalysisabstractWe release ManyNames v2 (MN v2), a verified version of an object naming dataset that contains dozens of valid names per object for 25K images.We analyze issues in the data collection method originally employed, standard in Language & Vision (L&V), and find that the main source of noise in the data comes from simulating a naming context solely from an image with a target object marked with a bounding box, which causes subjects to sometimes disagree regarding which object is the target.We also find that both the degree of this uncertainty in the original data and the amount of true naming variation in MN v2 differs substantially across object domains.We use MN v2 to analyze a popular L&V model and demonstrate its effectiveness on the task of object naming.However, our fine-grained analysis reveals that what appears to be human-like model behavior is not stable across domains, e.g., the model confuses people and clothing objects much more frequently than humans do.We also find that standard evaluations underestimate the actual effectiveness of the naming model: on the single-label names of the original dataset (Visual Genome), it obtains -27% accuracy points than on MN v2, that includes all valid object names. Carina Silberer, Sina Zarrieß, Matthijs Westera, Gemma Boleda |
COLING | 2 |
| 2020 | From "Before" to "After": Generating Natural Language Instructions from Image Pairs in a Simple Visual DomainabstractWhile certain types of instructions can be compactly expressed via images, there are situations where one might want to verbalise them, for example when directing someone.We investigate the task of Instruction Generation from Before/After Image Pairs which is to derive from images an instruction for effecting the implied change.For this, we make use of prior work on instruction following in a visual environment.We take an existing dataset, the BLOCKS data collected by Bisk et al. (2016) and investigate whether it is suitable for training an instruction generator as well.We find that it is, and investigate several simple baselines, taking these from the related task of image captioning.Through a series of experiments that simplify the task (by making image processing easier or completely side-stepping it; and by creating template-based targeted instructions), we investigate areas for improvement.We find that captioning models get some way towards solving the task, but have some difficulty with it, and future improvements must lie in the way the change is detected in the instruction. Robin Rojowiec, Jana Götze, Philipp Sadler, Henrik Voigt, Sina Zarrieß, David Schlangen |
INLG | 5 |
| 2020 | Object Naming in Language and Vision: A Survey and a New DatasetabstractPeople choose particular names for objects, such as dog or puppy for a given dog. Object naming has been studied in Psycholinguistics, but has received relatively little attention in Computational Linguistics. We review resources from Language and Vision that could be used to study object naming on a large scale, discuss their shortcomings, and create a new dataset that affords more opportunities for analysis and modeling. Our dataset, ManyNames, provides 36 name annotations for each of 25K objects in images selected from VisualGenome. We highlight the challenges involved and provide a preliminary analysis of the ManyNames data, showing that there is a high level of agreement in naming, on average. At the same time, the average number of name types associated with an object is much higher in our dataset than in existing corpora for Language and Vision, such that ManyNames provides a rich resource for studying phenomena like hierarchical variation (chihuahua vs. dog), which has been discussed at length in the theoretical literature, and other less well studied phenomena like cross-classification (cake vs. dessert). Carina Silberer, Sina Zarrieß, Gemma Boleda |
LREC | 2 |
| 2019 | Know What You Don't Know: Modeling a Pragmatic Speaker that Refers to Objects of Unknown CategoriesabstractZero-shot learning in Language & Vision is the task of correctly labelling (or naming) objects of novel categories.Another strand of work in L&V aims at pragmatically informative rather than "correct" object descriptions, e.g. in reference games.We combine these lines of research and model zero-shot reference games, where a speaker needs to successfully refer to a novel object in an image.Inspired by models of "rational speech acts", we extend a neural generator to become a pragmatic speaker reasoning about uncertain object categories.As a result of this reasoning, the generator produces fewer nouns and names of distractor categories as compared to a literal speaker.We show that this conversational strategy for dealing with novel objects often improves communicative success, in terms of resolution accuracy of an automatic listener. Sina Zarrieß, David Schlangen |
ACL (1) | 1 |
| 2019 | Sketch Me if You Can: Towards Generating Detailed Descriptions of Object Shape by Grounding in Images and DrawingsabstractHan T, Zarrieß S. Sketch Me if You Can: Towards Generating Detailed Descriptions of Object Shape by Grounding in Images and Drawings. In: Proceedings of the 12th International Conference on Natural Language Generation. Stroudsburg, PA, USA: Association for Computational Linguistics; 2019: 136-140. Sina Zarrieß |
INLG | 2 |
| 2019 | Tell Me More: A Dataset of Visual Scene Description SequencesabstractIlinykh N, Zarrieß S, Schlangen D. Tell Me More: A Dataset of Visual Scene Description Sequences. In: Proceedings of the 12th International Conference on Natural Language Generation. Stroudsburg, PA, USA: Association for Computational Linguistics; 2019: 152-157. Nikolai Ilinykh, Sina Zarrieß, David Schlangen |
INLG | 2 |
| 2019 | The Greennn Tree - Lengthening Position Influences Uncertainty PerceptionabstractBetz S, Zarrieß S, Székely É, Wagner P. The greennn tree - lengthening position influences uncertainty perception. In: Proceedings of Interspeech. 2019: 3990-3994. Simon Betz, Sina Zarrieß, Éva Székely, Petra Wagner |
INTERSPEECH | 2 |
| 2019 | Do Hesitations Facilitate Processing of Partially Defective System Utterances? An Exploratory Eye Tracking StudyabstractHaake K, Schimke S, Betz S, Zarrieß S. Do Hesitations Facilitate Processing of Partially Defective System Utterances? An Exploratory Eye Tracking Study. In: Proceedings of Interspeech. 2019: 1906-1910. Kristin Haake, Sarah Schimke, Simon Betz, Sina Zarrieß |
INTERSPEECH | 4 |
| 2018 | The Task Matters: Comparing Image Captioning and Task-Based Dialogical Image DescriptionabstractImage captioning models are typically trained on data that is collected from people who are asked to describe an image, without being given any further task context.As we argue here, this context independence is likely to cause problems for transferring to task settings in which image description is bound by task demands.We demonstrate that careful design of data collection is required to obtain image descriptions which are contextually bounded to a particular meta-level task.As a task, we use MeetUp!, a text-based communication game where two players have the goal of finding each other in a visual environment.To reach this goal, the players need to describe images representing their current location.We analyse a dataset from this domain and show that the nature of image descriptions found in MeetUp! is diverse, dynamic and rich with phenomena that are not present in descriptions obtained through a simple image captioning task, which we ran for comparison. Nikolai Ilinykh, Sina Zarrieß, David Schlangen |
INLG | 2 |
| 2018 | Decoding Strategies for Neural Referring Expression GenerationabstractRNN-based sequence generation is now widely used in NLP and NLG (natural language generation).Most work focusses on how to train RNNs, even though also decoding is not necessarily straightforward: previous work on neural MT found seq2seq models to radically prefer short candidates, and has proposed a number of beam search heuristics to deal with this.In this work, we assess decoding strategies for referring expression generation with neural models.Here, expression length is crucial: output should neither contain too much or too little information, in order to be pragmatically adequate.We find that most beam search heuristics developed for MT do not generalize well to referring expression generation (REG), and do not generally outperform greedy decoding.We observe that beam search heuristics for termination seem to override the model's knowledge of what a good stopping point is.Therefore, we also explore a recent approach called trainable decoding, which uses a small network to modify the RNN's hidden state for better decoding results.We find this approach to consistently outperform greedy decoding for REG. Sina Zarrieß, David Schlangen |
INLG | 1 |
| 2017 | Obtaining referential word meanings from visual and distributional information: Experiments on object namingabstractWe investigate object naming, which is an important sub-task of referring expression generation on real-world images.As opposed to mutually exclusive labels used in object recognition, object names are more flexible, subject to communicative preferences and semantically related to each other.Therefore, we investigate models of referential word meaning that link visual to lexical information which we assume to be given through distributional word embeddings.We present a model that learns individual predictors for object names that link visual and distributional aspects of word meaning during training.We show that this is particularly beneficial for zero-shot learning, as compared to projecting visual objects directly into the distributional space.In a standard object naming task, we find that different ways of combining lexical and visual information achieve very similar performance, though experiments on model combination suggest that they capture complementary aspects of referential meaning. Sina Zarrieß, David Schlangen |
ACL (1) | 1 |
| 2017 | Deriving continous grounded meaning representations from referentially structured multimodal contextsabstractCorpora of referring expressions paired with their visual referents are a good source for learning word meanings directly grounded in visual representations.Here, we explore additional ways of extracting from them word representations linked to multi-modal context: through expressions that refer to the same object, and through expressions that refer to different objects in the same scene.We show that continuous meaning representations derived from these contexts capture complementary aspects of similarity, even if not outperforming textual embeddings trained on very large amounts of raw text when tested on standard similarity benchmarks.We propose a new task for evaluating grounded meaning representations-detection of potentially co-referential phrases-and show that it requires precise denotational representations of attribute meanings, which our method provides.woman txt ref lady, girl, man, chick den lady, girl, women, blouse sit girl, guy, man, lady vis lady, girl, women, chick sidewalk txt ref pavement, ground, walkway, steps den street, sidewlak, walkway, pavement sit buildin, bldg, lamppost, street vis pavement, street, walkway, concrete grass txt Sina Zarrieß, David Schlangen |
EMNLP | 1 |
| 2017 | The Code2Text Challenge: Text Generation in Source LibrariesabstractWe propose a new shared task for tactical datato-text generation in the domain of source code libraries.Specifically, we focus on text generation of function descriptions from example software projects.Data is drawn from existing resources used for studying the related problem of semantic parser induction (Richardson and Kuhn, 2017b; Richardson and Kuhn, 2017a), and spans a wide variety of both natural languages and programming languages.In this paper, we describe these existing resources, which will serve as training and development data for the task, and discuss plans for building new independent test sets.1. Java Documentation * Returns the greater of two long values * / ... public static long max(long a, long b) 2. Python Documentation # from decimal.Context max(self, a, b): """Compares two values numerically and returns the maximum""" 3. aNALoGuE Challenge (Novikova and Rieser, 2016) MR input: name[Bibmbap House] food[French] priceRange[cheap], area[riverside] near[Clare Hall] NL output: Near Clare Hall, in the riverside area, Bibimbap serves French food in the price range cheap. Kyle Richardson 0001, Sina Zarrieß, Jonas Kuhn |
INLG | 2 |
| 2017 | Refer-iTTS: A System for Referring in Spoken Installments to Objects in Real-World ImagesabstractCurrent referring expression generation systems mostly deliver their output as one-shot, written expressions. We present on-going work on incremental generation of spoken expressions referring to objects in real-world images. This approach extends upon previous work using the words-as-classifier model for generation. We implement this generator in an incremental dialogue processing framework such that we can exploit an existing interface to incremental text-to-speech synthesis. Our system generates and synthesizes referring expressions while continuously observing non-verbal user reactions. Sina Zarrieß, Soledad López Gambino, David Schlangen |
INLG | 1 |
| 2017 | Increasing Recall of Lengthening Detection via Semi-Automatic ClassificationabstractBetz S, Voße J, Zarrieß S, Wagner P. Increasing Recall of Lengthening Detection via Semi-Automatic Classification. In: Proceedings of Interspeech. 2017: 1084-1088. Simon Betz, Jana Voße, Sina Zarrieß, Petra Wagner |
INTERSPEECH | 3 |
| 2017 | Beyond On-hold Messages: Conversational Time-buying in Task-oriented DialogueabstractA common convention in graphical user interfaces is to indicate a "wait state", for example while a program is preparing a response, through a changed cursor state or a progress bar.What should the analogue be in a spoken conversational system?To address this question, we set up an experiment in which a human information provider (IP) was given their information only in a delayed and incremental manner, which systematically created situations where the IP had the turn but could not provide task-related information.Our data analysis shows that 1) IPs bridge the gap until they can provide information by "re-purposing" a whole variety of task-and grounding-related communicative actions (e.g.echoing the user's request, signaling understanding, asserting partially relevant information), rather than being silent or explicitly asking for time (e.g."please wait"), and that 2) IPs combined these actions productively to ensure an ongoing conversation.These results, we argue, indicate that natural conversational interfaces should also be able to manage their time flexibly using a variety of conversational resources. Soledad López Gambino, Sina Zarrieß, David Schlangen |
SIGDIAL Conference | 2 |
| 2016 | Resolving References to Objects in Photographs using the Words-As-Classifiers ModelabstractA common use of language is to refer to visually present objects.Modelling it in computers requires modelling the link between language and perception.The "words as classifiers" model of grounded semantics views words as classifiers of perceptual contexts, and composes the meaning of a phrase through composition of the denotations of its component words.It was recently shown to perform well in a game-playing scenario with a small number of object types.We apply it to two large sets of real-world photographs that contain a much larger variety of object types and for which referring expressions are available.Using a pre-trained convolutional neural network to extract image region features, and augmenting these with positional information, we show that the model achieves performance competitive with the state of the art in a reference resolution task (given expression, find bounding box of its referent), while, as we argue, being conceptually simpler and more flexible. David Schlangen, Sina Zarrieß, Casey Kennington |
ACL (1) | 2 |
| 2016 | Easy Things First: Installments Improve Referring Expression Generation for Objects in PhotographsabstractResearch on generating referring expressions has so far mostly focussed on "oneshot reference", where the aim is to generate a single, discriminating expression.In interactive settings, however, it is not uncommon for reference to be established in "installments", where referring information is offered piecewise until success has been confirmed.We show that this strategy can also be advantageous in technical systems that only have uncertain access to object attributes and categories.We train a recently introduced model of grounded word meaning on a data set of REs for objects in images and learn to predict semantically appropriate expressions.In a human evaluation, we observe that users are sensitive to inadequate object names -which unfortunately are not unlikely to be generated from low-level visual input.We propose a solution inspired from human task-oriented interaction and implement strategies for avoiding and repairing semantically inaccurate words.We enhance a word-based REG with contextaware, referential installments and find that they substantially improve the referential success of the system. Sina Zarrieß, David Schlangen |
ACL (1) | 1 |
| 2016 | Towards Generating Colour Terms for Referents in Photographs: Prefer the Expected or the Unexpected?abstractColour terms have been a prime phenomenon for studying language grounding, though previous work focussed mostly on descriptions of simple objects or colour swatches.This paper investigates whether colour terms can be learned from more realistic and potentially noisy visual inputs, using a corpus of referring expressions to objects represented as regions in real-world images.We obtain promising results from combining a classifier that grounds colour terms in visual input with a recalibration model that adjusts probability distributions over colour terms according to contextual and object-specific preferences. Sina Zarrieß, David Schlangen |
INLG | 1 |
| 2016 | PentoRef: A Corpus of Spoken References in Task-oriented Dialogues
Sina Zarrieß, Julian Hough, Casey Kennington, Ramesh R. Manuvinakurike, David DeVault, Raquel Fernández, David Schlangen |
LREC | 1 |
| 2013 | Combining Referring Expression Generation and Surface Realization: A Corpus-Based Investigation of Architectures
Sina Zarrieß, Jonas Kuhn |
ACL (1) | 1 |
| 2012 | To what extent does sentence-internal realisation reflect discourse context? A study on word order
Sina Zarrieß, Aoife Cahill, Jonas Kuhn |
EACL | 1 |
| 2012 | Generating Non-Projective Word Order in Statistical Linearization
Bernd Bohnet, Anders Björkelund, Jonas Kuhn, Wolfgang Seeker, Sina Zarrieß |
EMNLP-CoNLL | 5 |
| 2012 | A Corpus-based Study of the German Recipient Passive
Patrick Ziering, Sina Zarrieß, Jonas Kuhn |
LREC | 2 |
| 2011 | Underspecifying and Predicting Voice for Surface Realisation Ranking
Sina Zarrieß, Aoife Cahill, Jonas Kuhn |
ACL | 1 |
| 2010 | Design and Development of Part-of-Speech-Tagging Resources for Wolof (Niger-Congo, spoken in Senegal)
Cheikh M. Bamba Dione, Jonas Kuhn, Sina Zarrieß |
LREC | 3 |