EDBT 2026 Demo / reviewers in the wild / expert
Vera Demberg
dblp:80/3304 · also Vera Demberg-Winterfors
· DBLP profile ↗
106ranked-venue papers
11as first author
63since 2021 · last 2026
0000-0002-8834-0020ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 100 · 10 first-author · 58 since 2021Applied, interdisciplinary, general and emerging computing · 29 · 3 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in VideoLMs for Multimodal Sarcasm Detection
Anisha Saha, Varsha Suresh, Timothy M. Hospedales, Vera Demberg |
LREC | 4 |
| 2026 | Human Label Variation in Implicit Discourse Relation Recognition
Frances Yung, Daniil Ignatev, Merel C. J. Scholman, Vera Demberg, Massimo Poesio |
LREC | 4 |
| 2025 | What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific PresentationsabstractTransforming recorded videos into concise and accurate textual summaries is a growing challenge in multimodal learning. This paper introduces VISTA, a dataset specifically designed for video-to-text summarization in scientific domains. VISTA contains 18,599 recorded AI conference presentations paired with their corresponding paper abstracts. We benchmark the performance of state-of-the-art large models and apply a plan-based framework to better capture the structured nature of abstracts. Both human and automated evaluations confirm that explicit planning enhances summary quality and factual consistency. However, a considerable gap remains between models and human performance, highlighting the challenges of our dataset. This study aims to pave the way for future research on scientific video-to-text summarization. Dongqi Liu 0001, Chenxi Whitehouse, Louis Mahon, Rohit Saxena, Zheng Zhao 0005, Yifu Qiu, Mirella Lapata, Vera Demberg |
ACL (1) | 9 |
| 2025 | Multimodal Pragmatic Jailbreak on Text-to-image ModelsabstractTong Liu, Zhixin Lai, Jiawen Wang, Gengyuan Zhang, Shuo Chen, Philip Torr, Vera Demberg, Volker Tresp, Jindong Gu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Tong Liu 0019, Zhixin Lai, Gengyuan Zhang, Shuo Chen 0014, Philip Torr 0001, Vera Demberg, Volker Tresp, Jindong Gu |
ACL (1) | 7 |
| 2025 | Enhancing Spoken Discourse Modeling in Language Models Using Gestural CuesabstractResearch in linguistics shows that non-verbal cues, such as gestures, play a crucial role in spoken discourse.For example, speakers perform hand gestures to indicate topic shifts, helping listeners identify transitions in discourse.In this work, we investigate whether the joint modeling of gestures using human motion sequences and language can improve spoken discourse modeling in language models.To integrate gestures into language models, we first encode 3D human motion sequences into discrete gesture tokens using a VQ-VAE.These gesture token embeddings are then aligned with text embeddings through feature alignment, mapping them into the text embedding space.To evaluate the gesture-aligned language model on spoken discourse, we construct text infilling tasks targeting three key discourse cues grounded in linguistic research: discourse connectives, stance markers, and quantifiers.Results show that incorporating gestures enhances marker prediction accuracy across the three tasks, highlighting the complementary information that gestures can offer in modeling spoken discourse.We view this work as an initial step toward leveraging non-verbal cues to advance spoken language modeling in language models. Varsha Suresh, Muhammad Hamza Mughal, Christian Theobalt, Vera Demberg |
ACL (1) | 4 |
| 2025 | An ACT-R model of resource-rational performance in a pragmatic reference game
John Duff, Alexandra Mayn, Vera Demberg |
CogSci | 3 |
| 2025 | People do not engage in ad-hoc reasoning about alternative messages when interacting with a literal speaker
Alexandra Mayn, John Duff, Natalia Bila, Vera Demberg |
CogSci | 4 |
| 2025 | Music-induced Positive Mood Stimulates Metaphor Production
Laura Pissani, Magdalena-Victoria Meiser, Vera Demberg |
CogSci | 3 |
| 2025 | Implicit Discourse Relation Classification For Nigerian PidginabstractNigerian Pidgin (NP) is an English-based creole language spoken by nearly 100 million people across Nigeria, and is still low-resource in NLP. In particular, there are currently no available discourse parsing tools, which, if available, would have the potential to improve various downstream tasks. Our research focuses on implicit discourse relation classification (IDRC) for NP, a task which, even in English, is not easily solved by prompting LLMs, but requires supervised training. % With this in mind, we have developed a framework for the task, which could also be used by researchers for other English-lexified languages. We systematically compare different approaches to the low resource IDRC task: in one approach, we use English IDRC tools directly on the NP text as well as on their English translations (followed by a back-projection of labels). In another approach, we create a synthetic discourse corpus for NP, in which we automatically translate the English discourse-annotated corpus PDTB to NP, project PDTB labels, and then train an NP IDR classifier. The latter approach of training a “native” NP classifier outperforms our baseline by 13.27% and 33.98% in f_{1} score for 4-way and 11-way classification, respectively. Muhammed Saeed, Peter Bourgonje, Vera Demberg |
COLING | 3 |
| 2025 | Retrieving Semantics from the Deep: an RAG Solution for Gesture SynthesisabstractNon-Verbal communication often comprises of semantically rich gestures that help convey the meaning of an utterance. Producing such semantic co-speech gestures has been a major challenge for the existing neural systems that can generate rhythmic beat gestures, but struggle to produce semantically meaningful gestures. Therefore, we present RAG-GESTURE, a diffusion-based gesture generation approach that leverages Retrieval Augmented Generation (RAG) to produce natural-looking and semantically rich gestures. Our neuro-explicit gesture generation approach is designed to produce semantic gestures grounded in interpretable linguistic knowledge. We achieve this by using explicit domain knowledge to retrieve exemplar motions from a database of co-speech gestures. Once retrieved, we then inject these semantic exemplar gestures into our diffusion-based gesture generation pipeline using DDIM inversion and retrieval guidance at the inference time without any need of training. Further, we propose a control paradigm for guidance, that allows the users to modulate the amount of influence each retrieval insertion has over the generated sequence. Our comparative evaluations demonstrate the validity of our approach against recent gesture generation approaches. The reader is urged to explore the results on our project page. Muhammad Hamza Mughal, Rishabh Dabral, Merel C. J. Scholman, Vera Demberg, Christian Theobalt |
CVPR | 4 |
| 2025 | MultiplEYE: Creating a multilingual eye-tracking-while-reading corpusabstractContains fulltext : 326363.pdf (Publisher’s version ) (Open Access) Deborah N. Jakobi, Maja Stegenwallner-Schütz, Nora Hollenstein, Cui Ding, Ramune Kaspere, Ana Matic Skoric, Eva Pavlinusic Vilus, Stefan Frank, Marie-Luise Müller, Kristine M. Jensen de López, Nik Kharlamov, Hanne B. Søndergaard Knudsen, Yevgeni Berzak, Ella Lion, Irina A. Sekerina, Cengiz Acartürk, Mohd Faizan Ansari, Katarzyna Harezlak, Pawel Kasprowski, Ana Bautista, Lisa Beinborn, Anna Bondar, Antonia Boznou, Leah Bradshaw, Jana Mara Hofmann, Thyra Krosness, Not Battesta Soliva, Anila Çepani, Kristina Cergol, Ana Dosen, Marijan Palmovic, Adelina Çerpja, Dalí Chirino, Jan Chromý, Vera Demberg, Iza Skrjanec, Nazik Dinçtopal Deniz, Inmaculada Fajardo, Mariola Giménez-Salvador, Xavier Mínguez-López, Maros Filip, Zigmunds Freibergs, Jessica Gomes, Andreia Janeiro, Paula Luegi, João Veríssimo, Sasho Gramatikov, Jana Hasenäcker, Alba Haveriku, Nelda Kote, Muhammad Mohsin Kamal, Hanna Kedzierska, Dorota Klimek-Jankowska, Sara Kosutar, Daniel Krakowczyk, Izabela Krejtz, Marta Lockiewicz, Kaidi Lõo, Jurgita Motiejuniene, Jamal Abdul Nasir, Johanne Sofie Krog Nedergård, Aysegül Özkan, Mikulás Preininger, Loredana Punga, David R. Reich, Chiara Tschirner, Spela Rot, Andreas Säuberli, Jordi Solé i Casals, Ekaterina Strati, Igor Svoboda, Evis Trandafili, Spyridoula Varlokosta, Mila Dimitrova-Vulchanova, Lena A. Jäger |
ETRA | 35 |
| 2025 | Born a Transformer - Always a Transformer? On the Effect of Pretraining on Architectural AbilitiesabstractTransformers have theoretical limitations in modeling certain sequence-to-sequence tasks, yet it remains largely unclear if these limitations play a role in large-scale pretrained LLMs, or whether LLMs might effectively overcome these constraints in practice due to the scale of both the models themselves and their pretraining data. We explore how these architectural constraints manifest after pretraining by studying a family of *retrieval* and *copying* tasks inspired by Liu et al. [2024a]. We use a recently proposed framework for studying length generalization [Huang et al., 2025] to provide guarantees for each of our settings. Empirically, we observe an *induction-versus-anti-induction asymmetry*, where pretrained models are better at retrieving tokens to the right (induction) rather than the left (anti-induction) of a query token. This asymmetry disappears upon targeted fine-tuning if length-generalization is guaranteed by theory. Mechanistic analysis reveals that this asymmetry is connected to the differences in the strength of induction versus anti-induction circuits within pretrained transformers. We validate our findings through practical experiments on real-world tasks demonstrating reliability risks. Our results highlight that pretraining selectively enhances certain transformer capabilities, but does not overcome fundamental length-generalization limits. Mayank Jobanputra, Yana Veitsman, Yash Raj Sarrof, Aleksandra Bakalova, Vera Demberg, Ellie Pavlick, Michael Hahn 0001 |
NeurIPS | 5 |
| 2025 | Synthetic Data Augmentation for Cross-domain Implicit Discourse Relation RecognitionabstractImplicit discourse relation recognition (IDRR) – the task of identifying the implicit coherence relation between two text spans – requires deep semantic understanding. Recent studies have shown that zero-/few-shot approaches significantly lag behind supervised models. However, LLMs may be useful for synthetic data augmentation, where LLMs generate a second argument following a specified coherence relation. We applied this approach in a cross-domain setting, generating discourse continuations using unlabelled target-domain data to adapt a base model which was trained on source-domain labelled data. Evaluations conducted on a large-scale test set revealed that different variations of the approach did not result in any significant improvements. We conclude that LLMs often fail to generate useful samples for IDRR, and emphasize the importance of considering both statistical significance and comparability when evaluating IDRR models. Frances Yung, Varsha Suresh, Zaynab Reza, Mansoor Ahmad, Vera Demberg |
SIGDIAL | 5 |
| 2025 | GestureCoach: Rehearsing for Engaging Talks with LLM-Driven Gesture RecommendationsabstractRehearse 3/10Why are we here, why are we alive?One key reason is that our ancestors on the savannas of Africa were really good at one thing…They weren't bigger than the animals they took down a lot of the time, they weren't faster than the animals they took down a lot of the time, but they were much better at banding together into groups and cooperating.…One key reason… Proactive Gesture Cues Current Slide Hover to Preview and Modify Gestures Presenter Notes with Gesture highlights Edit DeleteWhy were humans at hunting ?• Not bigger or faster than animals • Good at cooperationFigure 1: GestureCoach guides speakers to perform gestures while rehearsing their talk.A gesture recommendation model predicts text segments in the presenter notes that should be emphasized with gestures and retrieves relevant semantic gestures for each segment.During rehearsal, the system highlights the segments and proactively cues a video clip of the gesture by tracking users' speech, allowing them to integrate the gesture smoothly into their talk.Hovering over a segment allows users to preview and modify the associated gesture with alternate suggestions from the model. Ashwin Ram 0004, Varsha Suresh, Artin Saberpour, Vera Demberg, Jürgen Steimle |
UIST | 4 |
| 2025 | Explanatory Summarization with Discourse-Driven PlanningabstractAbstract Lay summaries for scientific documents typically include explanations to help readers grasp sophisticated concepts or arguments. However, current automatic summarization methods do not explicitly model explanations, which makes it difficult to align the proportion of explanatory content with human-written summaries. In this paper, we present a plan-based approach that leverages discourse frameworks to organize summary generation and guide explanatory sentences by prompting responses to the plan. Specifically, we propose two discourse-driven planning strategies, where the plan is conditioned as part of the input or part of the output prefix, respectively. Empirical experiments on three lay summarization datasets show that our approach outperforms existing state-of-the-art methods in terms of summary quality, and it enhances model robustness, controllability, and mitigates hallucination. The project information is available at https://dongqi.me/projects/ExpSum. Dongqi Liu 0001, Vera Demberg, Mirella Lapata |
Trans. Assoc. Comput. Linguistics | 3 |
| 2024 | Temperature-scaling surprisal estimates improve fit to human reading times - but does it do so for the "right reasons"?abstractA wide body of evidence shows that human language processing difficulty is predicted by the information-theoretic measure surprisal, a word's negative log probability in context.However, it is still unclear how to best estimate these probabilities needed for predicting human processing difficulty -while a long-standing belief held that models with lower perplexity would provide more accurate estimates of word predictability, and therefore lead to better reading time predictions, recent work has shown that for very large models, psycholinguistic predictive power decreases.One reason could be that language models might be more confident of their predictions than humans, because they have had exposure to several magnitudes more data.In this paper, we test what effect temperature-scaling of large language model (LLM) predictions has on surprisal estimates and their predictive power of reading times of English texts.Firstly, we show that calibration of large language models typically improves with model size, i.e. poorer calibration cannot account for poorer fit to reading times.Secondly, we find that temperature-scaling probabilities lead to a systematically better fit to reading times (up to 89% improvement in delta log likelihood), across several reading time corpora.Finally, we show that this improvement in fit is chiefly driven by words that are composed of multiple subword tokens. 1 Tong Liu 0019, Iza Skrjanec, Vera Demberg |
ACL (1) | 3 |
| 2024 | Interpreting implausible event descriptions under noise
Asya Achimova, Marjolein van Os, Vera Demberg, Martin V. Butz |
CogSci | 3 |
| 2024 | Uniform information density explains subject doubling in French
Yiming Liang, Pascal Amsili, Heather Burnett, Vera Demberg |
CogSci | 4 |
| 2024 | Adaptation to Speakers is modulated by working memory updating and theory of mind - a study investigating humor comprehension
Jia E. Loy, Vera Demberg |
CogSci | 2 |
| 2024 | What processing instructions do connectives provide? Modeling the facilitative effect of the connective
Marian Marchal, Merel C. J. Scholman, Ted Sanders, Vera Demberg |
CogSci | 4 |
| 2024 | Capable but not cooperative? Perceptions of ChatGPT as a pragmatic speaker
Alexandra Mayn, Jia E. Loy, Vera Demberg |
CogSci | 3 |
| 2024 | Retrieval-Augmented Modular Prompt Tuning for Low-Resource Data-to-Text GenerationabstractData-to-text (D2T) generation describes the task of verbalizing data, often given as attribute-value pairs. While this task is relevant for many different data domains beyond the traditionally well-explored tasks of weather forecasting, restaurant recommendations, and sports reporting, a major challenge to the applicability of data-to-text generation methods is typically data sparsity. For many applications, there is extremely little training data in terms of attribute-value inputs and target language outputs available for training a model. Given the sparse data setting, recently developed prompting methods seem most suitable for addressing D2T tasks since they do not require substantial amounts of training data, unlike finetuning approaches. However, prompt-based approaches are also challenging, as a) the design and search of prompts are non-trivial; and b) hallucination problems may occur because of the strong inductive bias of these models. In this paper, we propose a retrieval-augmented modular prompt tuning () method, which constructs prompts that fit the input data closely, thereby bridging the domain gap between the large-scale language model and the structured input data. Experiments show that our method generates texts with few hallucinations and achieves state-of-the-art performance on a dataset for drone handover message generation. Xudong Hong 0002, Mayank Jobanputra, Mattes Warning, Vera Demberg |
LREC/COLING | 5 |
| 2024 | Modeling Orthographic Variation Improves NLP Performance for Nigerian PidginabstractNigerian Pidgin is an English-derived contact language and is traditionally an oral language, spoken by approximately 100 million people. No orthographic standard has yet been adopted, and thus the few available Pidgin datasets that exist are characterised by noise in the form of orthographic variations. This contributes to under-performance of models in critical NLP tasks. The current work is the first to describe various types of orthographic variations commonly found in Nigerian Pidgin texts, and model this orthographic variation. The variations identified in the dataset form the basis of a phonetic-theoretic framework for word editing, which is used to generate orthographic variations to augment training data. We test the effect of this data augmentation on two critical NLP tasks: machine translation and sentiment analysis. The proposed variation generation framework augments the training data with new orthographic variants which are relevant for the test set but did not occur in the training set originally. Our results demonstrate the positive effect of augmenting the training data with a combination of real texts from other corpora as well as synthesized orthographic variation, resulting in performance improvements of 2.1 points in sentiment analysis and 1.4 BLEU points in translation to English. Pin-Jie Lin, Merel C. J. Scholman, Muhammed Saeed, Vera Demberg |
LREC/COLING | 4 |
| 2024 | SIGA: A Naturalistic NLI Dataset of English Scalar Implicatures with Gradable AdjectivesabstractMany utterances convey meanings that go beyond the literal meaning of a sentence. One class of such meanings is scalar implicatures, a phenomenon by which a speaker conveys the negation of a more informative utterance by producing a less informative utterance. This paper introduces a Natural Language Inference (NLI) dataset designed to investigate the ability of language models to interpret utterances with scalar implicatures. Our dataset is comprised of text extracted from the C4 English text corpus and annotated with both crowd-sourced and expert annotations. We evaluate NLI models based on DeBERTa to investigate 1) whether NLI models can learn to predict pragmatic inferences involving gradable adjectives and 2) whether models generalize to utterances involving unseen adjectives. We find that fine-tuning NLI models on our dataset significantly improves their performance to derive scalar implicatures, both for in-domain and for out-of domain examples. At the same time, we find that the investigated models still perform considerably worse on examples with scalar implicatures than on other types of NLI examples, highlighting that pragmatic inferences still pose challenges for current models. Rashid Nizamani, Sebastian Schuster 0001, Vera Demberg |
LREC/COLING | 3 |
| 2024 | SciNews: From Scholarly Complexities to Public Narratives - a Dataset for Scientific News Report GenerationabstractScientific news reports serve as a bridge, adeptly translating complex research articles into reports that resonate with the broader public. The automated generation of such narratives enhances the accessibility of scholarly insights. In this paper, we present a new corpus to facilitate this paradigm development. Our corpus comprises a parallel compilation of academic publications and their corresponding scientific news reports across nine disciplines. To demonstrate the utility and reliability of our dataset, we conduct an extensive analysis, highlighting the divergences in readability and brevity between scientific news narratives and academic manuscripts. We benchmark our dataset employing state-of-the-art text generation models. The evaluation process involves both automatic and human evaluation, which lays the groundwork for future explorations into the automated generation of scientific news reports. The dataset and code related to this work are available at https://dongqi.me/projects/SciNews. Dongqi Liu 0001, Yifan Wang 0019, Jia E. Loy, Vera Demberg |
LREC/COLING | 4 |
| 2024 | SpreadNaLa: A Naturalistic Code Generation Evaluation Dataset of Spreadsheet FormulasabstractAutomatic generation of code from natural language descriptions has emerged as one of the main use cases of large language models (LLMs). This has also led to a proliferation of datasets to track progress in the reliability of code generation models, including domains such as programming challenges and common data science tasks. However, existing datasets primarily target the use of code generation models to aid expert programmers in writing code. In this work, we consider a domain of code generation which is more frequently used by users without sophisticated programming skills: translating English descriptions to spreadsheet formulas that can be used to do everyday data processing tasks. We extract naturalistic instructions from StackOverflow posts and manually verify and standardize the corresponding spreadsheet formulas. We use this dataset to evaluate an off-the-shelf code generation model (GPT 3.5 text-davinci-003) as well as recently proposed pragmatic code generation procedures and find that Code Reviewer reranking (Zhang et al., 2022) performs best among the evaluated methods but still frequently generates formulas that differ from human-generated ones. Sebastian Schuster 0001, Ayesha Ansar, Om Agarwal, Vera Demberg |
LREC/COLING | 4 |
| 2024 | DiscoGeM 2.0: A Parallel Corpus of English, German, French and Czech Implicit Discourse RelationsabstractWe present DiscoGeM 2.0, a crowdsourced, parallel corpus of 12,834 implicit discourse relations, with English, German, French and Czech data. We propose and validate a new single-step crowdsourcing annotation method and apply it to collect new annotations in German, French and Czech. The corpus was constructed by having crowdsourced annotators choose a suitable discourse connective for each relation from a set of unambiguous candidates. Every instance was annotated by 10 workers. Our corpus hence represents the first multi-lingual resource that contains distributions of discourse interpretations for implicit relations. The results show that the connective insertion method of discourse annotation can be reliably extended to other languages. The resulting multi-lingual annotations also reveal that implicit relations inferred in one language may differ from those inferred in the translation, meaning the annotations are not always directly transferable. DiscoGem 2.0 promotes the investigation of cross-linguistic differences in discourse marking and could improve automatic discourse parsing applications. It is openly downloadable here: https://github.com/merelscholman/DiscoGeM. Frances Yung, Merel C. J. Scholman, Sárka Zikánová, Vera Demberg |
LREC/COLING | 4 |
| 2024 | RSA-Control: A Pragmatics-Grounded Lightweight Controllable Text Generation FrameworkabstractDespite significant advancements in natural language generation, controlling language models to produce texts with desired attributes remains a formidable challenge.In this work, we introduce RSA-Control, a training-free controllable text generation framework grounded in pragmatics.RSA-Control directs the generation process by recursively reasoning between imaginary speakers and listeners, enhancing the likelihood that target attributes are correctly interpreted by listeners amidst distractors.Additionally, we introduce a self-adjustable rationality parameter, which allows for automatic adjustment of control strength based on context.Our experiments, conducted with two task types and two types of language models, demonstrate that RSA-Control achieves strong attribute control while maintaining language fluency and content consistency.Our code is available at https://github.com/Ewanwong/RSA-Control. Yifan Wang 0019, Vera Demberg |
EMNLP | 2 |
| 2024 | RST-LoRA: A Discourse-Aware Low-Rank Adaptation for Long Document Abstractive SummarizationabstractFor long document summarization, discourse structure is important to discern the key content of the text and the differences in importance level between sentences.Unfortunately, the integration of rhetorical structure theory (RST) into parameter-efficient fine-tuning strategies for long document summarization remains unexplored.Therefore, this paper introduces RST-LoRA and proposes four RST-aware variants to explicitly incorporate RST into the LoRA model.Our empirical evaluation demonstrates that incorporating the type and uncertainty of rhetorical relations can complementarily enhance the performance of LoRA in summarization tasks.Furthermore, the best-performing variant we introduced outperforms the vanilla LoRA and full-parameter fine-tuning models, as confirmed by multiple automatic and human evaluations, and even surpasses previous stateof-the-art methods 1 . Dongqi Liu 0001, Vera Demberg |
NAACL-HLT | 2 |
| 2024 | Generalizing across Languages and Domains for Discourse Relation ClassificationabstractThe availability of corpora annotated for discourse relations is limited and discourse relation classification performance varies greatly depending on both language and domain.This is a problem for downstream applications that are intended for a language (i.e., not English) or a domain (i.e., not financial news) with comparatively low coverage for discourse annotations.In this paper, we experiment with a state-of-theart model for discourse relation classification, originally developed for English, extend it to a multi-lingual setting (testing on Italian, Portuguese and Turkish), and employ a simple, yet effective method to mark out-of-domain training instances.By doing so, we aim to contribute to better generalization and more robust discourse relation classification performance across both language and domain. Peter Bourgonje, Vera Demberg |
SIGDIAL | 2 |
| 2023 | Incorporating Distributions of Discourse Structure for Long Document Abstractive SummarizationabstractFor text summarization, the role of discourse structure is pivotal in discerning the core content of a text.Regrettably, prior studies on incorporating Rhetorical Structure Theory (RST) into transformer-based summarization models only consider the nuclearity annotation, thereby overlooking the variety of discourse relation types.This paper introduces the 'RSTformer', a novel summarization model that comprehensively incorporates both the types and uncertainty of rhetorical relations.Our RST-attention mechanism, rooted in document-level rhetorical structure, is an extension of the recently devised Longformer framework.Through rigorous evaluation, the model proposed herein exhibits significant superiority over state-of-theart models, as evidenced by its notable performance on several automatic metrics and human evaluation. 1 Dongqi Liu 0001, Yifan Wang 0019, Vera Demberg |
ACL (1) | 3 |
| 2023 | What inferences do people actually make upon encountering informationally redundant utterances? An individual differences study
Margarita Ryzhova, Alexandra Mayn, Vera Demberg |
CogSci | 3 |
| 2023 | Working memory updating modulates adaptation to speaker-specific use of uncertainty expressions
Sebastian Schuster 0001, Alexandra Mayn, Vera Demberg |
CogSci | 3 |
| 2023 | Word Familiarity Classification From a Single Trial Based on Eye-Movements. A Study in German and EnglishabstractIdentifying processing difficulty during reading due to unfamiliar words has promising applications in automatic text adaptation. We present a classification model that predicts whether a word is (un)known to the reader based on eye-movement measures. We examine German and English data and validate our model on unseen subjects and items achieving a high accuracy in both languages. Margarita Ryzhova, Iza Skrjanec, Nina Quach, Alice Virginia Chase, Emilia Ellsiepen, Vera Demberg |
ETRA | 6 |
| 2023 | Tackling Hallucinations in Neural Chart SummarizationabstractHallucinations in text generation occur when the system produces text that is not grounded in the input.In this work, we tackle the problem of hallucinations in neural chart summarization.Our analysis shows that the target side of chart summarization training datasets often contains additional information, leading to hallucinations.We propose a natural language inference (NLI) based method to preprocess the training data and show through human evaluation that our method significantly reduces hallucinations.We also found that shortening long-distance dependencies in the input sequence and adding chart-related information like title and legends improves the overall performance. Saad Obaid ul Islam, Iza Skrjanec, Ondrej Dusek, Vera Demberg |
INLG | 4 |
| 2023 | Expert-adapted language models improve the fit to reading timesabstractThe concept of surprisal refers to the predictability of a word based on its context. Surprisal is known to be predictive of human processing difficulty and is usually estimated by language models. However, because humans differ in their linguistic experience, they also differ in the actual processing difficulty they experience with a given word or sentence. We investigate whether models that are similar to the linguistic experience and background knowledge of a specific group of humans are better at predicting their reading times than a generic language model. We analyze reading times from the PoTeC corpus [15,27] of eye movements from biology and physics experts reading biology and physics texts. We find experts read in-domain texts faster than novices, especially domain-specific terms. Next, we train language models adapted to the biology and physics domains and show that surprisal obtained from these specialized models improves the fit to expert reading times above and beyond a generic language model. Iza Skrjanec, Frederik Yannick Broy, Vera Demberg |
KES | 3 |
| 2023 | Investigating Explicitation of Discourse Connectives in Translation using Automatic AnnotationsabstractDiscourse relations have different patterns of marking across different languages.As a result, discourse connectives are often added, omitted, or rephrased in translation.Prior work has shown a tendency for explicitation of discourse connectives, but such work was conducted using restricted sample sizes due to difficulty of connective identification and alignment.The current study exploits automatic methods to facilitate a large-scale study of connectives in English and German parallel texts.Our results based on over 300 types and 18000 instances of aligned connectives and an empirical approach to compare the cross-lingual specificity gap provide strong evidence of the Explicitation Hypothesis.We conclude that discourse relations are indeed more explicit in translation than texts written originally in the same language.Automatic annotations allow us to carry out translation studies of discourse relations on a large scale.Our methodology using relative entropy to study the specificity of connectives also provides more fine-grained insights into translation patterns. Frances Yung, Merel C. J. Scholman, Ekaterina Lapshinova-Koltunski, Christina Pollkläsener, Vera Demberg |
SIGDIAL | 5 |
| 2023 | Visual Writing Prompts: Character-Grounded Story Generation with Curated Image SequencesabstractAbstract Current work on image-based story generation suffers from the fact that the existing image sequence collections do not have coherent plots behind them. We improve visual story generation by producing a new image-grounded dataset, Visual Writing Prompts (VWP). VWP contains almost 2K selected sequences of movie shots, each including 5-10 images. The image sequences are aligned with a total of 12K stories which were collected via crowdsourcing given the image sequences and a set of grounded characters from the corresponding image sequence. Our new image sequence collection and filtering process has allowed us to obtain stories that are more coherent, diverse, and visually grounded compared to previous work. We also propose a character-based story generation model driven by coherence as a strong baseline. Evaluations show that our generated stories are more coherent, visually grounded, and diverse than stories generated with the current state-of-the-art model. Our code, image features, annotations and collected stories are available at https://vwprompt.github.io/. Xudong Hong 0002, Asad B. Sayeed, Khushboo Mehra, Vera Demberg, Bernt Schiele |
Trans. Assoc. Comput. Linguistics | 4 |
| 2023 | Design Choices for Crowdsourcing Implicit Discourse Relations: Revealing the Biases Introduced by Task DesignabstractAbstract Disagreement in natural language annotation has mostly been studied from a perspective of biases introduced by the annotators and the annotation frameworks. Here, we propose to analyze another source of bias—task design bias, which has a particularly strong impact on crowdsourced linguistic annotations where natural language is used to elicit the interpretation of lay annotators. For this purpose we look at implicit discourse relation annotation, a task that has repeatedly been shown to be difficult due to the relations’ ambiguity. We compare the annotations of 1,200 discourse relations obtained using two distinct annotation tasks and quantify the biases of both methods across four different domains. Both methods are natural language annotation tasks designed for crowdsourcing. We show that the task design can push annotators towards certain relations and that some discourse relation senses can be better elicited with one or the other annotation approach. We also conclude that this type of bias should be taken into account when training and testing models. Valentina Pyatkin, Frances Yung, Merel C. J. Scholman, Reut Tsarfaty, Ido Dagan, Vera Demberg |
Trans. Assoc. Comput. Linguistics | 6 |
| 2022 | Modeling atypicality inferences in pragmatic reasoning
Ekaterina Kravtchenko, Vera Demberg |
CogSci | 2 |
| 2022 | Partner effects and individual differences on perspective taking
Jia E. Loy, Vera Demberg |
CogSci | 2 |
| 2022 | Individual Differences in a Pragmatic Reference Game
Alexandra Mayn, Vera Demberg |
CogSci | 2 |
| 2022 | Pragmatics of Metaphor Revisited: Modeling the Role of Degree and Salience in Metaphor Understanding
Alexandra Mayn, Vera Demberg |
CogSci | 2 |
| 2022 | Pragmatic comprehension of implicatures - consistency within individuals across types and time
Margarita Ryzhova, Jia E. Loy, Vera Demberg |
CogSci | 3 |
| 2022 | Few-Shot Pidgin Text Adaptation via Contrastive Fine-TuningabstractThe surging demand for multilingual dialogue systems often requires a costly labeling process for each language addition. For low resource languages, human annotators are continuously tasked with the adaptation of resource-rich language utterances for each new domain. However, this prohibitive and impractical process can often be a bottleneck for low resource languages that are still without proper translation systems nor parallel corpus. In particular, it is difficult to obtain task-specific low resource language annotations for the English-derived creoles (e.g. Nigerian and Cameroonian Pidgin). To address this issue, we utilize the pretrained language models i.e. BART which has shown great potential in language generation/understanding – we propose to finetune the BART model to generate utterances in Pidgin by leveraging the proximity of the source and target languages, and utilizing positive and negative examples in constrastive training objectives. We collected and released the first parallel Pidgin-English conversation corpus in two dialogue domains and showed that this simple and effective technique is suffice to yield impressive results for English-to-Pidgin generation, which are two closely-related languages. Ernie Chang, Jesujoba O. Alabi, David Ifeoluwa Adelani, Vera Demberg |
COLING | 4 |
| 2022 | Programmable Annotation with Diversed Heuristics and Data DenoisingabstractNeural natural language generation (NLG) and understanding (NLU) models are costly and require massive amounts of annotated data to be competitive. Recent data programming frameworks address this bottleneck by allowing human supervision to be provided as a set of labeling functions to construct generative models that synthesize weak labels at scale. However, these labeling functions are difficult to build from scratch for NLG/NLU models, as they often require complex rule sets to be specified. To this end, we propose a novel data programming framework that can jointly construct labeled data for language generation and understanding tasks – by allowing the annotators to modify an automatically-inferred alignment rule set between sequence labels and text, instead of writing rules from scratch. Further, to mitigate the effect of poor quality labels, we propose a dually-regularized denoising mechanism for optimizing the NLU and NLG models. On two benchmarks we show that the framework can generate high-quality data that comes within a 1.48 BLEU and 6.42 slot F1 of the 100% human-labeled data (42k instances) with just 100 labeled data samples – outperforming benchmark annotation frameworks and other semi-supervised approaches. Ernie Chang, Alex Marin, Vera Demberg |
COLING | 3 |
| 2022 | Improving Zero-Shot Multilingual Text Generation via Iterative DistillationabstractThe demand for multilingual dialogue systems often requires a costly labeling process, where human translators derive utterances in low resource languages from resource rich language annotation. To this end, we explore leveraging the inductive biases for target languages learned by numerous pretrained teacher models by transferring them to student models via sequence-level knowledge distillation. By assuming no target language text, the both the teacher and student models need to learn from the target distribution in a few/zero-shot manner. On the MultiATIS++ benchmark, we explore the effectiveness of our proposed technique to derive the multilingual text for 6 languages, using only the monolingual English data and the pretrained models. We show that training on the synthetic multilingual generation outputs yields close performance to training on human annotations in both slot F1 and intent accuracy; the synthetic text also scores high in naturalness and correctness based on human evaluation. Ernie Chang, Alex Marin, Vera Demberg |
COLING | 3 |
| 2022 | Establishing Annotation Quality in Multi-label AnnotationsabstractIn many linguistic fields requiring annotated data, multiple interpretations of a single item are possible. Multi-label annotations more accurately reflect this possibility. However, allowing for multi-label annotations also affects the chance that two coders agree with each other. Calculating inter-coder agreement for multi-label datasets is therefore not trivial. In the current contribution, we evaluate different metrics for calculating agreement on multi-label annotations: agreement on the intersection of annotated labels, an augmented version of Cohen’s Kappa, and precision, recall and F1. We propose a bootstrapping method to obtain chance agreement for each measure, which allows us to obtain an adjusted agreement coefficient that is more interpretable. We demonstrate how various measures affect estimates of agreement on simulated datasets and present a case study of discourse relation annotations. We also show how the proportion of double labels, and the entropy of the label distribution, influences the measures outlined above and how a bootstrapped adjusted agreement can make agreement measures more comparable across datasets in multi-label scenarios. Marian Marchal, Merel C. J. Scholman, Frances Yung, Vera Demberg |
COLING | 4 |
| 2022 | Zero-shot Script ParsingabstractScript knowledge is useful to a variety of NLP tasks. However, existing resources only cover a small number of activities, limiting its practical usefulness. In this work, we propose a zero-shot learning approach to script parsing, the task of tagging texts with scenario-specific event and participant types, which enables us to acquire script knowledge without domain-specific annotations. We (1) learn representations of potential event and participant mentions by promoting cluster consistency according to the annotated data; (2) perform clustering on the event / participant candidates from unannotated texts that belongs to an unseen scenario. The model achieves 68.1/74.4 average F1 for event / participant parsing, respectively, outperforming a previous CRF model that, in contrast, has access to scenario-specific supervision. We also evaluate the model by testing on a different corpus, where it achieved 55.5/54.0 average F1 for event / participant parsing. Fangzhou Zhai, Vera Demberg, Alexander Koller |
COLING | 2 |
| 2022 | Logic-Guided Message Generation from Raw Real-Time Sensor DataabstractNatural language generation in real-time settings with raw sensor data is a challenging task. We find that formulating the task as an end-to-end problem leads to two major challenges in content selection – the sensor data is both redundant and diverse across environments, thereby making it hard for the encoders to select and reason on the data. We here present a new corpus for a specific domain that instantiates these properties. It includes handover utterances that an assistant for a semi-autonomous drone uses to communicate with humans during the drone flight. The corpus consists of sensor data records and utterances in 8 different environments. As a structured intermediary representation between data records and text, we explore the use of description logic (DL). We also propose a neural generation model that can alert the human pilot of the system state and environment in preparation of the handover of control. Ernie Chang, Alisa Kovtunova, Stefan Borgwardt, Vera Demberg, Kathryn Chapman, Hui-Syuan Yeh |
LREC | 4 |
| 2022 | DiscoGeM: A Crowdsourced Corpus of Genre-Mixed Implicit Discourse RelationsabstractWe present DiscoGeM, a crowdsourced corpus of 6,505 implicit discourse relations from three genres: political speech, literature, and encyclopedic texts. Each instance was annotated by 10 crowd workers. Various label aggregation methods were explored to evaluate how to obtain a label that best captures the meaning inferred by the crowd annotators. The results show that a significant proportion of discourse relations in DiscoGeM are ambiguous and can express multiple relation senses. Probability distribution labels better capture these interpretations than single labels. Further, the results emphasize that text genre crucially affects the distribution of discourse relations, suggesting that genre should be included as a factor in automatic relation classification. We make available the newly created DiscoGeM corpus, as well as the dataset with all annotator-level labels. Both the corpus and the dataset can facilitate a multitude of applications and research purposes, for example to function as training data to improve the performance of automatic discourse relation parsers, as well as facilitate research into non-connective signals of discourse relations. Merel C. J. Scholman, Tianai Dong, Frances Yung, Vera Demberg |
LREC | 4 |
| 2022 | Design Choices in Crowdsourcing Discourse Relation Annotations: The Effect of Worker Selection and TrainingabstractObtaining linguistic annotation from novice crowdworkers is far from trivial. A case in point is the annotation of discourse relations, which is a complicated task. Recent methods have obtained promising results by extracting relation labels from either discourse connectives (DCs) or question-answer (QA) pairs that participants provide. The current contribution studies the effect of worker selection and training on the agreement on implicit relation labels between workers and gold labels, for both the DC and the QA method. In Study 1, workers were not specifically selected or trained, and the results show that there is much room for improvement. Study 2 shows that a combination of selection and training does lead to improved results, but the method is cost- and time-intensive. Study 3 shows that a selection-only approach is a viable alternative; it results in annotations of comparable quality compared to annotations from trained participants. The results generalized over both the DC and QA method and therefore indicate that a selection-only approach could also be effective for other crowdsourced discourse annotation tasks. Merel C. J. Scholman, Valentina Pyatkin, Frances Yung, Ido Dagan, Reut Tsarfaty, Vera Demberg |
LREC | 6 |
| 2022 | Barch: an English Dataset of Bar Chart SummariesabstractWe present Barch, a new English dataset of human-written summaries describing bar charts. This dataset contains 47 charts based on a selection of 18 topics. Each chart is associated with one of the four intended messages expressed in the chart title. Using crowdsourcing, we collected around 20 summaries per chart, or one thousand in total. The text of the summaries is aligned with the chart data as well as with analytical inferences about the data drawn by humans. Our datasets is one of the first to explore the effect of intended messages on the data descriptions in chart summaries. Additionally, it lends itself well to the task of training data-driven systems for chart-to-text generation. We provide results on the performance of state-of-the-art neural generation models trained on this dataset and discuss the strengths and shortcomings of different models. Iza Skrjanec, Muhammad Salman Edhi, Vera Demberg |
LREC | 3 |
| 2022 | A Data-Driven Investigation of Noise-Adaptive Utterance Generation with Linguistic ModificationabstractIn noisy environments, speech can be hard to understand for humans. Spoken dialog systems can help to enhance the intelligibility of their output, either by modifying the speech synthesis (e.g., imitate Lombard speech) or by optimizing the language generation. We here focus on the second type of approach, by which an intended message is realized with words that are more intelligible in a specific noisy environment. By conducting a speech perception experiment, we created a dataset of 900 paraphrases in babble noise, perceived by native English speakers with normal hearing. We find that careful selection of paraphrases can improve intelligibility by 33% at SNR -5 dB. Our analysis of the data shows that the intelligibility differences between paraphrases are mainly driven by noise-robust acoustic cues. Furthermore, we propose an intelligibility-aware paraphrase ranking model, which outperforms baseline models with a relative improvement of 31.37% at SNR -5 dB. Anupama Chingacham, Vera Demberg, Dietrich Klakow |
SLT | 2 |
| 2021 | Pragmatics of Metaphor Revisited: Formalizing the Role of Typicality and Alternative Utterances in Metaphor Understanding
Alexandra Mayn, Vera Demberg |
CogSci | 2 |
| 2021 | Recognition of Minimal Pairs in (un)predictive Sentence Contexts in two Types of Noise
Marjolein van Os, Jutta Kray, Vera Demberg |
CogSci | 3 |
| 2021 | A bathtub by any other name: the reduction of German compounds in predictive contexts
Alessandra Zarcone, Vera Demberg |
CogSci | 2 |
| 2021 | Jointly Improving Language Understanding and Generation with Quality-Weighted Weak Supervision of Automatic LabelingabstractNeural natural language generation (NLG) and understanding (NLU) models are data-hungry and require massive amounts of annotated data to be competitive.Recent frameworks address this bottleneck with generative models that synthesize weak labels at scale, where a small amount of training labels are expertcurated and the rest of the data is automatically annotated.We follow that approach, by automatically constructing a large-scale weaklylabeled data with a fine-tuned GPT-2, and employ a semi-supervised framework to jointly train the NLG and NLU models.The proposed framework adapts the parameter updates to the models according to the estimated labelquality.On both the E2E and Weather benchmarks, we show that this weakly supervised training paradigm is an effective approach under low resource scenarios with as little as 10 data instances, and outperforming benchmark systems on both datasets when 100% of training data is used. Ernie Chang, Vera Demberg, Alex Marin |
EACL | 2 |
| 2021 | Neural Data-to-Text Generation with LM-based Text AugmentationabstractFor many new application domains for datato-text generation, the main obstacle in training neural models consists of a lack of training data.While usually large numbers of instances are available on the data side, often only very few text samples are available.To address this problem, we here propose a novel fewshot approach for this setting.Our approach automatically augments the data available for training by (i) generating new text samples based on replacing specific values by alternative ones from the same category, (ii) generating new text samples based on GPT-2, and (iii) proposing an automatic method for pairing the new text samples with data samples.As the text augmentation can introduce noise to the training data, we use cycle consistency as an objective, in order to make sure that a given data sample can be correctly reconstructed after having been formulated as text (and that text samples can be reconstructed from data).On both the E2E and WebNLG benchmarks, we show that this weakly supervised training paradigm is able to outperform fully supervised seq2seq models with less than 10% annotations.By utilizing all annotated data, our model can boost the performance of a standard seq2seq model by over 5 BLEU points, establishing a new state-of-the-art on both datasets. * Work done prior to joining Amazon.The Blue Spice is a restaurant that serves English cuisine. Ernie Chang, Xiaoyu Shen 0001, Vera Demberg, Hui Su |
EACL | 4 |
| 2021 | Does the Order of Training Samples Matter? Improving Neural Data-to-Text Generation with Curriculum LearningabstractRecent advancements in data-to-text generation largely take on the form of neural end-toend systems.Efforts have been dedicated to improving text generation systems by changing the order of training samples in a process known as curriculum learning.Past research on sequence-to-sequence learning showed that curriculum learning helps to improve both the performance and convergence speed.In this work, we delve into the same idea surrounding the training samples consisting of structured data and text pairs, where at each update, the curriculum framework selects training samples based on the model's competence.Specifically, we experiment with various difficulty metrics and put forward a soft edit distance metric for ranking training samples.Our benchmarks show faster convergence speed where training time is reduced by 38.7% and performance is boosted by 4.84 BLEU. Ernie Chang, Hui-Syuan Yeh, Vera Demberg |
EACL | 3 |
| 2021 | Why Do I Have to Take Over Control? Evaluating Safe Handovers with Advance Notice and Explanations in HADabstractIn highly automated driving (HAD), it is still an open question how machines can safely hand over control to humans, and if an advance notice with additional explanations can be beneficial in critical situations. Conceptually, use of formal methods from AI – description logic (DL) and automated planning – in order to more reliably predict when a handover is necessary, and to increase the advance notice for handovers by planning ahead at runtime, can provide a technological support for explanations using natural language generation. However, in this work we address only the user’s perspective with two contributions: First, we evaluate our concept in a driving simulator study (N=23) and find that an advance notice and spoken explanations were preferred over classical handover methods. Second, we propose a framework and an example test scenario specific to handovers that is based on the results of our study. Frederik Wiehr, Anke Hirsch, Lukas Schmitz, Nina Knieriemen, Antonio Krüger, Alisa Kovtunova, Stefan Borgwardt, Ernie Chang, Vera Demberg, Marcel Steinmetz, Jörg Hoffmann 0001 |
ICMI | 9 |
| 2021 | The SelectGen Challenge: Finding the Best Training Samples for Few-Shot Neural Text GenerationabstractWe propose a shared task on training instance selection for few-shot neural text generation.Large-scale pretrained language models have led to dramatic improvements in few-shot text generation.Nonetheless, almost all previous work simply applies random sampling to select the few-shot training instances.Little to no attention has been paid to the selection strategies and how they would affect model performance.The study of the selection strategy can help us to ( 1) make the most use of our annotation budget in downstream tasks and (2) better benchmark few-shot text generative models.We welcome submissions that present their selection strategies and the effects on the generation quality. Ernie Chang, Xiaoyu Shen 0001, Alex Marin, Vera Demberg |
INLG | 4 |
| 2021 | Exploring the Potential of Lexical Paraphrases for Mitigating Noise-Induced Comprehension ErrorsabstractListening in noisy environments can be difficult even for individuals with a normal hearing thresholds. The speech signal can be masked by noise, which may lead to word misperceptions on the side of the listener, and overall difficulty to understand the message. To mitigate hearing difficulties on listeners, a co-operative speaker utilizes voice modulation strategies like Lombard speech to generate noise-robust utterances, and similar solutions have been developed for speech synthesis systems. In this work, we propose an alternate solution of choosing noise-robust lexical paraphrases to represent an intended meaning. Our results show that lexical paraphrases differ in their intelligibility in noise. We evaluate the intelligibility of synonyms in context and find that choosing a lexical unit that is less risky to be misheard than its synonym introduced an average gain in comprehension of 37% at SNR -5 dB and 21% at SNR 0 dB for babble noise. Anupama Chingacham, Vera Demberg, Dietrich Klakow |
Interspeech | 2 |
| 2020 | Processing particularized pragmatic inferences under load
Margarita Ryzhova, Vera Demberg |
CogSci | 2 |
| 2020 | Story Generation with Rich DetailsabstractAutomatically generated stories should be not only coherent, but also interesting.Thus apart from realizing a story line, the text also have to include rich details to engage the readers.We propose a model that features two different generation components: (1) an outliner, which proceeds the main story line to establish global coherence; and (2) a detailer, which supplies relevant details to the story in a locally coherent manner.Human evaluation show that our model substantially improves the informativeness of generated text while retaining its coherence, outperforming a number of baselines. Fangzhou Zhai, Vera Demberg, Alexander Koller |
COLING | 2 |
| 2020 | Diverse and Relevant Visual Storytelling with Scene Graph EmbeddingsabstractA problem in automatically generated stories for image sequences is that they use overly generic vocabulary and phrase structure and fail to match the distributional characteristics of human-generated text. We address this problem by introducing explicit representations for objects and their relations by extracting scene graphs from the images. Utilizing an embedding of this scene graph enables our model to more explicitly reason over objects and their relations during story generation, compared to the global features from an object classifier used in previous work. We apply metrics that account for the diversity of words and phrases of generated stories as well as for reference to narratively-salient image features and show that our approach outperforms previous systems. Our experiments also indicate that our models obtain competitive results on reference-based metrics. Xudong Hong 0002, Rakshith Shetty, Asad B. Sayeed, Khushboo Mehra, Vera Demberg, Bernt Schiele |
CoNLL | 5 |
| 2019 | Next Sentence Prediction helps Implicit Discourse Relation Classification within and across DomainsabstractWei Shi, Vera Demberg. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Vera Demberg |
EMNLP/IJCNLP (1) | 2 |
| 2018 | Improving Variational Encoder-Decoders in Dialogue GenerationabstractVariational encoder-decoders (VEDs) have shown promising results in dialogue generation. However, the latent variable distributions are usually approximated by a much simpler model than the powerful RNN structure used for encoding and decoding, yielding the KL-vanishing problem and inconsistent training objective. In this paper, we separate the training step into two phases: The first phase learns to autoencode discrete texts into continuous embeddings, from which the second phase learns to generalize latent representations by reconstructing the encoded embedding. In this case, latent variables are sampled by transforming Gaussian noise through multi-layer perceptrons and are trained with a separate VED model, which has the potential of realizing a much more flexible distribution. We compare our model with current popular models and the experiment demonstrates substantial improvement in both metric-based and human evaluations. Xiaoyu Shen 0001, Hui Su, Shuzi Niu, Vera Demberg |
AAAI | 4 |
| 2018 | Toward Bayesian Synchronous Tree Substitution Grammars for Sentence PlanningabstractDeveloping conventional natural language generation systems requires extensive attention from human experts in order to craft complex sets of sentence planning rules.We propose a Bayesian nonparametric approach to learn sentence planning rules by inducing synchronous tree substitution grammars for pairs of text plans and morphosyntactically-specified dependency trees.Our system is able to learn rules which can be used to generate novel texts after training on small datasets. David M. Howcroft, Dietrich Klakow, Vera Demberg |
INLG | 3 |
| 2018 | A vision-grounded dataset for predicting typical locations for verbs
Nelson Mukuze, Anna Rohrbach, Vera Demberg, Bernt Schiele |
LREC | 3 |
| 2018 | Rollenwechsel-English: a large-scale semantic role corpus
Asad B. Sayeed, Pavel Shkadzko, Vera Demberg |
LREC | 3 |
| 2017 | Age differences in language comprehension during driving: Recovery from prediction errors is more effortful for older adults
Katja Häuser, Vera Demberg, Jutta Kray |
CogSci | 2 |
| 2017 | Psycholinguistic Models of Sentence Processing Improve Sentence Readability RankingabstractWhile previous research on readability has typically focused on document-level measures, recent work in areas such as natural language generation has pointed out the need of sentence-level readability measures.Much of psycholinguistics has focused for many years on processing measures that provide difficulty estimates on a word-by-word basis.However, these psycholinguistic measures have not yet been tested on sentence readability ranking tasks.In this paper, we use four psycholinguistic measures: idea density, surprisal, integration cost, and embedding depth to test whether these features are predictive of readability levels.We find that psycholinguistic features significantly improve performance by up to 3 percentage points over a standard document-level readability metric baseline. David M. Howcroft, Vera Demberg |
EACL (1) | 2 |
| 2017 | A Systematic Study of Neural Discourse Models for Implicit Discourse RelationabstractInferring implicit discourse relations in natural language text is the most difficult subtask in discourse parsing.Many neural network models have been proposed to tackle this problem.However, the comparison for this task is not unified, so we could hardly draw clear conclusions about the effectiveness of various architectures.Here, we propose neural network models that are based on feedforward and long-short term memory architecture and systematically study the effects of varying structures.To our surprise, the best-configured feedforward architecture outperforms LSTM-based model in most cases despite thorough tuning.Further, we compare our best feedforward system with competitive convolutional and recurrent networks and find that feedforward can actually be more effective.For the first time for this task, we compile and publish outputs from previous neural and nonneural systems to establish the standard for further comparison. Attapol Rutherford, Vera Demberg, Nianwen Xue |
EACL (1) | 2 |
| 2017 | Using Explicit Discourse Connectives in Translation for Implicit Discourse Relation ClassificationabstractImplicit discourse relation recognition is an extremely challenging task due to the lack of indicative connectives. Various neural network architectures have been proposed for this task recently, but most of them suffer from the shortage of labeled data. In this paper, we address this problem by procuring additional training data from parallel corpora: When humans translate a text, they sometimes add connectives (a process known as explicitation). We automatically back-translate it into an English connective and use it to infer a label with high confidence. We show that a training set several times larger than the original training set can be generated this way. With the extra labeled instances, we show that even a simple bidirectional Long Short-Term Memory Network can outperform the current state-of-the-art. Frances Yung, Raphaël Rubino, Vera Demberg |
IJCNLP(1) | 4 |
| 2017 | G-TUNA: a corpus of referring expressions in German, including duration informationabstractCorpora of referring expressions elicited from human participants in a controlled environment are an important resource for research on automatic referring expression generation.We here present G-TUNA, a new corpus of referring expressions for German.Using images of furniture as stimuli similarly to the TUNA and D-TUNA corpora, our corpus extends on these corpora by providing data collected in a simulated driving dual-task setting, and additionally provides exact duration annotations for the spoken referring expressions.This corpus will hence allow researchers to analyze the interaction between referring expression length and speech rate, under conditions where the listener is under high vs. low cognitive load. David M. Howcroft, Jorrig Vogels, Vera Demberg |
INLG | 3 |
| 2017 | The Extended SPaRKy Restaurant Corpus: Designing a Corpus with Variable Information Density
David M. Howcroft, Dietrich Klakow, Vera Demberg |
INTERSPEECH | 3 |
| 2017 | Modelling Semantic Expectation: Using Script Knowledge for Referent PredictionabstractRecent research in psycholinguistics has provided increasing evidence that humans predict upcoming content. Prediction also affects perception and might be a key to robustness in human language processing. In this paper, we investigate the factors that affect human prediction by building a computational model that can predict upcoming discourse referents based on linguistic knowledge alone vs. linguistic knowledge jointly with common-sense knowledge in the form of scripts. We find that script knowledge significantly improves model estimates of human predictions. In a second study, we test the highly controversial hypothesis that predictability influences referring expression type but do not find evidence for such an effect. Ashutosh Modi, Ivan Titov 0001, Vera Demberg, Asad B. Sayeed, Manfred Pinkal |
Trans. Assoc. Comput. Linguistics | 3 |
| 2016 | But vs. Although under the microscope
Fatemeh Torabi Asr, Vera Demberg |
CogSci | 2 |
| 2016 | From OpenCCG to AI Planning: Detecting Infeasible Edges in Sentence GenerationabstractThe search space in grammar-based natural language generation tasks can get very large, which is particularly problematic when generating long utterances or paragraphs. Using surface realization with OpenCCG as an example, we show that we can effectively detect partial solutions (edges) which cannot ultimately be part of a complete sentence because of their syntactic category. Formulating the completion of an edge into a sentence as finding a solution path in a large state-transition system, we demonstrate a connection to AI Planning which is concerned with this kind of problem. We design a compilation from OpenCCG into AI Planning allowing the detection of infeasible edges via AI Planning dead-end detection methods (proving the absence of a solution to the compilation). Our experiments show that this can filter out large fractions of infeasible edges in, and thus benefit the performance of, complex realization processes. Maximilian Schwenger, Álvaro Torralba, Jörg Hoffmann 0001, David M. Howcroft, Vera Demberg |
COLING | 5 |
| 2016 | Event participant modelling with neural networksabstractA common problem in cognitive modelling is lack of access to accurate broad-coverage models of event-level surprisal.As shown in, e.g., Bicknell et al. (2010), event-level knowledge does affect human expectations for verbal arguments.For example, the model should be able to predict that mechanics are likely to check tires, while journalists are more likely to check typos.Similarly, we would like to predict what locations are likely for playing football or playing flute in order to estimate the surprisal of actually-encountered locations.Furthermore, such a model can be used to provide a probability distribution over fillers for a thematic role which is not mentioned in the text at all.To this end, we train two neural network models (an incremental one and a non-incremental one) on large amounts of automatically rolelabelled text.Our models are probabilistic and can handle several roles at once, which also enables them to learn interactions between different role fillers.Evaluation shows a drastic improvement over current state-of-the-art systems on modelling human thematic fit judgements, and we demonstrate via a sentence similarity task that the system learns highly useful embeddings. Ottokar Tilk, Vera Demberg, Asad B. Sayeed, Dietrich Klakow, Stefan Thater |
EMNLP | 2 |
| 2016 | How can we adapt generation to the user's cognitive load?abstractAs language-based interaction becomes more ubiquitous and is used by in a larger and larger variety of different situations, the challenge for NLG systems is to not only convey a certain message correctly, but also do so in a way that is appropriate to the situation and the user. From various studies, we know that humans adapt the way they formulate their utterances to their conversational partners and may also change the way they say things as a function of the situation that the conversational partner is in (e.g. while talking to someone who is driving a car). Approaches from psycholinguistics (using information-theoretic measures as well as other complexity metrics) provide a way to formulate and quantify the demands that a certain formulation places on a hearer. In this talk, I will briefly survey ways of assessing human cognitive load in realistic settings, present current models of information density at the content level, and discuss the extent to which these measures have been found to drive choice of formulation in humans. Vera Demberg |
INLG | 1 |
| 2016 | Annotating Discourse Relations in Spoken Language: A Comparison of the PDTB and CCR Frameworks
Ines Rehbein, Merel C. J. Scholman, Vera Demberg |
LREC | 3 |
| 2016 | Improving event prediction by representing script participantsabstractAutomatically learning script knowledge has proved difficult, with previous work not or just barely beating a most-frequent baseline.Script knowledge is a type of world knowledge which can however be useful for various task in NLP and psycholinguistic modelling.We here propose a model that includes participant information (i.e., knowledge about which participants are relevant for a script) and show, on the Dinners from Hell corpus as well as the InScript corpus, that this knowledge helps us to significantly improve prediction performance on the narrative cloze task. Simon Ahrendt, Vera Demberg |
HLT-NAACL | 2 |
| 2015 | Vector-space calculation of semantic surprisal for predicting word pronunciation durationabstractAsad Sayeed, Stefan Fischer, Vera Demberg. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Asad B. Sayeed, Stefan Fischer 0008, Vera Demberg |
ACL (1) | 3 |
| 2015 | Semantically underinformative utterances trigger pragmatic inferences
Ekaterina Kravtchenko, Vera Demberg |
CogSci | 2 |
| 2015 | Improving unsupervised vector-space thematic fit evaluation via role-filler prototype clusteringabstractMost recent unsupervised methods in vector space semantics for assessing thematic fit (e.g.Erk, 2007;Baroni and Lenci, 2010;Sayeed and Demberg, 2014) create prototypical rolefillers without performing word sense disambiguation.This leads to a kind of sparsity problem: candidate role-fillers for different senses of the verb end up being measured by the same "yardstick", the single prototypical role-filler.In this work, we use three different feature spaces to construct robust unsupervised models of distributional semantics.We show that correlation with human judgements on thematic fit estimates can be improved consistently by clustering typical role-fillers and then calculating similarities of candidate rolefillers with these cluster centroids.The suggested methods can be used in any vector space model that constructs a prototype vector from a non-trivial set of typical vectors. Clayton Greenberg, Asad B. Sayeed, Vera Demberg |
HLT-NAACL | 3 |
| 2014 | Incremental and predictive discourse processing based on causal and concessive discourse markers: ERP studies on German and English
Heiner Drenhaus, Vera Demberg, Judith Köhne, Francesca Delogu |
CogSci | 2 |
| 2014 | Incremental Semantic Role Labeling with Tree Adjoining GrammarabstractWe introduce the task of incremental semantic role labeling (iSRL), in which semantic roles are assigned to incomplete input (sentence prefixes).iSRL is the semantic equivalent of incremental parsing, and is useful for language modeling, sentence completion, machine translation, and psycholinguistic modeling.We propose an iSRL system that combines an incremental TAG parser with a semantically enriched lexicon, a role propagation algorithm, and a cascade of classifiers.Our approach achieves an SRL Fscore of 78.38% on the standard CoNLL 2009 dataset.It substantially outperforms a strong baseline that combines gold-standard syntactic dependencies with heuristic role assignment, as well as a baseline based on Nivre's incremental dependency parser. Ioannis Konstas, Frank Keller, Vera Demberg, Mirella Lapata |
EMNLP | 3 |
| 2013 | Measuring linguistically-induced cognitive load during driving using the ConTRe taskabstractThis paper shows that fine-grained linguistic complexity has measurable effects on cognitive load with consequences for the design of in-car spoken dialogue systems. We used synthesized German sentences with grammatical ambiguities to test the additional workload caused by human sentence processing during driving. For the driving task, we used the Continuous Tracking and Reaction (ConTRe) task, which we believe is suitable for the measurement of the fine-grained effects of linguistically-related workload phenomena in automotive environments, as it provides millisecond-level driving deviation measurements on a continuous course. We applied the task in an eye-tracking environment, using a pupillometric measure of cognitive workload called the Index of Cognitive Activity (ICA). Vera Demberg, Asad B. Sayeed, Angela Castronovo, Christian Müller 0014 |
AutomotiveUI | 1 |
| 2013 | Pupillometry: the Index of Cognitive Activity in a dual-task study
Vera Demberg |
CogSci | 1 |
| 2013 | Integration Costs on Auxiliaries? - a self-paced reading study using WebExp
Vera Demberg |
CogSci | 1 |
| 2013 | The Index of Cognitive Activity as a Measure of Linguistic Processing
Vera Demberg, Evangelia Kiagia, Asad B. Sayeed |
CogSci | 1 |
| 2013 | Language and cognitive load in a dual task environment
Nikolaos Engonopoulos, Asad B. Sayeed, Vera Demberg |
CogSci | 3 |
| 2013 | The time-course of processing discourse connectives
Judith Köhne, Vera Demberg |
CogSci | 2 |
| 2013 | Identifying Predictive Collocations
Silas Weinbach, Vera Demberg |
CogSci | 2 |
| 2013 | Incremental, Predictive Parsing with Psycholinguistically Motivated Tree-Adjoining GrammarabstractPsycholinguistic research shows that key properties of the human sentence processor are incrementality, connectedness (partial structures contain no unattached nodes), and prediction (upcoming syntactic structure is anticipated). There is currently no broad-coverage parsing model with these properties, however. In this article, we present the first broad-coverage probabilistic parser for PLTAG, a variant of TAG that supports all three requirements. We train our parser on a TAG-transformed version of the Penn Treebank and show that it achieves performance comparable to existing TAG parsers that are incremental but not predictive. We also use our PLTAG model to predict human reading times, demonstrating a better fit on the Dundee eye-tracking corpus than a standard surprisal model. Vera Demberg, Frank Keller, Alexander Koller |
Comput. Linguistics | 1 |
| 2012 | Implicitness of Discourse Relations
Fatemeh Torabi Asr, Vera Demberg |
COLING | 2 |
| 2012 | Syntactic Surprisal Affects Spoken Word Duration in Conversational Contexts
Vera Demberg, Asad B. Sayeed, Philip Gorinski, Nikolaos Engonopoulos |
EMNLP-CoNLL | 1 |
| 2012 | German and English Treebanks and Lexica for Tree-Adjoining Grammars
Miriam Kaeshammer, Vera Demberg |
LREC | 2 |
| 2011 | A Strategy for Information Presentation in Spoken Dialog SystemsabstractIn spoken dialog systems, information must be presented sequentially, making it difficult to quickly browse through a large number of options. Recent studies have shown that user satisfaction is negatively correlated with dialog duration, suggesting that systems should be designed to maximize the efficiency of the interactions. Analysis of the logs of 2,000 dialogs between users and nine different dialog systems reveals that a large percentage of the time is spent on the information presentation phase, thus there is potentially a large pay-off to be gained from optimizing information presentation in spoken dialog systems. This article proposes a method that improves the efficiency of coping with large numbers of diverse options by selecting options and then structuring them based on a model of the user's preferences. This enables the dialog system to automatically determine trade-offs between alternative options that are relevant to the user and present these trade-offs explicitly. Multiple attractive options are thereby structured such that the user can gradually refine her request to find the optimal trade-off. To evaluate and challenge our approach, we conducted a series of experiments that test the effectiveness of the proposed strategy. Experimental results show that basing the content structuring and content selection process on a user model increases the efficiency and effectiveness of the user's interaction. Users complete their tasks more successfully and more quickly. Furthermore, user surveys revealed that participants found that the user-model based system presents complex trade-offs understandably and increases overall user satisfaction. The experiments also indicate that presenting users with a brief overview of options that do not fit their requirements significantly improves the user's overview of available options, also making them feel more confident in having been presented with all relevant options. Vera Demberg, Andi Winterboer, Johanna D. Moore |
Comput. Linguistics | 1 |
| 2010 | Syntactic and Semantic Factors in Processing Difficulty: An Integrated Measure
Jeff Mitchell 0001, Mirella Lapata, Vera Demberg, Frank Keller |
ACL | 3 |
| 2008 | Syntactic complexity induces explicit grounding in the Maptask corpus
Martin I. Tietze, Vera Demberg, Johanna D. Moore |
INTERSPEECH | 2 |
| 2007 | A Language-Independent Unsupervised Model for Morphological Segmentation
Vera Demberg |
ACL | 1 |
| 2007 | Phonological Constraints and Morphological Preprocessing for Grapheme-to-Phoneme Conversion
Vera Demberg, Helmut Schmid, Gregor Möhler |
ACL | 1 |
| 2006 | Information Presentation in Spoken Dialogue Systems
Vera Demberg, Johanna D. Moore |
EACL | 1 |