VLDB 2026 Research / reviewers in the wild / expert
David Schlangen
dblp:11/1189
· DBLP profile ↗
90ranked-venue papers
7as first author
23since 2021 · last 2025
0000-0002-2686-6887ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 82 · 7 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 9 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and AttitudesabstractRational speakers are supposed to know what they know and what they do not know, and to generate expressions matching the strength of evidence.In contrast, it is still a challenge for current large language models to generate corresponding utterances based on the assessment of facts and confidence in an uncertain real-world environment.While it has recently become popular to estimate and calibrate confidence of LLMs with verbalized uncertainty, what is lacking is a careful examination of the linguistic knowledge of uncertainty encoded in the latent space of LLMs.In this paper, we draw on typological frameworks of epistemic expressions to evaluate LLMs' knowledge of epistemic modality, using controlled stories.Our experiments show that the performance of LLMs in generating epistemic expressions is limited and not robust, and hence the expressions of uncertainty generated by LLMs are not always reliable.To build uncertainty-aware LLMs, it is necessary to enrich the semantic knowledge of epistemic modality in LLMs. Michael Vrazitulis, David Schlangen |
ACL (1) | 3 |
| 2025 | Using Game Play to Investigate Multimodal and Conversational Grounding in Large Multimodal ModelsabstractWhile the situation has improved for text-only models, it again seems to be the case currently that multimodal (text and image) models develop faster than ways to evaluate them. In this paper, we bring a recently developed evaluation paradigm from text models to multimodal models, namely evaluation through the goal-oriented game (self) play, complementing reference-based and preference-based evaluation. Specifically, we define games that challenge a model’s capability to represent a situation from visual information and align such representations through dialogue. We find that the largest closed models perform rather well on the games that we define, while even the best open-weight models struggle with them. On further analysis, we find that the exceptional deep captioning capabilities of the largest models drive some of the performance. There is still room to grow for both kinds of models, ensuring the continued relevance of the benchmark. Sherzod Hakimov, Yerkezhan Abdullayeva, Kushal Koshti, Antonia Schmidt, Yan Weiser, Anne Beyer, David Schlangen |
COLING | 7 |
| 2025 | Playpen: An Environment for Exploring Learning From Dialogue Game FeedbackabstractNicola Horst, Davide Mazzaccara, Antonia Schmidt, Michael Sullivan, Filippo Momentè, Luca Franceschetti, Philipp Sadler, Sherzod Hakimov, Alberto Testoni, Raffaella Bernardi, Raquel Fernández, Alexander Koller, Oliver Lemon, David Schlangen, Mario Giulianelli, Alessandro Suglia. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Nicola Horst, Davide Mazzaccara, Antonia Schmidt, Michael Sullivan, Filippo Momentè, Luca Franceschetti, Philipp Sadler, Sherzod Hakimov, Alberto Testoni, Raffaella Bernardi, Raquel Fernández, Alexander Koller, Oliver Lemon, David Schlangen, Mario Giulianelli, Alessandro Suglia |
EMNLP | 14 |
| 2025 | clem: todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System RealisationsabstractThe emergence of instruction-tuned large language models (LLMs) has advanced the field of dialogue systems, enabling both realistic user simulations and robust multi-turn conversational agents. However, existing research often evaluates these components in isolation, either focusing on a single user simulator or a specific system design, limiting the generalisability of insights across architectures and configurations. In this work, we propose clem:todd (chat-optimized LLMs for task-oriented dialogue systems development), a flexible framework for systematically evaluating dialogue systems under consistent conditions. clem:todd enables detailed benchmarking across combinations of user simulators and dialogue systems, whether existing models from literature or newly developed ones. To the best of our knowledge, clem:todd is the first evaluation framework for task-oriented dialogue systems that supports plug-and-play integration and ensures uniform datasets, evaluation metrics, and computational constraints. We showcase clem:todd’s flexibility by re-evaluating existing task-oriented dialogue systems within this unified setup and integrating three newly proposed dialogue systems into the same evaluation pipeline. Our results provide actionable insights into how architecture, scale, and prompting strategies affect dialogue performance, offering practical guidance for building efficient and effective conversational AI systems. Kranti Chalamalasetti, Sherzod Hakimov, David Schlangen |
SIGDIAL | 3 |
| 2024 | When Only Time Will Tell: Interpreting How Transformers Process Local Ambiguities Through the Lens of Restart-IncrementalityabstractIncremental models that process sentences one token at a time will sometimes encounter points where more than one interpretation is possible.Causal models are forced to output one interpretation and continue, whereas models that can revise may edit their previous output as the ambiguity is resolved.In this work, we look at how restart-incremental Transformers build and update internal states, in an effort to shed light on what processes cause revisions not viable in autoregressive models.We propose an interpretable way to analyse the incremental states, showing that their sequential structure encodes information on the garden path effect and its resolution.Our method brings insights on various bidirectional encoders for contextualised meaning representation and dependency parsing, contributing to show their advantage over causal models when it comes to revisions. 1 Brielen Madureira, Patrick Kahardipraja, David Schlangen |
ACL (1) | 3 |
| 2024 | Effects of Eye Movement Patterns and Scene-Object Relations on Description Production
Pelin Çelikkol, David Schlangen, Jochen Laubrock |
CogSci | 2 |
| 2024 | Learning Part-whole Hierarchies from the Sequence of Handwriting
David Schlangen, Dietrich Klakow |
CogSci | 2 |
| 2024 | Conceptual Pacts for Reference Resolution Using Small, Dynamically Constructed Language Models: A Study in Puzzle Building DialoguesabstractUsing Brennan and Clark’s theory of a Conceptual Pact, that when interlocutors agree on a name for an object, they are forming a temporary agreement on how to conceptualize that object, we present an extension to a simple reference resolver which simulates this process over time with different conversation pairs. In a puzzle construction domain, we model pacts with small language models for each referent which update during the interaction. When features from these pact models are incorporated into a simple bag-of-words reference resolver, the accuracy increases compared to using a standard pre-trained model. The model performs equally to a competitor using the same data but with exhaustive re-training after each prediction, while also being more transparent, faster and less resource-intensive. We also experiment with reducing the number of training interactions, and can still achieve reference resolution accuracies of over 80% in testing from observing a single previous interaction, over 20% higher than a pre-trained baseline. While this is a limited domain, we argue the model could be applicable to larger real-world applications in human and human-robot interaction and is an interpretable and transparent model. Julian Hough, Sina Zarrieß, Casey Kennington, David Schlangen, Massimo Poesio |
LREC/COLING | 4 |
| 2024 | Sharing the Cost of Success: A Game for Evaluating and Learning Collaborative Multi-Agent Instruction Giving and Following PoliciesabstractIn collaborative goal-oriented settings, the participants are not only interested in achieving a successful outcome, but do also implicitly negotiate the effort they put into the interaction (by adapting to each other). In this work, we propose a challenging interactive reference game that requires two players to coordinate on vision and language observations. The learning signal in this game is a score (given after playing) that takes into account the achieved goal and the players’ assumed efforts during the interaction. We show that a standard Proximal Policy Optimization (PPO) setup achieves a high success rate when bootstrapped with heuristic partner behaviors that implement insights from the analysis of human-human interactions. And we find that a pairing of neural partners indeed reduces the measured joint effort when playing together repeatedly. However, we observe that in comparison to a reasonable heuristic pairing there is still room for improvement—which invites further research in the direction of cost-sharing in collaborative interactions. Philipp Sadler, Sherzod Hakimov, David Schlangen |
LREC/COLING | 3 |
| 2024 | The Unreasonable Ineffectiveness of Nucleus Sampling on Mitigating Text MemorizationabstractThis work analyses the text memorization behavior of large language models (LLMs) when subjected to nucleus sampling.Stochastic decoding methods like nucleus sampling are typically applied to overcome issues such as monotonous and repetitive text generation, which are often observed with maximizationbased decoding techniques.We hypothesize that nucleus sampling might also reduce the occurrence of memorization patterns, because it could lead to the selection of tokens outside the memorized sequence.To test this hypothesis we create a diagnostic dataset with a known distribution of duplicates that gives us some control over the likelihood of memorization of certain parts of the training data.Our analysis of two GPT-Neo models fine-tuned on this dataset interestingly shows that (i) an increase of the nucleus size reduces memorization only modestly, and (ii) even when models do not engage in "hard" memorization -a verbatim reproduction of training samples -they may still display "soft" memorization whereby they generate outputs that echo the training data but without a complete one-by-one resemblance. Luka Borec, Philipp Sadler, David Schlangen |
INLG | 3 |
| 2024 | A Dialogue Game for Eliciting Balanced CollaborationabstractCollaboration is an integral part of human dialogue.Typical task-oriented dialogue games assign asymmetric roles to the participants, which limits their ability to elicit naturalistic roletaking in collaboration and its negotiation.We present a novel and simple online setup that favors balanced collaboration: a two-player 2D object placement game in which the players must negotiate the goal state themselves.We show empirically that human players exhibit a variety of role distributions, and that balanced collaboration improves task performance.We also present an LLM-based baseline agent which demonstrates that automatic playing of our game is an interesting challenge for artificial systems. Isidora Jeknic, David Schlangen, Alexander Koller |
SIGDIAL | 2 |
| 2024 | It Couldn't Help but Overhear: On the Limits of Modelling Meta-Communicative Grounding Acts with Supervised LearningabstractActive participation in a conversation is key to building common ground, since understanding is jointly tailored by producers and recipients.Overhearers are deprived of the privilege of performing grounding acts and can only conjecture about intended meanings.Still, data generation and annotation, modelling, training and evaluation of NLP dialogue models place reliance on the overhearing paradigm.How much of the underlying grounding processes are thereby forfeited?As we show, there is evidence pointing to the impossibility of properly modelling human meta-communicative acts with data-driven learning models.In this paper, we discuss this issue and provide a preliminary analysis on the variability of human decisions for requesting clarification.Most importantly, we wish to bring this topic back to the community's table, encouraging discussion on the consequences of having models designed to only "listen in". Brielen Madureira, David Schlangen |
SIGDIAL | 2 |
| 2023 | Revising with a Backward Glance: Regressions and Skips during Reading as Cognitive Signals for Revision Policies in Incremental ProcessingabstractIn NLP, incremental processors produce output in instalments, based on incoming prefixes of the linguistic input.Some tokens trigger revisions, causing edits to the output hypothesis, but little is known about why models revise when they revise.A policy that detects the time steps where revisions should happen can improve efficiency.Still, retrieving a suitable signal to train a revision policy is an open problem, since it is not naturally available in datasets.In this work, we investigate the appropriateness of regressions and skips in human reading eye-tracking data as signals to inform revision policies in incremental sequence labelling.Using generalised mixed-effects models, we find that the probability of regressions and skips by humans can potentially serve as useful predictors for revisions in BiLSTMs and Transformer models, with consistent results for various languages. Brielen Madureira, Pelin Çelikkol, David Schlangen |
CoNLL | 3 |
| 2023 | Instruction Clarification Requests in Multimodal Collaborative Dialogue Games: Tasks, and an Analysis of the CoDraw DatasetabstractIn visual instruction-following dialogue games, players can engage in repair mechanisms in face of an ambiguous or underspecified instruction that cannot be fully mapped to actions in the world.In this work, we annotate Instruction Clarification Requests (iCRs) in CoDraw, an existing dataset of interactions in a multimodal collaborative dialogue game.We show that it contains lexically and semantically diverse iCRs being produced self-motivatedly by players deciding to clarify in order to solve the task successfully.With 8.8k iCRs found in 9.9k dialogues, CoDraw-iCR (v1) is a large spontaneous iCR corpus, making it a valuable resource for data-driven research on clarification in dialogue.We then formalise and provide baseline models for two tasks: Determining when to make an iCR and how to recognise them, in order to investigate to what extent these tasks are learnable from data. Brielen Madureira, David Schlangen |
EACL | 2 |
| 2023 | Pento-DIARef: A Diagnostic Dataset for Learning the Incremental Algorithm for Referring Expression Generation from ExamplesabstractNLP tasks are typically defined extensionally through datasets containing example instantiations (e.g., pairs of image i and text t), but motivated intensionally through capabilities invoked in verbal descriptions of the task (e.g., "t is a description of i, for which the content of i needs to be recognised and understood").We present Pento-DIARef, a diagnostic dataset in a visual domain of puzzle pieces where referring expressions are generated by a wellknown symbolic algorithm (the "Incremental Algorithm"), which itself is motivated by appeal to a hypothesised capability (eliminating distractors through application of Gricean maxims).Our question then is whether the extensional description (the dataset) is sufficient for a neural model to pick up the underlying regularity and exhibit this capability given the simple task definition of producing expressions from visual inputs.We find that a model supported by a vision detection step and a targeted data generation scheme achieves an almost perfect BLEU@1 score and sentence accuracy, whereas simpler baselines do not. Philipp Sadler, David Schlangen |
EACL | 2 |
| 2023 | clembench: Using Game Play to Evaluate Chat-Optimized Language Models as Conversational AgentsabstractRecent work has proposed a methodology for the systematic evaluation of "Situated Language Understanding Agents"-agents that operate in rich linguistic and non-linguistic contexts-through testing them in carefully constructed interactive settings.Other recent work has argued that Large Language Models (LLMs), if suitably set up, can be understood as (simulators of) such agents.A connection suggests itself, which this paper explores: Can LLMs be evaluated meaningfully by exposing them to constrained game-like settings that are built to challenge specific capabilities?As a proof of concept, this paper investigates five interaction settings, showing that current chatoptimised LLMs are, to an extent, capable of following game-play instructions.Both this capability and the quality of the game play, measured by how well the objectives of the different games are met, follows the development cycle, with newer models generally performing better.The metrics even for the comparatively simple example games are far from being saturated, suggesting that the proposed instrument will remain to have diagnostic value.Our general framework for implementing and evaluating games with LLMs is available at https://github.com/clembench. Kranti Chalamalasetti, Jana Götze, Sherzod Hakimov, Brielen Madureira, Philipp Sadler, David Schlangen |
EMNLP | 6 |
| 2023 | TF-IDF based Scene-Object Relations Correlate With Visual AttentionabstractThe relative contribution of bottom-up and top-down attentional guidance is a central topic in vision research. Whereas attention is guided bottom-up by low-level saliency, top-down guidance involves the viewer’s knowledge and expectations accumulated throughout a lifetime. Here we explore the influence of high-level scene-object relations on viewing behavior. To assess top-down guidance, we score the relevance of linguistic object labels using methods from document analysis. Specifically, we computed the term frequency-inverse document frequency (TF-IDF), a statistic that reflects how important a term is to a document. We use object TF-IDF to measure how important a specific object is to a scene category and use these scores to predict eye movement distributions over scenes. Our results show that scene-specific objects are more likely to be fixated. Object TF-IDF had an effect partially independent of image saliency, suggesting that an object’s relevance for a scene category affects attention during scene perception. Pelin Çelikkol, Jochen Laubrock, David Schlangen |
ETRA | 3 |
| 2023 | The Road to Quality is Paved with Good Revisions: A Detailed Evaluation Methodology for Revision Policies in Incremental Sequence LabellingabstractIncremental dialogue model components produce a sequence of output prefixes based on incoming input.Mistakes can occur due to local ambiguities or to wrong hypotheses, making the ability to revise past outputs a desirable property that can be governed by a policy.In this work, we formalise and characterise edits and revisions in incremental sequence labelling and propose metrics to evaluate revision policies.We then apply our methodology to profile the incremental behaviour of three Transformerbased encoders in various tasks, paving the road for better revision policies. Brielen Madureira, Patrick Kahardipraja, David Schlangen |
SIGDIAL | 3 |
| 2022 | New or Old? Exploring How Pre-Trained Language Models Represent Discourse EntitiesabstractRecent research shows that pre-trained language models, built to generate text conditioned on some context, learn to encode syntactic knowledge to a certain degree. This has motivated researchers to move beyond the sentence-level and look into their ability to encode less studied discourse-level phenomena. In this paper, we add to the body of probing research by investigating discourse entity representations in large pre-trained language models in English. Motivated by early theories of discourse and key pieces of previous work, we focus on the information-status of entities as discourse-new or discourse-old. We present two probing models, one based on binary classification and another one on sequence labeling. The results of our experiments show that pre-trained language models do encode information on whether an entity has been introduced before or not in the discourse. However, this information alone is not sufficient to find the entities in a discourse, opening up interesting questions about the definition of entities for future work. Sharid Loáiciga, Anne Beyer, David Schlangen |
COLING | 3 |
| 2022 | The slurk Interaction Server Framework: Better Data for Better Dialog ModelsabstractThis paper presents the slurk software, a lightweight interaction server for setting up dialog data collections and running experiments. slurk enables a multitude of settings including text-based, speech and video interaction between two or more humans or humans and bots, and a multimodal display area for presenting shared or private interactive context. The software is implemented in Python with an HTML and JavaScript frontend that can easily be adapted to individual needs. It also provides a setup for pairing participants on common crowdworking platforms such as Amazon Mechanical Turk and some example bot scripts for common interaction scenarios. Jana Götze, Maike Paetzel-Prüsmann, Wencke Liermann, Tim Diekmann, David Schlangen |
LREC | 5 |
| 2021 | Space Efficient Context Encoding for Non-Task-Oriented Dialogue Generation with Graph Attention TransformerabstractFabian Galetzka, Jewgeni Rose, David Schlangen, Jens Lehmann. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Fabian Galetzka, Jewgeni Rose, David Schlangen, Jens Lehmann 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | Towards Incremental Transformers: An Empirical Analysis of Transformer Models for Incremental NLUabstractIncremental processing allows interactive systems to respond based on partial inputs, which is a desirable property e.g. in dialogue agents.The currently popular Transformer architecture inherently processes sequences as a whole, abstracting away the notion of time.Recent work attempts to apply Transformers incrementally via restart-incrementality by repeatedly feeding, to an unchanged model, increasingly longer input prefixes to produce partial outputs.However, this approach is computationally costly and does not scale efficiently for long sequences.In parallel, we witness efforts to make Transformers more efficient, e.g. the Linear Transformer (LT) with a recurrence mechanism.In this work, we examine the feasibility of LT for incremental NLU in English.Our results show that the recurrent LT model has better incremental performance and faster inference speed compared to the standard Transformer and LT with restartincrementality, at the cost of part of the nonincremental (full sequence) quality.We show that the performance drop can be mitigated by training the model to wait for right context before committing to an output and that training with input prefixes is beneficial for delivering correct partial outputs. Patrick Kahardipraja, Brielen Madureira, David Schlangen |
EMNLP (1) | 3 |
| 2021 | Is Incoherence Surprising? Targeted Evaluation of Coherence Prediction from Language ModelsabstractCoherent discourse is distinguished from a mere collection of utterances by the satisfaction of a diverse set of constraints, for example choice of expression, logical relation between denoted events, and implicit compatibility with world-knowledge. Do neural language models encode such constraints? We design an extendable set of test suites addressing different aspects of discourse and dialogue coherence. Unlike most previous coherence evaluation studies, we address specific linguistic devices beyond sentence order perturbations, which allow for a more fine-grained analysis of what constitutes coherence and what neural models trained on a language modelling objective are capable of encoding. Extending the targeted evaluation paradigm for neural language models (Marvin and Linzen, 2018) to phenomena beyond syntax, we show that this paradigm is equally suited to evaluate linguistic qualities that contribute to the notion of coherence. Anne Beyer, Sharid Loáiciga, David Schlangen |
NAACL-HLT | 3 |
| 2020 | Incremental Processing in the Age of Non-Incremental Encoders: An Empirical Assessment of Bidirectional Models for Incremental NLUabstractWhile humans process language incrementally, the best language encoders currently used in NLP do not.Both bidirectional LSTMs and Transformers assume that the sequence that is to be encoded is available in full, to be processed either forwards and backwards (BiL-STMs) or as a whole (Transformers).We investigate how they behave under incremental interfaces, when partial output must be provided based on partial input seen up to a certain time step, which may happen in interactive systems.We test five models on various NLU datasets and compare their performance using three incremental evaluation metrics.The results support the possibility of using bidirectional encoders in incremental mode while retaining most of their non-incremental quality.The "omni-directional" BERT model, which achieves better non-incremental performance, is impacted more by the incremental access.This can be alleviated by adapting the training regime (truncated training), or the testing procedure, by delaying the output until some right context is available or by incorporating hypothetical right contexts generated by a language model like GPT-2. Brielen Madureira, David Schlangen |
EMNLP (1) | 2 |
| 2020 | From "Before" to "After": Generating Natural Language Instructions from Image Pairs in a Simple Visual DomainabstractWhile certain types of instructions can be compactly expressed via images, there are situations where one might want to verbalise them, for example when directing someone.We investigate the task of Instruction Generation from Before/After Image Pairs which is to derive from images an instruction for effecting the implied change.For this, we make use of prior work on instruction following in a visual environment.We take an existing dataset, the BLOCKS data collected by Bisk et al. (2016) and investigate whether it is suitable for training an instruction generator as well.We find that it is, and investigate several simple baselines, taking these from the related task of image captioning.Through a series of experiments that simplify the task (by making image processing easier or completely side-stepping it; and by creating template-based targeted instructions), we investigate areas for improvement.We find that captioning models get some way towards solving the task, but have some difficulty with it, and future improvements must lie in the way the change is detected in the instruction. Robin Rojowiec, Jana Götze, Philipp Sadler, Henrik Voigt, Sina Zarrieß, David Schlangen |
INLG | 6 |
| 2020 | A Corpus of Controlled Opinionated and Knowledgeable Movie Discussions for Training Neural Conversation ModelsabstractFully data driven Chatbots for non-goal oriented dialogues are known to suffer from inconsistent behaviour across their turns, stemming from a general difficulty in controlling parameters like their assumed background personality and knowledge of facts. One reason for this is the relative lack of labeled data from which personality consistency and fact usage could be learned together with dialogue behaviour. To address this, we introduce a new labeled dialogue dataset in the domain of movie discussions, where every dialogue is based on pre-specified facts and opinions. We thoroughly validate the collected dialogue for adherence of the participants to their given fact and opinion profile, and find that the general quality in this respect is high. This process also gives us an additional layer of annotation that is potentially useful for training models. We introduce as a baseline an end-to-end trained self-attention decoder model trained on this data and show that it is able to generate opinionated responses that are judged to be natural and knowledgeable and show attentiveness. Fabian Galetzka, Chukwuemeka Uchenna Eneh, David Schlangen |
LREC | 3 |
| 2019 | Know What You Don't Know: Modeling a Pragmatic Speaker that Refers to Objects of Unknown CategoriesabstractZero-shot learning in Language & Vision is the task of correctly labelling (or naming) objects of novel categories.Another strand of work in L&V aims at pragmatically informative rather than "correct" object descriptions, e.g. in reference games.We combine these lines of research and model zero-shot reference games, where a speaker needs to successfully refer to a novel object in an image.Inspired by models of "rational speech acts", we extend a neural generator to become a pragmatic speaker reasoning about uncertain object categories.As a result of this reasoning, the generator produces fewer nouns and names of distractor categories as compared to a literal speaker.We show that this conversational strategy for dealing with novel objects often improves communicative success, in terms of resolution accuracy of an automatic listener. Sina Zarrieß, David Schlangen |
ACL (1) | 2 |
| 2019 | Tell Me More: A Dataset of Visual Scene Description SequencesabstractIlinykh N, Zarrieß S, Schlangen D. Tell Me More: A Dataset of Visual Scene Description Sequences. In: Proceedings of the 12th International Conference on Natural Language Generation. Stroudsburg, PA, USA: Association for Computational Linguistics; 2019: 152-157. Nikolai Ilinykh, Sina Zarrieß, David Schlangen |
INLG | 3 |
| 2019 | Can Neural Image Captioning be Controlled via Forced Attention?abstractLearned dynamic weighting of the conditioning signal (attention) has been shown to improve neural language generation in a variety of settings.The weights applied when generating a particular output sequence have also been viewed as providing a potentially explanatory insight into the internal workings of the generator.In this paper, we reverse the direction of this connection and ask whether through the control of the attention of the model we can control its output.Specifically, we take a standard neural image captioning model that uses attention, and fix the attention to predetermined areas in the image.We evaluate whether the resulting output is more likely to mention the class of the object in that area than the normally generated caption.We introduce three effective methods to control the attention and find that these are producing expected results in up to 27.43% of the cases. Philipp Sadler, Tatjana Scheffler, David Schlangen |
INLG | 3 |
| 2019 | From Explainability to Explanation: Using a Dialogue Setting to Elicit Annotations with JustificationsabstractDespite recent attempts in the field of explainable AI to go beyond black box prediction models, typically already the training data for supervised machine learning is collected in a manner that treats the annotator as a "black box", the internal workings of which remains unobserved.We present an annotation method where a task is given to a pair of annotators who collaborate on finding the best response.With this we want to shed light on the questions if the collaboration increases the quality of the responses and if this "thinking together" provides useful information in itself, as it at least partially reveals their reasoning steps.Furthermore, we expect that this setting puts the focus on explanation as a linguistic act, vs. explainability as a property of models.In a crowd-sourcing experiment, we investigated three different annotation tasks, each in a collaborative dialogical (two annotators) and monological (one annotator) setting.Our results indicate that our experiment elicits collaboration and that this collaboration increases the response accuracy.We see large differences in the annotators' behavior depending on the task.Similarly, we also observe that the dialog patterns emerging from the collaboration vary significantly with the task. Nazia Attari, Martin Heckmann, David Schlangen |
SIGdial | 3 |
| 2018 | Placing Objects in Gesture Space: Toward Incremental Interpretation of Multimodal Spatial DescriptionsabstractWhen describing routes not in the current environment, a common strategy is to anchor the description in configurations of salient landmarks, complementing the verbal descriptions by "placing" the non-visible landmarks in the gesture space. Understanding such multimodal descriptions and later locating the landmarks from real world is a challenging task for the hearer, who must interpret speech and gestures in parallel, fuse information from both modalities, build a mental representation of the description, and ground the knowledge to real world landmarks. In this paper, we model the hearer's task, using a multimodal spatial description corpus we collected. To reduce the variability of verbal descriptions, we simplified the setup to use simple objects as landmarks. We describe a real-time system to evaluate the separate and joint contribution of the modalities. We show that gestures not only help to improve the overall system performance, even if to a large extent they encode redundant information, but also result in earlier final correct interpretations. Being able to build and apply representations incrementally will be of use in more dialogical settings, we argue, where it can enable immediate clarification in cases of mismatch. Ting Han 0003, Casey Kennington, David Schlangen |
AAAI | 3 |
| 2018 | The Task Matters: Comparing Image Captioning and Task-Based Dialogical Image DescriptionabstractImage captioning models are typically trained on data that is collected from people who are asked to describe an image, without being given any further task context.As we argue here, this context independence is likely to cause problems for transferring to task settings in which image description is bound by task demands.We demonstrate that careful design of data collection is required to obtain image descriptions which are contextually bounded to a particular meta-level task.As a task, we use MeetUp!, a text-based communication game where two players have the goal of finding each other in a visual environment.To reach this goal, the players need to describe images representing their current location.We analyse a dataset from this domain and show that the nature of image descriptions found in MeetUp! is diverse, dynamic and rich with phenomena that are not present in descriptions obtained through a simple image captioning task, which we ran for comparison. Nikolai Ilinykh, Sina Zarrieß, David Schlangen |
INLG | 3 |
| 2018 | Decoding Strategies for Neural Referring Expression GenerationabstractRNN-based sequence generation is now widely used in NLP and NLG (natural language generation).Most work focusses on how to train RNNs, even though also decoding is not necessarily straightforward: previous work on neural MT found seq2seq models to radically prefer short candidates, and has proposed a number of beam search heuristics to deal with this.In this work, we assess decoding strategies for referring expression generation with neural models.Here, expression length is crucial: output should neither contain too much or too little information, in order to be pragmatically adequate.We find that most beam search heuristics developed for MT do not generalize well to referring expression generation (REG), and do not generally outperform greedy decoding.We observe that beam search heuristics for termination seem to override the model's knowledge of what a good stopping point is.Therefore, we also explore a recent approach called trainable decoding, which uses a small network to modify the RNN's hidden state for better decoding results.We find this approach to consistently outperform greedy decoding for REG. Sina Zarrieß, David Schlangen |
INLG | 2 |
| 2018 | A Corpus of Natural Multimodal Spatial Scene Descriptions
Ting Han 0003, David Schlangen |
LREC | 2 |
| 2017 | Obtaining referential word meanings from visual and distributional information: Experiments on object namingabstractWe investigate object naming, which is an important sub-task of referring expression generation on real-world images.As opposed to mutually exclusive labels used in object recognition, object names are more flexible, subject to communicative preferences and semantically related to each other.Therefore, we investigate models of referential word meaning that link visual to lexical information which we assume to be given through distributional word embeddings.We present a model that learns individual predictors for object names that link visual and distributional aspects of word meaning during training.We show that this is particularly beneficial for zero-shot learning, as compared to projecting visual objects directly into the distributional space.In a standard object naming task, we find that different ways of combining lexical and visual information achieve very similar performance, though experiments on model combination suggest that they capture complementary aspects of referential meaning. Sina Zarrieß, David Schlangen |
ACL (1) | 2 |
| 2017 | Joint, Incremental Disfluency Detection and Utterance Segmentation from SpeechabstractWe present the joint task of incremental disfluency detection and utterance segmentation and a simple deep learning system which performs it on transcripts and ASR results.We show how the constraints of the two tasks interact.Our joint-task system outperforms the equivalent individual task systems, provides competitive results and is suitable for future use in conversation agents in the psychiatric domain. Julian Hough, David Schlangen |
EACL (1) | 2 |
| 2017 | Deriving continous grounded meaning representations from referentially structured multimodal contextsabstractCorpora of referring expressions paired with their visual referents are a good source for learning word meanings directly grounded in visual representations.Here, we explore additional ways of extracting from them word representations linked to multi-modal context: through expressions that refer to the same object, and through expressions that refer to different objects in the same scene.We show that continuous meaning representations derived from these contexts capture complementary aspects of similarity, even if not outperforming textual embeddings trained on very large amounts of raw text when tested on standard similarity benchmarks.We propose a new task for evaluating grounded meaning representations-detection of potentially co-referential phrases-and show that it requires precise denotational representations of attribute meanings, which our method provides.woman txt ref lady, girl, man, chick den lady, girl, women, blouse sit girl, guy, man, lady vis lady, girl, women, chick sidewalk txt ref pavement, ground, walkway, steps den street, sidewlak, walkway, pavement sit buildin, bldg, lamppost, street vis pavement, street, walkway, concrete grass txt Sina Zarrieß, David Schlangen |
EMNLP | 2 |
| 2017 | It's Not What You Do, It's How You Do It: Grounding Uncertainty for a Simple RobotabstractFor effective HRI, robots must go beyond having good legibility of their intentions shown by their actions, but also ground the degree of uncertainty they have. We show how in simple robots which have spoken language understanding capacities, uncertainty can be communicated to users by principles of grounding in dialogue interaction even without natural language generation. We present a model which makes this possible for robots with limited communication channels beyond the execution of task actions themselves. We implement our model in a pick-and-place robot, and experiment with two strategies for grounding uncertainty. In an observer study, we show that participants observing interactions with the robot run by the two different strategies were able to infer the degree of understanding the robot had internally, and in the more uncertainty-expressive system, were also able to perceive the degree of internal uncertainty the robot had reliably. Julian Hough, David Schlangen |
HRI | 2 |
| 2017 | Temporal alignment using the incremental unit frameworkabstractWe propose a method for temporal alignment--a precondition of meaningful fusion--of multimodal systems, using the incremental unit dialogue system framework, which gives the system flexibility in how it handles alignment: either by delaying a modality for a specified amount of time, or by revoking (i.e., backtracking) processed information so multiple information sources can be processed jointly. We evaluate our approach in an offline experiment with multimodal data and find that using the incremental framework is flexible and shows promise as a solution to the problem of temporal alignment in multimodal systems. Casey Kennington, Ting Han 0003, David Schlangen |
ICMI | 3 |
| 2017 | Refer-iTTS: A System for Referring in Spoken Installments to Objects in Real-World ImagesabstractCurrent referring expression generation systems mostly deliver their output as one-shot, written expressions. We present on-going work on incremental generation of spoken expressions referring to objects in real-world images. This approach extends upon previous work using the words-as-classifier model for generation. We implement this generator in an incremental dialogue processing framework such that we can exploit an existing interface to incremental text-to-speech synthesis. Our system generates and synthesizes referring expressions while continuously observing non-verbal user reactions. Sina Zarrieß, Soledad López Gambino, David Schlangen |
INLG | 3 |
| 2017 | Towards Deep End-of-Turn Prediction for Situated Spoken Dialogue SystemsabstractThis work was supported by the Cluster of Excellence Cognitive Interaction Technology ‘CITEC’ (EXC 277) at Bielefeld University, funded by the German Research Foundation (DFG), and the DFG-funded DUEL project (grant SCHL 845/5-1). Angelika Maier, Julian Hough, David Schlangen |
INTERSPEECH | 3 |
| 2017 | The Intelligent Coaching Space: A Demonstration
Iwan de Kok, Felix Hülsmann, Thomas Waltemate, Cornelia Frank, Julian Hough, Thies Pfeiffer, David Schlangen, Thomas Schack, Mario Botsch, Stefan Kopp |
IVA | 7 |
| 2017 | Beyond On-hold Messages: Conversational Time-buying in Task-oriented DialogueabstractA common convention in graphical user interfaces is to indicate a "wait state", for example while a program is preparing a response, through a changed cursor state or a progress bar.What should the analogue be in a spoken conversational system?To address this question, we set up an experiment in which a human information provider (IP) was given their information only in a delayed and incremental manner, which systematically created situations where the IP had the turn but could not provide task-related information.Our data analysis shows that 1) IPs bridge the gap until they can provide information by "re-purposing" a whole variety of task-and grounding-related communicative actions (e.g.echoing the user's request, signaling understanding, asserting partially relevant information), rather than being silent or explicitly asking for time (e.g."please wait"), and that 2) IPs combined these actions productively to ensure an ongoing conversation.These results, we argue, indicate that natural conversational interfaces should also be able to manage their time flexibly using a variety of conversational resources. Soledad López Gambino, Sina Zarrieß, David Schlangen |
SIGDIAL Conference | 3 |
| 2017 | A simple generative model of incremental reference resolution for situated dialogue
Casey Kennington, David Schlangen |
Comput. Speech Lang. | 2 |
| 2016 | Resolving References to Objects in Photographs using the Words-As-Classifiers ModelabstractA common use of language is to refer to visually present objects.Modelling it in computers requires modelling the link between language and perception.The "words as classifiers" model of grounded semantics views words as classifiers of perceptual contexts, and composes the meaning of a phrase through composition of the denotations of its component words.It was recently shown to perform well in a game-playing scenario with a small number of object types.We apply it to two large sets of real-world photographs that contain a much larger variety of object types and for which referring expressions are available.Using a pre-trained convolutional neural network to extract image region features, and augmenting these with positional information, we show that the model achieves performance competitive with the state of the art in a reference resolution task (given expression, find bounding box of its referent), while, as we argue, being conceptually simpler and more flexible. David Schlangen, Sina Zarrieß, Casey Kennington |
ACL (1) | 1 |
| 2016 | Easy Things First: Installments Improve Referring Expression Generation for Objects in PhotographsabstractResearch on generating referring expressions has so far mostly focussed on "oneshot reference", where the aim is to generate a single, discriminating expression.In interactive settings, however, it is not uncommon for reference to be established in "installments", where referring information is offered piecewise until success has been confirmed.We show that this strategy can also be advantageous in technical systems that only have uncertain access to object attributes and categories.We train a recently introduced model of grounded word meaning on a data set of REs for objects in images and learn to predict semantically appropriate expressions.In a human evaluation, we observe that users are sensitive to inadequate object names -which unfortunately are not unlikely to be generated from low-level visual input.We propose a solution inspired from human task-oriented interaction and implement strategies for avoiding and repairing semantically inaccurate words.We enhance a word-based REG with contextaware, referential installments and find that they substantially improve the referential success of the system. Sina Zarrieß, David Schlangen |
ACL (1) | 2 |
| 2016 | "Look at Me!": Self-Interruptions as Attention Booster?abstractIn this paper we present results of an exploratory experiment investigating the effects of a contingently self-interrupting vs non-self-interrupting virtual agent who transmits information to a human interaction partner. In the experimental condition self-interruptions of the agent were triggered by an external event whereas in the control group the agent did not react to this event. We measured the effect of the agent's self-interruptions on human attention, memory performance and subjective ratings. In this paper we discuss the results with respect to the design of incremental human-agent dialogue modeling. Birte Richter, David Schlangen, Britta Wrede |
HAI | 2 |
| 2016 | Are you talking to me?: Improving the Robustness of Dialogue Systems in a Multi Party HRI Scenario by Incorporating Gaze Direction and Lip Movement of AttendeesabstractIn this paper, we present our humanoid robot "Meka", participating in a multi party human robot dialogue scenario. Active arbitration of the robot's attention based on multi-modal stimuli is utilised to observe persons which are outside of the robots field of view. We investigate the impact of this attention management and addressee recognition on the robot's capability to distinguish utterances directed at it from communication between humans. Based on the results of a user study, we show that mutual gaze at the end of an utterance, as a means of yielding a turn, is a substantial cue for addressee recognition. Verification of a speaker through the detection of lip movements can be used to further increase precision. Furthermore, we show that even a rather simplistic fusion of gaze and lip movement cues allows a considerable enhancement in addressee estimation, and can be altered to adapt to the requirements of a particular scenario. Viktor Richter, Birte Richter, Florian Lier, Sebastian Meyer zu Borgsen, David Schlangen, Franz Kummert, Sven Wachsmuth, Britta Wrede |
HAI | 5 |
| 2016 | Towards Generating Colour Terms for Referents in Photographs: Prefer the Expected or the Unexpected?abstractColour terms have been a prime phenomenon for studying language grounding, though previous work focussed mostly on descriptions of simple objects or colour swatches.This paper investigates whether colour terms can be learned from more realistic and potentially noisy visual inputs, using a corpus of referring expressions to objects represented as regions in real-world images.We obtain promising results from combining a classifier that grounds colour terms in visual input with a recalibration model that adjusts probability distributions over colour terms according to contextual and object-specific preferences. Sina Zarrieß, David Schlangen |
INLG | 2 |
| 2016 | How to Address Smart Homes with a Social Robot? A Multi-modal Corpus of User Interactions with an Intelligent Environment
Patrick Holthaus, Christian Leichsenring, Jasmin Bernotat, Viktor Richter, Marian Pohling, Birte Richter, Norman Köster, Sebastian Meyer zu Borgsen, René Zorn, Birte Schiffhauer, Kai Frederic Engelmann, Florian Lier, Simon Schulz, Philipp Cimiano, Friederike Eyssel, Thomas Hermann 0001, Franz Kummert, David Schlangen, Sven Wachsmuth, Petra Wagner, Britta Wrede, Sebastian Wrede 0001 |
LREC | 18 |
| 2016 | DUEL: A Multi-lingual Multimodal Dialogue Corpus for Disfluency, Exclamations and Laughter
Julian Hough, Laura E. de Ruiter, Simon Betz, Spyros Kousidis, David Schlangen, Jonathan Ginzburg |
LREC | 6 |
| 2016 | PentoRef: A Corpus of Spoken References in Task-oriented Dialogues
Sina Zarrieß, Julian Hough, Casey Kennington, Ramesh R. Manuvinakurike, David DeVault, Raquel Fernández, David Schlangen |
LREC | 7 |
| 2016 | Investigating Fluidity for Human-Robot Interaction with Real-time, Real-world Grounding StrategiesabstractWe present a simple real-time, real-world grounding framework, and a system which implements it in a simple robot, allowing investigation into different grounding strategies.We put particular focus on the grounding effects of non-linguistic task-related actions.We experiment with a trade-off between the fluidity of the grounding mechanism with the 'safety' of ensuring task success.The framework consists of a combination of interactive Harel statecharts and the Incremental Unit framework.We evaluate its in-robot implementation in a study with human users and find that in simple grounding situations, a model allowing greater fluidity is perceived to have better understanding of the user's speech. Julian Hough, David Schlangen |
SIGDIAL Conference | 2 |
| 2016 | Supporting Spoken Assistant Systems with a Graphical User Interface that Signals Incremental Understanding and Prediction StateabstractArguably, spoken dialogue systems are most often used not in hands/eyes-busy situations, but rather in settings where a graphical display is also available, such as a mobile phone.We explore the use of a graphical output modality for signalling incremental understanding and prediction state of the dialogue system.By visualising the current dialogue state and possible continuations of it as a simple tree, and allowing interaction with that visualisation (e.g., for confirmations or corrections), the system provides both feedback on past user actions and guidance on possible future ones, and it can span the continuum from slot filling to full prediction of user intent (such as GoogleNow).We evaluate our system with real users and report that they found the system intuitive and easy to use, and that incremental and adaptive settings enable users to accomplish more tasks. Casey Kennington, David Schlangen |
SIGDIAL Conference | 2 |
| 2016 | Real-Time Understanding of Complex Discriminative Scene DescriptionsabstractReal-world scenes typically have complex structure, and utterances about them consequently do as well.We devise and evaluate a model that processes descriptions of complex configurations of geometric shapes and can identify the described scenes among a set of candidates, including similar distractors.The model works with raw images of scenes, and by design can work word-by-word incrementally.Hence, it can be used in highly-responsive interactive and situated settings.Using a corpus of descriptions from game-play between human subjects (who found this to be a challenging task), we show that reconstruction of description structure in our system contributes to task success and supports the performance of the word-based model of grounded semantics that we use. Ramesh R. Manuvinakurike, Casey Kennington, David DeVault, David Schlangen |
SIGDIAL Conference | 4 |
| 2016 | Toward incremental dialogue act segmentation in fast-paced interactive dialogue systemsabstractIn this paper, we present and evaluate an approach to incremental dialogue act (DA) segmentation and classification.Our approach utilizes prosodic, lexico-syntactic and contextual features, and achieves an encouraging level of performance in offline corpus-based evaluation as well as in simulated human-agent dialogues.Our approach uses a pipeline of sequential processing steps, and we investigate the contribution of different processing steps to DA segmentation errors.We present our results using both existing and new metrics for DA segmentation.The incremental DA segmentation capability described here may help future systems to allow more natural speech from users and enable more natural patterns of interaction. Ramesh R. Manuvinakurike, Maike Paetzel-Prüsmann, Cheng Qu, David Schlangen, David DeVault |
SIGDIAL Conference | 4 |
| 2015 | Simple Learning and Compositional Application of Perceptually Grounded Word Meanings for Incremental Reference ResolutionabstractCasey Kennington, David Schlangen. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Casey Kennington, David Schlangen |
ACL (1) | 2 |
| 2015 | A Multimodal System for Real-Time Action Instruction in Motor Skill LearningabstractWe present a multimodal coaching system that supports online motor skill learning. In this domain, closed-loop interaction between the movements of the user and the action instructions by the system is an essential requirement. To achieve this, the actions of the user need to be measured and evaluated and the system must be able to give corrective instructions on the ongoing performance. Timely delivery of these instructions, particularly during execution of the motor skill by the user, is thus of the highest importance. Based on the results of an empirical study on motor skill coaching, we analyze the requirements for an interactive coaching system and present an architecture that combines motion analysis, dialogue management, and virtual human animation in a motion tracking and 3D virtual reality hardware setup. In a preliminary study we demonstrate that the current system is capable of delivering the closed-loop interaction that is required in the motor skill learning domain. Iwan de Kok, Julian Hough, Felix Hülsmann, Mario Botsch, David Schlangen, Stefan Kopp |
ICMI | 5 |
| 2015 | Micro-structure of disfluencies: basics for conversational speech synthesisabstractBetz S, Wagner P, Schlangen D. Micro-Structure of Disfluencies: Basics for Conversational Speech Synthesis. In: Interspeech 2015. 2015: 2222-2226. Simon Betz, Petra Wagner, David Schlangen |
INTERSPEECH | 3 |
| 2015 | Recurrent neural networks for incremental disfluency detectionabstractFor dialogue systems to become robust, they must be able to detect disfluencies accurately and with minimal latency.To meet this challenge, here we frame incremental disfluency detection as a word-by-word tagging task and, following their recent success in Spoken Language Understanding tasks, we test the performance of Recurrent Neural Networks (RNNs).We experiment with different inputs for RNNs to explore the effect of context on their ability to detect edit terms and repair disfluencies effectively.Although not eclipsing the state of the art in terms of utterance-final performance, RNNs achieve good detection results, requiring no feature engineering and using simple input vectors representing the incoming utterance as their training input.Furthermore, RNNs show very good incremental properties with low latency and very good output stability, surpassing previously reported results in these measures. Julian Hough, David Schlangen |
INTERSPEECH | 2 |
| 2015 | Incrementally Tracking Reference in Human/Human Dialogue Using Linguistic and Extra-Linguistic InformationabstractCasey Kennington, Ryu Iida, Takenobu Tokunaga, David Schlangen. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Casey Kennington, Ryu Iida, Takenobu Tokunaga, David Schlangen |
HLT-NAACL | 4 |
| 2014 | Better Driving and Recall When In-car Information Presentation Uses Situationally-Aware Incremental Speech Output GenerationabstractIt is established that driver distraction is the result of sharing cognitive resources between the primary task (driving) and any other secondary task. In the case of holding conversations, a human passenger who is aware of the driving conditions can choose to interrupt his speech in situations potentially requiring more attention from the driver, but in-car information systems typically do not exhibit such sensitivity. We have designed and tested such a system in a driving simulation environment. Unlike other systems, our system delivers information via speech (calendar entries with scheduled meetings) but is able to react to signals from the environment to interrupt when the driver needs to be fully attentive to the driving task and subsequently resume its delivery. Distraction is measured by a secondary short-term memory task. In both tasks, drivers perform significantly worse when the system does not adapt its speech, while they perform equally well to control conditions (no concurrent task) when the system intelligently interrupts and resumes. Casey Kennington, Spyros Kousidis, Timo Baumann, Hendrik Buschmeier, Stefan Kopp, David Schlangen |
AutomotiveUI | 6 |
| 2014 | Situated Incremental Natural Language Understanding using a Multimodal, Linguistically-driven Update Model
Casey Kennington, Spyros Kousidis, David Schlangen |
COLING | 3 |
| 2014 | A Multimodal In-Car Dialogue System That Tracks The Driver's AttentionabstractWhen a passenger speaks to a driver, he or she is co-located with the driver, is generally aware of the situation, and can stop speaking to allow the driver to focus on the driving task. In-car dialogue systems ignore these important aspects, making them more distracting than even cell-phone conversations. We developed and tested a "situationally-aware" dialogue system that can interrupt its speech when a situation which requires more attention from the driver is detected, and can resume when driving conditions return to normal. Furthermore, our system allows driver-controlled resumption of interrupted speech via verbal or visual cues (head nods). Over two experiments, we found that the situationally-aware spoken dialogue system improves driving performance and attention to the speech content, while driver-controlled speech resumption does not hinder performance in either of these two tasks Spyros Kousidis, Casey Kennington, Timo Baumann, Hendrik Buschmeier, Stefan Kopp, David Schlangen |
ICMI | 6 |
| 2014 | InproTKs: A Toolkit for Incremental Situated ProcessingabstractIn order to process incremental situated dialogue, it is necessary to accept information from various sensors, each tracking, in real-time, different aspects of the physical situation.We present extensions of the incremental processing toolkit IN-PROTK which make it possible to plug in such multimodal sensors and to achieve situated, real-time dialogue.We also describe a new module which enables the use in INPROTK of the Google Web Speech API, which offers speech recognition with a very large vocabulary and a wide choice of languages.We illustrate the use of these extensions with a description of two systems handling different situated settings. Casey Kennington, Spyros Kousidis, David Schlangen |
SIGDIAL Conference | 3 |
| 2014 | Situated incremental natural language understanding using Markov Logic Networks
Casey Kennington, David Schlangen |
Comput. Speech Lang. | 2 |
| 2013 | MINT.tools: tools and adaptors supporting acquisition, annotation and analysis of multimodal corporaabstractKousidis S, Pfeiffer T, Schlangen D. MINT.tools: Tools and Adaptors Supporting Acquisition, Annotation and Analysis of Multimodal Corpora. In: Proceedings of Interspeech 2013. ISCA; 2013. Spyros Kousidis, Thies Pfeiffer, David Schlangen |
INTERSPEECH | 3 |
| 2013 | A cross-linguistic study on turn-taking and temporal alignment in verbal interactionabstractKousidis S, Schlangen D, Skopeteas S. A cross-linguistic study on turn-taking and temporal alignment in verbal interaction. In: Proceedings of Interspeech 2013. 2013. Spyros Kousidis, David Schlangen, Stavros Skopeteas |
INTERSPEECH | 2 |
| 2013 | Open-ended, Extensible System Utterances Are Preferred, Even If They Require Filled Pauses
Timo Baumann, David Schlangen |
SIGDIAL Conference | 2 |
| 2013 | Interpreting Situated Dialogue Utterances: an Update Model that Uses Speech, Gaze, and Gesture Information
Casey Kennington, Spyros Kousidis, David Schlangen |
SIGDIAL Conference | 3 |
| 2013 | Investigating speaker gaze and pointing behaviour in human-computer interaction with the mint.tools collection
Spyros Kousidis, Casey Kennington, David Schlangen |
SIGDIAL Conference | 3 |
| 2012 | Joint Satisfaction of Syntactic and Pragmatic Constraints Improves Incremental Spoken Language Understanding
Andreas Peldszus, Okko Buß, Timo Baumann, David Schlangen |
EACL | 4 |
| 2012 | Evaluating Prosodic Processing for Incremental Speech SynthesisabstractBaumann T, Schlangen D. Evaluating Prosodic Processing for Incremental Speech Synthesis. In: Proceedings of Interspeech. 2012. Timo Baumann, David Schlangen |
INTERSPEECH | 2 |
| 2012 | Combining Incremental Language Generation and Incremental Speech Synthesis for Adaptive Information Presentation
Hendrik Buschmeier, Timo Baumann, Benjamin Dosch, Stefan Kopp, David Schlangen |
SIGDIAL Conference | 5 |
| 2012 | Markov Logic Networks for Situated Incremental Natural Language Understanding
Casey Kennington, David Schlangen |
SIGDIAL Conference | 2 |
| 2011 | Predicting the Micro-Timing of User Input for an Incremental Spoken Dialogue System that Completes a User's Ongoing Turn
Timo Baumann, David Schlangen |
SIGDIAL Conference | 2 |
| 2010 | Collaborating on Utterances with a Spoken Dialogue System Using an ISU-based Approach to Incremental Dialogue Management
Okko Buß, Timo Baumann, David Schlangen |
SIGDIAL Conference | 3 |
| 2010 | Comparing Local and Sequential Models for Statistical Incremental Natural Language Understanding
Silvan Heintze, Timo Baumann, David Schlangen |
SIGDIAL Conference | 3 |
| 2010 | Middleware for Incremental Processing in Conversational Agents
David Schlangen, Timo Baumann, Hendrik Buschmeier, Okko Buß, Stefan Kopp, Gabriel Skantze, Ramin Yaghoubzadeh |
SIGDIAL Conference | 1 |
| 2009 | A General, Abstract Model of Incremental Dialogue Processing
David Schlangen, Gabriel Skantze |
EACL | 1 |
| 2009 | Incremental Dialogue Processing in a Micro-Domain
Gabriel Skantze, David Schlangen |
EACL | 2 |
| 2009 | No sooner said than done? testing incrementality of semantic interpretations of spontaneous speechabstractIdeally, a spoken dialogue system should react without much delay to a user's utterance.Such a system would already select an object, for instance, before the user has finished her utterance about moving this particular object to a particular place.A prerequisite for such a prompt reaction is that semantic representations are built up on the fly and passed on to other modules.Few approaches to incremental semantics construction exist, and, to our knowledge, none of those has been systematically tested on a spontaneous speech corpus.In this paper, we develop measures to test empirically on transcribed spontaneous speech to what extent we can create semantic interpretation on the fly with an incremental semantic chunker that builds a frame semantics. Michaela Atterer, Timo Baumann, David Schlangen |
INTERSPEECH | 3 |
| 2009 | Evaluating the potential utility of ASR n-best lists for incremental spoken dialogue systemsabstractThe potential of using ASR n-best lists for dialogue systems has often been recognised (if less often realised): it is often the case that even when the top-ranked hypothesis is erroneous, a better one can be found at a lower rank.In this paper, we describe metrics for evaluating whether the same potential carries over to incremental dialogue systems, where ASR output is consumed and reacted upon while speech is still ongoing.We show that even small N can provide an advantage for semantic processing, at a cost of a computational overhead. Timo Baumann, Okko Buß, Michaela Atterer, David Schlangen |
INTERSPEECH | 4 |
| 2009 | Assessing and Improving the Performance of Speech Recognition for Incremental Systems
Timo Baumann, Michaela Atterer, David Schlangen |
HLT-NAACL | 3 |
| 2009 | TELIDA: A Package for Manipulation and Visualization of Timed Linguistic Data
Titus von der Malsburg, Timo Baumann, David Schlangen |
SIGDIAL Conference | 3 |
| 2009 | Incremental Reference Resolution: The Task, Metrics for Evaluation, and a Bayesian Filtering Model that is Sensitive to Disfluencies
David Schlangen, Timo Baumann, Michaela Atterer |
SIGDIAL Conference | 1 |
| 2007 | Speaking through a noisy channel - experiments on inducing clarification behaviour in human-human dialogueabstractWe report results of an experiment on inducing communication problems in human-human dialogue.We set up a voice-only cooperative task where we manipulated one channel by replacing (in real-time, at random points) all signal with noise.Altogether around 10% of the speaker's signal was thus removed.We found an increase in clarification requests of a form that has previously been hypothesised to be used mainly for clarifying acoustic problems.We also found a correlation between the percentage of an utterance being manipulated and the use of devices for pointing out error locations.From our findings, we derive a gold-standard policy for clarification behaviour. David Schlangen, Raquel Fernández |
INTERSPEECH | 1 |
| 2006 | From reaction to prediction: experiments with computational models of turn-takingabstractDeciding when to take (or not to take) the turn in a conversation is an important task.It has been stressed in the descriptive literature that such decisions must involve prediction, as they often seem to be made before a transition place has been reached.In computational systems, however, turn-taking is normally a reaction to parameters like pause length.In this paper, we report on experiments that try to bridge this gap.We describe an experiment (using controlled stimuli) that shows human performance at prediction of turn-taking decisions and then show that a model automatically induced from data can reach a similar level of performance.We then describe a series of experiments on spontaneous dialogue data where we combine pause thresholds with syntactic and prosodic information to make turn-taking decisions, successively reducing the pause threshold until reaction becomes prediction.All our classifiers improve significantly over the baselines; prediction however is shown to be the hardest task, and we discuss additional information sources that could improve it. David Schlangen |
INTERSPEECH | 1 |
| 2006 | Interaction in Task-Oriented Human-Human Dialogue: the Effects of Different turn-Taking PoliciesabstractIn human-human dialogue, the allocation of turns between the participants is normally managed smoothly, without the participants paying much attention to it. In contrast, for spoken dialogue systems turn allocation is a difficult task, and often technical restrictions are introduced to simplify it. In this paper we investigate, by comparing two experimentally collected corpora of human-human task oriented dialogue, what the consequences are of imposing one particular kind of restriction, namely that of using a simplex channel managed by push-to-talk (PTT). We found, as expected, a loss of interactivity in the PTT condition (fewer, longer turns; more silences), but surprisingly, no loss of efficiency; in fact, the subjects in the PTT condition were able to finish their task in roughly the same time, while using fewer words along the way. We analyse here the differences in the interaction patterns and the interplay of 'naturalness' and efficiency as relevant factors for practical system development. Raquel Fernández, Tatjana Lucht, Kepa Joseba Rodríguez, David Schlangen |
SLT | 4 |
| 2005 | Towards Finding and Fixing Fragments-Using ML to Identify Non-Sentential Utterances and their Antecedents in Multi-Party DialogueabstractNon-sentential utterances (e.g., short-answers as in "Who came to the party?"--- "Peter.") are pervasive in dialogue. As with other forms of ellipsis, the elided material is typically present in the context (e.g., the question that a short answer answers). We present a machine learning approach to the novel task of identifying fragments and their antecedents in multiparty dialogue. We compare the performance of several learning algorithms, using a mixture of structural and lexical features, and show that the task of identifying antecedents given a fragment can be learnt successfully (f(0.5) = .76); we discuss why the task of identifying fragments is harder (f(0.5) = .41) and finally report on a combined task (f(0.5) = .38). David Schlangen |
ACL | 1 |