VLDB 2026 Research / reviewers in the wild / expert
Svetlana Stoyanchev
dblp:99/7516
· DBLP profile ↗
29ranked-venue papers
13as first author
9since 2021 · last 2025
0000-0002-8079-1652ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 9 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Human ratings of LLM response generation in pair-programming dialogueabstractWe take first steps in exploring whether Large Language Models (LLMs) can be adapted to dialogic learning practices, specifically pair programming — LLMs have primarily been implemented as programming assistants, not fully exploiting their dialogic potential. We used new dialogue data from real pair-programming interactions between students, prompting state-of-the-art LLMs to assume the role of a student, when generating a response that continues the real dialogue. We asked human annotators to rate human and AI responses on the criteria through which we operationalise the LLMs’ suitability for educational dialogue: Coherence, Collaborativeness, and whether they appeared human. Results show model differences, with Llama-generated responses being rated similarly to human answers on all three criteria. Thus, for at least one of the models we investigated, the LLM utterance-level response generation appears to be suitable for pair-programming dialogue. Cecilia Domingo, Paul Piwek, Svetlana Stoyanchev, Michel Wermelinger, Kaustubh Adhikari, Rama Sanand Doddipatla |
INLG | 3 |
| 2024 | Semantic Map-based Generation of Navigation InstructionsabstractWe are interested in the generation of navigation instructions, either in their own right or as training material for robotic navigation task. In this paper, we propose a new approach to navigation instruction generation by framing the problem as an image captioning task using semantic maps as visual input. Conventional approaches employ a sequence of panorama images to generate navigation instructions. Semantic maps abstract away from visual details and fuse the information in multiple panorama images into a single top-down representation, thereby reducing computational complexity to process the input. We present a benchmark dataset for instruction generation using semantic maps, propose an initial model and ask human subjects to manually assess the quality of generated instructions. Our initial investigations show promise in using semantic maps for instruction generation instead of a sequence of panorama images, but there is vast scope for improvement. We release the code for data preparation and model training at https://github.com/chengzu-li/VLGen. Chengzu Li, Simone Teufel, Rama Sanand Doddipatla, Svetlana Stoyanchev |
LREC/COLING | 5 |
| 2024 | WHISMA: A Speech-LLM to Perform Zero-Shot Spoken Language UnderstandingabstractSpeech large language models (speech-LLMs) integrate speech and text-based foundation models to provide a unified framework for handling a wide range of downstream tasks. In this paper, we introduce WHISMA, a speech-LLM tailored for spoken language understanding (SLU) that demonstrates robust performance in various zero-shot settings. WHISMA combines the speech encoder from Whisper with the Llama-3 LLM, and is fine-tuned in a parameter-efficient manner on a comprehensive collection of SLU-related datasets. Our experiments show that WHISMA significantly improves the zero-shot slot filling performance on the SLURP benchmark, achieving a relative gain of 26.6% compared to the current state-of-the-art model. Furthermore, to evaluate WHISMA’s generalisation capabilities to unseen domains, we develop a new task-agnostic benchmark named SLU-GLUE. The evaluation results indicate that WHISMA outperforms an existing speech-LLM (Qwen-Audio) with a relative gain of 33.0%. Mohan Li, Cong-Thanh Do, Simon Keizer, Youmna Farag, Svetlana Stoyanchev, Rama Sanand Doddipatla |
SLT | 5 |
| 2024 | Entity Resolution in Situated Dialog With Unimodal and Multimodal TransformersabstractIn this work we address the entity resolution task for situated multimodal dialog investigating how a unimodal approach, which uses only textual information as input (representing visual attributes as text), compares to a multimodal system, which processes both text and visual information. We analyze two of the top performing models presented in the Tenth Dialog Systems Technology Challenge and propose modifications that enhance their performance on the multimodal coreference resolution task. We evaluate these approaches on in- and out-of-domain settings by training the models on the fashion domain and testing on the furniture domain, and vice-versa, to assess the generalizability of the models. Through systematic analysis, we show that while both systems achieve similar performance on in-domain scenarios, the multimodal system generalizes better to out-of-domain settings. A combination strategy of enhanced unimodal and multimodal systems achieves F1 = 0.80 (5% absolute gain compared to the best performing system). Finally, human performance on the same task is evaluated on a small subset, suggesting that the performance of the current automatic models is on par with people on this task. Alejandro Santorum Varela, Svetlana Stoyanchev, Simon Keizer, Rama Sanand Doddipatla, Kate M. Knill |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Combining Structured and Unstructured Knowledge in an Interactive Search Dialogue SystemabstractUsers of interactive search dialogue systems specify their preferences with natural language utterances.However, a schema-driven system is limited to handling the preferences that correspond to the predefined database content.In this work, we present a methodology for extending a schema-driven interactive search dialogue system with the ability to handle unconstrained user preferences.Using unsupervised semantic similarity metrics and text snippets associated with the search items, the system identifies suitable items for the user's unconstrained natural language query.In a crowd-sourced evaluation, the users were asked to chat with our extended restaurant search system.Based on objective metrics and subjective user ratings, we demonstrate the feasibility of using this unsupervised low latency approach to extend a schema-driven search dialogue system to handle unconstrained user preferences. Svetlana Stoyanchev, Suraj Pandey, Simon Keizer, Norbert Braunschweiler, Rama Sanand Doddipatla |
SIGDIAL | 1 |
| 2022 | Factors in Emotion Recognition With Deep Learning Models Using Speech and Text on Multiple CorporaabstractEmotion recognition performance of deep learning models is influenced by multiple factors such as acoustic condition, textual content, style of emotion expression (e.g. acted, natural), etc. In this paper, multiple factors are analysed by training and evaluating state-of-the-art deep learning models using the input modalities speech, text, and their combination across 6 emotional speech corpora. A novel deep learning model architecture is presented that further improves the state-of-the-art in multimodal emotion recognition with speech and text on the IEMOCAP corpus. Results from models trained on individual corpora show that combining speech and text improves performance only on corpora where the text of utterances varies across different emotions, while it reduced performance on corpora with fixed text expressed in different emotions, where the speech-only models performed better. Further, cross-corpus investigations are presented to understand the robustness to changing acoustic and textual content. Results show that models perform significantly better in matched conditions in particular single corpus models perform better than multi-corpus models, with the latter showing a tendency to be more robust to acoustic variations, while performance still depends on characteristics of both training corpora and test corpus. Norbert Braunschweiler, Rama Sanand Doddipatla, Simon Keizer, Svetlana Stoyanchev |
IEEE Signal Process. Lett. | 4 |
| 2021 | A Study on Cross-Corpus Speech Emotion Recognition and Data AugmentationabstractModels that can handle a wide range of speakers and acoustic conditions are essential in speech emotion recognition (SER). Often, these models tend to show mixed results when presented with speakers or acoustic conditions that were not visible during training. This paper investigates the impact of cross-corpus data complementation and data augmentation on the performance of SER models in matched (test-set from same corpus) and mismatched (test-set from different corpus) conditions. Investigations using six emotional speech corpora that include single and multiple speakers as well as variations in emotion style (acted, elicited, natural) and recording conditions are presented. Observations show that, as expected, models trained on single corpora perform best in matched conditions while performance decreases between 10-40% in mismatched conditions, depending on corpus specific features. Models trained on mixed corpora can be more stable in mismatched contexts, and the performance reductions range from 1 to 8% when compared with single corpus models in matched conditions. Data augmentation yields additional gains up to 4% and seem to benefit mismatched conditions more than matched ones. Norbert Braunschweiler, Rama Sanand Doddipatla, Simon Keizer, Svetlana Stoyanchev |
ASRU | 4 |
| 2021 | Dialogue Strategy Adaptation to New Action Sets Using Multi-Dimensional ModellingabstractA major bottleneck for building statistical spoken dialogue systems for new domains and applications is the need for large amounts of training data. To address this problem, we adopt the multi-dimensional approach to dialogue management and evaluate its potential for transfer learning. Specifically, we exploit pre-trained task-independent policies to speed up training for an extended task-specific action set, in which the single summary action for requesting a slot is replaced by multiple slot-specific request actions. Policy optimisation and evaluation experiments using an agenda-based user simulator show that with limited training data, much better performance levels can be achieved when using the proposed multi-dimensional adaptation method. We confirm this improvement in a crowd-sourced human user evaluation of our spoken dialogue system, comparing partially trained policies. The multi-dimensional system (with adaptation on limited training data in the target scenario) outperforms the one-dimensional baseline (without adaptation on the same amount of training data) by 7% perceived success rate. Simon Keizer, Norbert Braunschweiler, Svetlana Stoyanchev, Rama Sanand Doddipatla |
ASRU | 3 |
| 2021 | Action State Update Approach to Dialogue ManagementabstractUtterance interpretation is one of the main functions of a dialogue manager, which is the key component of a dialogue system. We propose the action state update approach (ASU) for utterance interpretation, featuring a statistically trained binary classifier used to detect dialogue state update actions in the text of a user utterance. Our goal is to interpret referring expressions in user input without a domain-specific natural language understanding component. For training the model, we use active learning to automatically select simulated training examples. With both user-simulated and interactive human evaluations, we show that the ASU approach successfully interprets user utterances in a dialogue system, including those with referring expressions. Svetlana Stoyanchev, Simon Keizer, Rama Sanand Doddipatla |
ICASSP | 1 |
| 2018 | Corpus and Annotation Towards NLU for Customer Ordering DialogsabstractOrdering products and services through virtual agents is possible but suffers limitations on the kind of ordering that is possible or on the naturalness of the conversation. We address these limitations by collecting a corpus of human-human dialogs in the food ordering domain. We create a food focused annotation scheme that is tailored for this corpus but customizable for other applications. After annotating the corpus, we find corpus characteristics that may make it more natural, such as complexity of food item mentions and use of multiple intent utterances. Furthermore, we train and evaluate preliminary statistical item and intent models using the annotated corpus. John Chen 0001, Rashmi Prasad, Svetlana Stoyanchev, Ethan Selfridge, Srinivas Bangalore, Michael Johnston |
SLT | 3 |
| 2016 | Evaluation of Semantic Dependency Labeling Across Domains
Svetlana Stoyanchev, Amanda Stent, Srinivas Bangalore |
AAAI | 1 |
| 2016 | Rapid Prototyping of Form-driven Dialogue Systems Using an Open-source FrameworkabstractMost human-machine communication for information access through speech, text and graphical interfaces are mediated by forms -i.e.lists of named fields.However, deploying form-filling dialogue systems still remains a challenging task due to the effort and skill required to author such systems.We describe an extension to the OpenDial framework that enables the rapid creation of functional dialogue systems by non-experts.The dialogue designer specifies the slots and their types as input and the tool generates a domain specification that drives a slot-filling dialogue system.The presented approach provides several benefits compared to traditional techniques based on flowcharts, such as the use of probabilistic reasoning and flexible grounding strategies. Svetlana Stoyanchev, Pierre Lison, Srinivas Bangalore |
SIGDIAL Conference | 1 |
| 2015 | Localized error detection for targeted clarification in a virtual assistantabstractWe propose a novel approach for addressing automatic speech recognition (ASR) and natural language understanding (NLU) errors in an interactive spoken dialog system using targeted clarification (TC). TC applies when a spoken utterance is partially recognized by focusing a clarification question on the misrecognized part of the utterance. A key component of TC is accurate detection of localized ASR and NLU errors in an utterance. In this work, we develop statistical models of presence and correctness for domain concepts within an ASR/NLU result and use these to drive a targeted clarification (TC) strategy. We evaluate the accuracy of the models and their effect on the dialog strategy in an interactive multimodal assistant. Svetlana Stoyanchev, Michael Johnston |
ICASSP | 1 |
| 2014 | Dialogue Act Modeling for Non-Visual Web AccessabstractSpeech-enabled dialogue systems have the potential to enhance the ease with which blind individuals can interact with the Web beyond what is possible with screen read-ers- the currently available assistive tech-nology which narrates the textual content on the screen and provides shortcuts to navigate the content. In this paper, we present a dialogue act model towards de-veloping a speech enabled browsing sys-tem. The model is based on the corpus data that was collected in a wizard-of-oz study with 24 blind individuals who were assigned a gamut of browsing tasks. The development of the model included exten-sive experiments with assorted feature sets and classifiers; the outcomes of the exper-iments and the analysis of the results are presented. 1 Vikas Ashok, Yevgen Borodin, Svetlana Stoyanchev, I. V. Ramakrishnan |
SIGDIAL Conference | 3 |
| 2014 | MVA: The Multimodal Virtual AssistantabstractMichael Johnston, John Chen, Patrick Ehlen, Hyuckchul Jung, Jay Lieske, Aarthi Reddy, Ethan Selfridge, Svetlana Stoyanchev, Brant Vasilieff, Jay Wilpon. Proceedings of the 15th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL). 2014. Michael Johnston, John Chen 0001, Patrick Ehlen, Hyuckchul Jung, Jay Lieske, Aarthi M. Reddy, Ethan Selfridge, Svetlana Stoyanchev, Brant Vasilieff, Jay G. Wilpon |
SIGDIAL Conference | 8 |
| 2014 | Detecting Inappropriate Clarification Requests in Spoken Dialogue SystemsabstractAlex Liu, Rose Sloan, Mei-Vern Then, Svetlana Stoyanchev, Julia Hirschberg, Elizabeth Shriberg. Proceedings of the 15th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL). 2014. Rose Sloan, Mei-Vern Then, Svetlana Stoyanchev, Julia Hirschberg, Elizabeth Shriberg |
SIGDIAL Conference | 4 |
| 2013 | "Can you give me another word for hyperbaric?": Improving speech translation using targeted clarification questionsabstractWe present a novel approach for improving communication success between users of speech-to-speech translation systems by automatically detecting errors in the output of automatic speech recognition (ASR) and statistical machine translation (SMT) systems. Our approach initiates system-driven targeted clarification about errorful regions in user input and repairs them given user responses. Our system has been evaluated by unbiased subjects in live mode, and results show improved success of communication between users of the system. Necip Fazil Ayan, Arindam Mandal, Michael W. Frandsen, Jing Zheng 0001, Peter Blasco, Andreas Kathol, Frédéric Béchet, Benoît Favre, Alex Marin, Tom Kwiatkowski, Mari Ostendorf, Luke Zettlemoyer, Philipp Salletmayr, Julia Hirschberg, Svetlana Stoyanchev |
ICASSP | 15 |
| 2013 | Exploring Features For Localized Detection of Speech Recognition Errors
Eli Pincus, Svetlana Stoyanchev, Julia Hirschberg |
SIGDIAL Conference | 2 |
| 2013 | Modelling Human Clarification Strategies
Svetlana Stoyanchev, Julia Hirschberg |
SIGDIAL Conference | 1 |
| 2012 | Fully Automated Generation of Question-Answer Pairs for Scripted Virtual Instruction
Pascal Kuyten, Timothy W. Bickmore, Svetlana Stoyanchev, Paul Piwek, Helmut Prendinger, Mitsuru Ishizuka |
IVA | 3 |
| 2012 | Localized detection of speech recognition errorsabstractWe address the problem of localized error detection in Automatic Speech Recognition (ASR) output. Localized error detection seeks to identify which particular words in a user's utterance have been misrecognized. Identifying misrecognized words permits one to create targeted clarification strategies for spoken dialogue systems, allowing the system to ask clarification questions targeting the particular type of misrecognition, in contrast to the “please repeat/rephrase” strategies used in most current dialogue systems. We present results of machine learning experiments using ASR confidence scores together with prosodic and syntactic features to predict whether 1) an utterance contains an error, and 2) whether a word in a misrecognized utterance is misrecognized. We show that by adding syntactic features to the ASR features when predicting misrecognized utterances the F-measure improves by 13.3% compared to using ASR features alone. By adding syntactic and prosodic features when predicting misrecognized words F-measure improves by 40%. Svetlana Stoyanchev, Philipp Salletmayr, Julia Hirschberg |
SLT | 1 |
| 2011 | Comparing Modes of Information Presentation: Text versus ECA and Single versus Two ECAs
Svetlana Stoyanchev, Paul Piwek, Helmut Prendinger |
IVA | 1 |
| 2011 | The CODA System for Monologue-to-Dialogue Generation
Svetlana Stoyanchev, Paul Piwek |
SIGDIAL Conference | 1 |
| 2010 | The First Question Generation Shared Task Evaluation Challenge
Vasile Rus, Brendan Wyse, Paul Piwek, Mihai C. Lintean, Svetlana Stoyanchev, Cristian Moldovan |
INLG | 5 |
| 2010 | Harvesting Re-usable High-level Rules for Expository Dialogue Generation
Svetlana Stoyanchev, Paul Piwek |
INLG | 1 |
| 2010 | Constructing the CODA Corpus: A Parallel Corpus of Monologues and Expository Dialogues
Svetlana Stoyanchev, Paul Piwek |
LREC | 1 |
| 2010 | Generating Expository Dialogue from Monologue: Motivation, Corpus and Preliminary Rules
Paul Piwek, Svetlana Stoyanchev |
HLT-NAACL | 2 |
| 2009 | Concept Form Adaptation in Human-Computer Dialog
Svetlana Stoyanchev, Amanda Stent |
SIGDIAL Conference | 1 |
| 2008 | Name-aware speech recognition for interactive question answeringabstractIn this work we show how interactivity in a voice-enabled question answering application may improve speech recognition. We allow the user to provide a target named entity before asking the question. Then we build a named entity specific language model using the documents containing the named entity. The question-specific model is obtained by merging the named entity specific model with the model built on a set of questions. We present a set of experiments using the TREC question set on the AQUAINT corpus. The question-specific language model is compared with the baseline model built by merging a model of the AQUAINT corpus and past TREC questions. The question-specific model achieves 32.2% reduction in word error rate from the baseline using the questions where pronominal references are resolved. Svetlana Stoyanchev, Gökhan Tür, Dilek Hakkani-Tür |
ICASSP | 1 |