Ondrej Dusek

dblp:126/8739 · DBLP profile ↗
← Back
48ranked-venue papers
9as first author
26since 2021 · last 2026
0000-0002-1415-1702ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 47 · 9 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Reasoning Gets Harder for LLMs Inside A Dialogue
abstract
Large Language Models (LLMs) achieve strong performance on many reasoning benchmarks, yet these evaluations typically focus on isolated tasks that differ from real-world usage in task-oriented dialogue (TOD).In this setting, LLMs must perform reasoning inherently while generating text and adhering to instructions on role, format, and style.This mismatch raises concerns about whether benchmark performance accurately reflects models' reasoning robustness in TOD setting.We investigate how framing reasoning tasks within TOD affects LLM performance by introducing BOUL-DER, a new dynamic benchmark covering eight travel-related tasks that require arithmetic, spatial, and temporal reasoning with both commonsense and formal aspects.Each problem is presented in both isolated and dialogue-based variants, enabling controlled comparison while mitigating data contamination.Experiments on eight LLMs reveal a substantial and consistent performance gap between isolated and dialogue settings.Through ablations and qualitative analysis, we show that this gap is largely driven by the multi-turn nature of dialogue, with additional effects from role conditioning and tool-use requirements.Our results highlight the need to evaluate LLM reasoning in realistic interactive scenarios.1
Ivan Kartác, Mateusz Lango, Ondrej Dusek
ACL (1)3
2026 HotelCheckSpan: A Benchmark Dataset for LLM Faithfulness
Patricia Schmidtová, Ondrej Dusek, Saad Mahamood
LREC2
2025 OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs
abstract
Large Language Models (LLMs) have demonstrated great potential as evaluators of NLG systems, allowing for high-quality, reference-free, and multi-aspect assessments. However, existing LLM-based metrics suffer from two major drawbacks: reliance on proprietary models to generate training data or perform evaluations, and a lack of fine-grained, explanatory feedback. We introduce OpeNLGauge, a fully open-source, reference-free NLG evaluation metric that provides accurate explanations based on individual error spans. OpeNLGauge is available as a two-stage ensemble of larger open-weight LLMs, or as a small fine-tuned evaluation model, with confirmed generalizability to unseen tasks, domains and aspects. Our extensive meta-evaluation shows that OpeNLGauge achieves competitive correlation with human judgments, outperforming state-of-the-art models on certain tasks while maintaining full reproducibility and providing explanations more than twice as accurate.
Ivan Kartác, Mateusz Lango, Ondrej Dusek
INLG3
2025 When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition
abstract
The increasing trust in large language models (LLMs), especially in the form of chatbots, is often undermined by the lack of their extrinsic evaluation. This holds particularly true in nutrition, where randomised controlled trials (RCTs) are the gold standard, and experts demand them for evidence-based deployment. LLMs have shown promising results in this field, but these are limited to intrinsic setups. We address this gap by running the first RCT involving LLMs for nutrition. We augment a rule-based chatbot with two LLM-based features: (1) message rephrasing for conversational variety and engagement, and (2) nutritional counselling through a fine-tuned model. In our seven-week RCT (n=81), we compare chatbot variants with and without LLM integration. We measure effects on dietary outcome, emotional well-being, and engagement. Despite our LLM-based features performing well in intrinsic evaluation, we find that they did not yield consistent benefits in real-world deployment. These results highlight critical gaps between intrinsic evaluations and real-world impact, emphasising the need for interdisciplinary, human-centred approaches.
Karen Jia-Hui Li, Simone Balloccu, Ondrej Dusek, Ehud Reiter
INLG3
2025 FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation
abstract
Table-to-text generation (insight generation from tables) is a challenging task that requires precision in analyzing the data. In addition, the evaluation of existing benchmarks is affected by contamination of Large Language Model (LLM) training data as well as domain imbalance. We introduce FreshTab, an on-the-fly table-to-text benchmark generation from Wikipedia, to combat the LLM data contamination problem and enable domain-sensitive evaluation. While non-English table-to-text datasets are limited, FreshTab collects datasets in different languages on demand (we experiment with German, Russian and French in addition to English). We find that insights generated by LLMs from recent tables collected by our method appear clearly worse by automatic metrics, but this does not translate into LLM and human evaluations. Domain effects are visible in all evaluations, showing that a domain-balanced benchmark is more challenging.
Kristýna Onderková, Ondrej Plátek, Zdenek Kasner, Ondrej Dusek
INLG4
2025 Do My Eyes Deceive Me? A Survey of Human Evaluations of Hallucinations in NLG
abstract
Hallucinations are one of the most pressing challenges for large language models (LLMs). While numerous methods have been proposed to detect and mitigate them automatically, human evaluation continues to serve as the gold standard. However, these human evaluations of hallucinations show substantial variation in definitions, terminology, and evaluation practices. In this paper, we survey 64 studies involving human evaluation of hallucination published between 2019 and 2024, to investigate how hallucinations are currently defined and assessed. Our analysis reveals a lack of consistency in definitions and exposes several concerning methodological shortcomings. Crucial details, such as evaluation guidelines, user interface design, inter-annotator agreement metrics, and annotator demographics, are frequently under-reported or omitted altogether.
Patrícia Schmidtová, Eduardo Calò, Simone Balloccu, Dimitra Gkatzia, Rudali Huidrom, Mateusz Lango, Fahime Same, Vilém Zouhar, Saad Mahamood, Ondrej Dusek
INLG10
2025 How (un)faithful are explainable LLM-based NLG metrics?
abstract
Explainable NLG metrics are becoming a popular research topic; however, the faithfulness of the explanations they provide is typically not evaluated. In this work, we propose a testbed for assessing the faithfulness of span-based metrics by performing controlled perturbations of their explanations and observing changes in the final score. We show that several popular LLM evaluators do not consistently produce faithful explanations.
Alex Terentowicz, Mateusz Lango, Ondrej Dusek
INLG3
2024 Beyond Traditional Benchmarks: Analyzing Behaviors of Open LLMs on Data-to-Text Generation
abstract
We analyze the behaviors of open large language models (LLMs) on the task of data-totext (D2T) generation, i.e., generating coherent and relevant text from structured data.To avoid the issue of LLM training data contamination with standard benchmarks, we design QUINTD -a tool for collecting novel structured data records from public APIs.We find that open LLMs (Llama 2, Mistral, and Zephyr) can generate fluent and coherent texts in zero-shot settings from data in common formats collected with QUINTD.However, we show that the semantic accuracy of the outputs is a major issue: both according to human annotators and our reference-free metric based on GPT-4, more than 80% of the outputs of open LLMs contain at least one semantic error.We publicly release the code, data, and model outputs. 1
Zdenek Kasner, Ondrej Dusek
ACL (1)2
2024 Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs
abstract
Simone Balloccu, Patrícia Schmidtová, Mateusz Lango, Ondrej Dusek. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Simone Balloccu, Patrícia Schmidtová, Mateusz Lango, Ondrej Dusek
EACL (1)4
2024 Multilingual Text Style Transfer: Datasets & Models for Indian Languages
abstract
Text style transfer (TST) involves altering the linguistic style of a text while preserving its style-independent content.This paper focuses on sentiment transfer, a popular TST subtask, across a spectrum of Indian languages: Hindi, Magahi, Malayalam, Marathi, Punjabi, Odia, Telugu, and Urdu, expanding upon previous work on English-Bangla sentiment transfer (Mukherjee et al., 2023a).We introduce dedicated datasets of 1,000 positive and 1,000 negative style-parallel sentences for each of these eight languages.We then evaluate the performance of various benchmark models categorized into parallel, non-parallel, cross-lingual, and shared learning approaches, including the Llama2 and GPT-3.5 large language models (LLMs).Our experiments highlight the significance of parallel data in TST and demonstrate the effectiveness of the Masked Style Filling (MSF) approach (Mukherjee et al., 2023a) in non-parallel techniques.Moreover, crosslingual and joint multilingual learning methods show promise, offering insights into selecting optimal models tailored to the specific language and task requirements.To the best of our knowledge, this work represents the first comprehensive exploration of the TST task as sentiment transfer across a diverse set of languages.
Sourabrata Mukherjee, Atul Kr. Ojha, Akanksha Bansal, Deepak Alok, John P. McCrae, Ondrej Dusek
INLG6
2024 Are Large Language Models Actually Good at Text Style Transfer?
abstract
We analyze the performance of large language models (LLMs) on Text Style Transfer (TST), specifically focusing on sentiment transfer and text detoxification across three languages: English, Hindi, and Bengali.Text Style Transfer involves modifying the linguistic style of a text while preserving its core content.We evaluate the capabilities of pre-trained LLMs using zero-shot and few-shot prompting as well as parameter-efficient finetuning on publicly available datasets.Our evaluation using automatic metrics, GPT-4 and human evaluations reveals that while some prompted LLMs perform well in English, their performance in on other languages (Hindi, Bengali) remains average.However, finetuning significantly improves results compared to zero-shot and fewshot prompting, making them comparable to previous state-of-the-art.This underscores the necessity of dedicated datasets and specialized models for effective TST.
Sourabrata Mukherjee, Atul Kr. Ojha, Ondrej Dusek
INLG3
2024 Automatic Metrics in Natural Language Generation: A survey of Current Evaluation Practices
abstract
Patricia Schmidtova, Saad Mahamood, Simone Balloccu, Ondrej Dusek, Albert Gatt, Dimitra Gkatzia, David M. Howcroft, Ondrej Platek, Adarsa Sivaprasad. Proceedings of the 17th International Natural Language Generation Conference. 2024.
Patrícia Schmidtová, Saad Mahamood, Simone Balloccu, Ondrej Dusek, Albert Gatt, Dimitra Gkatzia, David M. Howcroft, Ondrej Plátek, Adarsa Sivaprasad
INLG4
2024 Leveraging Large Language Models for Building Interpretable Rule-Based Data-to-Text Systems
abstract
We introduce a simple approach that uses a large language model (LLM) to automatically implement a fully interpretable rule-based data-to-text system in pure Python.Experimental evaluation on the WebNLG dataset showed that such a constructed system produces text of better quality (according to the BLEU and BLEURT metrics) than the same LLM prompted to directly produce outputs, and produces fewer hallucinations than a BART language model fine-tuned on the same data.Furthermore, at runtime, the approach generates text in a fraction of the processing time required by neural approaches, using only a single CPU.
Jedrzej Warczynski, Mateusz Lango, Ondrej Dusek
INLG3
2023 Mind the Labels: Describing Relations in Knowledge Graphs With Pretrained Models
abstract
Pretrained language models (PLMs) for data-totext (D2T) generation can use human-readable data labels such as column headings, keys, or relation names to generalize to out-of-domain examples.However, the models are wellknown in producing semantically inaccurate outputs if these labels are ambiguous or incomplete, which is often the case in D2T datasets.In this paper, we expose this issue on the task of descibing a relation between two entities.For our experiments, we collect a novel dataset for verbalizing a diverse set of 1,522 unique relations from three large-scale knowledge graphs (Wikidata, DBPedia, YAGO).We find that although PLMs for D2T generation expectedly fail on unclear cases, models trained with a large variety of relation labels are surprisingly robust in verbalizing novel, unseen relations.We argue that using data with a diverse set of clear and meaningful labels is key to training D2T generation systems capable of generalizing to novel domains. 1
Zdenek Kasner, Ioannis Konstas, Ondrej Dusek
EACL3
2023 Critic-Driven Decoding for Mitigating Hallucinations in Data-to-text Generation
abstract
Hallucination of text ungrounded in the input is a well-known problem in neural data-to-text generation.Many methods have been proposed to mitigate it, but they typically require altering model architecture or collecting additional data, and thus cannot be easily applied to an existing model.In this paper, we explore a new way to mitigate hallucinations by combining the probabilistic output of a generator language model (LM) with the output of a special "text critic" classifier, which guides the generation by assessing the match between the input data and the text generated so far.Our method does not need any changes to the underlying LM's architecture or training procedure and can thus be combined with any model and decoding operating on word probabilities.The critic does not need any additional training data, using the base LM's training data and synthetic negative examples.Our experimental results show that our method improves over the baseline on the WebNLG and OpenDialKG benchmarks.
Mateusz Lango, Ondrej Dusek
EMNLP2
2023 Tackling Hallucinations in Neural Chart Summarization
abstract
Hallucinations in text generation occur when the system produces text that is not grounded in the input.In this work, we tackle the problem of hallucinations in neural chart summarization.Our analysis shows that the target side of chart summarization training datasets often contains additional information, leading to hallucinations.We propose a natural language inference (NLI) based method to preprocess the training data and show through human evaluation that our method significantly reduces hallucinations.We also found that shortening long-distance dependencies in the input sequence and adding chart-related information like title and legends improves the overall performance.
Saad Obaid ul Islam, Iza Skrjanec, Ondrej Dusek, Vera Demberg
INLG3
2023 Leveraging Low-resource Parallel Data for Text Style Transfer
abstract
Text style transfer (TST) involves transforming a text into a desired style while approximately preserving its content.The biggest challenge in TST in the general lack of parallel data.Many existing approaches rely on complex models using substantial non-parallel data, with mixed results.In this paper, we leverage a pretrained BART language model with minimal parallel data and incorporate low-resource methods such as hyperparameter tuning, data augmentation, and self-training, which have not been explored in TST.We further include novel style-based rewards in the training loss.Through extensive experiments in sentiment transfer, a sub-task of TST, we demonstrate that our simple yet effective approaches achieve well-balanced results, surpassing non-parallel approaches and highlighting the usefulness of parallel data even in small amounts. 1
Sourabrata Mukherjee, Ondrej Dusek
INLG2
2023 Are Large Language Models All You Need for Task-Oriented Dialogue?
abstract
Instruction-finetuned large language models (LLMs) gained a huge popularity recently, thanks to their ability to interact with users through conversation.In this work, we aim to evaluate their ability to complete multi-turn tasks and interact with external databases in the context of established task-oriented dialogue benchmarks.We show that in explicit belief state tracking, LLMs underperform compared to specialized task-specific models.Nevertheless, they show some ability to guide the dialogue to a successful ending through their generated responses if they are provided with correct slot values.Furthermore, this ability improves with few-shot in-domain examples.
Vojtech Hudecek, Ondrej Dusek
SIGDIAL2
2022 Neural Pipeline for Zero-Shot Data-to-Text Generation
abstract
In data-to-text (D2T) generation, training on in-domain data leads to overfitting to the data representation and repeating training data noise.We examine how to avoid finetuning pretrained language models (PLMs) on D2T generation datasets while still taking advantage of surface realization capabilities of PLMs.Inspired by pipeline approaches, we propose to generate text by transforming single-item descriptions with a sequence of modules trained on generaldomain text-based operations: ordering, aggregation, and paragraph compression.We train PLMs for performing these operations on a synthetic corpus WIKIFLUENT which we build from English Wikipedia.Our experiments on two major triple-to-text datasets-WebNLG and E2E-show that our approach enables D2T generation from RDF triples in zero-shot settings.1
Zdenek Kasner, Ondrej Dusek
ACL (1)2
2022 A Unifying View On Task-oriented Dialogue Annotation
Vojtech Hudecek, Léon-Paul Schaub, Daniel Stancl, Patrick Paroubek, Ondrej Dusek
LREC5
2022 AARGH! End-to-end Retrieval-Generation for Task-Oriented Dialog
abstract
We introduce AARGH, an end-to-end taskoriented dialog system combining retrieval and generative approaches in a single model, aiming at improving dialog management and lexical diversity of outputs.The model features a new response selection method based on an action-aware training objective and a simplified single-encoder retrieval architecture which allow us to build an end-to-end retrievalenhanced generation model where retrieval and generation share most of the parameters.On the MultiWOZ dataset, we show that our approach produces more diverse outputs while maintaining or improving state tracking and context-to-response generation performance, compared to state-of-the-art baselines.
Tomás Nekvinda, Ondrej Dusek
SIGDIAL2
2022 The Seventh Workshop on Search-Oriented Conversational Artificial Intelligence (SCAI'22)
abstract
The goal of the seventh edition of SCAI (https://scai.info) is to bring together and further grow a community of researchers and practitioners interested in conversational systems for information access. The previous iterations of the workshop already demonstrated the breadth and multidisciplinarity inherent in the design and development of conversational search agents. The proposed shift from traditional web search to search interfaces enabled via human-like dialogue leads to a number of challenges, and although such challenges have received more attention in the recent years, there are many pending research questions that should be addressed by the information retrieval community and can largely benefit from a collaboration with other research fields, such as natural language processing, machine learning, human-computer interaction and dialogue systems. This workshop is intended as a platform enabling a continuous discussion of the major research challenges that surround the design of search-oriented conversational systems. This year, participants have the opportunity to meet in person and have more in-depth interactive discussions with a full-day onsite workshop.
Gustavo Penha, Svitlana Vakulenko, Ondrej Dusek, Leigh Clark, Vaishali Pal, Vaibhav Adlakha
SIGIR3
2021 Discovering Dialogue Slots with Weak Supervision
abstract
Vojtěch Hudeček, Ondřej Dušek, Zhou Yu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Vojtech Hudecek, Ondrej Dusek
ACL/IJCNLP (1)2
2021 AggGen: Ordering and Aggregating while Generating
abstract
Xinnuo Xu, Ondřej Dušek, Verena Rieser, Ioannis Konstas. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Xinnuo Xu, Ondrej Dusek, Verena Rieser, Ioannis Konstas
ACL/IJCNLP (1)2
2021 Text-in-Context: Token-Level Error Detection for Table-to-Text Generation
abstract
We present our Charles-UPF submission for the Shared Task on Evaluating Accuracy in Generated Texts at INLG 2021.Our system can detect the errors automatically using a combination of a rule-based natural language generation (NLG) system and pretrained language models (LMs).We first utilize a rulebased NLG system to generate sentences with facts that can be derived from the input.For each sentence we evaluate, we select a subset of facts which are relevant by measuring semantic similarity to the sentence in question.Finally, we finetune a pretrained language model on annotated data along with the relevant facts for fine-grained error detection.On the test set, we achieve 69% recall and 75% precision with a model trained on a mixture of human-annotated and synthetic data.
Zdenek Kasner, Simon Mille, Ondrej Dusek
INLG3
2021 Underreporting of errors in NLG output, and what to do about it
abstract
Emiel van Miltenburg, Miruna Clinciu, Ondřej Dušek, Dimitra Gkatzia, Stephanie Inglis, Leo Leppänen, Saad Mahamood, Emma Manning, Stephanie Schoch, Craig Thomson, Luou Wen. Proceedings of the 14th International Conference on Natural Language Generation. 2021.
Emiel van Miltenburg, Miruna-Adriana Clinciu, Ondrej Dusek, Dimitra Gkatzia, Stephanie Inglis, Leo Leppänen, Saad Mahamood, Emma Manning, Stephanie Schoch, Craig Thomson, Luou Wen
INLG3
2020 Fact-based Content Weighting for Evaluating Abstractive Summarisation
abstract
ive summarisation is notoriously hard to evaluate since standard word-overlap-based metrics are insufficient. We introduce a new evaluation metric which is based on fact-level content weighting, i.e. relating the facts of the document to the facts of the summary. We fol- low the assumption that a good summary will reflect all relevant facts, i.e. the ones present in the ground truth (human-generated refer- ence summary). We confirm this hypothe- sis by showing that our weightings are highly correlated to human perception and compare favourably to the recent manual highlight- based metric of Hardy et al. (2019).
Xinnuo Xu, Ondrej Dusek, Verena Rieser, Ioannis Konstas
ACL2
2020 Evaluating Semantic Accuracy of Data-to-Text Generation with Natural Language Inference
abstract
A major challenge in evaluating data-to-text (D2T) generation is measuring the semantic accuracy of the generated text, i.e. checking if the output text contains all and only facts supported by the input data.We propose a new metric for evaluating the semantic accuracy of D2T generation based on a neural model pretrained for natural language inference (NLI).We use the NLI model to check textual entailment between the input data and the output text in both directions, allowing us to reveal omissions or hallucinations.Input data are converted to text for NLI using trivial templates.Our experiments on two recent D2T datasets show that our metric can achieve high accuracy in identifying erroneous system outputs.
Ondrej Dusek, Zdenek Kasner
INLG1
2020 Data-to-Text Generation with Iterative Text Editing
abstract
We present a novel approach to data-to-text generation based on iterative text editing.Our approach maximizes the completeness and semantic accuracy of the output text while leveraging the abilities of recent pre-trained models for text editing (LASERTAGGER) and language modeling (GPT-2) to improve the text fluency.To this end, we first transform data items to text using trivial templates, and then we iteratively improve the resulting text by a neural model trained for the sentence fusion task.The output of the model is filtered by a simple heuristic and reranked with an offthe-shelf pre-trained language model.We evaluate our approach on two major data-to-text datasets (WebNLG, Cleaned E2E) and analyze its caveats and benefits.Furthermore, we show that our formulation of data-to-text generation opens up the possibility for zero-shot domain adaptation using a general-domain dataset for sentence fusion.
Zdenek Kasner, Ondrej Dusek
INLG2
2020 One Model, Many Languages: Meta-Learning for Multilingual Text-to-Speech
abstract
We introduce an approach to multilingual speech synthesis which uses the meta-learning concept of contextual parameter generation and produces natural-sounding multilingual speech using more languages and less training data than previous approaches. Our model is based on Tacotron 2 with a fully convolutional input text encoder whose weights are predicted by a separate parameter generator network. To boost voice cloning, the model uses an adversarial speaker classifier with a gradient reversal layer that removes speaker-specific information from the encoder. We arranged two experiments to compare our model with baselines using various levels of cross-lingual parameter sharing, in order to evaluate: (1) stability and performance when training on low amounts of data, (2) pronunciation accuracy and voice quality of code-switching synthesis. For training, we used the CSS10 dataset and our new small dataset based on Common Voice recordings in five languages. Our model is shown to effectively share information across languages and according to a subjective evaluation test, it produces more natural and accurate code-switching speech than the baselines.
Tomás Nekvinda, Ondrej Dusek
INTERSPEECH2
2020 SpeedySpeech: Efficient Neural Speech Synthesis
abstract
While recent neural sequence-to-sequence models have greatly improved the quality of speech synthesis, there has not been a system capable of fast training, fast inference and high-quality audio synthesis at the same time. We propose a student-teacher network capable of high-quality faster-than-real-time spectrogram synthesis, with low requirements on computational resources and fast training time. We show that self-attention layers are not necessary for generation of high quality audio. We utilize simple convolutional blocks with residual connections in both student and teacher networks and use only a single attention layer in the teacher model. Coupled with a MelGAN vocoder, our model's voice quality was rated significantly higher than Tacotron 2. Our model can be efficiently trained on a single GPU and can run in real time even on a CPU. We provide both our source code and audio samples in our GitHub repository.
Jan Vainer, Ondrej Dusek
INTERSPEECH2
2020 Evaluating the state-of-the-art of End-to-End Natural Language Generation: The E2E NLG challenge
abstract
This paper provides a comprehensive analysis of the first shared task on End-to-End Natural Language Generation (NLG) and identifies avenues for future research based on the results. This shared task aimed to assess whether recent end-to-end NLG systems can generate more complex output by learning from datasets containing higher lexical richness, syntactic complexity and diverse discourse phenomena. Introducing novel automatic and human metrics, we compare 62 systems submitted by 17 institutions, covering a wide range of approaches, including machine learning architectures – with the majority implementing sequence-to-sequence models (seq2seq) – as well as systems based on grammatical rules and templates. Seq2seq-based systems have demonstrated a great potential for NLG in the challenge. We find that seq2seq systems generally score high in terms of word-overlap metrics and human evaluations of naturalness – with the winning Slug system (Juraska et al., 2018) being seq2seq-based. However, vanilla seq2seq models often fail to correctly express a given meaning representation if they lack a strong semantic control mechanism applied during decoding. Moreover, seq2seq models can be outperformed by hand-engineered systems in terms of overall quality, as well as complexity, length and diversity of outputs. This research has influenced, inspired and motivated a number of recent studies outwith the original competition, which we also summarise as part of this paper.
Ondrej Dusek, Jekaterina Novikova, Verena Rieser
Comput. Speech Lang.1
2019 Semantic Noise Matters for Neural Natural Language Generation
abstract
Neural natural language generation (NNLG) systems are known for their pathological outputs, i.e. generating text which is unrelated to the input specification.In this paper, we show the impact of semantic noise on state-of-theart NNLG models which implement different semantic control mechanisms.We find that cleaned data can improve semantic correctness by up to 97%, while maintaining fluency.We also find that the most common error is omitting information, rather than hallucination.
Ondrej Dusek, David M. Howcroft, Verena Rieser
INLG1
2019 Neural Generation for Czech: Data and Baselines
abstract
We present the first dataset targeted at end-toend NLG in Czech in the restaurant domain, along with several strong baseline models using the sequence-to-sequence approach.While non-English NLG is under-explored in general, Czech, as a morphologically rich language, makes the task even harder: Since Czech requires inflecting named entities, delexicalization or copy mechanisms do not work out-ofthe-box and lexicalizing the generated outputs is non-trivial.In our experiments, we present two different approaches to this this problem: (1) using a neural language model to select the correct inflected form while lexicalizing, (2) a two-step generation setup: our sequence-to-sequence model generates an interleaved sequence of lemmas and morphological tags, which are then inflected by a morphological generator. Hledáte vhodnou restauraci na X-good_for_meal ? Do-you-look-for a-suitable restaurant for [breakfast]X-name najdete v oblasti X-area .[Baráčnická rychta] you-find in the-area [of-Malá Strana] Chcete najít restauraci, kde se dobře X-good_for_meal ?Do-you-want to-find a-restaurant where yourself well [you-will-have-breakfast]X-name je na X-area . [Baráčnická rychta] is in [Malá Strana]Malá Strana NNFS1-----A---- Malé Strany
Ondrej Dusek, Filip Jurcícek
INLG1
2019 Automatic Quality Estimation for Natural Language Generation: Ranting (Jointly Rating and Ranking)
abstract
We present a recurrent neural network based system for automatic quality estimation of natural language generation (NLG) outputs, which jointly learns to assign numerical ratings to individual outputs and to provide pairwise rankings of two different outputs.The latter is trained using pairwise hinge loss over scores from two copies of the rating network.We use learning to rank and synthetic data to improve the quality of ratings assigned by our system: we synthesise training pairs of distorted system outputs and train the system to rank the less distorted one higher.This leads to a 12% increase in correlation with human ratings over the previous benchmark.We also establish the state of the art on the dataset of relative rankings from the E2E NLG Challenge (Dušek et al., 2019), where synthetic data lead to a 4% accuracy increase over the base model.
Ondrej Dusek, Karin Sevegnani, Ioannis Konstas, Verena Rieser
INLG1
2019 User Evaluation of a Multi-dimensional Statistical Dialogue System
abstract
We present the first complete spoken dialogue system driven by a multi-dimensional statistical dialogue manager.This framework has been shown to substantially reduce data needs by leveraging domain-independent dimensions, such as social obligations or feedback, which (as we show) can be transferred between domains.In this paper, we conduct a user study and show that the performance of a multi-dimensional system, which can be adapted from a source domain, is equivalent to that of a one-dimensional baseline, which can only be trained from scratch.
Simon Keizer, Ondrej Dusek, Xingkun Liu, Verena Rieser
SIGdial2
2018 Better Conversations by Modeling, Filtering, and Optimizing for Coherence and Diversity
abstract
We present three enhancements to existing encoder-decoder models for open-domain conversational agents, aimed at effectively modeling coherence and promoting output diversity: (1) We introduce a measure of coherence as the GloVe embedding similarity between the dialogue context and the generated response, (2) we filter our training corpora based on the measure of coherence to obtain topically coherent and lexically diverse context-response pairs, (3) we then train a response generator using a conditional variational autoencoder model that incorporates the measure of coherence as a latent variable and uses a context gate to guarantee topical consistency with the context and promote lexical diversity.Experiments on the OpenSubtitles corpus show a substantial improvement over competitive neural models in terms of BLEU score as well as metrics of coherence and diversity.
Xinnuo Xu, Ondrej Dusek, Ioannis Konstas, Verena Rieser
EMNLP2
2018 Improving Context Modelling in Multimodal Dialogue Generation
abstract
In this work, we investigate the task of textual response generation in a multimodal task-oriented dialogue system.Our work is based on the recently released Multimodal Dialogue (MMD) dataset (Saha et al., 2017) in the fashion domain.We introduce a multimodal extension to the Hierarchical Recurrent Encoder-Decoder (HRED) model and show that this extension outperforms strong baselines in terms of text-based similarity metrics.We also showcase the shortcomings of current vision and language models by performing an error analysis on our system's output.
Shubham Agarwal 0001, Ondrej Dusek, Ioannis Konstas, Verena Rieser
INLG2
2018 Findings of the E2E NLG Challenge
abstract
This paper summarises the experimental setup and results of the first shared task on end-to-end (E2E) natural language generation (NLG) in spoken dialogue systems.Recent end-to-end generation systems are promising since they reduce the need for data annotation.However, they are currently limited to small, delexicalised datasets.The E2E NLG shared task aims to assess whether these novel approaches can generate better-quality output by learning from a dataset containing higher lexical richness, syntactic complexity and diverse discourse phenomena.We compare 62 systems submitted by 17 institutions, covering a wide range of approaches, including machine learning architectures -with the majority implementing sequence-to-sequence models (seq2seq) -as well as systems based on grammatical rules and templates.
Ondrej Dusek, Jekaterina Novikova, Verena Rieser
INLG1
2017 Why We Need New Evaluation Metrics for NLG
abstract
The majority of NLG evaluation relies on automatic metrics, such as BLEU.In this paper, we motivate the need for novel, system-and data-independent automatic evaluation methods: We investigate a wide range of metrics, including state-of-the-art word-based and novel grammar-based ones, and demonstrate that they only weakly reflect human judgements of system outputs as generated by data-driven, end-to-end NLG.We also show that metric performance is data-and system-specific.Nevertheless, our results also suggest that automatic metrics perform reliably at system-level and can support system development by finding cases where a system performs poorly.4 https://github.com/glampouras/JLOLS_NLG 5 Note that we use lexicalised versions of SFHOTEL and SFREST and a partially lexicalised version of BAGEL, where proper names and place names are replaced by placeholders ("X"), in correspondence with the outputs generated by the MR: inform(name=X, area=X, pricerange=moderate, type=restaurant) Reference: "X is a moderately priced restaurant in X."
Jekaterina Novikova, Ondrej Dusek, Amanda Cercas Curry, Verena Rieser
EMNLP2
2017 The E2E Dataset: New Challenges For End-to-End Generation
abstract
This paper describes the E2E data, a new dataset for training end-to-end, datadriven natural language generation systems in the restaurant domain, which is ten times bigger than existing, frequently used datasets in this area.The E2E dataset poses new challenges: (1) its human reference texts show more lexical richness and syntactic variation, including discourse phenomena; (2) generating from this set requires content selection.As such, learning from this dataset promises more natural, varied and less template-like system utterances.We also establish a baseline on this dataset, which illustrates some of the difficulties associated with this data.
Jekaterina Novikova, Ondrej Dusek, Verena Rieser
SIGDIAL Conference2
2016 A Context-aware Natural Language Generator for Dialogue Systems
abstract
We present a novel natural language generation system for spoken dialogue systems capable of entraining (adapting) to users' way of speaking, providing contextually appropriate responses.The generator is based on recurrent neural networks and the sequence-to-sequence approach.It is fully trainable from data which include preceding context along with responses to be generated.We show that the context-aware generator yields significant improvements over the baseline in both automatic metrics and a human pairwise preference test.
Ondrej Dusek, Filip Jurcícek
SIGDIAL Conference1
2015 Training a Natural Language Generator From Unaligned Data
abstract
Ondřej Dušek, Filip Jurčíček. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Ondrej Dusek, Filip Jurcícek
ACL (1)1
2014 Free English and Czech telephone speech corpus shared under the CC-BY-SA 3.0 license
Matej Korvas, Ondrej Plátek, Ondrej Dusek, Lukás Zilka, Filip Jurcícek
LREC3
2014 Multilingual Test Sets for Machine Translation of Search Queries for Cross-Lingual Information Retrieval in the Medical Domain
Zdenka Uresová, Jan Hajic 0001, Pavel Pecina, Ondrej Dusek
LREC4
2014 Alex: Bootstrapping a Spoken Dialogue System for a New Domain by Real Users
abstract
When deploying a spoken dialogue system in a new domain, one faces a situation where little to no data is available to train domain-specific statistical models. We describe our experience with bootstrapping a dialogue system for public transit and weather information in real-word deployment under public use. We proceeded incrementally, starting from a minimal system put on a toll-free telephone number to collect speech data. We were able to incorporate statistical modules trained on collected data ‐ in-domain speech recognition language models and spoken language understanding ‐ while simultaneously extending the domain, making use of automatically generated semantic annotation. Our approach shows that a successful system can be built with minimal effort and no in-domain data at hand.
Ondrej Dusek, Ondrej Plátek, Lukás Zilka, Filip Jurcícek
SIGDIAL Conference1
2014 Adaptation of machine translation for multilingual information retrieval in the medical domain
Pavel Pecina, Ondrej Dusek, Lorraine Goeuriot, Jan Hajic 0001, Jaroslava Hlavácová, Gareth J. F. Jones, Liadh Kelly, Johannes Leveling, David Marecek, Michal Novák 0001, Martin Popel, Rudolf Rosa, Ales Tamchyna, Zdenka Uresová
Artif. Intell. Medicine2
2012 The Joy of Parallelism with CzEng 1.0
Ondrej Bojar, Zdenek Zabokrtský, Ondrej Dusek, Petra Galuscáková, Martin Majlis, David Marecek, Jirka Marsík, Michal Novák 0001, Martin Popel, Ales Tamchyna
LREC3