VLDB 2026 Research / reviewers in the wild / expert
Michael White 0001
dblp:76/6763-1
· DBLP profile ↗
35ranked-venue papers
12as first author
9since 2021 · last 2026
0000-0002-3062-1719ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 12 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VISTA: Verification In Sequential Turn-based AssessmentabstractHallucination-defined here as generated statements unsupported or contradicted by available evidence or conversational context-remains a major obstacle to using conversational AI systems in settings that demand factual reliability.Existing metrics evaluate isolated responses or treat unverifiable content as errors, limiting their use for multi-turn dialogue.We introduce VISTA (Verification In Sequential Turn-based Assessment), a framework for evaluating conversational factuality via claim-level verification and sequential consistency tracking.VISTA decomposes each turn into atomic claims, verifies them against trusted sources and dialogue history, and categorizes unverifiable statements (subjective, contradicted, lacking evidence, or abstaining).Across eight large language models and four dialogue factuality benchmarks (AIS, BEGIN, FAITHDIAL, and FADE), VISTA substantially improves hallucination detection over FActScore and LLM-as-Judge baselines.Human evaluation confirms that VISTA's decomposition improves annotator agreement and reveals inconsistencies in existing benchmarks.Further analyses show that incorporating dialogue context into verification substantially improves contradiction detection, and that VISTA reliably identifies abstentions.By modeling factuality as a dynamic property of conversation, VISTA offers a more transparent, human-aligned measure of truthfulness in dialogue systems. Ashley Lewis, Andrew Perrault, Eric Fosler-Lussier, Michael White 0001 |
ACL (1) | 4 |
| 2024 | When is Tree Search Useful for LLM Planning? It Depends on the DiscriminatorabstractIn this paper, we examine how large language models (LLMs) solve multi-step problems under a language agent framework with three components: a generator, a discriminator, and a planning method.We investigate the practical utility of two advanced planning methods, iterative correction and tree search.We present a comprehensive analysis of how discrimination accuracy affects the overall performance of agents when using these two methods or a simpler method, re-ranking.Experiments on two tasks, text-to-SQL parsing and mathematical reasoning, show that: (1) advanced planning methods demand discriminators with at least 90% accuracy to achieve significant improvements over re-ranking; (2) current LLMs' discrimination abilities have not met the needs of advanced planning methods to achieve such improvements; (3) with LLM-based discriminators, advanced planning methods may not adequately balance accuracy and efficiency.For example, compared to the other two methods, tree search is at least 10-20 times slower but leads to negligible performance gains, which hinders its real-world applications.1 Ziru Chen, Michael White 0001, Raymond J. Mooney, Ali Payani, Yu Su 0001, Huan Sun 0001 |
ACL (1) | 2 |
| 2024 | A randomized prospective study of a hybrid rule- and data-driven virtual patientabstractAbstract Randomized prospective studies represent the gold standard for experimental design. In this paper, we present a randomized prospective study to validate the benefits of combining rule-based and data-driven natural language understanding methods in a virtual patient dialogue system. The system uses a rule-based pattern matching approach together with a machine learning (ML) approach in the form of a text-based convolutional neural network, combining the two methods with a simple logistic regression model to choose between their predictions for each dialogue turn. In an earlier, retrospective study, the hybrid system yielded a nearly 50% error reduction on our initial data, in part due to the differential performance between the two methods as a function of label frequency. Given these gains, and considering that our hybrid approach is unique among virtual patient systems, we compare the hybrid system to the rule-based system by itself in a randomized prospective study. We evaluate 110 unique medical student subjects interacting with the system over 5,296 conversation turns, to verify whether similar gains are observed in a deployed system. This prospective study broadly confirms the findings from the earlier one but also highlights important deficits in our training data. The hybrid approach still improves over either rule-based or ML approaches individually, even handling unseen classes with some success. However, we observe that live subjects ask more out-of-scope questions than expected. To better handle such questions, we investigate several modifications to the system combination component. These show significant overall accuracy improvements and modest F1 improvements on out-of-scope queries in an offline evaluation. We provide further analysis to characterize the difficulty of the out-of-scope problem that we have identified, as well as to suggest future improvements over the baseline we establish here. Adam Stiff, Michael White 0001, Eric Fosler-Lussier, Lifeng Jin, Evan Jaffe, Douglas Danforth |
Nat. Lang. Eng. | 2 |
| 2023 | Bootstrapping a Conversational Guide for Colonoscopy PrepabstractPulkit Arya, Madeleine Bloomquist, Subhankar Chakraborty, Andrew Perrault, William Schuler, Eric Fosler-Lussier, Michael White. Proceedings of the 24th Meeting of the Special Interest Group on Discourse and Dialogue. 2023. Pulkit Arya, Madeleine Bloomquist, Subhankar Chakraborty, Andrew Perrault, William Schuler, Eric Fosler-Lussier, Michael White 0001 |
SIGDIAL | 7 |
| 2022 | Generating Discourse Connectives with Pre-trained Language Models: Conditioning on Discourse Relations Helps Reconstruct the PDTBabstractWe report results of experiments using BART (Lewis et al., 2019) and the Penn Discourse Tree Bank (Webber et al., 2019) (PDTB) to generate texts with correctly realized discourse relations.We address a question left open by previous research (Yung et al., 2021;Ko and Li, 2020) concerning whether conditioning the model on the intended discourse relationwhich corresponds to adding explicit discourse relation information into the input to the model-improves its performance.Our results suggest that including discourse relation information in the input of the model significantly improves the consistency with which it produces a correctly realized discourse relation in the output.We compare our models' performance to known results concerning the discourse structures found in written text and their possible explanations in terms of discourse interpretation strategies hypothesized in the psycholinguistics literature.Our findings suggest that natural language generation models based on current pre-trained Transformers will benefit from infusion with discourse level information if they aim to construct discourses with the intended relations. Symon Jory Stevens-Guille, Aleksandre Maskharashvili, Michael White 0001 |
SIGDIAL | 4 |
| 2021 | Building Adaptive Acceptability Classifiers for Neural NLGabstractSoumya Batra, Shashank Jain, Peyman Heidari, Ankit Arun, Catharine Youngs, Xintong Li, Pinar Donmez, Shawn Mei, Shiunzu Kuo, Vikas Bhardwaj, Anuj Kumar, Michael White. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Soumya Batra, Shashank Jain, Peyman Heidari, Ankit Arun, Catharine Youngs, Pinar Donmez, Shawn Mei, Shiunzu Kuo, Vikas Bhardwaj, Michael White 0001 |
EMNLP (1) | 12 |
| 2021 | Self-Training for Compositional Neural NLG in Task-Oriented DialogueabstractNeural approaches to natural language generation in task-oriented dialogue have typically required large amounts of annotated training data to achieve satisfactory performance, especially when generating from compositional inputs.To address this issue, we show that selftraining enhanced with constrained decoding yields large gains in data efficiency on a conversational weather dataset that employs compositional meaning representations.In particular, our experiments indicate that self-training with constrained decoding can enable sequence-tosequence models to achieve satisfactory quality using vanilla decoding with five to ten times less data than with ordinary supervised baseline; moreover, by leveraging pretrained models, data efficiency can be increased further to fifty times.We confirm the main automatic results with human evaluations and show that they extend to an enhanced, compositional version of the E2E dataset.The end result is an approach that makes it possible to achieve acceptable performance on compositional NLG tasks using hundreds rather than tens of thousands of training samples. Symon Jory Stevens-Guille, Aleksandre Maskharashvili, Michael White 0001 |
INLG | 4 |
| 2021 | Neural Methodius Revisited: Do Discourse Relations Help with Pre-Trained Models Too?abstractRecent developments in natural language generation (NLG) have bolstered arguments in favor of re-introducing explicit coding of discourse relations in the input to neural models.In the Methodius corpus, a meaning representation (MR) is hierarchically structured and includes discourse relations.Meanwhile pre-trained language models have been shown to implicitly encode rich linguistic knowledge which provides an excellent resource for NLG.By virtue of synthesizing these lines of research, we conduct extensive experiments on the benefits of using pre-trained models and discourse relation information in MRs, focusing on the improvement of discourse coherence and correctness.We redesign the Methodius corpus; we also construct another Methodius corpus in which MRs are not hierarchically structured but flat.We report experiments on different versions of the corpora, which probe when, where, and how pre-trained models benefit from MRs with discourse relation information in them.We conclude that discourse relations significantly improve NLG when data is limited. Aleksandre Maskharashvili, Symon Jory Stevens-Guille, Michael White 0001 |
INLG | 4 |
| 2021 | Getting to Production with Few-shot Natural Language Generation ModelsabstractPeyman Heidari, Arash Einolghozati, Shashank Jain, Soumya Batra, Lee Callender, Ankit Arun, Shawn Mei, Sonal Gupta, Pinar Donmez, Vikas Bhardwaj, Anuj Kumar, Michael White. Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2021. Peyman Heidari, Arash Einolghozati, Shashank Jain, Soumya Batra, Lee Callender, Ankit Arun, Shawn Mei, Sonal Gupta, Pinar Donmez, Vikas Bhardwaj, Michael White 0001 |
SIGDIAL | 12 |
| 2020 | Neural NLG for Methodius: From RST Meaning Representations to TextsabstractWhile classic NLG systems typically made use of hierarchically structured content plans that included discourse relations as central components, more recent neural approaches have mostly mapped simple, flat inputs to texts without representing discourse relations explicitly.In this paper, we investigate whether it is beneficial to include discourse relations in the input to neural data-to-text generators for texts where discourse relations play an important role.To do so, we reimplement the sentence planning and realization components of a classic NLG system, Methodius, using LSTM sequence-to-sequence (seq2seq) models.We find that although seq2seq models can learn to generate fluent and grammatical texts remarkably well with sufficiently representative Methodius training data, they cannot learn to correctly express Methodius's SIMILARITY and CONTRAST comparisons unless the corresponding RST relations are included in the inputs.Additionally, we experiment with using self-training and reverse model reranking to better handle train/test data mismatches, and find that while these methods help reduce content errors, it remains essential to include discourse relations in the input to obtain optimal performance. Symon Jory Stevens-Guille, Aleksandre Maskharashvili, Amy Isard, Michael White 0001 |
INLG | 5 |
| 2019 | Constrained Decoding for Neural NLG from Compositional Representations in Task-Oriented DialogueabstractGenerating fluent natural language responses from structured semantic representations is a critical step in task-oriented conversational systems.Avenues like the E2E NLG Challenge have encouraged the development of neural approaches, particularly sequence-tosequence (Seq2Seq) models for this problem.The semantic representations used, however, are often underspecified, which places a higher burden on the generation model for sentence planning, and also limits the extent to which generated responses can be controlled in a live system.In this paper, we (1) propose using tree-structured semantic representations, like those used in traditional rule-based NLG systems, for better discourse-level structuring and sentence-level planning; (2) introduce a challenging dataset using this representation for the weather domain; (3) introduce a constrained decoding approach for Seq2Seq models that leverages this representation to improve semantic correctness; and (4) demonstrate promising results on our dataset and the E2E dataset. Anusha Balakrishnan, Jinfeng Rao, Kartikeya Upasani, Michael White 0001, Rajen Subba |
ACL (1) | 4 |
| 2019 | A Tree-to-Sequence Model for Neural NLG in Task-Oriented DialogabstractGenerating fluent natural language responses from structured semantic representations is a critical step in task-oriented conversational systems.Sequence-to-sequence models on flat meaning representations (MR) have been dominant in this task, for example in the E2E NLG Challenge.Previous work has shown that a tree-structured MR can improve the model for better discourse-level structuring and sentence-level planning.In this work, we propose a tree-to-sequence model that uses a tree-LSTM encoder to leverage the tree structures in the input MR, and further enhance the decoding by a structure-enhanced attention mechanism.In addition, we explore combining these enhancements with constrained decoding to improve semantic correctness.Our method not only shows significant improvements over standard seq2seq baselines, but also is more data-efficient and generalizes better to hard scenarios. Jinfeng Rao, Kartikeya Upasani, Anusha Balakrishnan, Michael White 0001, Rajen Subba |
INLG | 4 |
| 2016 | Enhancing PTB Universal Dependencies for Grammar-Based Surface RealizationabstractGrammar-based surface realizers require inputs compatible with their reversible, constraint-based grammars, including a proper representation of unbounded dependencies and coordination.In this paper, we report on progress towards creating realizer inputs along the lines of those used in the first surface realization shared task that satisfy this requirement.To do so, we augment the Universal Dependencies that result from running the Stanford Dependency Converter on the Penn Treebank with the unbounded and coordination dependencies in the CCGbank, since only the latter takes the Penn Treebank's trace information into account.An evaluation against gold standard dependencies shows that the enhanced dependencies have greatly enhanced recall with moderate precision.We conclude with a discussion of the implications of the work for a second realization shared task. David L. King, Michael White 0001 |
INLG | 2 |
| 2016 | A Corpus of Word-Aligned Asked and Anticipated Questions in a Virtual Patient Dialogue System
Ajda Gokcen, Evan Jaffe, Johnsey Erdmann, Michael White 0001, Douglas Danforth |
LREC | 4 |
| 2014 | That's Not What I Meant! Using Parsers to Avoid Structural Ambiguities in Generated TextabstractWe investigate whether parsers can be used for self-monitoring in surface realization in order to avoid egregious errors involving "vicious" ambiguities, namely those where the intended interpretation fails to be considerably more likely than alternative ones.Using parse accuracy in a simple reranking strategy for selfmonitoring, we find that with a stateof-the-art averaged perceptron realization ranking model, BLEU scores cannot be improved with any of the well-known Treebank parsers we tested, since these parsers too often make errors that human readers would be unlikely to make.However, by using an SVM ranker to combine the realizer's model score together with features from multiple parsers, including ones designed to make the ranker more robust to parsing mistakes, we show that significant increases in BLEU scores can be achieved.Moreover, via a targeted manual analysis, we demonstrate that the SVM reranker frequently manages to avoid vicious ambiguities, while its ranking errors tend to affect fluency much more often than adequacy. Manjuan Duan, Michael White 0001 |
ACL (1) | 2 |
| 2014 | Towards Surface Realization with CCGs Induced from DependenciesabstractWe present a novel algorithm for inducing Combinatory Categorial Grammars from dependency treebanks, along with initial experiments showing that it can be used to achieve competitive realization results using an enhanced version of the surface realization shared task data. 1 Michael White 0001 |
INLG | 1 |
| 2012 | Minimal Dependency Length in Realization Ranking
Michael White 0001, Rajakrishnan Rajkumar |
EMNLP-CoNLL | 1 |
| 2012 | The Surface Realisation Task: Recent Developments and Future Plans
Anya Belz, Bernd Bohnet, Simon Mille, Leo Wanner, Michael White 0001 |
INLG | 5 |
| 2012 | Shared Task Proposal: Syntactic Paraphrase Ranking
Michael White 0001 |
INLG | 1 |
| 2010 | Further Meta-Evaluation of Broad-Coverage Surface Realization
Dominic Espinosa, Rajakrishnan Rajkumar, Michael White 0001, Shoshana Berleant |
EMNLP | 3 |
| 2010 | Machine learning for text selection with expressive unit-selection voicesabstractWe show that a ranking model produced by machine learning outperforms two baselines when applied to the task of selecting texts for use in creating a unit-selection synthesis voice with good domain coverage. The model learns to predict the estimated utility of an utterance based on features relating it to the utterances selected so far and a corpus of target utterances. Our analyses indicate that our discriminative approach continues to work well even though the presence of rich prosodic and nonprosodic features significantly expands the search space beyond what has previously been handled by greedy methods. Index Terms: speech synthesis, unit selection, machine learning Dominic Espinosa, Michael White 0001, Eric Fosler-Lussier, Chris Brew |
INTERSPEECH | 2 |
| 2010 | Generating Tailored, Comparative Descriptions with Contextually Appropriate IntonationabstractGenerating responses that take user preferences into account requires adaptation at all levels of the generation process. This article describes a multi-level approach to presenting user-tailored information in spoken dialogues which brings together for the first time multi-attribute decision models, strategic content planning, surface realization that incorporates prosody prediction, and unit selection synthesis that takes the resulting prosodic structure into account. The system selects the most important options to mention and the attributes that are most relevant to choosing between them, based on the user model. Multiple options are selected when each offers a compelling trade-off. To convey these trade-offs, the system employs a novel presentation strategy which straightforwardly lends itself to the determination of information structure, as well as the contents of referring expressions. During surface realization, the prosodic structure is derived from the information structure using Combinatory Categorial Grammar in a way that allows phrase boundaries to be determined in a flexible, data-driven fashion. This approach to choosing pitch accents and edge tones is shown to yield prosodic structures with significantly higher acceptability than baseline prosody prediction models in an expert evaluation. These prosodic structures are then shown to enable perceptibly more natural synthesis using a unit selection voice that aims to produce the target tunes, in comparison to two baseline synthetic voices. An expert evaluation and f0 analysis confirm the superiority of the generator-driven intonation and its contribution to listeners' ratings. Michael White 0001, Robert A. J. Clark, Johanna D. Moore |
Comput. Linguistics | 1 |
| 2009 | Perceptron Reranking for CCG Realization
Michael White 0001, Rajakrishnan Rajkumar |
EMNLP | 1 |
| 2009 | Eye tracking for the online evaluation of prosody in speech synthesis: not so fast!abstractThis paper presents an eye-tracking experiment comparing the processing of different accent patterns in unit selection synthesis and human speech. The synthetic speech results failed to replicate the facilitative effect of contextually appropriate accent patterns found with human speech, while producing a more robust intonational garden-path effect with contextually inappropriate patterns, both of which could be due to processing delays seen with the synthetic speech. As the synthetic speech was of high quality, the results indicate that eye tracking holds promise as a highly sensitive and objective method for the online evaluation of prosody in speech synthesis. Index Terms: speech synthesis, evaluation, prosody, eye tracking, unit selection Michael White 0001, Rajakrishnan Rajkumar, Kiwako Ito, Shari R. Speer |
INTERSPEECH | 1 |
| 2008 | Hypertagging: Supertagging for Surface Realization with CCG
Dominic Espinosa, Michael White 0001, Dennis Mehay |
ACL | 2 |
| 2008 | Projecting Propbank Roles onto the CCGbank
Stephen A. Boxwell, Michael White 0001 |
LREC | 2 |
| 2006 | Learning to Say It Well: Reranking Realizations by Predicted Synthesis QualityabstractThis paper presents a method for adapting a language generator to the strengths and weaknesses of a synthetic voice, thereby improving the naturalness of synthetic speech in a spoken language dialogue system. The method trains a discriminative reranker to select paraphrases that are predicted to sound natural when synthesized. The ranker is trained on realizer and synthesizer features in supervised fashion, using human judgements of synthetic voice quality on a sample of the paraphrases representative of the generator's capability. Results from a cross-validation study indicate that discriminative paraphrase reranking can achieve substantial improvements in naturalness on average, ameliorating the problem of highly variable synthesis quality typically encountered with today's unit selection synthesizers. Crystal Nakatsu, Michael White 0001 |
ACL | 2 |
| 2006 | CCG Chart Realization from Disjunctive Inputs
Michael White 0001 |
INLG | 1 |
| 2005 | Multimodal Generation in the COMIC Dialogue System
Mary Ellen Foster, Michael White 0001, Andrea Setzer, Roberta Catizone |
ACL | 2 |
| 2004 | Reining in CCG Chart Realization
Michael White 0001 |
INLG | 1 |
| 1998 | EXEMPLARS: A Practical, Extensible Framework For Dynamic Text Generation
Michael White 0001, Ted Caldwell |
INLG | 1 |
| 1993 | The Imperfective Paradox and Trajectory-of-Motion EventsabstractIn the first part of the paper, I present a new treatment of THE IMPERFECTIVE PARADOX (Dowty 1979) for the restricted case of trajectory-of-motion events. This treatment extends and refines those of Moens and Steedman (1988) and Jackendoff (1991). In the second part, I describe an implemented algorithm based on this treatment which determines whether a specified sequence of such events is or is not possible under certain situationally supplied constraints and restrictive assumptions. Michael White 0001 |
ACL | 1 |
| 1993 | Delimitedness And Trajectory-Of-Motion Events
Michael White 0001 |
EACL | 1 |
| 1992 | On The Interpretation Of Natural Language Instructions
Barbara Di Eugenio, Michael White 0001 |
COLING | 2 |
| 1992 | Conceptual Structures And Ccc: Linking Theory And Incorporated Argument Adjuncts
Michael White 0001 |
COLING | 1 |