David Wan

dblp:17/4695 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 8 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 PrefixNLI: Detecting Factual Inconsistencies as Soon as They Arise
abstract
Natural Language Inference (NLI) models have been used in various ways to improve the factuality of LLM outputs.This is typically done by applying an NLI model to judge whether the model output is entailed from the supposed evidence, triggering some corrective actions, such as beam reranking at inference time or RL rewards during training.While NLI models are trained to detect factual inconsistencies over complete sentences, decisions in the common autoregressive generation architecture are made for each evolving text prefix, during decoding.Addressing this setting, we generalize the entailment detection task to apply over arbitrary text prefixes, and suggest its utility for improving generation faithfulness.Providing suitable evaluation and training datasets for this task, we train MiniTruePrefixes, a novel specialized model that better detects factual inconsistencies over text prefixes, outperforming comparable baseline NLI models by 5-14 F1 points in prefix-level entailment.We further demonstrate that integrating MiniTruePrefixes into a controlled decoding framework substantially improves factual consistency in abstractive summarization.When guided by Mini-TruePrefixes, LLaMA-3.2-3B-Instructmatches the faithfulness and runtime of the 8B model from the same model family, while using only half the memory.
Sapir Harary, Eran Hirsch, Aviv Slobodkin, David Wan, Mohit Bansal, Ido Dagan
ACL (1)4
2026 Localizing Factual Inconsistencies in Attributable Text Generation
abstract
Abstract There has been an increasing interest in detecting hallucinations in model-generated texts, both manually and automatically, at varying levels of granularity. However, most existing methods fail to precisely pinpoint the errors. In this work, we introduce QASemConsistency, a new formalism for localizing factual inconsistencies in attributable text generation, at a fine-grained level. Drawing inspiration from Neo-Davidsonian formal semantics, we propose decomposing the generated text into minimal predicate-argument level propositions, expressed as simple question-answer (QA) pairs, and assess whether each individual QA pair is supported by a trusted reference text. As each QA pair corresponds to a single semantic relation between a predicate and an argument, QASemConsistency effectively localizes the unsupported information. We first demonstrate the effectiveness of the QASemConsistency methodology for human annotation, by collecting crowdsourced annotations of granular consistency errors, while achieving a substantial inter-annotator agreement. This benchmark includes more than 3K instances spanning various tasks of attributable text generation. We also show that QASemConsistency yields factual consistency scores that correlate well with human judgments. Finally, we implement several methods for automatically detecting localized factual inconsistencies, with both supervised entailment models and LLMs.1
Arie Cattan, Paul Roit, Shiyue Zhang 0001, David Wan, Roee Aharoni, Idan Szpektor, Mohit Bansal, Ido Dagan
Trans. Assoc. Comput. Linguistics4
2025 LAQuer: Localized Attribution Queries in Content-grounded Generation
abstract
Eran Hirsch, Aviv Slobodkin, David Wan, Elias Stengel-Eskin, Mohit Bansal, Ido Dagan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Eran Hirsch, Aviv Slobodkin, David Wan, Elias Stengel-Eskin, Mohit Bansal, Ido Dagan
ACL (1)3
2025 MAMM-Refine: A Recipe for Improving Faithfulness in Generation with Multi-Agent Collaboration
abstract
David Wan, Justin Chen, Elias Stengel-Eskin, Mohit Bansal. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
David Wan, Justin Chih-Yao Chen, Elias Stengel-Eskin, Mohit Bansal
NAACL (Long Papers)1
2025 On Positional Bias of Faithfulness for Long-form Summarization
abstract
David Wan, Jesse Vig, Mohit Bansal, Shafiq Joty. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
David Wan, Jesse Vig, Mohit Bansal, Shafiq R. Joty
NAACL (Long Papers)1
2024 Contrastive Region Guidance: Improving Grounding in Vision-Language Models Without Training
David Wan, Jaemin Cho 0001, Elias Stengel-Eskin, Mohit Bansal
ECCV (79)1
2023 Extractive is not Faithful: An Investigation of Broad Unfaithfulness Problems in Extractive Summarization
abstract
The problems of unfaithful summaries have been widely discussed under the context of abstractive summarization.Though extractive summarization is less prone to the common unfaithfulness issues of abstractive summaries, does that mean extractive is equal to faithful?Turns out that the answer is no.In this work, we define a typology with five types of broad unfaithfulness problems (including and beyond not-entailment) that can appear in extractive summaries, including incorrect coreference, incomplete coreference, incorrect discourse, incomplete discourse, as well as other misleading information.We ask humans to label these problems out of 1600 English summaries produced by 16 diverse extractive systems.We find that 30% of the summaries have at least one of the five issues.To automatically detect these problems, we find that 5 existing faithfulness evaluation metrics for summarization have poor correlations with human judgment.To remedy this, we propose a new metric, EXTEVAL, that is designed for detecting unfaithful extractive summaries and is shown to have the best performance.We hope our work can increase the awareness of unfaithfulness problems in extractive summarization and help future work to evaluate and resolve these issues.1 * Equal contribution. 1 Our data and code are publicly available at https: //github.com/ZhangShiyue/extractive_is_ not_faithful. Document:(CNN) Most climbers who try don't succeed in summiting the 29,035-foot-high Mount Everest, the world's tallest peak.But they do leave their trash.Thousands of pounds of it.That's why an experienced climbing group from the Indian army plans to trek up the 8,850-meter mountain to pick up at least 4,000 kilograms (more than 8,000 pounds) of waste from the high-altitude camps, according to India Today.The mountain is part of the Himalaya mountain range on the border between Nepal and the Tibet region.The 34-member team plans to depart for Kathmandu on Saturday and start the ascent in mid-May.The upcoming trip marks the 50th anniversary of the first Indian team to scale Mount Everest [...]More than 200 climbers have died attempting to climb the peak, part of a UNESCO World Heritage Site.The Indian expedition isn't the first attempt to clean up the trash left by generations of hikers[...] Summary 1 (incorrect coreference): (CNN) Most climbers who try don't succeed in summiting the 29,035-foot-high Mount Everest, the world's tallest peak.That's why an experienced climbing group from the Indian army plans to trek up the 8,850-meter mountain to pick up at least 4,000 kilograms (more than 8,000 pounds) of waste from the high-altitude camps, according to India Today.[...] Summary 2 (incomplete coreference & incorrect discourse) : That's why an experienced climbing group from the Indian army plans to trek up the 8,850-meter mountain to pick up at least 4,000 kilograms More than 200 climbers have died to clean up the trash [...] Summary 3 (incomplete discourse & incomplete coreference): But they do leave their trash.Thousands of pounds of it.[...
Shiyue Zhang 0001, David Wan, Mohit Bansal
ACL (1)2
2023 Faithfulness-Aware Decoding Strategies for Abstractive Summarization
abstract
Despite significant progress in understanding and improving faithfulness in abstractive summarization, the question of how decoding strategies affect faithfulness is less studied.We present a systematic study of the effect of generation techniques such as beam search and nucleus sampling on faithfulness in abstractive summarization.We find a consistent trend where beam search with large beam sizes produces the most faithful summaries while nucleus sampling generates the least faithful ones.We propose two faithfulness-aware generation methods to further improve faithfulness over current generation techniques: (1) ranking candidates generated by beam search using automatic faithfulness metrics and (2) incorporating lookahead heuristics that produce a faithfulness score on the future summary.We show that both generation methods significantly improve faithfulness across two datasets as evaluated by four automatic faithfulness metrics and human evaluation.To reduce computational cost, we demonstrate a simple distillation approach that allows the model to generate faithful summaries with just greedy decoding.1
David Wan, Mengwen Liu, Kathy McKeown, Markus Dreyer, Mohit Bansal
EACL1
2023 HistAlign: Improving Context Dependency in Language Generation by Aligning with History
abstract
Language models (LMs) can generate hallucinations and incoherent outputs, which highlights their weak context dependency.Cache-LMs, which augment LMs with a memory of recent history, can increase context dependency and have shown remarkable performance in diverse language generation tasks.However, we find that even with training, the performance gain stemming from the cache component of current cache-LMs is suboptimal due to the misalignment between the current hidden states and those stored in the memory.In this work, we present HISTALIGN, a new training approach to ensure good cache alignment such that the model receives useful signals from the history.We first prove our concept on a simple and synthetic task where the memory is essential for correct predictions, and we show that the cache component of HISTALIGN is better aligned and improves overall performance.Next, we evaluate HISTALIGN on diverse downstream language generation tasks, including prompt continuation, abstractive summarization, and data-to-text.We demonstrate that HISTALIGN improves text coherence and faithfulness in open-ended and conditional generation settings, respectively.HISTALIGN is also generalizable across different model families, showcasing its strength in improving context dependency of LMs in diverse scenarios.1
David Wan, Shiyue Zhang 0001, Mohit Bansal
EMNLP1
2022 Constrained Regeneration for Cross-Lingual Query-Focused Extractive Summarization
abstract
Query-focused summaries of foreign-language, retrieved documents can help a user understand whether a document is actually relevant to the query term. A standard approach to this problem is to first translate the source documents and then perform extractive summarization to find relevant snippets. However, in a cross-lingual setting, the query term does not necessarily appear in the translations of relevant documents. In this work, we show that constrained machine translation and constrained post-editing can improve human relevance judgments by including a query term in a summary when its translation appears in the source document. We also present several strategies for selecting only certain documents for regeneration which yield further improvements
Elsbeth Turcan, David Wan, Faisal Ladhak, Petra Galuscáková, Sukanta Sen, Svetlana Tchistiakova, Weijia Xu, Marine Carpuat, Kenneth Heafield, Douglas W. Oard, Kathy McKeown
COLING2
2022 Evaluating and Improving Factuality in Multimodal Abstractive Summarization
abstract
Current metrics for evaluating factuality for abstractive document summarization have achieved high correlations with human judgment, but they do not account for the vision modality and thus are not adequate for visionand-language summarization.We propose CLIPBERTSCORE, a simple weighted combination of CLIPScore (Hessel et al., 2021) and BERTScore (Zhang* et al., 2020) to leverage the robustness and strong factuality detection performance between image-summary and document-summary, respectively.Next, due to the lack of meta-evaluation benchmarks to evaluate the quality of multimodal factuality metrics, we collect human judgments of factuality with respect to documents and images.We show that this simple combination of two metrics in the zero-shot setting achieves higher correlations than existing factuality metrics for document summarization, outperforms an existing multimodal summarization metric, and performs competitively with strong multimodal factuality metrics specifically fine-tuned for the task.Our thorough analysis demonstrates the robustness and high correlation of CLIP-BERTSCORE and its components on four factuality metric-evaluation benchmarks.Finally, we demonstrate two practical downstream applications of our CLIPBERTSCORE metric: for selecting important images to focus on during training, and as a reward for reinforcement learning to improve factuality of multimodal summary generation w.r.t automatic and human evaluation. 1
David Wan, Mohit Bansal
EMNLP1
2022 FactPEGASUS: Factuality-Aware Pre-training and Fine-tuning for Abstractive Summarization
abstract
We present FACTPEGASUS, an abstractive summarization model that addresses the problem of factuality during pre-training and finetuning: (1) We augment the sentence selection strategy of PEGASUS's (Zhang et al., 2020) pre-training objective to create pseudosummaries that are both important and factual;(2) We introduce three complementary components for fine-tuning.The corrector removes hallucinations present in the reference summary, the contrastor uses contrastive learning to better differentiate nonfactual summaries from factual ones, and the connector bridges the gap between the pre-training and finetuning for better transfer of knowledge.Experiments on three downstream tasks demonstrate that FACTPEGASUS substantially improves factuality evaluated by multiple automatic metrics and humans.Our thorough analysis suggests that FACTPEGASUS is more factual than using the original pre-training objective in zero-shot and few-shot settings, retains factual behavior more robustly than strong baselines, and does not rely entirely on becoming more extractive to improve factuality. 1
David Wan, Mohit Bansal
NAACL-HLT1
2021 Segmenting Subtitles for Correcting ASR Segmentation Errors
abstract
David Wan, Chris Kedzie, Faisal Ladhak, Elsbeth Turcan, Petra Galuscakova, Elena Zotkina, Zhengping Jiang, Peter Bell, Kathleen McKeown. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
David Wan, Chris Kedzie, Faisal Ladhak, Elsbeth Turcan, Petra Galuscáková, Elena Zotkina, Zhengping Jiang, Peter Bell 0001, Kathy McKeown
EACL1