VLDB 2026 Research / reviewers in the wild / expert
Philippe Laban
dblp:220/3590
· DBLP profile ↗
22ranked-venue papers
11as first author
21since 2021 · last 2025
0000-0001-9685-3961ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 8 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | KRISTEVA: Close Reading as a Novel Task for Benchmarking Interpretive ReasoningabstractPeiqi Sui, Juan Diego Rodriguez, Philippe Laban, J. Dean Murphy, Joseph P. Dexter, Richard Jean So, Samuel Baker, Pramit Chaudhuri. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Peiqi Sui, Juan Diego Rodriguez, Philippe Laban, Dean Murphy, Joseph P. Dexter, Richard Jean So, Samuel Baker, Pramit Chaudhuri |
ACL (1) | 3 |
| 2025 | Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through Edits
Tuhin Chakrabarty, Philippe Laban, Chien-Sheng Wu |
CHI | 2 |
| 2025 | BingoGuard: LLM Content Moderation Tools with Risk LevelsabstractMalicious content generated by large language models (LLMs) can pose varying degrees of harm.
Although existing LLM-based moderators can detect harmful content, they struggle to assess risk levels and may miss lower-risk outputs.
Accurate risk assessment allows platforms with different safety thresholds to tailor content filtering and rejection. In this paper, we introduce per-topic severity rubrics for 11 harmful topics and build BingoGuard, an LLM-based moderation system designed to predict both binary safety labels and severity levels.
To address the lack of annotations on levels of severity, we propose a scalable generate-then-filter framework that first generates responses across different severity levels and then filters out low-quality responses. Using this framework, we create BingoGuardTrain, a training dataset with 54,897 examples covering a variety of topics, response severity, styles, and BingoGuardTest, a test set with 988 examples explicitly labeled based on our severity rubrics that enables fine-grained analysis on model behaviors on different severity levels. Our BingoGuard-8B, trained on BingoGuardTrain, achieves the state-of-the-art performance on several moderation benchmarks, including WildGuardTest and HarmBench, as well as BingoGuardTest, outperforming best public models, WildGuard, by 4.3\%. Our analysis demonstrates that incorporating severity levels into training significantly enhances detection performance and enables the model to effectively gauge the severity of harmful responses. Warning: this paper includes red-teaming examples that may be harmful in nature. Fan Yin, Philippe Laban, Yilun Zhou, Yixin Mao, Vaibhav Vats, Linnea Ross, Divyansh Agarwal, Caiming Xiong, Chien-Sheng Wu |
ICLR | 2 |
| 2025 | CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic EnvironmentsabstractKung-Hsiang Huang, Akshara Prabhakar, Sidharth Dhawan, Yixin Mao, Huan Wang, Silvio Savarese, Caiming Xiong, Philippe Laban, Chien-Sheng Wu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kung-Hsiang Huang, Akshara Prabhakar, Sidharth Dhawan, Yixin Mao, Huan Wang 0016, Silvio Savarese, Caiming Xiong, Philippe Laban, Chien-Sheng Wu |
NAACL (Long Papers) | 8 |
| 2025 | Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question CoverageabstractKaige Xie, Philippe Laban, Prafulla Kumar Choubey, Caiming Xiong, Chien-Sheng Wu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kaige Xie, Philippe Laban, Prafulla Kumar Choubey, Caiming Xiong, Chien-Sheng Wu |
NAACL (Long Papers) | 2 |
| 2024 | Art or Artifice? Large Language Models and the False Promise of CreativityabstractResearchers have argued that large language models (LLMs) exhibit high-quality writing capabilities from blogs to stories. However, evaluating objectively the creativity of a piece of writing is challenging. Inspired by the Torrance Test of Creative Thinking (TTCT) [64], which measures creativity as a process, we use the Consensual Assessment Technique [3] and propose Torrance Test of Creative Writing (TTCW) to evaluate creativity as product. TTCW consists of 14 binary tests organized into the original dimensions of Fluency, Flexibility, Originality, and Elaboration. We recruit 10 creative writers and implement a human assessment of 48 stories written either by professional authors or LLMs using TTCW. Our analysis shows that LLM-generated stories pass 3-10X less TTCW tests than stories written by professionals. In addition, we explore the use of LLMs as assessors to automate the TTCW evaluation, revealing that none of the LLMs positively correlate with the expert assessments. Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, Chien-Sheng Wu |
CHI | 2 |
| 2024 | Summary of a Haystack: A Challenge to Long-Context LLMs and RAG SystemsabstractLLMs and RAG systems are now capable of handling millions of input tokens or more.However, evaluating the output quality of such systems on long-context tasks remains challenging, as tasks like Needle-in-a-Haystack lack complexity.In this work, we argue that summarization can play a central role in such evaluation.We design a procedure to synthesize Haystacks of documents, ensuring that specific insights repeat across documents.The "Summary of a Haystack" (SummHay) task then requires a system to process the Haystack and generate, given a query, a summary that identifies the relevant insights and precisely cites the source documents.Since we have precise knowledge of what insights should appear in a haystack summary and what documents should be cited, we implement a highly reproducible automatic evaluation that can score summaries on two aspects -Coverage and Citation.We generate Haystacks in two domains (conversation, news), and perform a large-scale evaluation of 14 LLMs and corresponding 50 RAG systems.Our findings indicate that SummHay is an open challenge for current systems, as even systems provided with an Oracle signal of document relevance lag our estimate of human performance (56%) by 10+ points on a Joint Score.Without a retriever, long-context LLMs like GPT-4o and Claude 3 Opus score below 20% on SummHay.We show SummHay can also be used to study enterprise RAG systems and position bias in long-context models.We hope future systems can equal and surpass human performance on SummHay. Philippe Laban, Alexander R. Fabbri, Caiming Xiong, Chien-Sheng Wu |
EMNLP | 1 |
| 2024 | MiniCheck: Efficient Fact-Checking of LLMs on Grounding DocumentsabstractRecognizing if LLM output can be grounded in evidence is central to many tasks in NLP: retrieval-augmented generation, summarization, document-grounded dialogue, and more.Current approaches to this kind of factchecking are based on verifying each piece of a model generation against potential evidence using an LLM.However, this process can be very computationally expensive, requiring many calls to a model to check a single response.In this work, we show how to build small fact-checking models that have GPT-4level performance but for 400x lower cost.We do this by constructing synthetic training data with GPT-4, which involves creating realistic yet challenging instances of factual errors via a structured generation procedure.Training on this data teaches models to check each fact in the claim and recognize synthesis of information across sentences.For evaluation, we unify datasets from recent work on factchecking and grounding LLM generations into a new benchmark, LLM-AGGREFACT.Our best system MiniCheck-FT5 (770M parameters) outperforms all systems of comparable size and reaches GPT-4 accuracy.We release LLM-AGGREFACT, code for data synthesis, and models. 1 Liyan Tang, Philippe Laban, Greg Durrett |
EMNLP | 2 |
| 2024 | Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News ArticlesabstractKung-Hsiang Huang, Philippe Laban, Alexander Fabbri, Prafulla Kumar Choubey, Shafiq Joty, Caiming Xiong, Chien-Sheng Wu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Kung-Hsiang Huang, Philippe Laban, Alexander R. Fabbri, Prafulla Kumar Choubey, Shafiq R. Joty, Caiming Xiong, Chien-Sheng Wu |
NAACL-HLT | 2 |
| 2024 | Beyond the Chat: Executable and Verifiable Text-Editing with LLMsabstractConversational interfaces powered by Large Language Models (LLMs) have recently become a popular way to obtain feedback during document editing. However, standard chat-based conversational interfaces cannot explicitly surface the editing changes that they suggest. To give the author more control when editing with an LLM, we present InkSync, an editing interface that suggests executable edits directly within the document being edited. Because LLMs are known to introduce factual errors, Inksync also supports a 3-stage approach to mitigate this risk: Warn authors when a suggested edit introduces new information, help authors Verify the new information’s accuracy through external search, and allow a third party to Audit with a-posteriori verification via a trace of all auto-generated content. Two usability studies confirm the effectiveness of InkSync’s components when compared to standard LLM-based chat interfaces, leading to more accurate and more efficient editing, and improved user experience. Philippe Laban, Jesse Vig, Marti A. Hearst, Caiming Xiong, Chien-Sheng Wu |
UIST | 1 |
| 2023 | SWiPE: A Dataset for Document-Level Simplification of Wikipedia PagesabstractPhilippe Laban, Jesse Vig, Wojciech Kryscinski, Shafiq Joty, Caiming Xiong, Chien-Sheng Wu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Philippe Laban, Jesse Vig, Wojciech Kryscinski, Shafiq R. Joty, Caiming Xiong, Chien-Sheng Wu |
ACL (1) | 1 |
| 2023 | Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error DetectorsabstractLiyan Tang, Tanya Goyal, Alex Fabbri, Philippe Laban, Jiacheng Xu, Semih Yavuz, Wojciech Kryscinski, Justin Rousseau, Greg Durrett. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Liyan Tang, Tanya Goyal, Alexander R. Fabbri, Philippe Laban, Jiacheng Xu 0001, Semih Yavuz, Wojciech Kryscinski, Justin F. Rousseau, Greg Durrett |
ACL (1) | 4 |
| 2023 | Did You Read the Instructions? Rethinking the Effectiveness of Task Definitions in Instruction LearningabstractLarge language models (LLMs) have shown impressive performance in following natural language instructions to solve unseen tasks.However, it remains unclear whether models truly understand task definitions and whether the human-written definitions are optimal.In this paper, we systematically study the role of task definitions in instruction learning.We first conduct an ablation analysis informed by human annotations to understand which parts of a task definition are most important, and find that model performance only drops substantially when removing contents describing the task output, in particular label information.Next, we propose an automatic algorithm to compress task definitions to a minimal supporting set of tokens, and find that 60% of tokens can be removed while maintaining or even improving model performance.Based on these results, we propose two strategies to help models better leverage task instructions: (1) providing only key information for tasks in a common structured format, and (2) adding a metatuning stage to help the model better understand the definitions.With these two strategies, we achieve a 4.2 Rouge-L improvement over 119 unseen test tasks. Fan Yin, Jesse Vig, Philippe Laban, Shafiq R. Joty, Caiming Xiong, Chien-Sheng Wu |
ACL (1) | 3 |
| 2023 | Designing and Evaluating Interfaces that Highlight News Coverage Diversity Using Discord QuestionsabstractModern news aggregators do the hard work of organizing a large news stream, creating collections for a given news story with tens of source options. This paper shows that navigating large source collections for a news story can be challenging without further guidance. In this work, we design three interfaces – the Annotated Article, the Recomposed Article, and the Question Grid – aimed at accompanying news readers in discovering coverage diversity while they read. A first usability study with 10 journalism experts confirms the designed interfaces all reveal coverage diversity and determine each interface’s potential use cases and audiences. In a second usability study, we developed and implemented a reading exercise with 95 novice news readers to measure exposure to coverage diversity. Results show that Annotated Article users are able to answer questions 34% more completely than with two existing interfaces while finding the interface equally easy to use. Philippe Laban, Chien-Sheng Wu, Lidiya Murakhovs'ka, Xiang 'Anthony' Chen, Caiming Xiong |
CHI | 1 |
| 2023 | SummEdits: Measuring LLM Ability at Factual Reasoning Through The Lens of SummarizationabstractPhilippe Laban, Wojciech Kryscinski, Divyansh Agarwal, Alexander Fabbri, Caiming Xiong, Shafiq Joty, Chien-Sheng Wu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Philippe Laban, Wojciech Kryscinski, Divyansh Agarwal, Alexander R. Fabbri, Caiming Xiong, Shafiq R. Joty, Chien-Sheng Wu |
EMNLP | 1 |
| 2023 | Marvista: Exploring the Design of a Human-AI Collaborative News Reading ToolabstractWe explore the design of Marvista—a human-AI collaborative tool that employs a suite of natural language processing models to provide end-to-end support for reading online news articles. Before reading an article, Marvista helps a user plan what to read by filtering text based on how much time one can spend and what questions one is interested to find out from the article. During reading, Marvista helps the user reflect on their understanding of each paragraph with AI-generated questions. After reading, Marvista generates an explainable human-AI summary that combines AI’s processing of the text, the user’s reading behavior, and user-generated data in the reading process. In contrast to prior work that offered (content-independent) interaction techniques or devices for reading, Marvista takes a human-AI collaborative approach that contributes text-specific guidance (content-aware) to support the entire reading process. Xiang 'Anthony' Chen, Chien-Sheng Wu, Lidiya Murakhovs'ka, Philippe Laban, Wenhao Liu 0003, Caiming Xiong |
ACM Trans. Comput. Hum. Interact. | 4 |
| 2022 | Near-Negative Distinction: Giving a Second Life to Human Evaluation DatasetsabstractPrecisely assessing the progress in natural language generation (NLG) tasks is challenging, and human evaluation to establish a preference in a model's output over another is often necessary.However, human evaluation is usually costly, difficult to reproduce, and non-reusable.In this paper, we propose a new and simple automatic evaluation method for NLG called Near-Negative Distinction (NND) that repurposes prior human annotations into NND tests.In an NND test, an NLG model must place a higher likelihood on a high-quality output candidate than on a near-negative candidate with a known error.Model performance is established by the number of NND tests a model passes, as well as the distribution over task-specific errors the model fails on.Through experiments on three NLG tasks (question generation, question answering, and summarization), we show that NND achieves a higher correlation with human judgments than standard NLG evaluation metrics.We then illustrate NND evaluation in four practical scenarios, for example performing fine-grain model analysis, or studying model training dynamics.Our findings suggest that NND can give a second life to human annotations and provide low-cost NLG evaluation. Philippe Laban, Chien-Sheng Wu, Wenhao Liu 0003, Caiming Xiong |
EMNLP | 1 |
| 2022 | NewsPod: Automatic and Interactive News PodcastsabstractNews podcasts are a popular medium to stay informed and dive deep into news topics. Today, most podcasts are handcrafted by professionals. In this work, we advance the state-of-the-art in automatically generated podcasts, making use of recent advances in natural language processing and text-to-speech technology. We present NewsPod, an automatically generated, interactive news podcast. The podcast is divided into segments, each centered on a news event, with each segment structured as a Question and Answer conversation, whose goal is to engage the listener. A key aspect of the design is the use of distinct voices for each role (questioner, responder), to better simulate a conversation. Another novel aspect of NewsPod allows listeners to interact with the podcast by asking their own questions and receiving automatically generated answers. We validate the soundness of this system design through two usability studies, focused on evaluating the narrative style and interactions with the podcast, respectively. We find that NewsPod is preferred over a baseline by participants, with 80% claiming they would use the system in the future. Philippe Laban, Elicia Ye, Srujay Korlakunta, John F. Canny, Marti A. Hearst |
IUI | 1 |
| 2022 | SummaC: Re-Visiting NLI-based Models for Inconsistency Detection in SummarizationabstractAbstract In the summarization domain, a key requirement for summaries is to be factually consistent with the input document. Previous work has found that natural language inference (NLI) models do not perform competitively when applied to inconsistency detection. In this work, we revisit the use of NLI for inconsistency detection, finding that past work suffered from a mismatch in input granularity between NLI datasets (sentence-level), and inconsistency detection (document level). We provide a highly effective and light-weight method called SummaCConv that enables NLI models to be successfully used for this task by segmenting documents into sentence units and aggregating scores between pairs of sentences. We furthermore introduce a new benchmark called SummaC (Summary Consistency) which consists of six large inconsistency detection datasets. On this dataset, SummaCConv obtains state-of-the-art results with a balanced accuracy of 74.4%, a 5% improvement compared with prior work. Philippe Laban, Tobias Schnabel, Paul N. Bennett, Marti A. Hearst |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | Keep It Simple: Unsupervised Simplification of Multi-Paragraph TextabstractPhilippe Laban, Tobias Schnabel, Paul Bennett, Marti A. Hearst. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Philippe Laban, Tobias Schnabel, Paul N. Bennett, Marti A. Hearst |
ACL/IJCNLP (1) | 1 |
| 2021 | News Headline Grouping as a Challenging NLU TaskabstractRecent progress in Natural Language Understanding (NLU) has seen the latest models outperform human performance on many standard tasks.These impressive results have led the community to introspect on dataset limitations, and iterate on more nuanced challenges.In this paper, we introduce the task of HeadLine Grouping (HLG) and a corresponding dataset (HLGD) consisting of 20,056 pairs of news headlines, each labeled with a binary judgement as to whether the pair belongs within the same group.On HLGD, human annotators achieve high performance of around 0.9 F-1, while current state-of-the art Transformer models only reach 0.75 F-1, opening the path for further improvements.We further propose a novel unsupervised Headline Generator Swap model for the task of HeadLine Grouping that achieves within 3 F-1 of the best supervised model.Finally, we analyze highperforming models with consistency tests, and find that models are not consistent in their predictions, revealing modeling limits of current architectures. Philippe Laban, Lucas Bandarkar, Marti A. Hearst |
NAACL-HLT | 1 |
| 2020 | The Summary Loop: Learning to Write Abstractive Summaries Without ExamplesabstractThis work presents a new approach to unsupervised abstractive summarization based on maximizing a combination of coverage and fluency for a given length constraint.It introduces a novel method that encourages the inclusion of key terms from the original document into the summary: key terms are masked out of the original document and must be filled in by a coverage model using the current generated summary.A novel unsupervised training procedure leverages this coverage model along with a fluency model to generate and score summaries.When tested on popular news summarization datasets, the method outperforms previous unsupervised methods by more than 2 R-1 points, and approaches results of competitive supervised methods.Our model attains higher levels of abstraction with copied passages roughly two times shorter than prior work, and learns to compress and merge sentences without supervision. Philippe Laban, Andrew Hsi, John F. Canny, Marti A. Hearst |
ACL | 1 |