EDBT 2026 Demo / reviewers in the wild / expert
Chien-Sheng Wu
dblp:180/5537 · also Chien-Sheng Jason Wu
· DBLP profile ↗
48ranked-venue papers
7as first author
31since 2021 · last 2026
0000-0002-5598-5324ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 6 first-author · 25 since 2021Human-computer interaction and ubiquitous computing · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GTA: Generating Long-horizon Tasks for Web Agents at ScaleabstractTenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou, Muhao Chen, Jonathan May, Chien-Sheng Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou, Muhao Chen 0001, Jonathan May, Chien-Sheng Wu |
ACL (1) | 7 |
| 2025 | Unanswerability Evaluation for Retrieval Augmented GenerationabstractExisting evaluation frameworks for retrievalaugmented generation (RAG) systems focus on answerable queries, but they overlook the importance of appropriately rejecting unanswerable requests.In this paper, we introduce UAEval4RAG, a comprehensive evaluation framework designed to evaluate whether RAG systems effectively handle unanswerable queries specific to a given knowledge base.We first define a taxonomy with six unanswerable categories, and UAEval4RAG automatically synthesizes diverse and challenging queries for any given knowledge base and evaluate the RAG systems with unanswered ratio and acceptable ratio metrics.We also conduct experiments with various RAG components and prompting strategies across four datasets, which reveals that due to varying knowledge distribution across datasets, no single configuration consistently delivers optimal performance on both answerable and unanswerable requests across different knowledge bases.Our findings highlight the critical role of component selection and prompt design in optimizing RAG systems to balance the accuracy of answerable queries with high rejection rates of unanswerable ones.UAEval4RAG provides valuable insights and tools for developing more robust and reliable RAG systems. Prafulla Kumar Choubey, Caiming Xiong, Chien-Sheng Wu |
ACL (1) | 4 |
| 2025 | Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through Edits
Tuhin Chakrabarty, Philippe Laban, Chien-Sheng Wu |
CHI | 3 |
| 2025 | Proactive Conversational Agents with Inner Thoughts
Xingyu Liu 0002, Shitao Fang, Weiyan Shi 0001, Chien-Sheng Wu, Takeo Igarashi, Xiang 'Anthony' Chen |
CHI | 4 |
| 2025 | ReGenesis: LLMs can Grow into Reasoning Generalists via Self-ImprovementabstractPost-training Large Language Models (LLMs) with explicit reasoning trajectories can enhance their reasoning abilities. However, acquiring such high-quality trajectory data typically demands meticulous supervision from humans or superior models, which can be either expensive or license-constrained. In this paper, we explore how far an LLM can improve its reasoning by self-synthesizing reasoning paths as training data without any additional supervision. Existing self-synthesizing methods, such as STaR, suffer from poor generalization to out-of-domain (OOD) reasoning tasks. We hypothesize it is due to that their self-synthesized reasoning paths are too task-specific, lacking general task-agnostic reasoning guidance. To address this, we propose **Reasoning Generalist via Self-Improvement (ReGenesis)**, a method to *self-synthesize reasoning paths as post-training data by progressing from abstract to concrete*. More specifically, ReGenesis self-synthesizes reasoning paths by converting general reasoning guidelines into task-specific ones, generating reasoning structures, and subsequently transforming these structures into reasoning paths, without the need for human-designed task-specific examples used in existing methods. We show that ReGenesis achieves superior performance on all in-domain and OOD settings tested compared to existing methods. For six OOD tasks specifically, while previous methods exhibited an average performance decrease of approximately 4.6% after post training, ReGenesis delivers around 6.1% performance improvement. We also conduct an in-depth analysis of our framework and show ReGenesis is effective across various language models and design choices. Congying Xia, Xinyi Yang 0002, Caiming Xiong, Chien-Sheng Wu, Chen Xing |
ICLR | 5 |
| 2025 | BingoGuard: LLM Content Moderation Tools with Risk LevelsabstractMalicious content generated by large language models (LLMs) can pose varying degrees of harm.
Although existing LLM-based moderators can detect harmful content, they struggle to assess risk levels and may miss lower-risk outputs.
Accurate risk assessment allows platforms with different safety thresholds to tailor content filtering and rejection. In this paper, we introduce per-topic severity rubrics for 11 harmful topics and build BingoGuard, an LLM-based moderation system designed to predict both binary safety labels and severity levels.
To address the lack of annotations on levels of severity, we propose a scalable generate-then-filter framework that first generates responses across different severity levels and then filters out low-quality responses. Using this framework, we create BingoGuardTrain, a training dataset with 54,897 examples covering a variety of topics, response severity, styles, and BingoGuardTest, a test set with 988 examples explicitly labeled based on our severity rubrics that enables fine-grained analysis on model behaviors on different severity levels. Our BingoGuard-8B, trained on BingoGuardTrain, achieves the state-of-the-art performance on several moderation benchmarks, including WildGuardTest and HarmBench, as well as BingoGuardTest, outperforming best public models, WildGuard, by 4.3\%. Our analysis demonstrates that incorporating severity levels into training significantly enhances detection performance and enables the model to effectively gauge the severity of harmful responses. Warning: this paper includes red-teaming examples that may be harmful in nature. Fan Yin, Philippe Laban, Yilun Zhou, Yixin Mao, Vaibhav Vats, Linnea Ross, Divyansh Agarwal, Caiming Xiong, Chien-Sheng Wu |
ICLR | 10 |
| 2025 | SiReRAG: Indexing Similar and Related Information for Multihop ReasoningabstractIndexing is an important step towards strong performance in retrieval-augmented generation (RAG) systems. However, existing methods organize data based on either semantic similarity (similarity) or related information (relatedness), but do not cover both perspectives comprehensively. Our analysis reveals that modeling only one perspective results in insufficient knowledge synthesis, leading to suboptimal performance on complex tasks requiring multihop reasoning. In this paper, we propose SiReRAG, a novel RAG indexing approach that explicitly considers both similar and related information. On the similarity side, we follow existing work and explore some variances to construct a similarity tree based on recursive summarization. On the relatedness side, SiReRAG extracts propositions and entities from texts, groups propositions via shared entities, and generates recursive summaries to construct a relatedness tree. We index and flatten both similarity and relatedness trees into a unified retrieval pool. Our experiments demonstrate that SiReRAG consistently outperforms state-of-the-art indexing methods on three multihop datasets (MuSiQue, 2WikiMultiHopQA, and HotpotQA), with an average 1.9% improvement in F1 scores. As a reasonably efficient solution, SiReRAG enhances existing reranking methods significantly, with up to 7.8% improvement in average F1 scores. Our code is available at https://github.com/SalesforceAIResearch/SiReRAG. Prafulla Kumar Choubey, Alexander R. Fabbri, Gabriel Bernadett-Shapiro, Rui Zhang 0037, Prasenjit Mitra 0001, Caiming Xiong, Chien-Sheng Wu |
ICLR | 8 |
| 2025 | CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic EnvironmentsabstractKung-Hsiang Huang, Akshara Prabhakar, Sidharth Dhawan, Yixin Mao, Huan Wang, Silvio Savarese, Caiming Xiong, Philippe Laban, Chien-Sheng Wu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kung-Hsiang Huang, Akshara Prabhakar, Sidharth Dhawan, Yixin Mao, Huan Wang 0016, Silvio Savarese, Caiming Xiong, Philippe Laban, Chien-Sheng Wu |
NAACL (Long Papers) | 9 |
| 2025 | ReIFE: Re-evaluating Instruction-Following EvaluationabstractYixin Liu, Kejian Shi, Alexander Fabbri, Yilun Zhao, PeiFeng Wang, Chien-Sheng Wu, Shafiq Joty, Arman Cohan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yixin Liu 0003, Kejian Shi, Alexander R. Fabbri, Yilun Zhao 0001, Peifeng Wang, Chien-Sheng Wu, Shafiq R. Joty, Arman Cohan |
NAACL (Long Papers) | 6 |
| 2025 | Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question CoverageabstractKaige Xie, Philippe Laban, Prafulla Kumar Choubey, Caiming Xiong, Chien-Sheng Wu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kaige Xie, Philippe Laban, Prafulla Kumar Choubey, Caiming Xiong, Chien-Sheng Wu |
NAACL (Long Papers) | 5 |
| 2024 | Art or Artifice? Large Language Models and the False Promise of CreativityabstractResearchers have argued that large language models (LLMs) exhibit high-quality writing capabilities from blogs to stories. However, evaluating objectively the creativity of a piece of writing is challenging. Inspired by the Torrance Test of Creative Thinking (TTCT) [64], which measures creativity as a process, we use the Consensual Assessment Technique [3] and propose Torrance Test of Creative Writing (TTCW) to evaluate creativity as product. TTCW consists of 14 binary tests organized into the original dimensions of Fluency, Flexibility, Originality, and Elaboration. We recruit 10 creative writers and implement a human assessment of 48 stories written either by professional authors or LLMs using TTCW. Our analysis shows that LLM-generated stories pass 3-10X less TTCW tests than stories written by professionals. In addition, we explore the use of LLMs as assessors to automate the TTCW evaluation, revealing that none of the LLMs positively correlate with the expert assessments. Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, Chien-Sheng Wu |
CHI | 5 |
| 2024 | Summary of a Haystack: A Challenge to Long-Context LLMs and RAG SystemsabstractLLMs and RAG systems are now capable of handling millions of input tokens or more.However, evaluating the output quality of such systems on long-context tasks remains challenging, as tasks like Needle-in-a-Haystack lack complexity.In this work, we argue that summarization can play a central role in such evaluation.We design a procedure to synthesize Haystacks of documents, ensuring that specific insights repeat across documents.The "Summary of a Haystack" (SummHay) task then requires a system to process the Haystack and generate, given a query, a summary that identifies the relevant insights and precisely cites the source documents.Since we have precise knowledge of what insights should appear in a haystack summary and what documents should be cited, we implement a highly reproducible automatic evaluation that can score summaries on two aspects -Coverage and Citation.We generate Haystacks in two domains (conversation, news), and perform a large-scale evaluation of 14 LLMs and corresponding 50 RAG systems.Our findings indicate that SummHay is an open challenge for current systems, as even systems provided with an Oracle signal of document relevance lag our estimate of human performance (56%) by 10+ points on a Joint Score.Without a retriever, long-context LLMs like GPT-4o and Claude 3 Opus score below 20% on SummHay.We show SummHay can also be used to study enterprise RAG systems and position bias in long-context models.We hope future systems can equal and surpass human performance on SummHay. Philippe Laban, Alexander R. Fabbri, Caiming Xiong, Chien-Sheng Wu |
EMNLP | 4 |
| 2024 | Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News ArticlesabstractKung-Hsiang Huang, Philippe Laban, Alexander Fabbri, Prafulla Kumar Choubey, Shafiq Joty, Caiming Xiong, Chien-Sheng Wu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Kung-Hsiang Huang, Philippe Laban, Alexander R. Fabbri, Prafulla Kumar Choubey, Shafiq R. Joty, Caiming Xiong, Chien-Sheng Wu |
NAACL-HLT | 7 |
| 2024 | Beyond the Chat: Executable and Verifiable Text-Editing with LLMsabstractConversational interfaces powered by Large Language Models (LLMs) have recently become a popular way to obtain feedback during document editing. However, standard chat-based conversational interfaces cannot explicitly surface the editing changes that they suggest. To give the author more control when editing with an LLM, we present InkSync, an editing interface that suggests executable edits directly within the document being edited. Because LLMs are known to introduce factual errors, Inksync also supports a 3-stage approach to mitigate this risk: Warn authors when a suggested edit introduces new information, help authors Verify the new information’s accuracy through external search, and allow a third party to Audit with a-posteriori verification via a trace of all auto-generated content. Two usability studies confirm the effectiveness of InkSync’s components when compared to standard LLM-based chat interfaces, leading to more accurate and more efficient editing, and improved user experience. Philippe Laban, Jesse Vig, Marti A. Hearst, Caiming Xiong, Chien-Sheng Wu |
UIST | 5 |
| 2023 | SWiPE: A Dataset for Document-Level Simplification of Wikipedia PagesabstractPhilippe Laban, Jesse Vig, Wojciech Kryscinski, Shafiq Joty, Caiming Xiong, Chien-Sheng Wu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Philippe Laban, Jesse Vig, Wojciech Kryscinski, Shafiq R. Joty, Caiming Xiong, Chien-Sheng Wu |
ACL (1) | 6 |
| 2023 | Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human EvaluationabstractYixin Liu, Alex Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir Radev. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yixin Liu 0003, Alexander R. Fabbri, Pengfei Liu 0003, Yilun Zhao 0001, Linyong Nan, Ruilin Han, Simeng Han, Shafiq R. Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir R. Radev |
ACL (1) | 9 |
| 2023 | Socratic Pretraining: Question-Driven Pretraining for Controllable SummarizationabstractIn long document controllable summarization, where labeled data is scarce, pretrained models struggle to adapt to the task and effectively respond to user queries.In this paper, we introduce SOCRATIC pretraining, a question-driven, unsupervised pretraining objective specifically designed to improve controllability in summarization tasks.By training a model to generate and answer relevant questions in a given context, SOCRATIC pretraining enables the model to more effectively adhere to user-provided queries and identify relevant content to be summarized.We demonstrate the effectiveness of this approach through extensive experimentation on two summarization domains, short stories and dialogue, and multiple control strategies: keywords, questions, and factoid QA pairs.Our pretraining method relies only on unlabeled documents and a question generation system and outperforms pre-finetuning approaches that use additional supervised data.Furthermore, our results show that SOCRATIC pretraining cuts task-specific labeled data requirements in half, is more faithful to userprovided queries, and achieves state-of-the-art performance on QMSum and SQuALITY.Joseph L Fleiss.1971.Measuring nominal scale agreement among many raters. Artidoro Pagnoni, Alexander R. Fabbri, Wojciech Kryscinski, Chien-Sheng Wu |
ACL (1) | 4 |
| 2023 | Did You Read the Instructions? Rethinking the Effectiveness of Task Definitions in Instruction LearningabstractLarge language models (LLMs) have shown impressive performance in following natural language instructions to solve unseen tasks.However, it remains unclear whether models truly understand task definitions and whether the human-written definitions are optimal.In this paper, we systematically study the role of task definitions in instruction learning.We first conduct an ablation analysis informed by human annotations to understand which parts of a task definition are most important, and find that model performance only drops substantially when removing contents describing the task output, in particular label information.Next, we propose an automatic algorithm to compress task definitions to a minimal supporting set of tokens, and find that 60% of tokens can be removed while maintaining or even improving model performance.Based on these results, we propose two strategies to help models better leverage task instructions: (1) providing only key information for tasks in a common structured format, and (2) adding a metatuning stage to help the model better understand the definitions.With these two strategies, we achieve a 4.2 Rouge-L improvement over 119 unseen test tasks. Fan Yin, Jesse Vig, Philippe Laban, Shafiq R. Joty, Caiming Xiong, Chien-Sheng Wu |
ACL (1) | 6 |
| 2023 | Designing and Evaluating Interfaces that Highlight News Coverage Diversity Using Discord QuestionsabstractModern news aggregators do the hard work of organizing a large news stream, creating collections for a given news story with tens of source options. This paper shows that navigating large source collections for a news story can be challenging without further guidance. In this work, we design three interfaces – the Annotated Article, the Recomposed Article, and the Question Grid – aimed at accompanying news readers in discovering coverage diversity while they read. A first usability study with 10 journalism experts confirms the designed interfaces all reveal coverage diversity and determine each interface’s potential use cases and audiences. In a second usability study, we developed and implemented a reading exercise with 95 novice news readers to measure exposure to coverage diversity. Results show that Annotated Article users are able to answer questions 34% more completely than with two existing interfaces while finding the interface equally easy to use. Philippe Laban, Chien-Sheng Wu, Lidiya Murakhovs'ka, Xiang 'Anthony' Chen, Caiming Xiong |
CHI | 2 |
| 2023 | SummEdits: Measuring LLM Ability at Factual Reasoning Through The Lens of SummarizationabstractPhilippe Laban, Wojciech Kryscinski, Divyansh Agarwal, Alexander Fabbri, Caiming Xiong, Shafiq Joty, Chien-Sheng Wu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Philippe Laban, Wojciech Kryscinski, Divyansh Agarwal, Alexander R. Fabbri, Caiming Xiong, Shafiq R. Joty, Chien-Sheng Wu |
EMNLP | 7 |
| 2023 | Towards Interpretable and Efficient Automatic Reference-Based Summarization EvaluationabstractYixin Liu, Alexander Fabbri, Yilun Zhao, Pengfei Liu, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir Radev. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Yixin Liu 0003, Alexander R. Fabbri, Yilun Zhao 0001, Pengfei Liu 0003, Shafiq R. Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir R. Radev |
EMNLP | 6 |
| 2023 | Model ensemble instead of prompt fusion: a sample-specific knowledge transfer method for few-shot prompt tuning
Chen Xing, Prafulla Kumar Choubey, Chien-Sheng Wu, Caiming Xiong |
ICLR | 4 |
| 2023 | Marvista: Exploring the Design of a Human-AI Collaborative News Reading ToolabstractWe explore the design of Marvista—a human-AI collaborative tool that employs a suite of natural language processing models to provide end-to-end support for reading online news articles. Before reading an article, Marvista helps a user plan what to read by filtering text based on how much time one can spend and what questions one is interested to find out from the article. During reading, Marvista helps the user reflect on their understanding of each paragraph with AI-generated questions. After reading, Marvista generates an explainable human-AI summary that combines AI’s processing of the text, the user’s reading behavior, and user-generated data in the reading process. In contrast to prior work that offered (content-independent) interaction techniques or devices for reading, Marvista takes a human-AI collaborative approach that contributes text-specific guidance (content-aware) to support the entire reading process. Xiang 'Anthony' Chen, Chien-Sheng Wu, Lidiya Murakhovs'ka, Philippe Laban, Wenhao Liu 0003, Caiming Xiong |
ACM Trans. Comput. Hum. Interact. | 2 |
| 2022 | DialFact: A Benchmark for Fact-Checking in DialogueabstractFact-checking is an essential tool to mitigate the spread of misinformation and disinformation.We introduce the task of fact-checking in dialogue, which is a relatively unexplored area.We construct DIALFACT, a testing benchmark dataset of 22,245 annotated conversational claims, paired with pieces of evidence from Wikipedia.There are three sub-tasks in DIALFACT: 1) Verifiable claim detection task distinguishes whether a response carries verifiable factual information; 2) Evidence retrieval task retrieves the most relevant Wikipedia snippets as evidence; 3) Claim verification task predicts a dialogue response to be supported, refuted, or not enough information.We found that existing fact-checking models trained on non-dialogue data like FEVER (Thorne et al., 2018) fail to perform well on our task, and thus, we propose a simple yet data-efficient solution to effectively improve fact-checking performance in dialogue.We point out unique challenges in DIALFACT such as handling the colloquialisms, coreferences and retrieval ambiguities in the error analysis to shed light on future research in this direction 1 .Dialogue Context: I have family in Ireland!Have you ever been there?Evidence: Ireland is an island in the North Atlantic.Non-Verifiable Response: I haven't been but want to!Verifiable Supported Response: I haven't.It is an island in the north Atlantic right?Verifiable Refuted Response: I haven't been.Isn't it somewhere in north Pacific?Verifiable NEI Response: I haven't been.I heard it's the most popular tourist location in Europe! Prakhar Gupta, Chien-Sheng Wu, Wenhao Liu 0003, Caiming Xiong |
ACL (1) | 2 |
| 2022 | QAConv: Question Answering on Informative ConversationsabstractThis paper introduces QAConv, 1 , a new question answering (QA) dataset that uses conversations as a knowledge source.We focus on informative conversations, including business emails, panel discussions, and work channels.Unlike open-domain and task-oriented dialogues, these conversations are usually long, complex, asynchronous, and involve strong domain knowledge.In total, we collect 34,608 QA pairs from 10,259 selected conversations with both human-written and machinegenerated questions.We use a question generator and a dialogue summarizer as auxiliary tools to collect and recommend questions.The dataset has two testing scenarios: chunk mode and full mode, depending on whether the grounded partial conversation is provided or retrieved.Experimental results show that stateof-the-art pretrained QA systems have limited zero-shot performance and tend to predict our questions as unanswerable.Our dataset provides a new training and evaluation testbed to facilitate QA on conversations research. Chien-Sheng Wu, Andrea Madotto, Wenhao Liu 0003, Pascale Fung, Caiming Xiong |
ACL (1) | 1 |
| 2022 | Conformal Predictor for Improving Zero-Shot Text Classification EfficiencyabstractPre-trained language models (PLMs) have been shown effective for zero-shot (0shot) text classification.0shot models based on natural language inference (NLI) and next sentence prediction (NSP) employ cross-encoder architecture and infer by making a forward pass through the model for each label-text pair separately.This increases the computational cost to make inferences linearly in the number of labels.In this work, we improve the efficiency of such cross-encoder-based 0shot models by restricting the number of likely labels using another fast base classifier-based conformal predictor (CP) calibrated on samples labeled by the 0shot model.Since a CP generates prediction sets with coverage guarantees, it reduces the number of target labels without excluding the most probable label based on the 0shot model.We experiment with three intent and two topic classification datasets.With a suitable CP for each dataset, we reduce the average inference time for NLI-and NSP-based models by 25.6% and 22.2% respectively, without dropping performance below the predefined error rate of 1%. Prafulla Kumar Choubey, Yu Bai 0017, Chien-Sheng Wu, Wenhao Liu 0003, Nazneen Fatema Rajani |
EMNLP | 3 |
| 2022 | Improving Factual Consistency in Summarization with Compression-Based Post-EditingabstractState-of-the-art summarization models still struggle to be factually consistent with the input text.A model-agnostic way to address this problem is post-editing the generated summaries.However, existing approaches typically fail to remove entity errors if a suitable input entity replacement is not available or may insert erroneous content.In our work, we focus on removing extrinsic entity errors, or entities not in the source, to improve consistency while retaining the summary's essential information and form.We propose to use sentence-compression data to train the post-editing model to take a summary with extrinsic entity errors marked with special tokens and output a compressed, well-formed summary with those errors removed.We show that this model improves factual consistency while maintaining ROUGE, improving entity precision by up to 30% on XSum, and that this model can be applied on top of another post-editor, improving entity precision by up to a total of 38%.We perform an extensive comparison of post-editing approaches that demonstrate trade-offs between factual consistency, informativeness, and grammaticality, and we analyze settings where posteditors show the largest improvements. Alexander R. Fabbri, Prafulla Kumar Choubey, Jesse Vig, Chien-Sheng Wu, Caiming Xiong |
EMNLP | 4 |
| 2022 | Near-Negative Distinction: Giving a Second Life to Human Evaluation DatasetsabstractPrecisely assessing the progress in natural language generation (NLG) tasks is challenging, and human evaluation to establish a preference in a model's output over another is often necessary.However, human evaluation is usually costly, difficult to reproduce, and non-reusable.In this paper, we propose a new and simple automatic evaluation method for NLG called Near-Negative Distinction (NND) that repurposes prior human annotations into NND tests.In an NND test, an NLG model must place a higher likelihood on a high-quality output candidate than on a near-negative candidate with a known error.Model performance is established by the number of NND tests a model passes, as well as the distribution over task-specific errors the model fails on.Through experiments on three NLG tasks (question generation, question answering, and summarization), we show that NND achieves a higher correlation with human judgments than standard NLG evaluation metrics.We then illustrate NND evaluation in four practical scenarios, for example performing fine-grain model analysis, or studying model training dynamics.Our findings suggest that NND can give a second life to human annotations and provide low-cost NLG evaluation. Philippe Laban, Chien-Sheng Wu, Wenhao Liu 0003, Caiming Xiong |
EMNLP | 2 |
| 2022 | UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language ModelsabstractTianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang, Noah A. Smith, Luke Zettlemoyer, Tao Yu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Tianbao Xie, Chen Henry Wu, Peng Shi 0010, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong 0005, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao 0002, Dragomir R. Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang 0037, Noah A. Smith, Luke Zettlemoyer, Tao Yu 0009 |
EMNLP | 7 |
| 2022 | QAFactEval: Improved QA-Based Factual Consistency Evaluation for SummarizationabstractAlexander Fabbri, Chien-Sheng Wu, Wenhao Liu, Caiming Xiong. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Alexander R. Fabbri, Chien-Sheng Wu, Wenhao Liu 0003, Caiming Xiong |
NAACL-HLT | 2 |
| 2021 | GraPPa: Grammar-Augmented Pre-Training for Table Semantic Parsing
Tao Yu 0009, Chien-Sheng Wu, Xi Victoria Lin, Bailin Wang, Yi Chern Tan, Xinyi Yang 0002, Dragomir R. Radev, Richard Socher, Caiming Xiong |
ICLR | 2 |
| 2020 | Explicit Memory Tracker with Coarse-to-Fine Reasoning for Conversational Machine ReadingabstractYifan Gao, Chien-Sheng Wu, Shafiq Joty, Caiming Xiong, Richard Socher, Irwin King, Michael Lyu, Steven C.H. Hoi. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Yifan Gao 0001, Chien-Sheng Wu, Shafiq R. Joty, Caiming Xiong, Richard Socher, Irwin King, Michael R. Lyu, Steven C. H. Hoi |
ACL | 2 |
| 2020 | Discern: Discourse-Aware Entailment Reasoning Network for Conversational Machine ReadingabstractYifan Gao, Chien-Sheng Wu, Jingjing Li, Shafiq Joty, Steven C.H. Hoi, Caiming Xiong, Irwin King, Michael Lyu. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Yifan Gao 0001, Chien-Sheng Wu, Jingjing Li 0007, Shafiq R. Joty, Steven C. H. Hoi, Caiming Xiong, Irwin King, Michael R. Lyu |
EMNLP (1) | 2 |
| 2020 | TOD-BERT: Pre-trained Natural Language Understanding for Task-Oriented DialogueabstractThe underlying difference of linguistic patterns between general text and task-oriented dialogue makes existing pre-trained language models less useful in practice.In this work, we unify nine human-human and multi-turn task-oriented dialogue datasets for language modeling.To better model dialogue behavior during pre-training, we incorporate user and system tokens into the masked language modeling.We propose a contrastive objective function to simulate the response selection task.Our pre-trained task-oriented dialogue BERT (TOD-BERT) outperforms strong baselines like BERT on four downstream taskoriented dialogue applications, including intention recognition, dialogue state tracking, dialogue act prediction, and response selection.We also show that TOD-BERT has a stronger few-shot ability that can mitigate the data scarcity problem for task-oriented dialogue. Chien-Sheng Wu, Steven C. H. Hoi, Richard Socher, Caiming Xiong |
EMNLP (1) | 1 |
| 2020 | Probing Task-Oriented Dialogue Representation from Language ModelsabstractThis paper investigates pre-trained language models to find out which model intrinsically carries the most informative representation for task-oriented dialogue tasks.We approach the problem from two aspects: supervised classifier probe and unsupervised mutual information probe.We fine-tune a feed-forward layer as the classifier probe on top of a fixed pretrained language model with annotated labels in a supervised way.Meanwhile, we propose an unsupervised mutual information probe to evaluate the mutual dependence between a real clustering and a representation clustering.The goals of this empirical paper are to 1) investigate probing techniques, especially from the unsupervised mutual information aspect, 2) provide guidelines of pre-trained language model selection for the dialogue research community, 3) find insights of pre-training factors for dialogue application that may be the key to success. Chien-Sheng Wu, Caiming Xiong |
EMNLP (1) | 1 |
| 2020 | Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language InferenceabstractJianguo Zhang, Kazuma Hashimoto, Wenhao Liu, Chien-Sheng Wu, Yao Wan, Philip Yu, Richard Socher, Caiming Xiong. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Jianguo Zhang 0005, Kazuma Hashimoto, Wenhao Liu 0003, Chien-Sheng Wu, Yao Wan 0001, Philip S. Yu, Richard Socher, Caiming Xiong |
EMNLP (1) | 4 |
| 2020 | Getting To Know You: User Attribute Extraction from DialoguesabstractUser attributes provide rich and useful information for user understanding, yet structured and easy-to-use attributes are often sparsely populated. In this paper, we leverage dialogues with conversational agents, which contain strong suggestions of user information, to automatically extract user attributes. Since no existing dataset is available for this purpose, we apply distant supervision to train our proposed two-stage attribute extractor, which surpasses several retrieval and generation baselines on human evaluation. Meanwhile, we discuss potential applications (e.g., personalized recommendation and dialogue systems) of such extracted user attributes, and point out current limitations to cast light on future work. Chien-Sheng Wu, Andrea Madotto, Zhaojiang Lin, Peng Xu 0008, Pascale Fung |
LREC | 1 |
| 2020 | A Simple Language Model for Task-Oriented DialogueabstractTask-oriented dialogue is often decomposed into three tasks: understanding user input, deciding actions, and generating a response. While such decomposition might suggest a dedicated model for each sub-task, we find a simple, unified approach leads to state-of-the-art performance on the MultiWOZ dataset. SimpleTOD is a simple approach to task-oriented dialogue that uses a single, causal language model trained on all sub-tasks recast as a single sequence prediction problem. This allows SimpleTOD to fully leverage transfer learning from pre-trained, open domain, causal language models such as GPT-2. SimpleTOD improves over the prior state-of-the-art in joint goal accuracy for dialogue state tracking, and our analysis reveals robustness to noisy annotations in this setting. SimpleTOD also improves the main metrics used to evaluate action decisions and response generation in an end-to-end setting: inform rate by 8.1 points, success rate by 9.7 points, and combined score by 7.2 points. Ehsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz, Richard Socher |
NeurIPS | 3 |
| 2019 | Personalizing Dialogue Agents via Meta-LearningabstractExisting personalized dialogue models use human designed persona descriptions to improve dialogue consistency.Collecting such descriptions from existing dialogues is expensive and requires hand-crafted feature designs.In this paper, we propose to extend Model-Agnostic Meta-Learning (MAML) (Finn et al., 2017) to personalized dialogue learning without using any persona descriptions.Our model learns to quickly adapt to new personas by leveraging only a few dialogue samples collected from the same user, which is fundamentally different from conditioning the response on the persona descriptions.Empirical results on Persona-chat dataset (Zhang et al., 2018) indicate that our solution outperforms non-metalearning baselines using automatic evaluation metrics, and in terms of human-evaluated fluency and consistency. Andrea Madotto, Zhaojiang Lin, Chien-Sheng Wu, Pascale Fung |
ACL (1) | 3 |
| 2019 | Transferable Multi-Domain State Generator for Task-Oriented Dialogue SystemsabstractOver-dependence on domain ontology and lack of knowledge sharing across domains are two practical and yet less studied problems of dialogue state tracking.Existing approaches generally fall short in tracking unknown slot values during inference and often have difficulties in adapting to new domains.In this paper, we propose a TRAnsferable Dialogue statE generator (TRADE) that generates dialogue states from utterances using a copy mechanism, facilitating knowledge transfer when predicting (domain, slot, value) triplets not encountered during training.Our model is composed of an utterance encoder, a slot gate, and a state generator, which are shared across domains.Empirical results demonstrate that TRADE achieves state-of-the-art joint goal accuracy of 48.62% for the five domains of Mul-tiWOZ, a human-human dialogue dataset.In addition, we show its transferring ability by simulating zero-shot and few-shot dialogue state tracking for unseen domains.TRADE achieves 60.58% joint goal accuracy in one of the zero-shot domains, and is able to adapt to few-shot cases without forgetting already trained domains.* Work partially done while the first author was an intern at Salesforce Research.Usr: I am looking for a cheap restaurant in the centre of the Chien-Sheng Wu, Andrea Madotto, Ehsan Hosseini-Asl, Caiming Xiong, Richard Socher, Pascale Fung |
ACL (1) | 1 |
| 2019 | Code-Switched Language Models Using Neural Based Synthetic Data from Parallel SentencesabstractTraining code-switched language models is difficult due to lack of data and complexity in the grammatical structure.Linguistic constraint theories have been used for decades to generate artificial code-switching sentences to cope with this issue.However, this require external word alignments or constituency parsers that create erroneous results on distant languages.We propose a sequence-to-sequence model using a copy mechanism to generate code-switching data by leveraging parallel monolingual translations from a limited source of code-switching data.The model learns how to combine words from parallel sentences and identifies when to switch one language to the other.Moreover, it captures code-switching constraints by attending and aligning the words in inputs, without requiring any external knowledge.Based on experimental results, the language model trained with the generated sentences achieves state-of-theart performance and improves end-to-end automatic speech recognition. Genta Indra Winata, Andrea Madotto, Chien-Sheng Wu, Pascale Fung |
CoNLL | 3 |
| 2019 | Clickbait? Sensational Headline Generation with Auto-tuned Reinforcement LearningabstractPeng Xu, Chien-Sheng Wu, Andrea Madotto, Pascale Fung. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Peng Xu 0008, Chien-Sheng Wu, Andrea Madotto, Pascale Fung |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Global-to-local Memory Pointer Networks for Task-Oriented Dialogue
Chien-Sheng Wu, Richard Socher, Caiming Xiong |
ICLR (Poster) | 1 |
| 2018 | Mem2Seq: Effectively Incorporating Knowledge Bases into End-to-End Task-Oriented Dialog SystemsabstractEnd-to-end task-oriented dialog systems usually suffer from the challenge of incorporating knowledge bases.In this paper, we propose a novel yet simple end-toend differentiable model called memoryto-sequence (Mem2Seq) to address this issue.Mem2Seq is the first neural generative model that combines the multihop attention over memories with the idea of pointer network.We empirically show how Mem2Seq controls each generation step, and how its multi-hop attention mechanism helps in learning correlations between memories.In addition, our model is quite general without complicated taskspecific designs.As a result, we show that Mem2Seq can be trained faster and attain the state-of-the-art performance on three different task-oriented dialog datasets. Andrea Madotto, Chien-Sheng Wu, Pascale Fung |
ACL (1) | 2 |
| 2018 | Improving Large-Scale Fact-Checking using Decomposable Attention Models and Lexical TaggingabstractFact-checking of textual sources needs to effectively extract relevant information from large knowledge bases.In this paper, we extend an existing pipeline approach to better tackle this problem.We propose a neural ranker using a decomposable attention model that dynamically selects sentences to achieve promising improvement in evidence retrieval F1 by 38.80%, with (×65) speedup compared to a TF-IDF method.Moreover, we incorporate lexical tagging methods into our pipeline framework to simplify the tasks and render the model more generalizable.As a result, our framework achieves promising performance on a large-scale fact extraction and verification dataset with speedup. Nayeon Lee, Chien-Sheng Wu, Pascale Fung |
EMNLP | 2 |
| 2018 | End-to-End Dynamic Query Memory Network for Entity-Value Independent Task-Oriented DialogabstractIn this paper, we propose an end-to-end Dynamic Query Memory Network (DQMemNN) with a delexicalization mechanism for task-oriented dialog systems. The added dynamic component enables memory networks to capture the dialog's sequential dependencies by using a context-based query. Besides, the delexicalization mechanism reduces learning complexity and it alleviates the out-of-vocabulary entity problems. Experiments show that DQMemNN outperforms original end-to-end memory network models on bAbI full-dialog task by 3.1 % per-response and 39.3% per-dialog accuracy. In addition, the proposed framework achieves a promising average per-response accuracy of 99.7% and per-dialog accuracy of 97.8% without hand-crafted rules and features. Chien-Sheng Wu, Andrea Madotto, Genta Indra Winata, Pascale Fung |
ICASSP | 1 |
| 2016 | Towards Empathetic Human-Robot Interactions
Pascale Fung, Dario Bertero, Yan Wan 0004, Anik Dey, Ricky Ho Yin Chan, Farhad Bin Siddique, Yang Yang 0130, Chien-Sheng Wu, Ruixi Lin |
CICLing (2) | 8 |
| 2016 | Real-Time Speech Emotion and Sentiment Recognition for Interactive Dialogue SystemsabstractIn this paper, we describe our approach of enabling an interactive dialogue system to recognize user emotion and sentiment in realtime.These modules allow otherwise conventional dialogue systems to have "empathy" and answer to the user while being aware of their emotion and intent.Emotion recognition from speech previously consists of feature engineering and machine learning where the first stage causes delay in decoding time.We describe a CNN model to extract emotion from raw speech input without feature engineering.This approach even achieves an impressive average of 65.7% accuracy on six emotion categories, a 4.5% improvement when compared to the conventional feature based SVM classification.A separate, CNN-based sentiment analysis module recognizes sentiments from speech recognition results, with 82.5 Fmeasure on human-machine dialogues when trained with out-of-domain data. Dario Bertero, Farhad Bin Siddique, Chien-Sheng Wu, Yan Wan 0004, Ricky Ho Yin Chan, Pascale Fung |
EMNLP | 3 |