VLDB 2026 Research / reviewers in the wild / expert
Lu Wang 0008
dblp:49/3800-8
· DBLP profile ↗
70ranked-venue papers
9as first author
40since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 68 · 8 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CASPER in the Machine: Insights into Character Variety in LLM-Generated StoriesabstractAs LLM-generated text is increasingly used, especially in fictional domains, we explore how much LLM-generated stories differ from human-written stories.In this work, we focus on characters.We borrow definitions from narratology to analyze 8 intricate category-pairs of character, such as stylization and wholeness.These category-pairs consider more than just basic characteristics.They assess how characters are portrayed within their stories.After automatically inferring categories of characters within both LLM and human-written stories, we compare and contrast these two sets of stories.We consider the following overarching questions: (1) Do LLMs and human-written stories have similar characters? and (2) Do LLMs generate stories with a variety of characters?Our analysis includes research questions that focus on stories generated by popular LLMs and recently published human-written stories.We describe a number of interesting similarities, differences and key takeaways.1 Anneliese Brei, Abhisheik Sharma, Nicholas Sanaie, Lu Wang 0008, Snigdha Chaturvedi |
ACL (1) | 4 |
| 2026 | Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM OverthinkingabstractXinliang Frederick Zhang, Anhad Mohananey, Alexandra Chronopoulou, Pinelopi Papalampidi, Somit Gupta, Tsendsuren Munkhdalai, Lu Wang, Shyam Upadhyay. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xinliang Frederick Zhang, Anhad Mohananey, Alexandra Chronopoulou, Pinelopi Papalampidi, Somit Gupta, Tsendsuren Munkhdalai, Lu Wang 0008, Shyam Upadhyay |
ACL (1) | 7 |
| 2025 | FactBench: A Dynamic Benchmark for In-the-Wild Language Model Factuality EvaluationabstractThe rapid adoption of language models (LMs) across diverse applications has raised concerns about their factuality, i.e., their consistency with real-world facts.We introduce VERIFY, an evidence-based evaluation pipeline that measures LMs' factuality in real-world user interactions.VERIFY considers the verifiability of LM-generated content and categorizes content units as Supported, Unsupported, or Undecidable based on Web-retrieved evidence.Importantly, factuality judgment by VERIFY more strongly correlates with human evaluations than existing methods.Using VER-IFY, we identify "hallucination prompts," i.e., those that frequently elicit factual errors in LM responses.These prompts form FACTBENCH, a dataset of 1K prompts spanning 150 topics and tiered into Easy, Moderate, and Hard prompts.We benchmark widely-used openweight and proprietary LMs from six families, yielding three key findings: (i) LMs' factual precision declines from Easy to Hard prompts, (ii) factuality does not necessarily improve with scale; Llama3.1-405B-Instructperforms comparably to or worse than its 70B variant, and (iii) Gemini1.5-Proshows a notably higher refusal rate, with over-refusal in 25% of cases. Farima Fatahi Bayat, Sheza Munir, Lu Wang 0008 |
ACL (1) | 4 |
| 2025 | Efficient Ensemble for Fine-tuning Language Models on Multiple DatasetsabstractThis paper develops an ensemble method for fine-tuning a language model to multiple datasets.Existing methods, such as quantized LoRA (QLoRA), are efficient when adapting to a single dataset.When training on multiple datasets of different tasks, a common setup in practice, it remains unclear how to design an efficient adaptation for fine-tuning language models.We propose to use an ensemble of multiple smaller adapters instead of a single adapter per task.We design an efficient algorithm that partitions n datasets into m groups, where m is typically much smaller than n in practice, and train one adapter for each group before taking a weighted combination to form the ensemble.The algorithm leverages a first-order approximation property of low-rank adaptation to quickly obtain the fine-tuning performances of dataset combinations since methods like LoRA stay close to the base model.Hence, we use the gradients of the base model to estimate its behavior during fine-tuning.Empirically, this approximation holds with less than 1% error on models with up to 34 billion parameters, leading to an estimation of true fine-tuning performances under 5% error while speeding up computation compared to base fine-tuning by 105 times.When applied to fine-tune Llama and GPT models on ten text classification tasks, our approach provides up to 10% higher average test accuracy over QLoRA, with only 9% more FLOPs.On a Llama model with 34 billion parameters, an ensemble of QLoRA increases test accuracy by 3% compared to QLoRA, with only 8% more FLOPs. Ziniu Zhang, Lu Wang 0008, Hongyang R. Zhang |
ACL (1) | 3 |
| 2025 | On Many-Shot In-Context Learning for Long-Context EvaluationabstractMany-shot in-context learning (ICL) has emerged as a unique setup to both utilize and test the ability of large language models to handle long context.This paper delves into long-context language model (LCLM) evaluation through many-shot ICL.We first ask: what types of ICL tasks benefit from additional demonstrations, and how effective are they in evaluating LCLMs?We find that classification and summarization tasks show performance improvements with additional demonstrations, while translation and reasoning tasks do not exhibit clear trends.Next, we investigate the extent to which different tasks necessitate retrieval versus global context understanding.We develop metrics to categorize ICL tasks into two groups: (i) similar-sample learning (SSL): tasks where retrieval of the most similar examples is sufficient for good performance, and (ii) all-sample learning (ASL): tasks that necessitate a deeper comprehension of all examples in the prompt.Lastly, we introduce a new many-shot ICL benchmark built on existing ICL tasks, MANYICLBENCH, to characterize model's ability on both fronts and benchmark 12 LCLMs using MANYICLBENCH.We find that while state-of-the-art models demonstrate good performance up to 64k tokens in SSL tasks, many models experience significant performance drops at only 16k tokens in ASL tasks. Kaijian Zou, Muhammad Khalifa, Lu Wang 0008 |
ACL (1) | 3 |
| 2025 | SYNC: A Synthetic Long-Context Understanding Benchmark for Controlled Comparisons of Model CapabilitiesabstractRecently, researchers have turned to synthetic tasks for evaluating long-context capabilities of large language models (LLMs) , as they offer more flexibility than realistic benchmarks in scaling both input length and dataset size.However, existing synthetic tasks typically target narrow skill sets such as retrieving information from massive input, limiting their ability to comprehensively assess model capabilities.Furthermore, existing benchmarks often pair each task with a different input context, creating confounding factors that prevent fair crosstask comparison.To address these limitations, we introduce SYNC, a new evaluation suite of synthetic tasks spanning domains including graph understanding and translation.Each domain includes three tasks designed to test a wide range of capabilities-from retrieval, to multi-hop tracking, and to global context understanding that that requires chain-of-thought (CoT) reasoning.Crucially, all tasks share the same context, enabling controlled comparisons of model performance.We evaluate 14 LLMs on SYNC and observe substantial performance drops on more challenging tasks, underscoring the benchmark's difficulty.Additional experiments highlight the necessity of CoT reasoning and demonstrate that SYNC poses a robust challenge for future models. Shuyang Cao, Kaijian Zou, Lu Wang 0008 |
EMNLP | 3 |
| 2025 | Structured Moral Reasoning in Language Models: A Value-Grounded Evaluation FrameworkabstractLarge language models (LLMs) are increasingly deployed in domains requiring moral understanding, yet their reasoning often remains shallow, and misaligned with human reasoning (Jiang et al., 2021).Unlike humans, whose moral reasoning integrates contextual trade-offs, value systems, and ethical theories, LLMs often rely on surface patterns, leading to biased decisions in morally and ethically complex scenarios.To address this gap, we present a value-grounded framework for evaluating and distilling structured moral reasoning in LLMs.We benchmark 12 open-source models across four moral datasets using a taxonomy of prompts grounded in value systems, ethical theories, and cognitive reasoning strategies.Our evaluation is guided by four questions: (1) Does reasoning improve LLM decision-making over direct prompting?(2) Which types of value/ethical frameworks most effectively guide LLM reasoning?(3) Which cognitive reasoning strategies lead to better moral performance?(4) Can small-sized LLMs acquire moral competence through distillation?We find that prompting with explicit moral structure consistently improves accuracy and coherence, with first-principles reasoning and Schwartz's + care-ethics scaffolds yielding the strongest gains.Furthermore, our supervised distillation approach transfers moral competence from large to small models without additional inference cost.Together, our results offer a scalable path toward interpretable and value-grounded models. Mohna Chakraborty, Lu Wang 0008, David Jurgens |
EMNLP | 2 |
| 2025 | Answer Convergence as a Signal for Early Stopping in ReasoningabstractChain-of-thought (CoT) prompting enhances reasoning in large language models (LLMs) but often leads to verbose and redundant outputs, thus increasing inference cost.We hypothesize that many reasoning steps are unnecessary for producing correct answers.To investigate this, we start with a systematic study to examine what is the minimum reasoning required for a model to reach a stable decision.We find that on reasoning tasks like math, models typically converge to their final answers after 60% of the reasoning steps, suggesting substantial redundancy in the remaining content.Based on these insights, we propose three inference-time strategies to improve efficiency: (1) early stopping via answer consistency, (2) boosting the probability of generating end-of-reasoning signals, and (3) a supervised method that learns when to stop based on internal activations.Experiments across five benchmarks and five open-weights LLMs show that our methods significantly reduce token usage with little or no accuracy drop.In particular, on NaturalQuestions, Answer Consistency reduces tokens by over 40% while further improving accuracy.Our work underscores the importance of cost-effective reasoning methods that operate at inference time, offering practical benefits for real-world applications.1 Lu Wang 0008 |
EMNLP | 2 |
| 2025 | VeriFact: Enhancing Long-Form Factuality Evaluation with Refined Fact Extraction and Reference FactsabstractLarge language models (LLMs) excel at generating long-form responses, but evaluating their factuality remains challenging due to complex inter-sentence dependencies within the generated facts.Prior solutions predominantly follow a decompose-decontextualize-verify pipeline but often fail to capture essential context and miss key relational facts.In this paper, we introduce VERIFACT, a factuality evaluation framework designed to enhance fact extraction by identifying and resolving incomplete and missing facts to support more accurate verification results.Moreover, we introduce FACTR-BENCH 1 , a benchmark that evaluates both precision and recall in long-form model responses, whereas prior work primarily focuses on precision.FACTRBENCH provides reference fact sets from advanced LLMs and human-written answers, enabling recall assessment.Empirical evaluations show that VERIFACT significantly enhances fact completeness and preserves complex facts with critical relational information, resulting in more accurate factuality evaluation.Benchmarking various open-and close-weight LLMs on FACTRBENCH indicate that larger models within same model family improve precision and recall, but high precision does not always correlate with high recall, underscoring the importance of comprehensive factuality assessment.1.There is a demand for gold as jewelry 2. The price of gold could drop by 20-50% or more.3. Copper and silver are competitive with gold in terms of price.Anthropic. Sheza Munir, Yiyang Gu, Lu Wang 0008 |
EMNLP | 5 |
| 2025 | Unstructured Evidence Attribution for Long Context Query Focused SummarizationabstractLarge language models (LLMs) are capable of generating coherent summaries from very long contexts given a user query, and extracting and citing evidence spans helps improve the trustworthiness of these summaries.Whereas previous work has focused on evidence citation with fixed levels of granularity (e.g.sentence, paragraph, document, etc.), we propose to extract unstructured (i.e., spans of any length) evidence in order to acquire more relevant and consistent evidence than in the fixed granularity case.We show how existing systems struggle to copy and properly cite unstructured evidence, which also tends to be "lost-in-the-middle".To help models perform this task, we create the Summaries with Unstructured Evidence Text dataset (SUnsET), a synthetic dataset generated using a novel pipeline, which can be used as training supervision for unstructured evidence summarization.We demonstrate across 5 LLMs and 4 datasets spanning human written, synthetic, single, and multi-document settings that LLMs adapted with SUnsET generate more relevant and factually consistent evidence with their summaries, extract evidence from more diverse locations in their context, and can generate more relevant and consistent summaries than baselines with no fine-tuning and fixed granularity evidence.We release SUnsET and our generation code to the public.1 Dustin Wright 0001, Zain Muhammad Mujahid, Lu Wang 0008, Isabelle Augenstein, David Jurgens |
EMNLP | 3 |
| 2025 | PRIME: Large Language Model Personalization with Cognitive Dual-Memory and Personalized Thought ProcessabstractLarge language model (LLM) personalization aims to align model outputs with individuals' unique preferences and opinions.While recent efforts have implemented various personalization methods, a unified theoretical framework that can systematically understand the drivers of effective personalization is still lacking.In this work, we integrate the well-established cognitive dual-memory model into LLM personalization, by mirroring episodic memory to historical user engagements and semantic memory to long-term, evolving user beliefs.Specifically, we systematically investigate memory instantiations and introduce a unified framework, PRIME, using episodic and semantic memory mechanisms.We further augment PRIME with a novel personalized thinking capability inspired by the slow thinking strategy.Moreover, recognizing the absence of suitable benchmarks, we introduce a dataset using Change My View (CMV) from Reddit 1 , specifically designed to evaluate long-context personalization.Extensive experiments validate PRIME's effectiveness across both longand short-context scenarios.Further analysis confirms that PRIME effectively captures dynamic personalization beyond mere popularity biases. Xinliang Frederick Zhang, Nick Beauchamp, Lu Wang 0008 |
EMNLP | 3 |
| 2025 | Linear-Time Demonstration Selection for In-Context Learning via Gradient EstimationabstractThis paper introduces an algorithm to select demonstration examples for in-context learning of a query set.Given a set of n examples, how can we quickly select k out of n to best serve as the conditioning for downstream inference?This problem has broad applications in prompt tuning and chain-of-thought reasoning.Since model weights remain fixed during in-context learning, previous work has sought to design methods based on the similarity of token embeddings.This work proposes a new approach based on gradients of the output taken in the input embedding space.Our approach estimates model outputs through a first-order approximation using the gradients.Then, we apply this estimation to multiple randomly sampled subsets.Finally, we aggregate the sampled subset outcomes to form an influence score for each demonstration, and select k most relevant examples.This procedure only requires precomputing model outputs and gradients once, resulting in a linear-time algorithm relative to model and training set sizes.Extensive experiments across various models and datasets validate the efficiency of our approach.We show that the gradient estimation procedure yields approximations of full inference with less than 1% error across six datasets.This allows us to scale up subset selection that would otherwise run full inference by up to 37.7× on models with up to 34 billion parameters, and outperform existing selection methods based on input embeddings by 11% on average. Ziniu Zhang, Zhenshuo Zhang, Lu Wang 0008, Jennifer G. Dy, Hongyang R. Zhang |
EMNLP | 4 |
| 2025 | MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?abstractWe introduce MLRC-Bench, a benchmark designed to quantify how effectively language agents can tackle challenging Machine Learning (ML) Research Competitions, with a focus on open research problems that demand novel methodologies.Unlike prior work, e.g., AI Scientist, which evaluates the end-to-end agentic pipeline by using LLM-as-a-judge, MLRC-Bench measures the key steps of proposing and implementing novel research methods and evaluates them with rigorous protocol and objective metrics.Our curated suite of 7 competition tasks reveals significant challenges for LLM agents. Even the best-performing tested agent (gemini-exp-1206 under MLAB) closes only 9.3% of the gap between baseline and top human participant scores.Furthermore, our analysis reveals a misalignment between the LLM-judged innovation and their actual performance on cutting-edge ML research problems. MLRC-Bench is a dynamic benchmark, which is designed to continually grow with new ML competitions to encourage rigorous and objective evaluations of AI’s research capabilities. Our leaderboard and code are publicly available at https://huggingface.co/spaces/launch/MLRC_Bench. Muhammad Khalifa, Shitanshu Bhushan, Grant D. Murphy, Lajanugen Logeswaran, Jaekyeom Kim, Moontae Lee, Honglak Lee, Lu Wang 0008 |
NeurIPS | 9 |
| 2024 | MIDGARD: Self-Consistency Using Minimum Description Length for Structured Commonsense ReasoningabstractWe study the task of conducting structured reasoning as generating a reasoning graph from natural language input using large language models (LLMs).Previous approaches have explored various prompting schemes, yet they suffer from error propagation due to the autoregressive nature and single-pass-based decoding, which lack error correction capability.Additionally, relying solely on a single sample may result in the omission of true nodes and edges.To counter this, we draw inspiration from self-consistency (SC), which involves sampling a diverse set of reasoning chains and taking the majority vote as the final answer.To tackle the substantial challenge of applying SC on generated graphs, we propose MIDGARD (MInimum Description length Guided Aggregation of Reasoning in Directed acyclic graph) that leverages Minimum Description Length (MDL)-based formulation to identify consistent properties among the different graph samples generated by an LLM.This formulation helps reject properties that appear in only a few samples, which are likely to be erroneous, while enabling the inclusion of missing elements without compromising precision.Our method demonstrates superior performance than comparisons across various structured reasoning tasks, including argument structure extraction, explanation graph generation, inferring dependency relations among actions for everyday tasks, and semantic graph generation from natural texts. Inderjeet Nair, Lu Wang 0008 |
ACL (1) | 2 |
| 2024 | Analyzing Occupational Distribution Representation in Japanese Language ModelsabstractRecent advances in large language models (LLMs) have enabled users to generate fluent and seemingly convincing text. However, these models have uneven performance in different languages, which is also associated with undesirable societal biases toward marginalized populations. Specifically, there is relatively little work on Japanese models, despite it being the thirteenth most widely spoken language. In this work, we first develop three Japanese language prompts to probe LLMs’ understanding of Japanese names and their association between gender and occupations. We then evaluate a variety of English, multilingual, and Japanese models, correlating the models’ outputs with occupation statistics from the Japanese Census Bureau from the last 100 years. Our findings indicate that models can associate Japanese names with the correct gendered occupations when using constrained decoding. However, with sampling or greedy decoding, Japanese language models have a preference for a small set of stereotypically gendered occupations, and multilingual models, though trained on Japanese, are not always able to understand Japanese prompts. Katsumi Ibaraki, Winston Wu, Lu Wang 0008, Rada Mihalcea |
LREC/COLING | 3 |
| 2024 | Enhancing Language Model Factuality via Activation-Based Confidence Calibration and Guided DecodingabstractCalibrating language models (LMs) aligns their generation confidence with the actual likelihood of answer correctness, which can inform users about LMs' reliability and mitigate hallucinated content.However, prior calibration methods, such as self-consistency-based and logit-based approaches, are either limited in inference-time efficiency or fall short of providing informative signals.Moreover, simply filtering out low-confidence responses reduces the LM's helpfulness when the answers are correct.Therefore, effectively using calibration techniques to enhance an LM's factuality remains an unsolved challenge.In this paper, we first propose an activation-based calibration method, ACTCAB, which trains a linear layer on top of the LM's last-layer activations that can better capture the representations of knowledge.Built on top of ACTCAB, we further propose CODEC, a confidence-guided decoding strategy to elicit truthful answers with high confidence from LMs.By evaluating on five popular QA benchmarks, ACTCAB achieves superior calibration performance than all competitive baselines, e.g., by reducing the average expected calibration error (ECE) score by up to 39%.Further experiments on CODEC show consistent improvements in several LMs' factuality on challenging QA datasets, such as TruthfulQA, highlighting the value of confidence signals in enhancing the factuality. 1 Farima Fatahi Bayat, Lu Wang 0008 |
EMNLP | 3 |
| 2024 | Closing the Loop: Learning to Generate Writing Feedback via Language Model Simulated Student RevisionsabstractProviding feedback is widely recognized as crucial for refining students' writing skills.Recent advances in language models (LMs) have made it possible to automatically generate feedback that is actionable and well-aligned with humanspecified attributes.However, it remains unclear whether the feedback generated by these models is truly effective in enhancing the quality of student revisions.Moreover, prompting LMs with a precise set of instructions to generate feedback is nontrivial due to the lack of consensus regarding the specific attributes that can lead to improved revising performance.To address these challenges, we propose PROF that PROduces Feedback via learning from LM simulated student revisions.PROF aims to iteratively optimize the feedback generator by directly maximizing the effectiveness of students' overall revising performance as simulated by LMs.Focusing on an economic essay assignment, we empirically test the efficacy of PROF and observe that our approach not only surpasses a variety of baseline methods in effectiveness of improving students' writing but also demonstrates enhanced pedagogical values, even though it was not explicitly trained for this aspect. Inderjeet Nair, Jiaye Tan, Xiaotian Su 0001, Anne Gere, Xu Wang 0016, Lu Wang 0008 |
EMNLP | 6 |
| 2024 | LitCab: Lightweight Language Model Calibration over Short- and Long-form ResponsesabstractA model is considered well-calibrated when its probability estimate aligns with the actual likelihood of the output being correct. Calibrating language models (LMs) is crucial, as it plays a vital role in detecting and mitigating hallucinations of LMs as well as building more trustworthy models. However, standard calibration techniques may not be suited for LM calibration. For instance, post-processing methods such as temperature scaling do not reorder the candidate generations. On the other hand, training-based methods require fine-tuning the entire model, which is impractical for LMs of large scale. We present LitCab, a lightweight calibration mechanism consisting of a single linear layer that takes the input text representation and predicts a bias term, which is then added to the LM output logits. LitCab improves model calibration by only adding < 2% of the original model parameters. For evaluation, we construct CaT, a benchmark consisting of eight text generation tasks, covering responses ranging from short phrases to paragraphs. We test LitCab with Llama2-7B, where it improves calibration across all tasks, reducing the average ECE score by as large as 30%. We further conduct a comprehensive evaluation with multiple popular open-sourced LMs from GPT and LLaMA families, yielding the following key findings: (i) Larger models within the same family exhibit better calibration on tasks with short generation tasks, but not necessarily for longer ones. (ii) GPT-family models show superior calibration compared to LLaMA, Llama2, and Vicuna models, despite having much fewer parameters. (iii) Fine-tuning pretrained model (e.g., LLaMA) with samples of limited purpose (e.g., conversations) may lead to worse calibration, highlighting the importance of fine-tuning setups for calibrating LMs. Muhammad Khalifa, Lu Wang 0008 |
ICLR | 3 |
| 2024 | AWESOME: GPU Memory-constrained Long Document Summarization using Memory Mechanism and Global Salient ContentabstractLong document summarization systems are critical for domains with lengthy and jargonladen text, yet they present significant challenges to researchers and developers with limited computing resources.Existing solutions mainly focus on efficient attentions or divideand-conquer strategies.The former reduces theoretical time complexity, but is still memoryheavy.The latter methods sacrifice global context, leading to uninformative and incoherent summaries.This work aims to leverage the memory-efficient nature of divide-and-conquer methods while preserving global context.Concretely, our framework AWESOME uses two novel mechanisms: (1) External memory mechanisms track previously encoded document segments and their corresponding summaries, to enhance global document understanding and summary coherence.(2) Global salient content is further identified beforehand to augment each document segment to support its summarization.Extensive experiments on diverse genres of text, including government reports, meeting transcripts, screenplays, scientific papers, and novels, show that AWESOME produces summaries with improved informativeness, faithfulness, and coherence than competitive baselines on longer documents, while having a smaller GPU memory footprint.Encoder Shuyang Cao, Lu Wang 0008 |
NAACL-HLT | 2 |
| 2024 | PELMS: Pre-training for Effective Low-Shot Multi-Document SummarizationabstractJoseph Peper, Wenzhao Qiu, Lu Wang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Joseph Peper, Wenzhao Qiu, Lu Wang 0008 |
NAACL-HLT | 3 |
| 2024 | MOKA: Moral Knowledge Augmentation for Moral Event ExtractionabstractXinliang Frederick Zhang, Winston Wu, Nick Beauchamp, Lu Wang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xinliang Frederick Zhang, Winston Wu, Nick Beauchamp, Lu Wang 0008 |
NAACL-HLT | 4 |
| 2023 | Few-shot Reranking for Multi-hop QA via Language Model PromptingabstractWe study few-shot reranking for multi-hop QA (MQA) with open-domain questions.To alleviate the need for a large number of labeled question-document pairs for retriever training, we propose PROMPTRANK, which relies on language model prompting for multi-hop path reranking.PROMPTRANK first constructs an instruction-based prompt that includes a candidate document path and then computes the relevance score between a given question and the path based on the conditional likelihood of the question given the path prompt according to a language model.PROMPTRANK yields strong retrieval performance on HotpotQA with only 128 training examples compared to state-of-theart methods trained on thousands of examples -73.6 recall@10 by PROMPTRANK vs. 77.8 by PathRetriever (Asai et al., 2020) and 77.5 by multi-hop dense retrieval (Xiong et al., 2021). Muhammad Khalifa, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, Lu Wang 0008 |
ACL (1) | 5 |
| 2023 | ReadingQuizMaker: A Human-NLP Collaborative System that Supports Instructors to Design High-Quality Reading Quiz QuestionsabstractDespite that reading assignments are prevalent, methods to encourage students to actively read are limited. We propose a system ReadingQuizMaker that supports instructors to conveniently design high-quality questions to help students comprehend readings. ReadingQuizMaker adapts to instructors’ natural workflows of creating questions, while providing NLP-based process-oriented support. ReadingQuizMaker enables instructors to decide when and which NLP models to use, select the input to the models, and edit the outcomes. In an evaluation study, instructors found the resulting questions to be comparable to their previously designed quizzes. Instructors praised ReadingQuizMaker for its ease of use, and considered the NLP suggestions to be satisfying and helpful. We compared ReadingQuizMaker with a control condition where instructors were given automatically generated questions to edit. Instructors showed a strong preference for the human-AI teaming approach provided by ReadingQuizMaker. Our findings suggest the importance of giving users control and showing an immediate preview of AI outcomes when providing AI support. Xinyi Lu 0004, Simin Fan, Jessica Houghton, Lu Wang 0008, Xu Wang 0016 |
CHI | 4 |
| 2023 | All Things Considered: Detecting Partisan Events from News Media with Cross-Article ComparisonabstractPublic opinion is shaped by the information news media provide, and that information in turn may be shaped by the ideological preferences of media outlets.But while much attention has been devoted to media bias via overt ideological language or topic selection, a more unobtrusive way in which the media shape opinion is via the strategic inclusion or omission of partisan events that may support one side or the other.We develop a latent variable-based framework to predict the ideology of news articles by comparing multiple articles on the same story and identifying partisan events whose inclusion or omission reveals ideology.Our experiments first validate the existence of partisan event selection, and then show that article alignment and cross-document comparison detect partisan events and article ideology better than competitive baselines.Our results reveal the high-level form of media bias, which is present even among mainstream media with strong norms of objectivity and nonpartisanship. Yujian Liu, Xinliang Frederick Zhang, Kaijian Zou, Ruihong Huang, Nick Beauchamp, Lu Wang 0008 |
EMNLP | 6 |
| 2023 | Cross-Cultural Analysis of Human Values, Morals, and Biases in Folk TalesabstractFolk tales are strong cultural and social influences in children's lives, and they are known to teach morals and values.However, existing studies on folk tales are largely limited to European tales.In our study, we compile a large corpus of over 1,900 tales originating from 27 diverse cultures across six continents.Using a range of lexicons and correlation analyses, we examine how human values, morals, and gender biases are expressed in folk tales across cultures.We discover differences between cultures in prevalent values and morals, as well as cross-cultural trends in problematic gender biases.Furthermore, we find trends of reduced value expression when examining public-domain fiction stories, extrinsically validate our analyses against the multicultural Schwartz Survey of Cultural Values, and find traditional gender biases associated with values, morals, and agency.This largescale cross-cultural study of folk tales paves the way for future studies on how literature influences and reflects cultural norms. Winston Wu, Lu Wang 0008, Rada Mihalcea |
EMNLP | 2 |
| 2023 | Merging Generated and Retrieved Knowledge for Open-Domain QAabstractOpen-domain question answering (QA) systems are often built with retrieval modules.However, retrieving passages from a given source is known to suffer from insufficient knowledge coverage.Alternatively, prompting large language models (LLMs) to generate contextual passages based on their parametric knowledge has been shown to improve QA performance.Yet, LLMs tend to "hallucinate" content that conflicts with the retrieved knowledge.Based on the intuition that answers supported by both sources are more likely to be correct, we propose COMBO, a Compatibility-Oriented knowledge Merging for Better Open-domain QA framework, to effectively leverage the two sources of information.Concretely, we match LLM-generated passages with retrieved counterparts into compatible pairs, based on discriminators trained with silver compatibility labels.Then a Fusionin-Decoder-based (Izacard and Grave, 2021b) reader model handles passage pairs to arrive at the final answer.Experiments show that COMBO outperforms competitive baselines on three out of four tested open-domain QA benchmarks.Further analysis reveals that our proposed framework demonstrates greater efficacy in scenarios with a higher degree of knowledge conflicts. Muhammad Khalifa, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, Lu Wang 0008 |
EMNLP | 6 |
| 2023 | General then Personal: Decoupling and Pre-training for Personalized Headline GenerationabstractAbstract Personalized Headline Generation aims to generate unique headlines tailored to users’ browsing history. In this task, understanding user preferences from click history and incorporating them into headline generation pose challenges. Existing approaches typically rely on predefined styles as control codes, but personal style lacks explicit definition or enumeration, making it difficult to leverage traditional techniques. To tackle these challenges, we propose General Then Personal (GTP), a novel framework comprising user modeling, headline generation, and customization. We train the framework using tailored designs that emphasize two central ideas: (a) task decoupling and (b) model pre-training. With the decoupling mechanism separating the task into generation and customization, two mechanisms, i.e., information self-boosting and mask user modeling, are further introduced to facilitate the training and text control. Additionally, we introduce a new evaluation metric to address existing limitations. Extensive experiments conducted on the PENS dataset, considering both zero-shot and few-shot scenarios, demonstrate that GTP outperforms state-of-the-art methods. Furthermore, ablation studies and analysis emphasize the significance of decoupling and pre-training. Finally, the human evaluation validates the effectiveness of our approaches.1 Yun-Zhu Song, Yi-Syuan Chen, Lu Wang 0008, Hong-Han Shuai |
Trans. Assoc. Comput. Linguistics | 3 |
| 2022 | HIBRIDS: Attention with Hierarchical Biases for Structure-aware Long Document SummarizationabstractDocument structure is critical for efficient information consumption.However, it is challenging to encode it efficiently into the modern Transformer architecture.In this work, we present HIBRIDS, which injects Hierarchical Biases foR Incorporating Document Structure into the calculation of attention scores.We further present a new task, hierarchical questionsummary generation, for summarizing salient content in the source document into a hierarchy of questions and summaries, where each follow-up question inquires about the content of its parent question-summary pair.We also annotate a new dataset with 6, 153 questionsummary hierarchies labeled on long government reports.Experiment results show that our model produces better question-summary hierarchies than comparisons on both hierarchy quality and content coverage, a finding also echoed by human judges.Additionally, our model improves the generation of longform summaries from lengthy government reports and Wikipedia articles, as measured by ROUGE scores. Shuyang Cao, Lu Wang 0008 |
ACL (1) | 2 |
| 2022 | Sentence-level Media Bias Analysis Informed by Discourse StructuresabstractAs polarization continues to rise among both the public and the news media, increasing attention has been devoted to detecting media bias.Most recent work in the NLP community, however, identify bias at the level of individual articles.However, each article itself comprises multiple sentences, which vary in their ideological bias.In this paper, we aim to identify sentences within an article that can illuminate and explain the overall bias of the entire article.We show that understanding the discourse role of a sentence in telling a news story, as well as its relation with nearby sentences, can reveal the ideological leanings of an author even when the sentence itself appears merely neutral.In particular, we consider using a functional news discourse structure and PDTB discourse relations to inform bias sentence identification, and distill the auxiliary knowledge from the two types of discourse structure into our bias sentence identification system.Experimental results on benchmark datasets show that incorporating both the global functional discourse structure and local rhetorical discourse relations can effectively increase the recall of bias sentence identification by 8.27% -8.62%, as well as increase the precision by 2.82% -3.48% 1 . Yuanyuan Lei 0001, Ruihong Huang, Lu Wang 0008, Nick Beauchamp |
EMNLP | 3 |
| 2022 | Late Fusion with Triplet Margin Objective for Multimodal Ideology Prediction and AnalysisabstractPrior work on ideology prediction has largely focused on single modalities, i.e., text or images.In this work, we introduce the task of multimodal ideology prediction, where a model predicts binary or five-point scale ideological leanings, given a text-image pair with political content.We first collect five new large-scale datasets with English documents and images along with their ideological leanings, covering news articles from a wide range of mainstream media in US and social media posts from Reddit and Twitter.We conduct in-depth analyses on news articles and reveal differences in image content and usage across the political spectrum.Furthermore, we perform extensive experiments and ablation studies, demonstrating the effectiveness of targeted pretraining objectives on different model components.Our bestperforming model, a late-fusion architecture pretrained with a triplet objective over multimodal content, outperforms the state-of-the-art text-only model by almost 4% and a strong multimodal baseline with no pretraining by over 3%. Changyuan Qiu, Winston Wu, Xinliang Frederick Zhang, Lu Wang 0008 |
EMNLP | 4 |
| 2022 | Generative Entity-to-Entity Stance Detection with Knowledge Graph AugmentationabstractStance detection is typically framed as predicting the sentiment in a given text towards a target entity.However, this setup overlooks the importance of the source entity, i.e., who is expressing the opinion.In this paper, we emphasize the need for studying interactions among entities when inferring stances.We first introduce a new task, entity-to-entity (E2E) stance detection, which primes models to identify entities in their canonical names and discern stances jointly.To support this study, we curate a new dataset with 10,619 annotations labeled at the sentence-level from news articles of different ideological leanings.We present a novel generative framework to allow the generation of canonical names for entities as well as stances among them.We further enhance the model with a graph encoder to summarize entity activities and external knowledge surrounding the entities.Experiments show that our model outperforms strong comparisons by large margins.Further analyses demonstrate the usefulness of E2E stance detection for understanding media quotation and stance landscape, as well as inferring entity ideology. Xinliang Frederick Zhang, Nick Beauchamp, Lu Wang 0008 |
EMNLP | 3 |
| 2022 | Towards Process-Oriented, Modular, and Versatile Question Generation that Meets Educational NeedsabstractNLP-powered automatic question generation (QG) techniques carry great pedagogical potential of saving educators' time and benefiting student learning.Yet, QG systems have not been widely adopted in classrooms to date.In this work, we aim to pinpoint key impediments and investigate how to improve the usability of automatic QG techniques for educational purposes by understanding how instructors construct questions and identifying touch points to enhance the underlying NLP models.We perform an in-depth need finding study with 11 instructors across 7 different universities, and summarize their thought processes and needs when creating questions.While instructors show great interests in using NLP systems to support question design, none of them has used such tools in practice.They resort to multiple sources of information, ranging from domain knowledge to students' misconceptions, all of which missing from today's QG systems.We argue that building effective human-NLP collaborative QG systems that emphasize instructor control and explainability is imperative for real-world adoption.We call for QG systems to provide process-oriented support, use modular design, and handle diverse sources of input. Xu Wang 0016, Simin Fan, Jessica Houghton, Lu Wang 0008 |
NAACL-HLT | 4 |
| 2021 | Controllable Open-ended Question Generation with A New Question Type OntologyabstractShuyang Cao, Lu Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Shuyang Cao, Lu Wang 0008 |
ACL/IJCNLP (1) | 2 |
| 2021 | DYPLOC: Dynamic Planning of Content Using Mixed Language Models for Text GenerationabstractXinyu Hua, Ashwin Sreevatsa, Lu Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xinyu Hua, Ashwin Sreevatsa, Lu Wang 0008 |
ACL/IJCNLP (1) | 3 |
| 2021 | Learning To Segment Actions From Visual and Language Instructions via Differentiable Weak Sequence AlignmentabstractWe address the problem of unsupervised localization of task-relevant actions (key-steps) and feature learning in instructional videos using both visual and language instructions. Our key observation is that the sequences of visual and linguistic key-steps are weakly aligned: there is an ordered one-to-one correspondence between most visual and language key-steps, while some key-steps in one modality are absent in the other. To recover the two sequences, we develop an ordered prototype learning module, which extracts visual and linguistic prototypes representing key-steps. To find weak alignment and perform feature learning, we develop a differentiable weak sequence alignment (DWSA) method that finds ordered one-to-one matching between sequences while allowing some items in a sequence to stay unmatched. We develop an efficient forward and backward algorithm for computing the alignment and the loss derivative with respect to parameters of visual and language feature learning modules. By experiments on two instructional video datasets, we show that our method significantly improves the state of the art. Yuhan Shen, Lu Wang 0008, Ehsan Elhamifar |
CVPR | 2 |
| 2021 | CLIFF: Contrastive Learning for Improving Faithfulness and Factuality in Abstractive SummarizationabstractWe study generating abstractive summaries that are faithful and factually consistent with the given articles.A novel contrastive learning formulation is presented, which leverages both reference summaries, as positive training data, and automatically generated erroneous summaries, as negative training data, to train summarization systems that are better at distinguishing between them.We further design four types of strategies for creating negative samples, to resemble errors made commonly by two state-of-the-art models, BART and PEGASUS, found in our new human annotations of summary errors.Experiments on XSum and CNN/Daily Mail show that our contrastive learning framework is robust across datasets and models.It consistently produces more factual summaries than strong comparisons with post error correction, entailmentbased reranking, and unlikelihood training, according to QA-based factuality evaluation.Human judges echo the observation and find that our model summaries correct more errors.REFERENCE: A "rare" short-eared owl found emaciated in Flintshire is now recuperating well, the RSPCA have said.SWAPENT: Flintshire → Bettisfield ⇒ A "rare" short-eared owl found emaciated in Bettisfield is now recuperating well, the RSPCA have said.MASKENT: A "rare" short-eared owl found emaciated in [MASK] is now recuperating well, the RSPCA have said.⇒ A "rare" short-eared owl found emaciated in a field in South Yorkshire is now recuperating well, the RSPCA have said.MASKREL: A "rare" short-eared owl found [MASK] in [MASK] is now recuperating well, the RSPCA have said.⇒ A "rare" short-eared owl found dead in London is now recuperating well, the RSPCA have said.REGENENT: A "rare" short-eared owl found emaciated in ⇒ A "rare" short-eared owl found emaciated in Nottinghamshire is now at a wildlife centre to recover.REGENREL: A "rare" short-eared owl found ⇒ A "rare" short-eared owl found in the grounds of a former coal mine is being cared for by the RSPCA in Somerset.SYSLOWCON: An injured golden owl found in a former coal mine in Lancashire is being cared for by the RSPCA. Shuyang Cao, Lu Wang 0008 |
EMNLP (1) | 2 |
| 2021 | Attention Head Masking for Inference Time Content Selection in Abstractive SummarizationabstractHow can we effectively inform content selection in Transformer-based abstractive summarization models?In this work, we present a simple-yet-effective attention head masking technique, which is applied on encoderdecoder attentions to pinpoint salient content at inference time.Using attention head masking, we are able to reveal the relation between encoder-decoder attentions and content selection behaviors of summarization models.We then demonstrate its effectiveness on three document summarization datasets based on both in-domain and cross-domain settings.Importantly, our models outperform prior state-ofthe-art models on CNN/Daily Mail and New York Times datasets.Moreover, our inferencetime masking technique is also data-efficient, requiring less than 20% of the training samples to outperform BART fine-tuned on the full CNN/DailyMail dataset. Shuyang Cao, Lu Wang 0008 |
NAACL-HLT | 2 |
| 2021 | Inference Time Style Control for SummarizationabstractHow to generate summaries of different styles without requiring corpora in the target styles, or training separate models?We present two novel methods that can be deployed during summary decoding on any pre-trained Transformer-based summarization model.(1) Decoder state adjustment instantly modifies decoder final states with externally trained style scorers, to iteratively refine the output against a target style.(2) Word unit prediction constrains the word usage to impose strong lexical control during generation.In experiments of summarizing with simplicity control, automatic evaluation and human judges both find our models producing outputs in simpler languages while still informative.We also generate news headlines with various ideological leanings, which can be distinguished by humans with a reasonable probability. Shuyang Cao, Lu Wang 0008 |
NAACL-HLT | 2 |
| 2021 | Efficient Attentions for Long Document SummarizationabstractLuyang Huang, Shuyang Cao, Nikolaus Parulian, Heng Ji, Lu Wang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Luyang Huang, Shuyang Cao, Nikolaus Nova Parulian, Heng Ji 0001, Lu Wang 0008 |
NAACL-HLT | 5 |
| 2021 | Controllable Summarization with Constrained Markov Decision ProcessabstractAbstract We study controllable text summarization, which allows users to gain control on a particular attribute (e.g., length limit) of the generated summaries. In this work, we propose a novel training framework based on Constrained Markov Decision Process (CMDP), which conveniently includes a reward function along with a set of constraints, to facilitate better summarization control. The reward function encourages the generation to resemble the human-written reference, while the constraints are used to explicitly prevent the generated summaries from violating user-imposed requirements. Our framework can be applied to control important attributes of summarization, including length, covered entities, and abstractiveness, as we devise specific constraints for each of these aspects. Extensive experiments on popular benchmarks show that our CMDP framework helps generate informative summaries while complying with a given attribute’s requirement.1 Hou Pong Chan, Lu Wang 0008, Irwin King |
Trans. Assoc. Comput. Linguistics | 2 |
| 2020 | Copy or Rewrite: Hybrid Summarization with Hierarchical Reinforcement LearningabstractJointly using the extractive and abstractive summarization methods can combine their complementary advantages, generating both informative and concise summary. Existing methods that adopt an extract-then-abstract strategy have achieved impressive results, yet they suffer from the information loss in the abstraction step because they compress all the selected sentences without distinguish. Especially when the whole sentence is summary-worthy, salient content would be lost by compression. To address this problem, we propose HySum, a hybrid framework for summarization that can flexibly switch between copying sentence and rewriting sentence according to the degree of redundancy. In this way, our approach can effectively combine the advantages of two branches of summarization, juggling informativity and conciseness. Moreover, we based on Hierarchical Reinforcement Learning, propose an end-to-end reinforcing method to bridge together the extraction module and rewriting module, which can enhance the cooperation between them. Automatic evaluation shows that our approach significantly outperforms the state-of-the-arts on the CNN/DailyMail corpus. Human evaluation also demonstrates that our generated summaries are more informative and concise than popular models. Liqiang Xiao, Lu Wang 0008, Hao He 0007, Yaohui Jin |
AAAI | 2 |
| 2020 | Discourse as a Function of Event: Profiling Discourse Structure in News Articles around the Main EventabstractUnderstanding discourse structures of news articles is vital to effectively contextualize the occurrence of a news event.To enable computational modeling of news structures, we apply an existing theory of functional discourse structure for news articles that revolves around the main event and create a human-annotated corpus of 802 documents spanning over four domains and three media sources.Next, we propose several documentlevel neural-network models to automatically construct news content structures.Finally, we demonstrate that incorporating system predicted news structures yields new state-of-theart performance for event coreference resolution.The news documents we annotated are openly available and the annotations are publicly released for future research 1 . Prafulla Kumar Choubey, Aaron Lee, Ruihong Huang, Lu Wang 0008 |
ACL | 4 |
| 2020 | Knowledge Graph-Augmented Abstractive Summarization with Semantic-Driven Cloze RewardabstractSequence-to-sequence models for abstractive summarization have been studied extensively, yet the generated summaries commonly suffer from fabricated content, and are often found to be near-extractive.We argue that, to address these issues, the summarizer should acquire semantic interpretation over input, e.g., via structured representation, to allow the generation of more informative summaries.In this paper, we present ASGARD, a novel framework for Abstractive Summarization with Graph-Augmentation and semantic-driven RewarD.We propose the use of dual encoders-a sequential document encoder and a graphstructured encoder-to maintain the global context and local characteristics of entities, complementing each other.We further design a reward based on a multiple choice cloze test to drive the model to better capture entity interactions.Results show that our models produce significantly higher ROUGE scores than a variant without knowledge graph as input on both New York Times and CNN/Daily Mail datasets.We also obtain better or comparable performance compared to systems that are finetuned from large pretrained language models.Human judges further rate our model outputs as more informative and containing fewer unfaithful errors.Input Article of New York Times: John M. Fabrizi, the mayor of Bridgeport, admitted on Tuesday that he had used cocaine and abused alcohol while in office.Mr. Fabrizi, who was appointed mayor in 2003 after the former mayor, Joseph P. Ganim, went to prison on corruption charges, said he had sought help for his drug problem about 18 months ago and that he had not used drugs since.About four months ago, he added, he stopped drinking alcohol. Luyang Huang, Lingfei Wu 0001, Lu Wang 0008 |
ACL | 3 |
| 2020 | Dynamic Online Conversation RecommendationabstractTrending topics in social media content evolve over time, and it is therefore crucial to understand social media users and their interpersonal communications in a dynamic manner.In this research we study dynamic online conversation recommendation, to help users engage in conversations that satisfy their evolving interests.Different from works in conversation recommendation which assume static user interests, our model captures the temporal aspects of user interests.Moreover, our model can cater for cold start problem where conversations are new and unseen in training.We propose a neural architecture to analyze changes of user interactions and interests over time, whose result is used to predict which discussions the users are likely to enter.We conduct experiments on large-scale collections of Reddit conversations.Results on three subreddits show that our model significantly outperforms state-of-the-art models based on static assumption of user interests.We further evaluate performance in cold start, and observe consistently better performance by our model when considering various degrees of sparsity of user's chatting history and conversation contexts.Lastly, our analysis also confirms the change of user interests.This further justify the advantage and efficacy of our model. Xingshan Zeng, Jing Li 0049, Lu Wang 0008, Zhiming Mao, Kam-Fai Wong |
ACL | 3 |
| 2020 | PAIR: Planning and Iterative Refinement in Pre-trained Transformers for Long Text GenerationabstractPre-trained Transformers have enabled impressive breakthroughs in generating long and fluent text, yet their outputs are often "rambling" without coherently arranged content.In this work, we present a novel content-controlled text generation framework, PAIR, with planning and iterative refinement, which is built upon a large model, BART.We first adapt the BERT model to automatically construct the content plans, consisting of keyphrase assignments and their corresponding sentence-level positions.The BART model is employed for generation without modifying its structure.We then propose a refinement algorithm to gradually enhance the generation quality within the sequence-tosequence framework.Evaluation with automatic metrics shows that adding planning consistently improves the generation quality on three distinct domains, with an average of 20 BLEU points and 12 METEOR points improvements.In addition, human judges rate our system outputs to be more relevant and coherent than comparisons without planning. Xinyu Hua, Lu Wang 0008 |
EMNLP (1) | 2 |
| 2020 | Modeling Content Importance for Summarization with Pre-trained Language ModelsabstractModeling content importance is an essential yet challenging task for summarization.Previous work is mostly based on statistical methods that estimate word-level salience, which does not consider semantics and larger context when quantifying importance.It is thus hard for these methods to generalize to semantic units of longer text spans.In this work, we apply information theory on top of pretrained language models and define the concept of importance from the perspective of information amount.It considers both the semantics and context when evaluating the importance of each semantic unit.With the help of pre-trained language models, it can easily generalize to different kinds of semantic units (n-grams or sentences).Experiments on CNN/Daily Mail and New York Times datasets demonstrate that our method can better model the importance of content than prior work based on F1 and ROUGE scores. Liqiang Xiao, Lu Wang 0008, Hao He 0007, Yaohui Jin |
EMNLP (1) | 2 |
| 2020 | Temporal Logic Point ProcessesabstractWe propose a modeling framework for event data and aim to answer questions such as \emph{when} and \emph{why} the next event would happen. Our proposed model excels in small data regime with the ability to incorporate domain knowledge in terms of logic rules. We model the dynamics of the event starts and ends via intensity function with the structures informed by a set of first-order temporal logic rules. Using the softened representation of temporal relations, and a weighted combination of logic rules, our probabilistic model can deal with uncertainty in events. Furthermore, many well-known point processes (e.g., Hawkes process, self-correcting point process) can be interpreted as special cases of our model given simple temporal logic rules. Our model, therefore, riches the family of point processes. We derive a maximum likelihood estimation procedure for our model and show that it can lead to accurate predictions when data are sparse and domain knowledge is critical. Shuang Li 0002, Lu Wang 0008, Xiaofu Chang, Xuqin Liu, Yao Xie 0002, Yuan Qi 0001 |
ICML | 2 |
| 2019 | Neural Keyphrase Generation via Reinforcement Learning with Adaptive RewardsabstractGenerating keyphrases that summarize the main points of a document is a fundamental task in natural language processing.Although existing generative models are capable of predicting multiple keyphrases for an input document as well as determining the number of keyphrases to generate, they still suffer from the problem of generating too few keyphrases.To address this problem, we propose a reinforcement learning (RL) approach for keyphrase generation, with an adaptive reward function that encourages a model to generate both sufficient and accurate keyphrases.Furthermore, we introduce a new evaluation method that incorporates name variations of the ground-truth keyphrases using the Wikipedia knowledge base.Thus, our evaluation method can more robustly evaluate the quality of predicted keyphrases.Extensive experiments on five real-world datasets of different scales demonstrate that our RL approach consistently and significantly improves the performance of the state-of-the-art generative models with both conventional and new evaluation methods. Document: DCE MRI data analysis for cancer area classification.The paper aims at improving the support of medical researchers in the context of in-vivo cancer imaging… The proposed approach is based on a three-step procedure: i) robust feature extraction from raw time-intensity curves, ii) voxel segmentation, and iii) voxel classification based on a learning-by-example approach… Finally, in the third step, a support vector machine (SVM) is trained to classify voxels according to the labels obtained by the clustering phase… Keyphrase labels: svm; Hou Pong Chan, Wang Chen 0001, Lu Wang 0008, Irwin King |
ACL (1) | 3 |
| 2019 | Argument Generation with Retrieval, Planning, and RealizationabstractAutomatic argument generation is an appealing but challenging task.In this paper, we study the specific problem of counterargument generation, and present a novel framework, CANDELA.It consists of a powerful retrieval system and a novel two-step generation model, where a text planning decoder first decides on the main talking points and a proper language style for each sentence, then a content realization decoder reflects the decisions and constructs an informative paragraph-level argument.Furthermore, our generation model is empowered by a retrieval system indexed with 12 million articles collected from Wikipedia and popular English news media, which provides access to highquality content with diversity.Automatic evaluation on a large-scale dataset collected from Reddit shows that our model yields significantly higher BLEU, ROUGE, and METEOR scores than the state-of-the-art and non-trivial comparisons.Human evaluation further indicates that our system arguments are more appropriate for refutation and richer in content. Xinyu Hua, Lu Wang 0008 |
ACL (1) | 3 |
| 2019 | BIGPATENT: A Large-Scale Dataset for Abstractive and Coherent SummarizationabstractMost existing text summarization datasets are compiled from the news domain, where summaries have a flattened discourse structure.In such datasets, summary-worthy content often appears in the beginning of input articles.Moreover, large segments from input articles are present verbatim in their respective summaries.These issues impede the learning and evaluation of systems that can understand an article's global content structure as well as produce abstractive summaries with high compression ratio.In this work, we present a novel dataset, BIGPATENT, consisting of 1.3 million records of U.S. patent documents along with human written abstractive summaries.Compared to existing summarization datasets, BIGPATENT has the following properties: i) summaries contain a richer discourse structure with more recurring entities, ii) salient content is evenly distributed in the input, and iii) lesser and shorter extractive fragments are present in the summaries.Finally, we train and evaluate baselines and popular learning models on BIGPATENT to shed light on new challenges and motivate future directions for summarization research. Eva Sharma, Chen Li 0037, Lu Wang 0008 |
ACL (1) | 3 |
| 2019 | Jointly Learning Semantic Parser and Natural Language Generator via Dual Information MaximizationabstractSemantic parsing aims to transform natural language (NL) utterances into formal meaning representations (MRs), whereas an NL generator achieves the reverse: producing a NL description for some given MRs.Despite this intrinsic connection, the two tasks are often studied separately in prior work.In this paper, we model the duality of these two tasks via a joint learning framework, and demonstrate its effectiveness of boosting the performance on both tasks.Concretely, we propose the method of dual information maximization (DIM) to regularize the learning process, where DIM empirically maximizes the variational lower bounds of expected joint distributions of NL and MRs.We further extend DIM to a semisupervision setup (SEMIDIM), which leverages unlabeled data of both tasks.Experiments on three datasets of dialogue management and code generation (and summarization) show that performance on both semantic parsing and NL generation can be consistently improved by DIM, in both supervised and semi-supervised setups 1 . Hai Ye, Lu Wang 0008 |
ACL (1) | 3 |
| 2019 | Joint Effects of Context and User History for Predicting Online Conversation Re-entriesabstractAs the online world continues its exponential growth, interpersonal communication has come to play an increasingly central role in opinion formation and change.In order to help users better engage with each other online, we study a challenging problem of re-entry prediction foreseeing whether a user will come back to a conversation they once participated in.We hypothesize that both the context of the ongoing conversations and the users' previous chatting history will affect their continued interests in future engagement.Specifically, we propose a neural framework with three main layers, each modeling context, user history, and interactions between them, to explore how the conversation context and user chatting history jointly result in their re-entry behavior.We experiment with two large-scale datasets collected from Twitter and Reddit.Results show that our proposed framework with biattention achieves an F1 score of 61.1 on Twitter conversations, outperforming the state-ofthe-art methods from previous work. * Jing Li is the corresponding author.…… H 1 : Is there literally no one on twitter who wants to talk about LET ME IN with me? :( H 2 : I think the change in overall tone was enough to let LMI stand on it's own.Love Giacchino's score too. Xingshan Zeng, Jing Li 0049, Lu Wang 0008, Kam-Fai Wong |
ACL (1) | 3 |
| 2019 | In Plain Sight: Media Bias Through the Lens of Factual ReportingabstractLisa Fan, Marshall White, Eva Sharma, Ruisi Su, Prafulla Kumar Choubey, Ruihong Huang, Lu Wang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Lisa Fan, Marshall White, Eva Sharma, Ruisi Su, Prafulla Kumar Choubey, Ruihong Huang, Lu Wang 0008 |
EMNLP/IJCNLP (1) | 7 |
| 2019 | Sentence-Level Content Planning and Style Specification for Neural Text GenerationabstractXinyu Hua, Lu Wang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xinyu Hua, Lu Wang 0008 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | An Entity-Driven Framework for Abstractive SummarizationabstractEva Sharma, Luyang Huang, Zhe Hu, Lu Wang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Eva Sharma, Luyang Huang, Lu Wang 0008 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Neural Conversation Recommendation with Online Interaction ModelingabstractXingshan Zeng, Jing Li, Lu Wang, Kam-Fai Wong. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xingshan Zeng, Jing Li 0049, Lu Wang 0008, Kam-Fai Wong |
EMNLP/IJCNLP (1) | 3 |
| 2018 | Neural Argument Generation Augmented with Externally Retrieved EvidenceabstractHigh quality arguments are essential elements for human reasoning and decision-making processes.However, effective argument construction is a challenging task for both human and machines.In this work, we study a novel task on automatically generating arguments of a different stance for a given statement.We propose an encoder-decoder style neural network-based argument generation model enriched with externally retrieved evidence from Wikipedia.Our model first generates a set of talking point phrases as intermediate representation, followed by a separate decoder producing the final argument based on both input and the keyphrases.Experiments on a large-scale dataset collected from Reddit show that our model constructs arguments with more topicrelevant content than a popular sequence-tosequence generation model according to both automatic evaluation and human assessments. Xinyu Hua, Lu Wang 0008 |
ACL (1) | 2 |
| 2018 | Semi-Supervised Learning for Neural Keyphrase GenerationabstractWe study the problem of generating keyphrases that summarize the key points for a given document.While sequence-to-sequence (seq2seq) models have achieved remarkable performance on this task (Meng et al., 2017), model training often relies on large amounts of labeled data, which is only applicable to resource-rich domains.In this paper, we propose semi-supervised keyphrase generation methods by leveraging both labeled data and large-scale unlabeled samples for learning.Two strategies are proposed.First, unlabeled documents are first tagged with synthetic keyphrases obtained from unsupervised keyphrase extraction methods or a selflearning algorithm, and then combined with labeled samples for training.Furthermore, we investigate a multi-task learning framework to jointly learn to generate keyphrases as well as the titles of the articles.Experimental results show that our semi-supervised learning-based methods outperform a state-of-the-art model trained with labeled data only. Hai Ye, Lu Wang 0008 |
EMNLP | 2 |
| 2018 | Microblog Conversation Recommendation via Joint Modeling of Topics and DiscourseabstractXingshan Zeng, Jing Li, Lu Wang, Nicholas Beauchamp, Sarah Shugars, Kam-Fai Wong. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Xingshan Zeng, Jing Li 0049, Lu Wang 0008, Nick Beauchamp, Sarah Shugars, Kam-Fai Wong |
NAACL-HLT | 3 |
| 2017 | Joint Modeling of Content and Discourse Relations in DialoguesabstractWe present a joint modeling approach to identify salient discussion points in spoken meetings as well as to label the discourse relations between speaker turns.A variation of our model is also discussed when discourse relations are treated as latent variables.Experimental results on two popular meeting corpora show that our joint model can outperform state-of-the-art approaches for both phrasebased content selection and discourse relation prediction tasks.We also evaluate our model on predicting the consistency among team members' understanding of their group decisions.Classifiers trained with features constructed from our model achieve significant better predictive performance than the state-of-the-art.D: Three different types of batteries.Um can either use a hand dynamo, or the kinetic type ones, you know that they use in watches, or else uh a solar powered one.B: Um the bat uh the battery for a a watch wouldn't require a lot of power, would be my one query.Is a kinetic one going to be able to supply enough power?D: Yeah, I don't think it would.C: Yeah.D: We should probably just use conventional batteries. Kechen Qin, Lu Wang 0008, Joseph Kim |
ACL (1) | 2 |
| 2017 | Weakly-Guided User Stance Prediction via Joint Modeling of Content and Social InteractionabstractSocial media websites have become a popular outlet for online users to express their opinions on controversial issues, such as gun control and abortion. Understanding users' stances and their arguments is a critical task for policy-making process and public deliberation. Existing methods rely on large amounts of human annotation for predicting stance on issues of interest, which is expensive and hard to scale to new problems. In this work, we present a weakly-guided user stance modeling framework which simultaneously considers two types of information: what do you say (via stance-based content generative model) and how do you behave (via social interaction-based graph regularization). We experiment with two types of social media data: news comments and discussion forum posts. Our model uniformly outperforms a logistic regression-based supervised method on stance-based link prediction for unseen users on news comments. Our method also achieves better or comparable stance prediction performance for discussion forum users, when compared with state-of-the-art supervised systems. Meanwhile, separate word distributions are learned for users of opposite stances. This potentially helps with better understanding and interpretation of conflicting arguments for controversial issues. Yizhou Sun, Lu Wang 0008, Yupeng Gu |
CIKM | 3 |
| 2017 | Winning on the Merits: The Joint Effects of Content and Style on Debate OutcomesabstractDebate and deliberation play essential roles in politics and government, but most models presume that debates are won mainly via superior style or agenda control. Ideally, however, debates would be won on the merits, as a function of which side has the stronger arguments. We propose a predictive model of debate that estimates the effects of linguistic features and the latent persuasive strengths of different topics, as well as the interactions between the two. Using a dataset of 118 Oxford-style debates, our model’s combination of content (as latent topics) and style (as linguistic features) allows us to predict audience-adjudicated winners with 74% accuracy, significantly outperforming linguistic features alone (66%). Our model finds that winning sides employ stronger arguments, and allows us to identify the linguistic features associated with strong or weak arguments. Lu Wang 0008, Nick Beauchamp, Sarah Shugars, Kechen Qin |
Trans. Assoc. Comput. Linguistics | 1 |
| 2016 | Neural Network-Based Abstract Generation for Opinions and ArgumentsabstractWe study the problem of generating abstractive summaries for opinionated text. We propose an attention-based neural network model that is able to absorb information from multiple text units to construct informative, concise, and fluent summaries. An importance-based sampling method is designed to allow the encoder to integrate information from an important subset of input. Automatic evaluation indicates that our system outperforms state-of-the-art abstractive and extractive summarization systems on two newly collected datasets of movie reviews and arguments. Our system summaries are also rated as more informative and grammatical in human evaluation. Lu Wang 0008, Wang Ling |
HLT-NAACL | 1 |
| 2015 | Socially-Informed Timeline Generation for Complex EventsabstractExisting timeline generation systems for complex events consider only information from traditional media, ignoring the rich social context provided by user-generated content that reveals representative public interests or insightful opinions. We instead aim to generate socially-informed timelines that contain both news article summaries and selected user comments. We present an optimization framework designed to balance topical cohesion between the article and comment summaries along with their informativeness and coverage of the event. Automatic evaluations on real-world datasets that cover four complex events show that our system produces more informative timelines than state-of-the-art systems. In human evaluation, the associated comment summaries are furthermore rated more insightful than editor's picks and comments ranked highly by users. Lu Wang 0008, Claire Cardie, Galen Marchetti |
HLT-NAACL | 1 |
| 2014 | Query-Focused Opinion Summarization for User-Generated Content
Lu Wang 0008, Hema Raghavan, Claire Cardie, Vittorio Castelli |
COLING | 1 |
| 2014 | Leveraging semantic web search and browse sessions for multi-turn spoken dialog systemsabstractTraining statistical dialog models in spoken dialog systems (SDS) requires large amounts of annotated data. The lack of scalable methods for data mining and annotation poses a significant hurdle for state-of-the-art statistical dialog managers. This paper presents an approach that directly leverage billions of web search and browse sessions to overcome this hurdle. The key insight is that task completion through web search and browse sessions is (a) predictable and (b) generalizes to spoken dialog task completion. The new method automatically mines behavioral search and browse patterns from web logs and translates them into spoken dialog models. We experiment with naturally occurring spoken dialogs and large scale web logs. Our session-based models outperform the state-of-the-art method for entity extraction task in SDS. We also achieve better performance for both entity and relation extraction on web search queries when compared with nontrivial baselines. Lu Wang 0008, Larry Heck, Dilek Hakkani-Tür |
ICASSP | 1 |
| 2013 | Domain-Independent Abstract Generation for Focused Meeting Summarization
Lu Wang 0008, Claire Cardie |
ACL (1) | 1 |
| 2013 | A Sentence Compression Based Framework to Query-Focused Multi-Document Summarization
Lu Wang 0008, Hema Raghavan, Vittorio Castelli, Radu Florian, Claire Cardie |
ACL (1) | 1 |
| 2012 | Unsupervised Topic Modeling Approaches to Decision Summarization in Spoken Meetings
Lu Wang 0008, Claire Cardie |
SIGDIAL Conference | 1 |
| 2012 | Focused Meeting Summarization via Unsupervised Relation Extraction
Lu Wang 0008, Claire Cardie |
SIGDIAL Conference | 1 |