Jackie Chi Kit Cheung

dblp:00/9012 · also Jackie C. K. Cheung, Jackie CK Cheung · DBLP profile ↗
← Back
77ranked-venue papers
10as first author
35since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 75 · 10 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Identifying and Analyzing Performance-Critical Tokens in Large Language Models
abstract
In-context learning (ICL) has emerged as an effective solution for few-shot learning with large language models (LLMs). However, how LLMs leverage demonstrations to specify a task and learn a corresponding computational function through ICL is underexplored. Drawing from the way humans learn from content-label mappings in demonstrations, we categorize the tokens in an ICL prompt into content, stopword, and template tokens. Our goal is to identify the types of tokens whose representations directly influence LLM's performance, a property we refer to as being performance-critical. By ablating representations from the attention of the test example, we find that the representations of informative content tokens have less influence on performance compared to template and stopword tokens, which contrasts with the human attention to informative words. We give evidence that the representations of performance-critical tokens aggregate information from the content tokens. Moreover, we demonstrate experimentally that lexical meaning, repetition, and structural cues are the main distinguishing characteristics of these tokens. Our work sheds light on how LLMs learn to perform tasks from demonstrations and deepens our understanding of the roles different types of tokens play in LLMs.
Yu Bai 0018, Heyan Huang, Cesare Spinoso Di Piano, Sanxing Chen, Marc-Antoine Rondeau, Yang Gao 0016, Jackie Chi Kit Cheung
AAAI7
2025 Error Diversity Matters: An Error-Resistant Ensemble Method for Unsupervised Dependency Parsing
abstract
We address unsupervised dependency parsing by building an ensemble of diverse existing models through post hoc aggregation of their output dependency parse structures. We observe that these ensembles often suffer from low robustness against weak ensemble components due to error accumulation. To tackle this problem, we propose an efficient ensemble-selection approach that considers error diversity and avoids error accumulation. Results demonstrate that our approach outperforms each individual model as well as previous ensemble techniques. Additionally, our experiments show that the proposed ensemble-selection method significantly enhances the performance and robustness of our ensemble, surpassing previously proposed strategies, which have not accounted for error diversity.
Behzad Shayegh, Hobie H.-B. Lee, Xiaodan Zhu 0001, Jackie Chi Kit Cheung, Lili Mou
AAAI4
2025 Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMs
abstract
The widespread success of large language models (LLMs) on NLP benchmarks has been accompanied by concerns that LLMs function primarily as stochastic parrots that reproduce texts similar to what they saw during pre-training, often erroneously.But what is the nature of their errors, and do these errors exhibit any regularities?In this work, we examine irrelevant context hallucinations, in which models integrate misleading contextual cues into their predictions.Through behavioral analysis, we show that these errors result from a structured yet flawed mechanism that we term class-based (mis)generalization, in which models combine abstract class cues with features extracted from the query or context to derive answers.Furthermore, mechanistic interpretability experiments on Llama-3, Mistral, and Pythia across 39 factual recall relation types reveal that this behavior is reflected in the model's internal computations: (i) abstract class representations are constructed in lower layers before being refined into specific answers in higher layers, (ii) feature selection is governed by two competing circuits -one prioritizing direct query-based reasoning, the other incorporating contextual cues -whose relative influences determine the final output.Our findings provide a more nuanced perspective on the stochastic parrot argument: through form-based training, LLMs can exhibit generalization leveraging abstractions, albeit in unreliable ways based on contextual cues -what we term stochastic chameleons. 1 * Equal contribution.
Ziling Cheng, Meng Cao 0003, Marc-Antoine Rondeau, Jackie Chi Kit Cheung
ACL (1)4
2025 (RSA)²: A Rhetorical-Strategy-Aware Rational Speech Act Framework for Figurative Language Understanding
abstract
Figurative language (e.g., irony, hyperbole, understatement) is ubiquitous in human communication, resulting in utterances where the literal and the intended meanings do not match. The Rational Speech Act (RSA) framework, which explicitly models speaker intentions, is the most widespread theory of probabilistic pragmatics, but existing implementations are either unable to account for figurative expressions or require modeling the implicit motivations for using figurative language (e.g., to express joy or annoyance) in a setting-specific way. In this paper, we introduce the Rhetorical-Strategy-Aware RSA (RSA)² framework which models figurative language use by considering a speaker’s employed rhetorical strategy. We show that (RSA)² enables human-compatible interpretations of non-literal utterances without modeling a speaker’s motivations for being non-literal. Combined with LLMs, it achieves state-of-the-art performance on the ironic split of PragMega+, a new irony interpretation dataset introduced in this study.
Cesare Spinoso Di Piano, David Eric Austin, Pablo Piantanida, Jackie Chi Kit Cheung
ACL (1)4
2025 Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation
abstract
Final-answer-based metrics are commonly used for evaluating large language models (LLMs) on math word problems, often taken as proxies for reasoning ability. However, such metrics conflate two distinct sub-skills: abstract formulation (capturing mathematical relationships using expressions) and arithmetic computation (executing the calculations). Through a disentangled evaluation on GSM8K and SVAMP, we find that the final-answer accuracy of Llama-3 and Qwen2.5 (1B-32B) without CoT is overwhelmingly bottlenecked by the arithmetic computation step and not by the abstract formulation step. Contrary to the common belief, we show that CoT primarily aids in computation, with limited impact on abstract formulation. Mechanistically, we show that these two skills are composed conjunctively even in a single forward pass without any reasoning steps via an abstract-then-compute mechanism: models first capture problem abstractions, then handle computation. Causal patching confirms these abstractions are present, transferable, composable, and precede computation. These behavioural and mechanistic findings highlight the need for disentangled evaluation to accurately assess LLM reasoning and to guide future improvements.
Ziling Cheng, Meng Cao 0003, Leila Pishdad, Yanshuai Cao, Jackie Chi Kit Cheung
EMNLP5
2025 Collaborative Rational Speech Act: Pragmatic Reasoning for Multi-Turn Dialog
abstract
As AI systems take on collaborative roles, they must reason about shared goals and beliefs-not just generate fluent language.The Rational Speech Act (RSA) framework offers a principled approach to pragmatic reasoning, but existing extensions face challenges in scaling to multi-turn, collaborative scenarios.In this paper, we introduce Collaborative Rational Speech Act (CRSA), an information-theoretic (IT) extension of RSA that models multi-turn dialog by optimizing a gain function adapted from rate-distortion theory.This gain is an extension of the gain model that is maximized in the original RSA model but takes into account the scenario in which both agents in a conversation have private information and produce utterances conditioned on the dialog.We demonstrate the effectiveness of CRSA on referential games and template-based doctor-patient dialogs in the medical domain.Empirical results show that CRSA yields more consistent, interpretable, and collaborative behavior than existing baselines, paving the way for more pragmatically competent language agents.
Lautaro Estienne, Gabriel Ben Zenou, Nona Naderi, Jackie Chi Kit Cheung, Pablo Piantanida
EMNLP4
2025 Beyond the Safety Bundle: Auditing the Helpful and Harmless Dataset
abstract
Khaoula Chehbouni, Jonathan Colaço Carr, Yash More, Jackie CK Cheung, Golnoosh Farnadi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Khaoula Chehbouni, Jonathan Colaço Carr, Yash More, Jackie Chi Kit Cheung, Golnoosh Farnadi
NAACL (Long Papers)4
2025 Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
abstract
Evaluating natural language generation (NLG) systems remains a core challenge, further complicated by the rise of general-purpose large language models (LLMs). Recently, large language models as judges (LLJs) have emerged as a scalable, cost-effective alternative to traditional metrics, but their validity remains underexplored. This position paper argues that the current enthusiasm around LLJs may be premature, as their adoption has outpaced rigorous scrutiny of their reliability and validity as evaluators. Drawing on measurement theory from the social sciences, we identify and critically assess four core assumptions underlying the use of LLJs: their ability to act as proxies for human judgment, their capabilities as evaluators, their scalability, and their cost-effectiveness. We examine how each of these assumptions may be challenged by the inherent limitations of LLMs, LLJs, or current practices in NLG evaluation. To ground our analysis, we explore three applications of LLJs at various stages of the machine learning pipeline: text summarization, data annotation and safety alignment. Finally, we highlight the need for more responsible evaluation practices in LLJs evaluation, to ensure that their growing role in the field supports, rather than undermines, progress in NLG.
Khaoula Chehbouni, Mohammed Haddou, Jackie Chi Kit Cheung, Golnoosh Farnadi
NeurIPS3
2025 Learning Task-Agnostic Representations through Multi-Teacher Distillation
abstract
Casting complex inputs into tractable representations is a critical step across various fields. Diverse embedding models emerge from differences in architectures, loss functions, input modalities and datasets, each capturing unique aspects of the input. Multi-teacher distillation leverages this diversity to enrich representations but often remains tailored to specific tasks. We introduce a task-agnostic framework based on a ``majority vote" objective function. We demonstrate that this function is bounded by the mutual information between the student and the teachers' embeddings, leading to a task-agnostic distillation loss that eliminates dependence on task-specific labels or prior knowledge. Comprehensive evaluations across text, vision models, and molecular modeling show that our method effectively leverages teacher diversity, resulting in representations enabling better performance for a wide range of downstream tasks such as classification, clustering, or regression. Additionally, we train and release state-of-the-art embedding models, enhancing downstream performance in various modalities.
Philippe Formont, Maxime Darrin, Banafsheh Karimian, Eric Granger, Jackie Chi Kit Cheung, Ismail Ben Ayed, Mohammadhadi Shateri, Pablo Piantanida
NeurIPS5
2024 Unsupervised Layer-Wise Score Aggregation for Textual OOD Detection
abstract
Out-of-distribution (OOD) detection is a rapidly growing field due to new robustness and security requirements driven by an increased number of AI-based systems. Existing OOD textual detectors often rely on anomaly scores (\textit{e.g.}, Mahalanobis distance) computed on the embedding output of the last layer of the encoder. In this work, we observe that OOD detection performance varies greatly depending on the task and layer output. More importantly, we show that the usual choice (the last layer) is rarely the best one for OOD detection and that far better results can be achieved, provided that an oracle selects the best layer. We propose a data-driven, unsupervised method to leverage this observation to combine layer-wise anomaly scores. In addition, we extend classical textual OOD benchmarks by including classification tasks with a more significant number of classes (up to 150), which reflects more realistic settings. On this augmented benchmark, we show that the proposed post-aggregation methods achieve robust and consistent results comparable to using the best layer according to an oracle while removing manual feature selection altogether.
Maxime Darrin, Guillaume Staerman, Eduardo Dadalto Câmara Gomes, Jackie Chi Kit Cheung, Pablo Piantanida, Pierre Colombo
AAAI4
2024 How Teachers Can Use Large Language Models and Bloom's Taxonomy to Create Educational Quizzes
abstract
Question generation (QG) is a natural language processing task with an abundance of potential benefits and use cases in the educational domain. In order for this potential to be realized, QG systems must be designed and validated with pedagogical needs in mind. However, little research has assessed or designed QG approaches with the input of real teachers or students. This paper applies a large language model-based QG approach where questions are generated with learning goals derived from Bloom's taxonomy. The automatically generated questions are used in multiple experiments designed to assess how teachers use them in practice. The results demonstrate that teachers prefer to write quizzes with automatically generated questions, and that such quizzes have no loss in quality compared to handwritten versions. Further, several metrics indicate that automatically generated questions can even improve the quality of the quizzes created, showing the promise for large scale use of QG in the classroom setting.
Sabina Elkins, Ekaterina Kochmar, Jackie Chi Kit Cheung, Iulian Serban
AAAI3
2024 GLIMPSE: Pragmatically Informative Multi-Document Summarization for Scholarly Reviews
abstract
Scientific peer review is essential for the quality of academic publications.However, the increasing number of paper submissions to conferences has strained the reviewing process.This surge poses a burden on area chairs who have to carefully read an ever-growing volume of reviews and discern each reviewer's main arguments as part of their decision process.In this paper, we introduce GLIMPSE, a summarization method designed to offer a concise yet comprehensive overview of scholarly reviews.Unlike traditional consensus-based methods, GLIMPSE extracts both common and unique opinions from the reviews.We introduce novel uniqueness scores based on the Rational Speech Act framework to identify relevant sentences in the reviews.Our method aims to provide a pragmatic glimpse into all reviews, offering a balanced perspective on their opinions.Our experimental results with both automatic metrics and human evaluation show that GLIMPSE generates more discriminative summaries than baseline methods in terms of human evaluation while achieving comparable performance with these methods in terms of automatic metrics.
Maxime Darrin, Ines Arous, Pablo Piantanida, Jackie Chi Kit Cheung
ACL (1)4
2024 COSMIC: Mutual Information for Task-Agnostic Summarization Evaluation
abstract
Assessing the quality of summarizers poses significant challenges-gold summaries are hard to obtain and their suitability depends on the use context of the summarization system.Who is the user of the system, and what do they intend to do with the summary?In response, we propose a novel task-oriented evaluation approach that assesses summarizers based on their capacity to produce summaries while preserving task outcomes.We theoretically establish both a lower and upper bound on the expected error rate of these tasks, which depends on the mutual information between source texts and generated summaries.We introduce COSMIC, a practical implementation of this metric, and demonstrate its strong correlation with human judgment-based metrics, as well as its effectiveness in predicting downstream task performance.Comparative analyses against established metrics like BERTScore and ROUGE highlight the competitive performance of COSMIC.
Maxime Darrin, Philippe Formont, Jackie Chi Kit Cheung, Pablo Piantanida
ACL (1)3
2024 ECBD: Evidence-Centered Benchmark Design for NLP
abstract
Yu Lu Liu, Su Lin Blodgett, Jackie Cheung, Q. Vera Liao, Alexandra Olteanu, Ziang Xiao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yu Lu Liu, Su Lin Blodgett, Jackie Chi Kit Cheung, Qingzi Vera Liao, Alexandra Olteanu, Ziang Xiao
ACL (1)3
2024 A Controlled Reevaluation of Coreference Resolution Models
abstract
All state-of-the-art coreference resolution (CR) models involve finetuning a pretrained language model. Whether the superior performance of one CR model over another is due to the choice of language model or other factors, such as the task-specific architecture, is difficult or impossible to determine due to lack of a standardized experimental setup. To resolve this ambiguity, we systematically evaluate five CR models and control for certain design decisions including the pretrained language model used by each. When controlling for language model size, encoder-based CR models outperform more recent decoder-based models in terms of both accuracy and inference speed. Surprisingly, among encoder-based CR models, more recent models are not always more accurate, and the oldest CR model that we test generalizes the best to out-of-domain textual genres. We conclude that controlling for the choice of language model reduces most, but not all, of the increase in F1 score reported in the past five years.
Ian Porada, Xiyuan Zou, Jackie Chi Kit Cheung
LREC/COLING3
2024 CItruS: Chunked Instruction-aware State Eviction for Long Sequence Modeling
abstract
Long sequence modeling has gained broad interest as large language models (LLMs) continue to advance.Recent research has identified that a large portion of hidden states within the key-value caches of Transformer models can be discarded (also termed evicted) without affecting the perplexity performance in generating long sequences.However, we show that these methods, despite preserving perplexity performance, often drop information that is important for solving downstream tasks, a problem which we call information neglect.To address this issue, we introduce Chunked Instruction-aware State Eviction (CItruS), a novel modeling technique that integrates the attention preferences useful for a downstream task into the eviction process of hidden states.In addition, we design a method for chunked sequence processing to further improve efficiency.Our training-free method exhibits superior performance on long sequence comprehension and retrieval tasks over several strong baselines under the same memory budget, while preserving language modeling perplexity.The code and data have been released at https: //github.com/ybai-nlp/CItruS.
Yu Bai 0018, Xiyuan Zou, Heyan Huang, Sanxing Chen, Marc-Antoine Rondeau, Yang Gao 0016, Jackie Chi Kit Cheung
EMNLP7
2024 Ensemble Distillation for Unsupervised Constituency Parsing
abstract
We investigate the unsupervised constituency parsing task, which organizes words and phrases of a sentence into a hierarchical structure without using linguistically annotated data. We observe that existing unsupervised parsers capture different aspects of parsing structures, which can be leveraged to enhance unsupervised parsing performance. To this end, we propose a notion of "tree averaging," based on which we further propose a novel ensemble method for unsupervised parsing. To improve inference efficiency, we further distill the ensemble knowledge into a student model; such an ensemble-then-distill process is an effective approach to mitigate the over-smoothing problem existing in common multi-teacher distilling methods. Experiments show that our method surpasses all previous approaches, consistently demonstrating its effectiveness and robustness across various runs, with different ensemble components, and under domain-shift conditions.
Behzad Shayegh, Yanshuai Cao, Xiaodan Zhu 0001, Jackie Chi Kit Cheung, Lili Mou
ICLR4
2024 Successor Features for Efficient Multi-Subject Controlled Text Generation
abstract
While large language models (LLMs) have achieved impressive performance in generating fluent and realistic text, controlling the generated text so that it exhibits properties such as safety, factuality, and non-toxicity remains challenging. Existing decoding-based controllable text generation methods are static in terms of the dimension of control; if the target subject is changed, they require new training. Moreover, it can quickly become prohibitive to concurrently control multiple subjects. To address these challenges, we first show that existing methods can be framed as a reinforcement learning problem, where an action-value function estimates the likelihood of a desired attribute appearing in the generated text. Then, we introduce a novel approach named SF-Gen, which leverages the concept of successor features to decouple the dynamics of LLMs from task-specific rewards. By employing successor features, our method proves to be memory-efficient and computationally efficient for both training and decoding, especially when dealing with multiple target subjects. To the best of our knowledge, our research represents the first application of successor features in text generation. In addition to its computational efficiency, the resultant language produced by our method is comparable to the SOTA (and outperforms baselines) in both control measures as well as language quality, which we demonstrate through a series of experiments in various controllable text generation tasks.
Meng Cao 0003, Mehdi Fatemi, Jackie Chi Kit Cheung, Samira Shabanian
ICML3
2024 When is an Embedding Model More Promising than Another?
abstract
Embedders play a central role in machine learning, projecting any object into numerical representations that can, in turn, be leveraged to perform various downstream tasks. The evaluation of embedding models typically depends on domain-specific empirical approaches utilizing downstream tasks, primarily because of the lack of a standardized framework for comparison. However, acquiring adequately large and representative datasets for conducting these assessments is not always viable and can prove to be prohibitively expensive and time-consuming. In this paper, we present a unified approach to evaluate embedders. First, we establish theoretical foundations for comparing embedding models, drawing upon the concepts of sufficiency and informativeness. We then leverage these concepts to devise a tractable comparison criterion (information sufficiency), leading to a task-agnostic and self-supervised ranking procedure. We demonstrate experimentally that our approach aligns closely with the capability of embedding models to facilitate various downstream tasks in both natural language processing and molecular biology. This effectively offers practitioners a valuable tool for prioritizing model trials.
Maxime Darrin, Philippe Formont, Ismail Ben Ayed, Jackie Chi Kit Cheung, Pablo Piantanida
NeurIPS4
2024 Do LLMs Build World Representations? Probing Through the Lens of State Abstraction
abstract
How do large language models (LLMs) encode the state of the world, including the status of entities and their relations, as described by a text? While existing work directly probes for a complete state of the world, our research explores whether and how LLMs abstract this world state in their internal representations. We propose a new framework for probing for world representations through the lens of state abstraction theory from reinforcement learning, which emphasizes different levels of abstraction, distinguishing between general abstractions that facilitate predicting future states and goal-oriented abstractions that guide the subsequent actions to accomplish tasks. To instantiate this framework, we design a text-based planning task, where an LLM acts as an agent in an environment and interacts with objects in containers to achieve a specified goal state. Our experiments reveal that fine-tuning as well as advanced pre-training strengthens LLM-built representations' tendency of maintaining goal-oriented abstractions during decoding, prioritizing task completion over recovery of the world's state and dynamics.
Zichao Li 0003, Yanshuai Cao, Jackie Chi Kit Cheung
NeurIPS3
2023 The KITMUS Test: Evaluating Knowledge Integration from Multiple Sources
abstract
Akshatha Arodi, Martin Pömsl, Kaheer Suleman, Adam Trischler, Alexandra Olteanu, Jackie Chi Kit Cheung. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Akshatha Arodi, Martin Pömsl, Kaheer Suleman, Adam Trischler, Alexandra Olteanu, Jackie Chi Kit Cheung
ACL (1)6
2023 Systematic Rectification of Language Models via Dead-end Analysis
Meng Cao 0003, Mehdi Fatemi, Jackie Chi Kit Cheung, Samira Shabanian
ICLR3
2022 Hallucinated but Factual! Inspecting the Factuality of Hallucinations in Abstractive Summarization
abstract
State-of-the-art abstractive summarization systems often generate hallucinations; i.e., content that is not directly inferable from the source text.Despite being assumed incorrect, we find that much hallucinated content is factual, namely consistent with world knowledge.These factual hallucinations can be beneficial in a summary by providing useful background information.In this work, we propose a novel detection approach that separates factual from non-factual hallucinations of entities.Our method utilizes an entity's prior and posterior probabilities according to pre-trained and finetuned masked language models, respectively.Empirical results suggest that our approach outperforms five baselines and strongly correlates with human judgments.Furthermore, we show that our detector, when used as a reward signal in an off-line reinforcement learning (RL) algorithm, significantly improves the factuality of summaries while maintaining the level of abstractiveness. 1
Meng Cao 0003, Yue Dong 0002, Jackie Chi Kit Cheung
ACL (1)3
2022 Characterizing Idioms: Conventionality and Contingency
abstract
Idioms are unlike most phrases in two important ways.First, words in an idiom have non-canonical meanings.Second, the noncanonical meanings of words in an idiom are contingent on the presence of other words in the idiom.Linguistic theories differ on whether these properties depend on one another, as well as whether special theoretical machinery is needed to accommodate idioms.We define two measures that correspond to the properties above, and we implement them using BERT (Devlin et al., 2019) and XLNet (Yang et al., 2019).We show that English idioms fall at the expected intersection of the two dimensions, but that the dimensions themselves are not correlated.Our results suggest that special machinery to handle idioms may not be warranted.
Michaela Socolof, Jackie Chi Kit Cheung, Michael Wagner 0019, Timothy J. O'Donnell
ACL (1)2
2022 Source-summary Entity Aggregation in Abstractive Summarization
abstract
In a text, entities mentioned earlier can be referred to in later discourse by a more general description. For example, Celine Dion and Justin Bieber can be referred to by Canadian singers or celebrities. In this work, we study this phenomenon in the context of summarization, where entities from a source text are generalized in the summary. We call such instances source-summary entity aggregations. We categorize these aggregations into two types and analyze them in the Cnn/Dailymail corpus, showing that they are reasonably frequent. We then examine how well three state-of-the-art summarization systems can generate such aggregations within summaries. We also develop techniques to encourage them to generate more aggregations. Our results show that there is significant room for improvement in producing semantically correct aggregations.
José Ángel González, Annie Louis, Jackie Chi Kit Cheung
COLING3
2022 Investigating the Performance of Transformer-Based NLI Models on Presuppositional Inferences
abstract
Presuppositions are assumptions that are taken for granted by an utterance, and identifying them is key to a pragmatic interpretation of language. In this paper, we investigate the capabilities of transformer models to perform NLI on cases involving presupposition. First, we present simple heuristics to create alternative “contrastive” test cases based on the ImpPres dataset and investigate the model performance on those test cases. Second, to better understand how the model is making its predictions, we analyze samples from sub-datasets of ImpPres and examine model performance on them. Overall, our findings suggest that NLI-trained transformer models seem to be exploiting specific structural and lexical cues as opposed to performing some kind of pragmatic reasoning.
Jad Kabbara, Jackie Chi Kit Cheung
COLING2
2022 A Multifaceted Framework to Evaluate Evasion, Content Preservation, and Misattribution in Authorship Obfuscation Techniques
abstract
Authorship obfuscation techniques have commonly been evaluated based on their ability to hide the author's identity (evasion) while preserving the content of the original text.However, to avoid overstating the systems' effectiveness, evasion detection must be evaluated using competitive identification techniques in settings that mimic real-life scenarios, and the outcomes of the content-preservation evaluation have to be interpretable by potential users of these obfuscation tools.Motivated by recent work on cross-topic authorship identification and content preservation in summarization, we re-evaluate different authorship obfuscation techniques on detection evasion and content preservation.Furthermore, we propose a new information-theoretic measure to characterize the misattribution harm that can be caused by detection evasion.Our results 1 reveal key weaknesses in state-of-the-art obfuscation techniques and a surprisingly competitive effectiveness from a back-translation baseline in all evaluation aspects.
Malik H. Altakrori, Thomas Scialom, Benjamin C. M. Fung, Jackie Chi Kit Cheung
EMNLP4
2022 Learning with Rejection for Abstractive Text Summarization
abstract
State-of-the-art abstractive summarization systems frequently hallucinate content that is not supported by the source document, mainly due to noise in the training dataset.Existing methods opt to drop the noisy samples or tokens from the training set entirely, reducing the effective training set size and creating an artificial propensity to copy words from the source.In this work, we propose a training objective for abstractive summarization based on rejection learning, in which the model learns whether or not to reject potentially noisy tokens.We further propose a regularized decoding objective that penalizes non-factual candidate summaries during inference by using the rejection probability learned during training.We show that our method considerably improves the factuality of generated summaries in automatic and human evaluations when compared to five baseline models, and that it does so while increasing the abstractiveness of the generated summaries.1 Source: (...) Chris Cox, the university's director of development, said the research centre would translate discoveries made in the laboratory into new treatments.He said it would house 150 additional researchers who will be developing new ideas and treatments.Research will focus on radiation therapy, lung cancer, women's cancers, melanoma and haematological oncology.The centre is the result of a partnership between The University of Manchester, The Christie NHS Foundation Trust and Cancer Research UK.The Christie's chief executive Caroline Shaw said the funding would "help facilitate groundbreaking research right here in Manchester".(...)
Meng Cao 0003, Yue Dong 0002, Jackie Chi Kit Cheung
EMNLP4
2022 Does Pre-training Induce Systematic Inference? How Masked Language Models Acquire Commonsense Knowledge
abstract
Transformer models pre-trained with a maskedlanguage-modeling objective (e.g., BERT) encode commonsense knowledge as evidenced by behavioral probes; however, the extent to which this knowledge is acquired by systematic inference over the semantics of the pretraining corpora is an open question.To answer this question, we selectively inject verbalized knowledge into the pre-training minibatches of BERT and evaluate how well the model generalizes to supported inferences after pre-training on the injected knowledge.We find generalization does not improve over the course of pre-training BERT from scratch, suggesting that commonsense knowledge is acquired from surface-level, co-occurrence patterns rather than induced, systematic reasoning.
Ian Porada, Alessandro Sordoni, Jackie Chi Kit Cheung
NAACL-HLT3
2021 Deep Discourse Analysis for Generating Personalized Feedback in Intelligent Tutor Systems
abstract
We explore creating automated, personalized feedback in an intelligent tutoring system (ITS). Our goal is to pinpoint correct and incorrect concepts in student answers in order to achieve better student learning gains. Although automatic methods for providing personalized feedback exist, they do not explicitly inform students about which concepts in their answers are correct or incorrect. Our approach involves decomposing students answers using neural discourse segmentation and classification techniques. This decomposition yields a relational graph over all discourse units covered by the reference solutions and student answers. We use this inferred relational graph structure and a neural classifier to match student answers with reference solutions and generate personalized feedback. Although the process is completely automated and data-driven, the personalized feedback generated is highly contextual, domain-aware and effectively targets each student's misconceptions and knowledge gaps. We test our method in a dialogue-based ITS and demonstrate that our approach results in high-quality feedback and significantly improved student learning gains.
Matt Grenander, Robert Belfer, Ekaterina Kochmar, Iulian Serban, François St-Hilaire, Jackie Chi Kit Cheung
AAAI6
2021 ADEPT: An Adjective-Dependent Plausibility Task
abstract
Ali Emami, Ian Porada, Alexandra Olteanu, Kaheer Suleman, Adam Trischler, Jackie Chi Kit Cheung. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Ali Emami, Ian Porada, Alexandra Olteanu, Kaheer Suleman, Adam Trischler, Jackie Chi Kit Cheung
ACL/IJCNLP (1)6
2021 Optimizing Deeper Transformers on Small Datasets
abstract
Peng Xu, Dhruv Kumar, Wei Yang, Wenjie Zi, Keyi Tang, Chenyang Huang, Jackie Chi Kit Cheung, Simon J.D. Prince, Yanshuai Cao. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Dhruv Kumar 0005, Wei Yang 0017, Wenjie Zi, Keyi Tang, Chenyang Huang 0001, Jackie Chi Kit Cheung, Simon Prince, Yanshuai Cao
ACL/IJCNLP (1)7
2021 Discourse-Aware Unsupervised Summarization for Long Scientific Documents
abstract
We propose an unsupervised graph-based ranking model for extractive summarization of long scientific documents.Our method assumes a two-level hierarchical graph representation of the source document, and exploits asymmetrical positional cues to determine sentence importance.Results on the PubMed and arXiv datasets show that our approach 1 outperforms strong unsupervised baselines by wide margins in automatic metrics and human evaluation.In addition, it achieves performance comparable to many state-of-the-art supervised approaches which are trained on hundreds of thousands of examples.These results suggest that patterns in the discourse structure are a strong signal for determining importance in scientific articles.
Yue Dong 0002, Andrei Romascanu, Jackie Chi Kit Cheung
EACL3
2021 Modeling Event Plausibility with Consistent Conceptual Abstraction
abstract
Ian Porada, Kaheer Suleman, Adam Trischler, Jackie Chi Kit Cheung. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Ian Porada, Kaheer Suleman, Adam Trischler, Jackie Chi Kit Cheung
NAACL-HLT4
2021 TIE: A Framework for Embedding-based Incremental Temporal Knowledge Graph Completion
abstract
Reasoning in a temporal knowledge graph (TKG) is a critical task for information retrieval and semantic search. It is particularly challenging when the TKG is updated frequently. The model has to adapt to changes in the TKG for efficient training and inference while preserving its performance on historical knowledge. Recent work approaches TKG completion (TKGC) by augmenting the encoder-decoder framework with a time-aware encoding function. However, naively fine-tuning the model at every time step using these methods does not address the problems of 1) catastrophic forgetting, 2) the model's inability to identify the change of facts (e.g., the change of the political affiliation and end of a marriage), and 3) the lack of training efficiency. To address these challenges, we present the Time-aware Incremental Embedding (TIE) framework, which combines TKG representation learning, experience replay, and temporal regularization. We introduce a set of metrics that characterizes the intransigence of the model and propose a constraint that associates the deleted facts with negative labels.
Jiapeng Wu, Yishi Xu, Yingxue Zhang 0001, Chen Ma 0001, Mark Coates, Jackie Chi Kit Cheung
SIGIR6
2020 An Analysis of Dataset Overlap on Winograd-Style Tasks
abstract
The Winograd Schema Challenge (WSC) and variants inspired by it have become important benchmarks for common-sense reasoning (CSR).Model performance on the WSC has quickly progressed from chance-level to near-human using neural language models trained on massive corpora.In this paper, we analyze the effects of varying degrees of overlap between these training corpora and the test instances in WSC-style tasks.We find that a large number of test instances overlap considerably with the corpora on which state-of-the-art models are (pre)trained, and that a significant drop in classification accuracy occurs when we evaluate models on instances with minimal overlap.Based on these results, we develop the KNOWREF-60K dataset, which consists of over 60k pronoun disambiguation problems scraped from web data.KNOWREF-60K is the largest corpus to date for WSC-style common-sense reasoning and exhibits a significantly lower proportion of overlaps with current pretraining corpora.
Ali Emami, Kaheer Suleman, Adam Trischler, Jackie Chi Kit Cheung
COLING4
2020 Learning Efficient Task-Specific Meta-Embeddings with Word Prisms
abstract
Word embeddings are trained to predict word cooccurrence statistics, which leads them to possess different lexical properties (syntactic, semantic, etc.) depending on the notion of context defined at training time.These properties manifest when querying the embedding space for the most similar vectors, and when used at the input layer of deep neural networks trained to solve downstream NLP problems.Meta-embeddings combine multiple sets of differently trained word embeddings, and have been shown to successfully improve intrinsic and extrinsic performance over equivalent models which use just one set of source embeddings.We introduce word prisms: a simple and efficient meta-embedding method that learns to combine source embeddings according to the task at hand.Word prisms learn orthogonal transformations to linearly combine the input source embeddings, which allows them to be very efficient at inference time.We evaluate word prisms in comparison to other meta-embedding methods on six extrinsic evaluations and observe that word prisms offer improvements in performance on all tasks. 1
Konstantinos C. Tsiolis, Kian Kenyon-Dean, Jackie Chi Kit Cheung
COLING4
2020 Factual Error Correction for Abstractive Summarization Models
abstract
Neural abstractive summarization systems have achieved promising progress, thanks to the availability of large-scale datasets and models pre-trained with self-supervised methods.However, ensuring the factual consistency of the generated summaries for abstractive summarization systems is a challenge.We propose a post-editing corrector module to address this issue by identifying and correcting factual errors in generated summaries.The neural corrector model is pre-trained on artificial examples that are created by applying a series of heuristic transformations on reference summaries.These transformations are inspired by an error analysis of state-of-the-art summarization model outputs.Experimental results show that our model is able to correct factual errors in summaries generated by other neural summarization models and outperforms previous models on factual consistency evaluation on the CNN/DailyMail dataset.We also find that transferring from artificial error correction to downstream settings is still very challenging 1 .Article: Jerusalem (CNN)The flame of remembrance burns in Jerusalem, and a song of memory haunts Valerie Braham as it never has before.(...) "Now I truly understand everyone who has lost a loved one," Braham said.Her husband, Philippe Braham, was one of 17 people killed in January's terror attacks in Paris.He was in a kosher supermarket when a gunman stormed in, killing four people, all of them Jewish.(...) Original: Valerie braham was one of 17 people killed in january's terror attacks in paris.(inconsistent) Corrected: Philippe braham was one of 17 people killed in january's terror attacks in paris.(consistent) Article: (...) Thursday's attack by al-Shabaab militants killed 147 people, including 142 students, three security officers and two university security personnel.The attack left 104 people injured, including 19 who are in critical condition, Nkaissery said.(...) Original: 147 people, including 142 students, are in critical condition.(inconsistent) Corrected: 19 people, including 142 students, are in critical condition.(inconsistent) Article: (CNN) Officer Michael Slager's five-year career with the North Charleston Police Department in South Carolina ended after he resorted to deadly force following a routine traffic stop.(...) His back is to Slager, who, from a few yards away, raises his gun and fires.Slager is now charged with murder.The FBI is involved in the investigation of the slaying of the father of four.(...) Original: Slager is now charged with murder.(consistent) Corrected: Michael Slager is now charged with murder.(consistent) Article: (CNN)The announcement this year of a new, original Dr. Seuss book sent a wave of nostalgic giddiness across Twitter, and months before publication, the number of pre-orders for "What Pet Should I Get?" continues to climb.(...) It features the spirited siblings from the beloved classic "One Fish Two Fish Red Fish Blue Fish" and is believed to have been written between 1958 and 1962.(...) Original: Seuss book sent a wave of nostalgic giddiness across twitter.(consistent) Corrected: "One Fish Two Fish Red Fish Blue Fish" book sent a wave of nostalgic giddiness across twitter.(inconsistent)
Meng Cao 0003, Yue Dong 0002, Jiapeng Wu, Jackie Chi Kit Cheung
EMNLP (1)4
2020 Multi-Fact Correction in Abstractive Text Summarization
abstract
Pre-trained neural abstractive summarization systems have dominated extractive strategies on news summarization performance, at least in terms of ROUGE.However, systemgenerated abstractive summaries often face the pitfall of factual inconsistency: generating incorrect facts with respect to the source text.To address this challenge, we propose Span-Fact, a suite of two factual correction models that leverages knowledge learned from question answering models to make corrections in system-generated summaries via span selection.Our models employ single or multimasking strategies to either iteratively or autoregressively replace entities in order to ensure semantic consistency w.r.t. the source text, while retaining the syntactic structure of summaries generated by abstractive summarization models.Experiments show that our models significantly boost the factual consistency of system-generated summaries without sacrificing summary quality in terms of both automatic metrics and human evaluation.* *Most of this work was done when the first author was an intern at Microsoft.CNNDM Source (CNN) About a quarter of a million Australian homes and businesses have no power after a "once in a decade" storm battered Sydney and nearby areas.About 4,500 people
Yue Dong 0002, Shuohang Wang, Zhe Gan, Yu Cheng 0001, Jackie Chi Kit Cheung, Jingjing Liu 0001
EMNLP (1)5
2020 TESA: A Task in Entity Semantic Aggregation for Abstractive Summarization
abstract
Human-written texts contain frequent generalizations and semantic aggregation of content. In a document, they may refer to a pair of named entities such as ‘London’ and ‘Paris’ with different expressions: “the major cities”, “the capital cities” and “two European cities”. Yet generation, especially, abstractive summarization systems have so far focused heavily on paraphrasing and simplifying the source content, to the exclusion of such semantic abstraction capabilities. In this paper, we present a new dataset and task aimed at the semantic aggregation of entities. TESA contains a dataset of 5.3K crowd-sourced entity aggregations of Person, Organization, and Location named entities. The aggregations are document-appropriate, meaning that they are produced by annotators to match the situational context of a given news article from the New York Times. We then build baseline models for generating aggregations given a tuple of entities and document context. We finetune on TESA an encoder-decoder language model and compare it with simpler classification methods based on linguistically informed features. Our quantitative and qualitative evaluations show reasonable performance in making a choice from a given list of expressions, but free-form expressions are understandably harder to generate and evaluate.
Clément Jumel, Annie Louis, Jackie Chi Kit Cheung
EMNLP (1)3
2020 Deconstructing word embedding algorithms
abstract
Word embeddings are reliable feature representations of words used to obtain high quality results for various NLP applications.Uncontextualized word embeddings are used in many NLP tasks today, especially in resourcelimited settings where high memory capacity and GPUs are not available.Given the historical success of word embeddings in NLP, we propose a retrospective on some of the most well-known word embedding algorithms.In this work, we deconstruct Word2vec, GloVe, and others, into a common form, unveiling some of the common conditions that seem to be required for making performant word embeddings.We believe that the theoretical findings in this paper can provide a basis for more informed development of future models.
Kian Kenyon-Dean, Edward Newell, Jackie Chi Kit Cheung
EMNLP (1)3
2020 TeMP: Temporal Message Passing for Temporal Knowledge Graph Completion
abstract
Inferring missing facts in temporal knowledge graphs (TKGs) is a fundamental and challenging task.Previous works have approached this problem by augmenting methods for static knowledge graphs to leverage time-dependent representations.However, these methods do not explicitly leverage multi-hop structural information and temporal facts from recent time steps to enhance their predictions.Additionally, prior work does not explicitly address the temporal sparsity and variability of entity distributions in TKGs.We propose the Temporal Message Passing (TeMP) framework to address these challenges by combining graph neural networks, temporal dynamics models, data imputation and frequency-based gating techniques.Experiments 1 on standard TKG tasks show that our approach provides substantial gains compared to the previous state of the art, achieving a 10.7% average relative improvement in Hits@10 across three standard benchmarks.Our analysis also reveals important sources of variability both within and across TKG datasets, and we introduce several simple but strong baselines that outperform the prior state of the art in certain settings.
Jiapeng Wu, Meng Cao 0003, Jackie Chi Kit Cheung, William L. Hamilton
EMNLP (1)3
2020 On Variational Learning of Controllable Representations for Text without Supervision
abstract
The variational autoencoder (VAE) can learn the manifold of natural images on certain datasets, as evidenced by meaningful interpolating or extrapolating in the continuous latent space. However, on discrete data such as text, it is unclear if unsupervised learning can discover similar latent space that allows controllable manipulation. In this work, we find that sequence VAEs trained on text fail to properly decode when the latent codes are manipulated, because the modified codes often land in holes or vacant regions in the aggregated posterior latent space, where the decoding network fails to generalize. Both as a validation of the explanation and as a fix to the problem, we propose to constrain the posterior mean to a learned probability simplex, and performs manipulation within this simplex. Our proposed method mitigates the latent vacancy problem and achieves the first success in unsupervised learning of controllable representations for text. Empirically, our method outperforms unsupervised baselines and strong supervised approaches on text style transfer, and is capable of performing more flexible fine-grained control over text generation than existing methods.
Jackie Chi Kit Cheung, Yanshuai Cao
ICML2
2020 Learning Lexical Subspaces in a Distributional Vector Space
abstract
In this paper, we propose LexSub, a novel approach towards unifying lexical and distributional semantics. We inject knowledge about lexical-semantic relations into distributional word embeddings by defining subspaces of the distributional vector space in which a lexical relation should hold. Our framework can handle symmetric attract and repel relations (e.g., synonymy and antonymy, respectively), as well as asymmetric relations (e.g., hypernymy and meronomy). In a suite of intrinsic benchmarks, we show that our model outperforms previous approaches on relatedness tasks and on hypernymy classification and detection, while being competitive on word similarity tasks. It also outperforms previous systems on extrinsic classification tasks that benefit from exploiting lexical relational cues. We perform a series of analyses to understand the behaviors of our model. 1 Code available at https://github.com/aishikchakraborty/LexSub .
Kushal Arora, Aishik Chakraborty, Jackie Chi Kit Cheung
Trans. Assoc. Comput. Linguistics3
2019 Contextualized Non-Local Neural Networks for Sequence Learning
abstract
Recently, a large number of neural mechanisms and models have been proposed for sequence learning, of which selfattention, as exemplified by the Transformer model, and graph neural networks (GNNs) have attracted much attention. In this paper, we propose an approach that combines and draws on the complementary strengths of these two methods. Specifically, we propose contextualized non-local neural networks (CN3), which can both dynamically construct a task-specific structure of a sentence and leverage rich local dependencies within a particular neighbourhood.Experimental results on ten NLP tasks in text classification, semantic matching, and sequence labelling show that our proposed model outperforms competitive baselines and discovers task-specific dependency structures, thus providing better interpretability to users.
Pengfei Liu 0003, Shuaichen Chang, Xuanjing Huang 0001, Jackie Chi Kit Cheung
AAAI5
2019 Learning Multi-Task Communication with Message Passing for Sequence Learning
abstract
We present two architectures for multi-task learning with neural sequence models. Our approach allows the relationships between different tasks to be learned dynamically, rather than using an ad-hoc pre-defined structure as in previous work. We adopt the idea from message-passing graph neural networks, and propose a general graph multi-task learning framework in which different tasks can communicate with each other in an effective and interpretable way. We conduct extensive experiments in text classification and sequence labelling to evaluate our approach on multi-task learning and transfer learning. The empirical results show that our models not only outperform competitive baselines, but also learn interpretable and transferable patterns across tasks.
Pengfei Liu 0003, Jie Fu 0001, Yue Dong 0002, Xipeng Qiu, Jackie Chi Kit Cheung
AAAI5
2019 Generating Character Descriptions for Automatic Summarization of Fiction
abstract
Summaries of fictional stories allow readers to quickly decide whether or not a story catches their interest. A major challenge in automatic summarization of fiction is the lack of standardized evaluation methodology or high-quality datasets for experimentation. In this work, we take a bottomup approach to this problem by assuming that story authors are uniquely qualified to inform such decisions. We collect a dataset of one million fiction stories with accompanying author-written summaries from Wattpad, an online story sharing platform. We identify commonly occurring summary components, of which a description of the main characters is the most frequent, and elicit descriptions of main characters directly from the authors for a sample of the stories. We propose two approaches to generate character descriptions, one based on ranking attributes found in the story text, the other based on classifying into a list of pre-defined attributes. We find that the classification-based approach performs the best in predicting character descriptions.
Jackie Chi Kit Cheung, Joel Oren
AAAI2
2019 EditNTS: An Neural Programmer-Interpreter Model for Sentence Simplification through Explicit Editing
abstract
We present the first sentence simplification model that learns explicit edit operations (ADD, DELETE, and KEEP) via a neural programmer-interpreter approach.Most current neural sentence simplification systems are variants of sequence-to-sequence models adopted from machine translation.These methods learn to simplify sentences as a byproduct of the fact that they are trained on complex-simple sentence pairs.By contrast, our neural programmer-interpreter is directly trained to predict explicit edit operations on targeted parts of the input sentence, resembling the way that humans might perform simplification and revision.Our model outperforms previous state-of-the-art neural sentence simplification models (without external knowledge) by large margins on three benchmark text simplification corpora in terms of SARI (+0.95 WikiLarge, +1.89 WikiSmall, +1.41 Newsela), and is judged by humans to produce overall better and simpler output sentences 1 .
Yue Dong 0002, Zichao Li 0003, Mehdi Rezagholizadeh, Jackie Chi Kit Cheung
ACL (1)4
2019 The KnowRef Coreference Corpus: Removing Gender and Number Cues for Difficult Pronominal Anaphora Resolution
abstract
We introduce a new benchmark for coreference resolution and NLI, KnowRef, that targets common-sense understanding and world knowledge. Previous coreference resolution tasks can largely be solved by exploiting the number and gender of the antecedents, or have been handcrafted and do not reflect the diversity of naturally occurring text. We present a corpus of over 8,000 annotated text passages with ambiguous pronominal anaphora. These instances are both challenging and realistic. We show that various coreference systems, whether rule-based, feature-rich, or neural, perform significantly worse on the task than humans, who display high inter-annotator agreement. To explain this performance gap, we show empirically that state-of-the art models often fail to capture context, instead relying on the gender or number of candidate antecedents to make a decision. We then use problem-specific insights to propose a data-augmentation trick called antecedent switching to alleviate this tendency in models. Finally, we show that antecedent switching yields promising results on other tasks as well: we use it to achieve state-of-the-art results on the GAP coreference task.
Ali Emami, Paul Trichelair, Adam Trischler, Kaheer Suleman, Hannes Schulz, Jackie Chi Kit Cheung
ACL (1)6
2019 A Cross-Domain Transferable Neural Coherence Model
abstract
Peng Xu, Hamidreza Saghir, Jin Sung Kang, Teng Long, Avishek Joey Bose, Yanshuai Cao, Jackie Chi Kit Cheung. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Hamidreza Saghir, Jin Sung Kang, Joey Bose, Yanshuai Cao, Jackie Chi Kit Cheung
ACL (1)7
2019 Referring Expression Generation Using Entity Profiles
abstract
Meng Cao, Jackie Chi Kit Cheung. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Meng Cao 0003, Jackie Chi Kit Cheung
EMNLP/IJCNLP (1)2
2019 Countering the Effects of Lead Bias in News Summarization via Multi-Stage Training and Auxiliary Losses
abstract
Matt Grenander, Yue Dong, Jackie Chi Kit Cheung, Annie Louis. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Matt Grenander, Yue Dong 0002, Jackie Chi Kit Cheung, Annie Louis
EMNLP/IJCNLP (1)3
2019 How Reasonable are Common-Sense Reasoning Tasks: A Case-Study on the Winograd Schema Challenge and SWAG
abstract
Paul Trichelair, Ali Emami, Adam Trischler, Kaheer Suleman, Jackie Chi Kit Cheung. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Paul Trichelair, Ali Emami, Adam Trischler, Kaheer Suleman, Jackie Chi Kit Cheung
EMNLP/IJCNLP (1)5
2018 Let's do it "again": A First Computational Approach to Detecting Adverbial Presupposition Triggers
abstract
We introduce the task of predicting adverbial presupposition triggers such as also and again.Solving such a task requires detecting recurring or similar events in the discourse context, and has applications in natural language generation tasks such as summarization and dialogue systems.We create two new datasets for the task, derived from the Penn Treebank and the Annotated English Gigaword corpora, as well as a novel attention mechanism tailored to this task.Our attention mechanism augments a baseline recurrent neural network without the need for additional trainable parameters, minimizing the added computational cost of our mechanism.We demonstrate that our model statistically outperforms a number of baselines, including an LSTM-based language model.
Andre Cianflone, Yulan Feng, Jad Kabbara, Jackie Chi Kit Cheung
ACL (1)4
2018 BanditSum: Extractive Summarization as a Contextual Bandit
abstract
In this work, we propose a novel method for training neural networks to perform singledocument extractive summarization without heuristically-generated extractive labels.We call our approach BANDITSUM as it treats extractive summarization as a contextual bandit (CB) problem, where the model receives a document to summarize (the context), and chooses a sequence of sentences to include in the summary (the action).A policy gradient reinforcement learning algorithm is used to train the model to select sequences of sentences that maximize ROUGE score.We perform a series of experiments demonstrating that BANDITSUM is able to achieve ROUGE scores that are better than or comparable to the state-of-the-art for extractive summarization, and converges using significantly fewer update steps than competing approaches.In addition, we show empirically that BANDIT-SUM performs significantly better than competing approaches when good summary sentences appear late in the source document.
Yue Dong 0002, Yikang Shen, Eric Crawford, Herke van Hoof, Jackie Chi Kit Cheung
EMNLP5
2018 A Knowledge Hunting Framework for Common Sense Reasoning
abstract
We introduce an automatic system that achieves state-of-the-art results on the Winograd Schema Challenge (WSC), a common sense reasoning task that requires diverse, complex forms of inference and knowledge.Our method uses a knowledge hunting module to gather text from the web, which serves as evidence for candidate problem resolutions.Given an input problem, our system generates relevant queries to send to a search engine, then extracts and classifies knowledge from the returned results and weighs them to make a resolution.Our approach improves F1 performance on the full WSC by 0.21 over the previous best and represents the first system to exceed 0.5 F1.We further demonstrate that the approach is competitive on the Choice of Plausible Alternatives (COPA) task, which suggests that it is generally applicable.
Ali Emami, Noelia De La Cruz, Adam Trischler, Kaheer Suleman, Jackie Chi Kit Cheung
EMNLP5
2018 A Hierarchical Neural Attention-based Text Classifier
abstract
Deep neural networks have been displaying superior performance over traditional supervised classifiers in text classification.They learn to extract useful features automatically when sufficient amount of data is presented.However, along with the growth in the number of documents comes the increase in the number of categories, which often results in poor performance of the multiclass classifiers.In this work, we use external knowledge in the form of topic category taxonomies to aide the classification by introducing a deep hierarchical neural attention-based classifier.Our model performs better than or comparable to state-of-the-art hierarchical models at significantly lower computational cost while maintaining high interpretability.
Koustuv Sinha, Yue Dong 0002, Jackie Chi Kit Cheung, Derek Ruths
EMNLP3
2018 Constructing a Lexicon of Relational Nouns
Edward Newell, Jackie Chi Kit Cheung
LREC2
2017 World Knowledge for Reading Comprehension: Rare Entity Prediction with Hierarchical LSTMs Using External Descriptions
abstract
Humans interpret texts with respect to some background information, or world knowledge, and we would like to develop automatic reading comprehension systems that can do the same.In this paper, we introduce a task and several models to drive progress towards this goal.In particular, we propose the task of rare entity prediction: given a web document with several entities removed, models are tasked with predicting the correct missing entities conditioned on the document context and the lexical resources.This task is challenging due to the diversity of language styles and the extremely large number of rare entities.We propose two recurrent neural network architectures which make use of external knowledge in the form of entity descriptions.Our experiments show that our hierarchical LSTM model performs significantly better at the rare entity prediction task than those that do not make use of external resources.
Emmanuel Bengio, Ryan Lowe, Jackie Chi Kit Cheung, Doina Precup
EMNLP4
2017 Nifty Assignments
abstract
I suspect that students learn more from our programming assignments than from our much sweated-over lectures, with their slide transitions, clip art, and joke attempts. A great assignment is deliberate about where the student hours go, concentrating the student's attention on material that is interesting and useful. The best assignments solve a problem that is topical and entertaining, providing motivation for the whole stack of work. Unfortunately, creating great programming assignments is both time consuming and error prone. The Nifty Assignments special session is all about promoting and sharing the ideas and ready-to-use materials of successful assignments.
Nick Parlante, Julie Zelenski, Dave Feinberg, Kunal Mishra, Josh Hug, Kevin Wayne, Michael Guerzhoy, Jackie Chi Kit Cheung, François Pitt
SIGCSE8
2017 Predicting Success in Goal-Driven Human-Human Dialogues
abstract
In goal-driven dialogue systems, success is often defined based on a structured definition of the goal.This requires that the dialogue system be constrained to handle a specific class of goals and that there be a mechanism to measure success with respect to that goal.However, in many human-human dialogues the diversity of goals makes it infeasible to define success in such a way.To address this scenario, we consider the task of automatically predicting success in goal-driven human-human dialogues using only the information communicated between participants in the form of text.We build a dataset from stackoverflow.com which consists of exchanges between two users in the technical domain where groundtruth success labels are available.We then propose a turn-based hierarchical neural network model that can be used to predict success without requiring a structured goal definition.We show this model outperforms rule-based heuristics and other baselines as it is able to detect patterns over the course of a dialogue and capture notions such as gratitude.
Michael Noseworthy, Jackie Chi Kit Cheung, Joelle Pineau
SIGDIAL Conference2
2016 Predicting sentential semantic compatibility for aggregation in text-to-text generation
abstract
We examine the task of aggregation in the context of text-to-text generation. We introduce a new aggregation task which frames the process as grouping input sentence fragments into clusters that are to be expressed as a single output sentence. We extract datasets for this task from a corpus using an automatic extraction process. Based on the results of a user study, we develop two gold-standard clusterings and corresponding evaluation methods for each dataset. We present a hierarchical clustering framework for predicting aggregation decisions on this task, which outperforms several baselines and can serve as a reference in future work.
Victor Chenal, Jackie Chi Kit Cheung
COLING2
2016 Capturing Pragmatic Knowledge in Article Usage Prediction using LSTMs
abstract
We examine the potential of recurrent neural networks for handling pragmatic inferences involving complex contextual cues for the task of article usage prediction. We train and compare several variants of Long Short-Term Memory (LSTM) networks with an attention mechanism. Our model outperforms a previous state-of-the-art system, achieving up to 96.63% accuracy on the WSJ/PTB corpus. In addition, we perform a series of analyses to understand the impact of various model choices. We find that the gain in performance can be attributed to the ability of LSTMs to pick up on contextual cues, both local and further away in distance, and that the model is able to solve cases involving reasoning about coreference and synonymy. We also show how the attention mechanism contributes to the interpretability of the model’s effectiveness.
Jad Kabbara, Yulan Feng, Jackie Chi Kit Cheung
COLING3
2016 Verb Phrase Ellipsis Resolution Using Discriminative and Margin-Infused Algorithms
Kian Kenyon-Dean, Jackie Chi Kit Cheung, Doina Precup
EMNLP2
2015 Indicative Tweet Generation: An Extractive Summarization Problem?
abstract
Social media such as Twitter have become an important method of communication, with potential opportunities for NLG to facilitate the generation of social media content.We focus on the generation of indicative tweets that contain a link to an external web page.While it is natural and tempting to view the linked web page as the source text from which the tweet is generated in an extractive summarization setting, it is unclear to what extent actual indicative tweets behave like extractive summaries.We collect a corpus of indicative tweets with their associated articles and investigate to what extent they can be derived from the articles using extractive methods.We also consider the impact of the formality and genre of the article.Our results demonstrate the limits of viewing indicative tweet generation as extractive summarization, and point to the need for the development of a methodology for tweet generation that is sensitive to genre-specific issues.
Priya Sidhaye, Jackie Chi Kit Cheung
EMNLP2
2014 Unsupervised Sentence Enhancement for Automatic Summarization
abstract
We present sentence enhancement as a novel technique for text-to-text genera-tion in abstractive summarization. Com-pared to extraction or previous approaches to sentence fusion, sentence enhancement increases the range of possible summary sentences by allowing the combination of dependency subtrees from any sentence from the source text. Our experiments in-dicate that our approach yields summary sentences that are competitive with a sen-tence fusion baseline in terms of con-tent quality, but better in terms of gram-maticality, and that the benefit of sen-tence enhancement relies crucially on an event coreference resolution algorithm us-ing distributional semantics. We also consider how text-to-text generation ap-proaches to summarization can be ex-tended beyond the source text by exam-ining how human summary writers incor-porate source-text-external elements into their summary sentences. 1
Jackie Chi Kit Cheung, Gerald Penn
EMNLP1
2013 Probabilistic Domain Modelling With Contextualized Distributional Semantic Vectors
Jackie Chi Kit Cheung, Gerald Penn
ACL (1)1
2013 Towards Robust Abstractive Multi-Document Summarization: A Caseframe Analysis of Centrality and Domain
Jackie Chi Kit Cheung, Gerald Penn
ACL (1)1
2013 Probabilistic Frame Induction
Jackie Chi Kit Cheung, Hoifung Poon, Lucy Vanderwende
HLT-NAACL1
2013 Multi-Document Summarization of Evaluative Text
abstract
In many decision‐making scenarios, people can benefit from knowing what other people's opinions are. As more and more evaluative documents are posted on the Web, summarizing these useful resources becomes a critical task for many organizations and individuals. This paper presents a framework for summarizing a corpus of evaluative documents about a single entity by a natural language summary. We propose two summarizers: an extractive summarizer and an abstractive one. As an additional contribution, we show how our abstractive summarizer can be modified to generate summaries tailored to a model of the user preferences that is solidly grounded in decision theory and can be effectively elicited from users. We have tested our framework in three user studies. In the first one, we compared the two summarizers. They performed equally well relative to each other quantitatively, while significantly outperforming a baseline standard approach to multidocument summarization. Trends in the results as well as qualitative comments from participants suggest that the summarizers have different strengths and weaknesses. After this initial user study, we realized that the diversity of opinions expressed in the corpus (i.e., its controversiality) might play a critical role in comparing abstraction versus extraction. To clearly pinpoint the role of controversiality, we ran a second user study in which we controlled for the degree of controversiality of the corpora that were summarized for the participants. The outcome of this study indicates that for evaluative text abstraction tends to be more effective than extraction, particularly when the corpus is controversial. In the third user study we assessed the effectiveness of our user tailoring strategy. The results of this experiment confirm that user tailored summaries are more informative than untailored ones.
Giuseppe Carenini, Jackie Chi Kit Cheung, Adam Pauls
Comput. Intell.2
2012 Evaluating Distributional Models of Semantics for Syntactically Invariant Inference
Jackie Chi Kit Cheung, Gerald Penn
EACL1
2012 Unsupervised Detection of Downward-Entailing Operators By Maximizing Classification Certainty
Jackie Chi Kit Cheung, Gerald Penn
EACL1
2012 Sequence clustering and labeling for unsupervised query intent discovery
abstract
One popular form of semantic search observed in several modern search engines is to recognize query patterns that trigger instant answers or domain-specific search, producing semantically enriched search results. This often requires understanding the query intent in addition to the meaning of the query terms in order to access structured data sources. A major challenge in intent understanding is to construct a domain-dependent schema and to annotate search queries based on such a schema, a process that to date has required much manual annotation effort. We present an unsupervised method for clustering queries with similar intent and for producing a pattern consisting of a sequence of semantic concepts and/or lexical items for each intent. Furthermore, we leverage the discovered intent patterns to automatically annotate a large number of queries beyond those used in clustering. We evaluated our method on 10 selected domains, discovering over 1400 intent patterns and automatically annotating 125K (and potentially many more) queries. We found that over 90% of patterns and 80% of instance annotations tested are judged to be correct by a majority of annotators.
Jackie Chi Kit Cheung
WSDM1
2010 Entity-Based Local Coherence Modelling Using Topological Fields
Jackie Chi Kit Cheung, Gerald Penn
ACL1
2010 Utilizing Extra-Sentential Context for Parsing
Jackie Chi Kit Cheung, Gerald Penn
EMNLP1
2009 Topological Field Parsing of German
Jackie Chi Kit Cheung, Gerald Penn
ACL/IJCNLP1
2008 Extractive vs. NLG-based Abstractive Summarization of Evaluative Text: The Effect of Corpus Controversiality
Giuseppe Carenini, Jackie Chi Kit Cheung
INLG2