VLDB 2026 Research / reviewers in the wild / expert
Claire Cardie
dblp:c/ClaireCardie
· DBLP profile ↗
114ranked-venue papers
16as first author
20since 2021 · last 2026
0000-0002-2061-6094ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 106 · 16 first-author · 19 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HAPO: Training Language Models to Reason Concisely via History-Aware Policy OptimizationabstractWhile scaling the length of responses at test-time has been shown to markedly improve the reasoning abilities and performance of large language models (LLMs), it often results in verbose outputs and increases inference cost. Prior approaches for efficient test-time scaling, typically using universal budget constraints or query-level length optimization, do not leverage historical information from previous encounters with the same problem during training. We hypothesize that this limits their ability to progressively make solutions more concise over time. To address this, we present History-Aware Policy Optimization (HAPO), which keeps track of a history state (e.g., the minimum length over previously generated correct responses) for each problem. HAPO employs a novel length reward function based on this history state to incentivize the discovery of correct solutions that are more concise than those previously found. Crucially, this reward structure avoids overly penalizing shorter incorrect responses with the goal of facilitating exploration towards more efficient solutions. By combining this length reward with a correctness reward, HAPO jointly optimizes for correctness and efficiency. We use HAPO to train DeepSeek-R1-Distill-Qwen-1.5B, DeepScaleR-1.5B-Preview, and Qwen-2.5-1.5B-Instruct, and evaluate HAPO on several math benchmarks that span various difficulty levels. Experiment results demonstrate that HAPO effectively induces LLMs’ concise reasoning abilities, producing length reductions of 33-59% with accuracy drops of only 2-5%. Claire Cardie |
AAAI | 3 |
| 2026 | MedSimAI: Simulation and Formative Feedback Generation to Enhance Deliberate Practice in Medical EducationabstractMedical education faces challenges in providing scalable, consistent clinical skills training. Simulation with standardized patients (SPs) develops communication and diagnostic skills, but remains resource-intensive and variable in feedback quality. Existing AI-based tools show promise yet often lack comprehensive assessment frameworks, evidence of clinical impact, and integration of self-regulated learning (SRL) principles. Through a multi-phase co-design process with medical education experts, we developed MedSimAI, an AI-powered simulation platform that enables deliberate practice through interactive patient encounters with immediate, structured feedback. Leveraging large language models, MedSimAI generates realistic clinical interactions and provides automated assessments aligned with validated evaluation frameworks. In a multi-institutional deployment (410 students; 1,024 encounters across three medical schools), 59.5% engaged in repeated practice. At one site, mean Objective Structured Clinical Examination (OSCE) history-taking scores rose from 82.8 to 88.8 (p < 0.001, d = 0.75), while a second site’s pilot showed no significant change. Automated scoring achieved 87% accuracy in identifying proficiency thresholds on the Master Interview Rating Scale (MIRS). Mixed-effects analyses revealed institution and case effects. Thematic analysis of 840 learner reflections highlighted challenges in missed items, organization, review-of-systems, and empathy. These findings position MedSimAI as a scalable formative platform for history-taking and communication, motivating staged curriculum integration and realism enhancements for advanced learners. Yann Hicke, Jadon Geathers, Kellen Vu, Justin Sewell, Claire Cardie, Jaideep Talwalkar, Dennis L. Shung, Anyanate Gwendolyne Jack, Susannah Cornes, MacKenzi Preston, René F. Kizilcec |
LAK | 5 |
| 2025 | Commit0: Library Generation from ScratchabstractWith the goal of benchmarking generative systems beyond expert software development ability, we introduce Commit0, a benchmark that challenges AI agents to write libraries from scratch. Agents are provided with a specification document outlining the library’s API as well as a suite of interactive unit tests, with the goal of producing an implementation of this API accordingly. The implementation is validated through running these unit tests. As a benchmark, Commit0 is designed to move beyond static one-shot code generation towards agents that must process long-form natural language specifications, adapt to multi-stage feedback, and generate code with complex dependencies. Commit0 also offers an interactive environment where models receive static analysis and execution feedback on the code they generate. Our experiments demonstrate that while current agents can pass some unit tests, none can yet fully reproduce full libraries. Results also show that interactive feedback is quite useful for models to generate code that passes more unit tests, validating the benchmarks that facilitate its use. We publicly release the benchmark, the interactive environment, and the leaderboard. Celine Lee, Justin T. Chiu, Claire Cardie, Matthias Gallé, Alexander M. Rush |
ICLR | 5 |
| 2025 | Are Triggers Needed for Document-Level Event Extraction?abstractAbstract Most existing work on event extraction has focused on sentence-level texts and presumes the identification of a trigger-span—a word or phrase in the input that evokes the occurrence of an event of interest. Event arguments are then extracted with respect to the trigger. Indeed, triggers are treated as integral to, and trigger detection as an essential component of, event extraction. In this paper, we provide the first investigation of the role of triggers for the more difficult and much less studied task of document-level event extraction. We analyze their usefulness in multiple end-to-end and pipelined transformer-based event extraction models for three document-level event extraction datasets, measuring performance using triggers of varying quality (human-annotated, LLM-generated, keyword-based, and random). We find that whether or not systems benefit from explicitly extracting triggers depends both on dataset characteristics (i.e., the typical number of events per document) and task-specific information available during extraction (i.e., natural language event schemas). Perhaps surprisingly, we also observe that the mere existence of triggers in the input, even random ones, is important for prompt-based in-context learning approaches to the task. Shaden Shaar, Wayne Chen, Maitreyi Chatterjee, Barry Wang, Claire Cardie |
Trans. Assoc. Comput. Linguistics | 6 |
| 2024 | I Could've Asked That: Reformulating Unanswerable QuestionsabstractWhen seeking information from unfamiliar documents, users frequently pose questions that cannot be answered by the documents.While existing large language models (LLMs) identify these unanswerable questions, they do not assist users in reformulating their questions, thereby reducing their overall utility.We curate COULDASK, an evaluation benchmark composed of existing and new datasets for document-grounded question answering, specifically designed to study reformulating unanswerable questions.We evaluate stateof-the-art open-source and proprietary LLMs on COULDASK.The results demonstrate the limited capabilities of these models in reformulating questions.Specifically, GPT-4 and Llama2-7B successfully reformulate questions only 26% and 12% of the time, respectively.Error analysis shows that 62% of the unsuccessful reformulations stem from the models merely rephrasing the questions or even generating identical questions.We publicly release the benchmark 1 and the code to reproduce the experiments 2 . Claire Cardie, Alexander M. Rush |
EMNLP | 3 |
| 2024 | WildChat: 1M ChatGPT Interaction Logs in the WildabstractChatbots such as GPT-4 and ChatGPT are now serving millions of users. Despite their widespread use, there remains a lack of public datasets showcasing how these tools are used by a population of users in practice. To bridge this gap, we offered free access to ChatGPT for online users in exchange for their affirmative, consensual opt-in to anonymously collect their chat transcripts and request headers. From this, we compiled WildChat, a corpus of 1 million user-ChatGPT conversations, which consists of over 2.5 million interaction turns. We compare WildChat with other popular user-chatbot interaction datasets, and find that our dataset offers the most diverse user prompts, contains the largest number of languages, and presents the richest variety of potentially toxic use-cases for researchers to study. In addition to timestamped chat transcripts, we enrich the dataset with demographic data, including state, country, and hashed IP addresses, alongside request headers. This augmentation allows for more detailed analysis of user behaviors across different geographical regions and temporal dimensions. Finally, because it captures a broad range of use cases, we demonstrate the dataset’s potential utility in fine-tuning instruction-following models. WildChat is released at https://wildchat.allen.ai under AI2 ImpACT Licenses. Xiang Ren 0001, Jack Hessel, Claire Cardie, Yejin Choi 0001, Yuntian Deng |
ICLR | 4 |
| 2023 | Abductive Commonsense Reasoning Exploiting Mutually Exclusive ExplanationsabstractAbductive reasoning aims to find plausible explanations for an event.This style of reasoning is critical for commonsense tasks where there are often multiple plausible explanations.Existing approaches for abductive reasoning in natural language processing (NLP) often rely on manually generated annotations for supervision; however, such annotations can be subjective and biased.Instead of using direct supervision, this work proposes an approach for abductive commonsense reasoning that exploits the fact that only a subset of explanations is correct for a given context.The method uses posterior regularization to enforce a mutual exclusion constraint, encouraging the model to learn the distinction between fluent explanations and plausible ones.We evaluate our approach on a diverse set of abductive reasoning datasets; experimental results show that our approach outperforms or is comparable to directly applying pretrained language models in a zeroshot manner and other knowledge-augmented zero-shot methods. Justin T. Chiu, Claire Cardie, Alexander M. Rush |
ACL (1) | 3 |
| 2023 | End-to-end Case-Based Reasoning for Commonsense Knowledge Base CompletionabstractPretrained language models have been shown to store knowledge in their parameters and have achieved reasonable performance in commonsense knowledge base completion (CKBC) tasks.However, CKBC is knowledge-intensive and it is reported that pretrained language models' performance in knowledge-intensive tasks are limited because of their incapability of accessing and manipulating knowledge.As a result, we hypothesize that providing retrieved passages that contain relevant knowledge as additional input to the CKBC task will improve performance.In particular, we draw insights from Case-Based Reasoning (CBR) -which aims to solve a new problem by reasoning with retrieved relevant cases, and investigate the direct application of it to CKBC.On two benchmark datasets, we demonstrate through automatic and human evaluations that our End-to-end Case-Based Reasoning Framework (ECBRF) generates more valid knowledge than the state-of-the-art COMET model for CKBC in both the fully supervised and few-shot settings.From the perspective of CBR, our framework addresses a fundamental question on whether CBR methodology can be utilized to improve deep learning models. Zonglin Yang 0001, Xinya Du, Erik Cambria, Claire Cardie |
EACL | 4 |
| 2023 | Hop, Union, Generate: Explainable Multi-hop Reasoning without Rationale SupervisionabstractExplainable multi-hop question answering (QA) not only predicts answers but also identifies rationales, i. e. subsets of input sentences used to derive the answers.This problem has been extensively studied under the supervised setting, where both answer and rationale annotations are given.Because rationale annotations are expensive to collect and not always available, recent efforts have been devoted to developing methods that do not rely on supervision for rationales.However, such methods have limited capacities in modeling interactions between sentences, let alone reasoning across multiple documents.This work proposes a principled, probabilistic approach for training explainable multi-hop QA systems without rationale supervision.Our approach performs multi-hop reasoning by explicitly modeling rationales as sets, enabling the model to capture interactions between documents and sentences within a document.Experimental results show that our approach is more accurate at selecting rationales than the previous methods, while maintaining similar accuracy in predicting answers. Justin T. Chiu, Claire Cardie, Alexander M. Rush |
EMNLP | 3 |
| 2022 | Improving Machine Reading Comprehension with Contextualized Commonsense KnowledgeabstractTo perform well on a machine reading comprehension (MRC) task, machine readers usually require commonsense knowledge that is not explicitly mentioned in the given documents.This paper aims to extract a new kind of structured knowledge from scripts and use it to improve MRC.We focus on scripts as they contain rich verbal and nonverbal messages, and two relevant messages originally conveyed by different modalities during a short time period may serve as arguments of a piece of commonsense knowledge as they function together in daily communications.To save human efforts to name relations, we propose to represent relations implicitly by situating such an argument pair in a context and call it contextualized knowledge.To use the extracted knowledge to improve MRC, we compare several fine-tuning strategies to use the weakly-labeled MRC data constructed based on contextualized knowledge and further design a teacher-student paradigm with multiple teachers to facilitate the transfer of knowledge in weakly-labeled MRC data.Experimental results show that our paradigm outperforms other methods that use weaklylabeled data and improves a state-of-the-art baseline by 4.3% in accuracy on a Chinese multiple-choice MRC dataset C 3 , wherein most of the questions require unstated prior knowledge.We also seek to transfer the knowledge to other tasks by simply adapting the resulting student reader, yielding a 2.9% improvement in F1 on a relation extraction dataset DialogRE, demonstrating the potential usefulness of the knowledge for non-MRC tasks that require document comprehension.Interior.Runaway office.Day. Kai Sun 0006, Dian Yu 0001, Jianshu Chen, Dong Yu 0001, Claire Cardie |
ACL (1) | 5 |
| 2022 | Automatic Error Analysis for Document-level Information ExtractionabstractAliva Das, Xinya Du, Barry Wang, Kejian Shi, Jiayuan Gu, Thomas Porter, Claire Cardie. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Aliva Das, Xinya Du, Barry Wang, Kejian Shi, Jiayuan Gu, Thomas Porter, Claire Cardie |
ACL (1) | 7 |
| 2022 | Faithful or Extractive? On Mitigating the Faithfulness-Abstractiveness Trade-off in Abstractive SummarizationabstractDespite recent progress in abstractive summarization, systems still suffer from faithfulness errors.While prior work has proposed models that improve faithfulness, it is unclear whether the improvement comes from an increased level of extractiveness of the model outputs as one naive way to improve faithfulness is to make summarization models more extractive.In this work, we present a framework for evaluating the effective faithfulness of summarization systems, by generating a faithfulnessabstractiveness trade-off curve that serves as a control at different operating points on the abstractiveness spectrum.We then show that the baseline system as well as recently proposed methods for improving faithfulness, fail to consistently improve over the control at the same level of abstractiveness.Finally, we learn a selector to identify the most faithful and abstractive summary for a given document, and show that this system can attain higher faithfulness scores in human evaluations while being more abstractive than the baseline system on two datasets.Moreover, we show that our system is able to achieve a better faithfulnessabstractiveness trade-off than the control at the same level of abstractiveness. Faisal Ladhak, Esin Durmus, He He 0001, Claire Cardie, Kathy McKeown |
ACL (1) | 4 |
| 2022 | Visual Prompt Tuning
Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge J. Belongie, Bharath Hariharan, Ser-Nam Lim |
ECCV (33) | 4 |
| 2022 | BeSt: The Belief and Sentiment CorpusabstractWe present the BeSt corpus, which records cognitive state: who believes what (i.e., factuality), and who has what sentiment towards what. This corpus is inspired by similar source-and-target corpora, specifically MPQA and FactBank. The corpus comprises two genres, newswire and discussion forums, in three languages, Chinese (Mandarin), English, and Spanish. The corpus is distributed through the LDC. Jennifer Tracey, Owen Rambow, Claire Cardie, Adam Dalton 0001, Hoa Trang Dang, Mona T. Diab, Bonnie J. Dorr, Louise Guthrie, Magdalena Markowska, Smaranda Muresan, Vinodkumar Prabhakaran, Samira Shaikh, Tomek Strzalkowski |
LREC | 3 |
| 2022 | Compositional Task-Oriented Parsing as Abstractive Question AnsweringabstractWenting Zhao, Konstantine Arkoudas, Weiqi Sun, Claire Cardie. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Konstantine Arkoudas, Claire Cardie |
NAACL-HLT | 4 |
| 2021 | Intentonomy: A Dataset and Study Towards Human Intent UnderstandingabstractAn image is worth a thousand words, conveying information that goes beyond the mere visual content therein. In this paper, we study the intent behind social media images with an aim to analyze how visual information can facilitate recognition of human intent. Towards this goal, we introduce an intent dataset, Intentonomy, comprising 14K images covering a wide range of everyday scenes. These images are manually annotated with 28 intent categories derived from a social psychology taxonomy. We then systematically study whether, and to what extent, commonly used visual information, i.e., object and context, contribute to human motive understanding. Based on our findings, we conduct further study to quantify the effect of attending to object and context classes as well as textual information in the form of hashtags when training an intent classifier. Our results quantitatively and qualitatively shed light on how visual and textual information can produce observable effects when predicting intent.1 Menglin Jia, Zuxuan Wu, Austin Reiter, Claire Cardie, Serge J. Belongie, Ser-Nam Lim |
CVPR | 4 |
| 2021 | GRIT: Generative Role-filler Transformers for Document-level Event Entity ExtractionabstractWe revisit the classic problem of documentlevel role-filler entity extraction (REE) for template filling.We argue that sentence-level approaches are ill-suited to the task and introduce a generative transformer-based encoderdecoder framework (GRIT) that is designed to model context at the document level: it can make extraction decisions across sentence boundaries; is implicitly aware of noun phrase coreference structure, and has the capacity to respect cross-role dependencies in the template structure.We evaluate our approach on the MUC-4 dataset, and show that our model performs substantially better than prior work.We also show that our modeling choices contribute to model performance, e.g., by implicitly capturing linguistic knowledge such as recognizing coreferent entity mentions. Xinya Du, Alexander M. Rush, Claire Cardie |
EACL | 3 |
| 2021 | Exploring Visual Engagement Signals for Representation LearningabstractVisual engagement in social media platforms comprises interactions with photo posts including comments, shares, and likes. In this paper, we leverage such Visual Engagement clues as supervisory signals for representation learning. However, learning from engagement signals is non-trivial as it is not clear how to bridge the gap between low-level visual information and high-level social interactions. We present VisE, a weakly supervised learning approach, which maps social images to pseudo labels derived by clustered engagement signals. We then study how models trained in this way benefit subjective downstream computer vision tasks such as emotion recognition or political bias detection. Through extensive studies, we empirically demonstrate the effectiveness of VisE across a diverse set of classification tasks beyond the scope of conventional recognition1. Menglin Jia, Zuxuan Wu, Austin Reiter, Claire Cardie, Serge J. Belongie, Ser-Nam Lim |
ICCV | 4 |
| 2021 | Template Filling with Generative TransformersabstractTemplate filling is generally tackled by a pipeline of two separate supervised systemsone for role-filler extraction and another for template/event recognition.Since pipelines consider events in isolation, they can suffer from error propagation.We introduce a framework based on end-to-end generative transformers for this task (i.e., GTT).It naturally models the dependence between entities both within a single event and across the multiple events described in a document.Experiments demonstrate that this framework substantially outperforms pipeline-based approaches, and other neural end-to-end baselines that do not model between-event dependencies.We further show that our framework specifically improves performance on documents containing multiple events. Xinya Du, Alexander M. Rush, Claire Cardie |
NAACL-HLT | 3 |
| 2021 | Adding Chit-Chat to Enhance Task-Oriented DialoguesabstractKai Sun, Seungwhan Moon, Paul Crook, Stephen Roller, Becka Silvert, Bing Liu, Zhiguang Wang, Honglei Liu, Eunjoon Cho, Claire Cardie. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Kai Sun 0006, Seungwhan Moon, Paul A. Crook, Stephen Roller, Becka Silvert, Zhiguang Wang, Eunjoon Cho, Claire Cardie |
NAACL-HLT | 10 |
| 2020 | Interpreting Pretrained Contextualized Representations via Reductions to Static EmbeddingsabstractContextualized representations (e.g.ELMo, BERT) have become the default pretrained representations for downstream NLP applications.In some settings, this transition has rendered their static embedding predecessors (e.g.Word2Vec, GloVe) obsolete.As a side-effect, we observe that older interpretability methods for static embeddings -while more mature than those available for their dynamic counterparts -are underutilized in studying newer contextualized representations.Consequently, we introduce simple and fully general methods for converting from contextualized representations to static lookup-table embeddings which we apply to 5 popular pretrained models and 9 sets of pretrained weights.Our analysis of the resulting static embeddings notably reveals that pooling over many contexts significantly improves representational quality under intrinsic evaluation.Complementary to analyzing representational quality, we consider social biases encoded in pretrained representations with respect to gender, race/ethnicity, and religion and find that bias is encoded disparately across pretrained models and internal layers even for models that share the same training data.Concerningly, we find dramatic inconsistencies between social bias estimators for word embeddings. Rishi Bommasani, Kelly Davis, Claire Cardie |
ACL | 3 |
| 2020 | Document-Level Event Role Filler Extraction using Multi-Granularity Contextualized EncodingabstractFew works in the literature of event extraction have gone beyond individual sentences to make extraction decisions.This is problematic when the information needed to recognize an event argument is spread across multiple sentences.We argue that document-level event extraction is a difficult task since it requires a view of a larger context to determine which spans of text correspond to event role fillers.We first investigate how end-toend neural sequence models (with pre-trained language model representations) perform on document-level role filler extraction, as well as how the length of context captured affects the models' performance.To dynamically aggregate information captured by neural representations learned at different levels of granularity (e.g., the sentence-and paragraph-level), we propose a novel multi-granularity reader.We evaluate our models on the MUC-4 event extraction dataset, and show that our best system performs substantially better than prior work.We also report findings on the relationship between context length and neural model performance on the task. Xinya Du, Claire Cardie |
ACL | 2 |
| 2020 | Dialogue-Based Relation ExtractionabstractWe present the first human-annotated dialoguebased relation extraction (RE) dataset Dialo-gRE, aiming to support the prediction of relation(s) between two arguments that appear in a dialogue.We further offer DialogRE as a platform for studying cross-sentence RE as most facts span multiple sentences.We argue that speaker-related information plays a critical role in the proposed task, based on an analysis of similarities and differences between dialogue-based and traditional RE tasks.Considering the timeliness of communication in a dialogue, we design a new metric to evaluate the performance of RE methods in a conversational setting and investigate the performance of several representative RE methods on DialogRE.Experimental results demonstrate that a speaker-aware extension on the best-performing model leads to gains in both the standard and conversational evaluation settings.DialogRE is available at https:// dataset.org/dialogre/. Dian Yu 0001, Kai Sun 0006, Claire Cardie, Dong Yu 0001 |
ACL | 3 |
| 2020 | Fashionpedia: Ontology, Segmentation, and an Attribute Localization Dataset
Menglin Jia, Mengyun Shi, Mikhail Sirotenko, Yin Cui, Claire Cardie, Bharath Hariharan, Hartwig Adam, Serge J. Belongie |
ECCV (1) | 5 |
| 2020 | Intrinsic Evaluation of Summarization DatasetsabstractHigh quality data forms the bedrock for building meaningful statistical models in NLP.Consequently, data quality must be evaluated either during dataset construction or post hoc.Almost all popular summarization datasets are drawn from natural sources and do not come with inherent quality assurance guarantees.In spite of this, data quality has gone largely unquestioned for many recent summarization datasets.We perform the first large-scale evaluation of summarization datasets by introducing 5 intrinsic metrics and applying them to 10 popular datasets.We find that data usage in recent summarization research is sometimes inconsistent with the underlying properties of the datasets employed.Further, we discover that our metrics can serve the additional purpose of being inexpensive heuristics for detecting generically low quality examples. Rishi Bommasani, Claire Cardie |
EMNLP (1) | 2 |
| 2020 | Event Extraction by Answering (Almost) Natural QuestionsabstractThe problem of event extraction requires detecting the event trigger and extracting its corresponding arguments.Existing work in event argument extraction typically relies heavily on entity recognition as a preprocessing/concurrent step, causing the well-known problem of error propagation.To avoid this issue, we introduce a new paradigm for event extraction by formulating it as a question answering (QA) task that extracts the event arguments in an end-to-end manner.Empirical results demonstrate that our framework outperforms prior methods substantially; in addition, it is capable of extracting event arguments for roles not seen at training time (i.e., in a zeroshot learning setting).1 Xinya Du, Claire Cardie |
EMNLP (1) | 2 |
| 2020 | Exploring the Role of Argument Structure in Online Debate PersuasionabstractOnline debate forums provide users a platform to express their opinions on controversial topics while being exposed to opinions from diverse set of viewpoints.Existing work in Natural Language Processing (NLP) has shown that linguistic features extracted from the debate text and features encoding the characteristics of the audience are both critical in persuasion studies.In this paper, we aim to further investigate the role of discourse structure of the arguments from online debates in their persuasiveness.In particular, we use the factor graph model to obtain features for the argument structure of debates from an online debating platform and incorporate these features to an LSTM-based model to predict the debater that makes the most convincing arguments.We find that incorporating argument structure features play an essential role in achieving the better predictive performance in assessing the persuasiveness of the arguments in online debates. Jialu Li 0001, Esin Durmus, Claire Cardie |
EMNLP (1) | 3 |
| 2020 | Investigating Prior Knowledge for Challenging Chinese Machine Reading ComprehensionabstractMachine reading comprehension tasks require a machine reader to answer questions relevant to the given document. In this paper, we present the first free-form multiple-Choice Chinese machine reading Comprehension dataset (C 3 ), containing 13,369 documents (dialogues or more formally written mixed-genre texts) and their associated 19,577 multiple-choice free-form questions collected from Chinese-as-a-second-language examinations. We present a comprehensive analysis of the prior knowledge (i.e., linguistic, domain-specific, and general world knowledge) needed for these real-world problems. We implement rule-based and popular neural methods and find that there is still a significant performance gap between the best performing model (68.5%) and human readers (96.0%), especiallyon problems that require prior knowledge. We further study the effects of distractor plausibility and data augmentation based on translated relevant datasets for English on model performance. We expect C 3 to present great challenges to existing systems as answering 86.8% of questions requires both knowledge within and beyond the accompanying document, and we hope that C 3 can serve as a platform to study how to leverage various kinds of prior knowledge to better understand a given written or orally oriented text. C 3 is available at https://dataset.org/c3/ . Kai Sun 0006, Dian Yu 0001, Dong Yu 0001, Claire Cardie |
Trans. Assoc. Comput. Linguistics | 4 |
| 2019 | Keeping Notes: Conditional Natural Language Generation with a Scratchpad EncoderabstractWe introduce the Scratchpad Mechanism, a novel addition to the sequence-to-sequence (seq2seq) neural network architecture and demonstrate its effectiveness in improving the overall fluency of seq2seq models for natural language generation tasks.By enabling the decoder at each time step to write to all of the encoder output layers, Scratchpad can employ the encoder as a "scratchpad" memory to keep track of what has been generated so far and thereby guide future generation.We evaluate Scratchpad in the context of three well-studied natural language generation tasks -Machine Translation, Question Generation, and Text Summarization -and obtain stateof-the-art or comparable performance on standard datasets for each task.Qualitative assessments in the form of human judgements (question generation), attention visualization (MT), and sample output (summarization) provide further evidence of the ability of Scratchpad to generate fluent and expressive output. Ryan Y. Benmalek, Madian Khabsa, Suma Desu, Claire Cardie, Michele Banko |
ACL (1) | 4 |
| 2019 | Multi-Source Cross-Lingual Model Transfer: Learning What to ShareabstractModern NLP applications have enjoyed a great boost utilizing neural networks models.Such deep neural models, however, are not applicable to most human languages due to the lack of annotated training data for various NLP tasks.Cross-lingual transfer learning (CLTL) is a viable method for building NLP models for a low-resource target language by leveraging labeled data from other (source) languages.In this work, we focus on the multilingual transfer setting where training data in multiple source languages is leveraged to further boost target language performance.Unlike most existing methods that rely only on language-invariant features for CLTL, our approach coherently utilizes both languageinvariant and language-specific features at instance level.Our model leverages adversarial networks to learn language-invariant features, and mixture-of-experts models to dynamically exploit the similarity between the target language and each individual source language 1 .This enables our model to learn effectively what to share between various languages in the multilingual setup.Moreover, when coupled with unsupervised multilingual embeddings, our model can operate in a zero-resource setting where neither target language training data nor cross-lingual resources are available.Our model achieves significant performance gains over prior art, as shown in an extensive set of experiments over multiple text classification and sequence tagging tasks including a large-scale industry dataset. Xilun Chen 0002, Ahmed Awadallah 0001, Hany Hassan, Wei Wang 0238, Claire Cardie |
ACL (1) | 5 |
| 2019 | A Corpus for Modeling User and Language Effects in Argumentation on Online DebatingabstractExisting argumentation datasets have succeeded in allowing researchers to develop computational methods for analyzing the content, structure and linguistic features of argumentative text.They have been much less successful in fostering studies of the effect of "user" traits -characteristics and beliefs of the participants -on the debate/argument outcome as this type of user information is generally not available.This paper presents a dataset of 78, 376 debates generated over a 10-year period along with surprisingly comprehensive participant profiles.We also complete an example study using the dataset to analyze the effect of selected user traits on the debate outcome in comparison to the linguistic features typically employed in studies of this kind. Esin Durmus, Claire Cardie |
ACL (1) | 2 |
| 2019 | Determining Relative Argument Specificity and Stance for Complex Argumentative StructuresabstractSystems for automatic argument generation and debate require the ability to (1) determine the stance of any claims employed in the argument and (2) assess the specificity of each claim relative to the argument context.Existing work on understanding claim specificity and stance, however, has been limited to the study of argumentative structures that are relatively shallow, most often consisting of a single claim that directly supports or opposes the argument thesis.In this paper, we tackle these tasks in the context of complex arguments on a diverse set of topics.In particular, our dataset consists of manually curated argument trees for 741 controversial topics covering 95,312 unique claims; lines of argument are generally of depth 2 to 6.We find that as the distance between a pair of claims increases along the argument path, determining the relative specificity of a pair of claims becomes easier and determining their relative stance becomes harder. Esin Durmus, Faisal Ladhak, Claire Cardie |
ACL (1) | 3 |
| 2019 | The Role of Pragmatic and Discourse Context in Determining Argument ImpactabstractEsin Durmus, Faisal Ladhak, Claire Cardie. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Esin Durmus, Faisal Ladhak, Claire Cardie |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Modeling the Factors of User Success in Online DebateabstractDebate is a process that gives individuals the opportunity to express, and to be exposed to, diverging viewpoints on controversial issues; and the existence of online debating platforms makes it easier for individuals to participate in debates and obtain feedback on their debating skills. But understanding the factors that contribute to a user's success in debate is complicated: while success depends, in part, on the characteristics of the language they employ, it is also important to account for the degree to which their beliefs and personal traits are compatible with that of the audience. Friendships and previous interactions among users on the platform may further influence success. Esin Durmus, Claire Cardie |
WWW | 2 |
| 2019 | DREAM: A Challenge Dataset and Models for Dialogue-Based Reading ComprehensionabstractWe present DREAM, the first dialogue-based multiple-choice reading comprehension data set. Collected from English as a Foreign Language examinations designed by human experts to evaluate the comprehension level of Chinese learners of English, our data set contains 10,197 multiple-choice questions for 6,444 dialogues. In contrast to existing reading comprehension data sets, DREAM is the first to focus on in-depth multi-turn multi-party dialogue understanding. DREAM is likely to present significant challenges for existing reading comprehension systems: 84% of answers are non-extractive, 85% of questions require reasoning beyond a single sentence, and 34% of questions also involve commonsense knowledge. We apply several popular neural reading comprehension models that primarily exploit surface information within the text and find them to, at best, just barely outperform a rule-based approach. We next investigate the effects of incorporating dialogue structure and different kinds of general world knowledge into both rule-based and (neural and non-neural) machine learning-based reading comprehension models. Experimental results on the DREAM data set show the effectiveness of dialogue structure and general world knowledge. DREAM is available at https://dataset.org/dream/ . Kai Sun 0006, Dian Yu 0001, Jianshu Chen, Dong Yu 0001, Yejin Choi 0001, Claire Cardie |
Trans. Assoc. Comput. Linguistics | 6 |
| 2018 | Harvesting Paragraph-level Question-Answer Pairs from WikipediaabstractWe study the task of generating from Wikipedia articles question-answer pairs that cover content beyond a single sentence.We propose a neural network approach that incorporates coreference knowledge via a novel gating mechanism.Compared to models that only take into account sentence-level information (Heilman and Smith, 2010; Du et al., 2017;Zhou et al., 2017), we find that the linguistic knowledge introduced by the coreference representation aids question generation significantly, producing models that outperform the current state-of-theart.We apply our system (composed of an answer span extraction system and the passage-level QG system) to the 10,000 top-ranking Wikipedia articles and create a corpus of over one million questionanswer pairs.We also provide a qualitative analysis for this large-scale generated corpus from Wikipedia. Xinya Du, Claire Cardie |
ACL (1) | 2 |
| 2018 | Unsupervised Multilingual Word EmbeddingsabstractMultilingual Word Embeddings (MWEs) represent words from multiple languages in a single distributional vector space.Unsupervised MWE (UMWE) methods acquire multilingual embeddings without cross-lingual supervision, which is a significant advantage over traditional supervised approaches and opens many new possibilities for low-resource languages.Prior art for learning UMWEs, however, merely relies on a number of independently trained Unsupervised Bilingual Word Embeddings (UBWEs) to obtain multilingual embeddings.These methods fail to leverage the interdependencies that exist among many languages.To address this shortcoming, we propose a fully unsupervised framework for learning MWEs 1 that directly exploits the relations between all language pairs.Our model substantially outperforms previous approaches in the experiments on multilingual word translation and cross-lingual word similarity.In addition, our model even beats supervised approaches trained with cross-lingual resources. Xilun Chen 0002, Claire Cardie |
EMNLP | 2 |
| 2018 | Towards Dynamic Computation Graphs via Sparse Latent StructureabstractDeep NLP models benefit from underlying structures in the data-e.g., parse treestypically extracted using off-the-shelf parsers.Recent attempts to jointly learn the latent structure encounter a tradeoff: either make factorization assumptions that limit expressiveness, or sacrifice end-to-end differentiability.Using the recently proposed SparseMAP inference, which retrieves a sparse distribution over latent structures, we propose a novel approach for end-to-end learning of latent structure predictors jointly with a downstream predictor.To the best of our knowledge, our method is the first to enable unrestricted dynamic computation graph construction from the global latent structure, while maintaining differentiability.2016).Assigning the conjunction as head instead seems preferable in a Child-Sum TreeLSTM. Vlad Niculae, André F. T. Martins, Claire Cardie |
EMNLP | 3 |
| 2018 | SparseMAP: Differentiable Sparse Structured InferenceabstractStructured prediction requires searching over a combinatorial number of structures. To tackle it, we introduce SparseMAP, a new method for sparse structured inference, together with corresponding loss functions. SparseMAP inference is able to automatically select only a few global structures: it is situated between MAP inference, which picks a single structure, and marginal inference, which assigns probability mass to all structures, including implausible ones. Importantly, SparseMAP can be computed using only calls to a MAP oracle, hence it is applicable even to problems where marginal inference is intractable, such as linear assignment. Moreover, thanks to the solution sparsity, gradient backpropagation is efficient regardless of the structure. SparseMAP thus enables us to augment deep neural networks with generic and sparse structured hidden layers. Experiments in dependency parsing and natural language inference reveal competitive accuracy, improved interpretability, and the ability to capture natural language ambiguities, which is attractive for pipeline systems. Vlad Niculae, André F. T. Martins, Mathieu Blondel, Claire Cardie |
ICML | 4 |
| 2018 | A Corpus of eRulemaking User Comments for Measuring Evaluability of Arguments
Joonsuk Park, Claire Cardie |
LREC | 2 |
| 2018 | Multinomial Adversarial Networks for Multi-Domain Text ClassificationabstractMany text classification tasks are known to be highly domain-dependent.Unfortunately, the availability of training data can vary drastically across domains.Worse still, for some domains there may not be any annotated data at all.In this work, we propose a multinomial adversarial network 1 (MAN) to tackle this real-world problem of multi-domain text classification (MDTC) in which labeled data may exist for multiple domains, but in insufficient amounts to train effective classifiers for one or more of the domains.We provide theoretical justifications for the MAN framework, proving that different instances of MANs are essentially minimizers of various f-divergence metrics (Ali and Silvey, 1966) among multiple probability distributions.MANs are thus a theoretically sound generalization of traditional adversarial networks that discriminate over two distributions.More specifically, for the MDTC task, MAN learns features that are invariant across multiple domains by resorting to its ability to reduce the divergence among the feature distributions of each domain.We present experimental results showing that MANs significantly outperform the prior art on the MDTC task.We also show that MANs achieve state-of-the-art performance for domains with no labeled data. Xilun Chen 0002, Claire Cardie |
NAACL-HLT | 2 |
| 2018 | Exploring the Role of Prior Beliefs for Argument PersuasionabstractEsin Durmus, Claire Cardie. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Esin Durmus, Claire Cardie |
NAACL-HLT | 2 |
| 2018 | Nested Named Entity Recognition RevisitedabstractArzoo Katiyar, Claire Cardie. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Arzoo Katiyar, Claire Cardie |
NAACL-HLT | 2 |
| 2018 | Adversarial Deep Averaging Networks for Cross-Lingual Sentiment ClassificationabstractIn recent years great success has been achieved in sentiment classification for English, thanks in part to the availability of copious annotated resources. Unfortunately, most languages do not enjoy such an abundance of labeled data. To tackle the sentiment classification problem in low-resource languages without adequate annotated data, we propose an Adversarial Deep Averaging Network (ADAN 1 ) to transfer the knowledge learned from labeled data on a resource-rich source language to low-resource languages where only unlabeled data exist. ADAN has two discriminative branches: a sentiment classifier and an adversarial language discriminator. Both branches take input from a shared feature extractor to learn hidden representations that are simultaneously indicative for the classification task and invariant across languages. Experiments on Chinese and Arabic sentiment classification demonstrate that ADAN significantly outperforms state-of-the-art systems. Xilun Chen 0002, Yu Sun 0020, Ben Athiwaratkun, Claire Cardie, Kilian Q. Weinberger |
Trans. Assoc. Comput. Linguistics | 4 |
| 2017 | Learning to Ask: Neural Question Generation for Reading ComprehensionabstractWe study automatic question generation for sentences from text passages in reading comprehension.We introduce an attention-based sequence learning model for the task and investigate the effect of encoding sentence-vs.paragraph-level information.In contrast to all previous work, our model does not rely on hand-crafted rules or a sophisticated NLP pipeline; it is instead trainable end-to-end via sequenceto-sequence learning.Automatic evaluation results show that our system significantly outperforms the state-of-the-art rule-based system.In human evaluations, questions generated by our system are also rated as being more natural (i.e., grammaticality, fluency) and as more difficult to answer (in terms of syntactic and lexical divergence from the original text and reasoning needed to answer). Xinya Du, Junru Shao, Claire Cardie |
ACL (1) | 3 |
| 2017 | Going out on a limb: Joint Extraction of Entity Mentions and Relations without Dependency TreesabstractWe present a novel attention-based recurrent neural network for joint extraction of entity mentions and relations. We show that attention along with long short term memory (LSTM) network can extract semantic relations between entity mentions without having access to dependency trees. Experiments on Automatic Content Extraction (ACE) corpora show that our model significantly outperforms feature-based joint model by Li and Ji (2014). We also compare our model with an end-to-end tree-based LSTM model (SPTree) by Miwa and Bansal (2016) and show that our model performs within 1% on entity mentions and 2% on relations. Our fine-grained analysis also shows that our model performs significantly better on Agent-Artifact relations, while SPTree performs better on Physical and Part-Whole relations. Arzoo Katiyar, Claire Cardie |
ACL (1) | 2 |
| 2017 | Argument Mining with Structured SVMs and RNNsabstractWe propose a novel factor graph model for argument mining, designed for settings in which the argumentative relations in a document do not necessarily form a tree structure.(This is the case in over 20% of the web comments dataset we release.)Our model jointly learns elementary unit type classification and argumentative relation prediction.Moreover, our model supports SVM and RNN parametrizations, can enforce structure constraints (e.g., transitivity), and can express dependencies between adjacent relations and propositions.Our approaches outperform unstructured baselines in both web comments and argumentative essay datasets. Vlad Niculae, Joonsuk Park, Claire Cardie |
ACL (1) | 3 |
| 2017 | Identifying Where to Focus in Reading Comprehension for Neural Question GenerationabstractA first step in the task of automatically generating questions for testing reading comprehension is to identify questionworthy sentences, i.e. sentences in a text passage that humans find it worthwhile to ask questions about.We propose a hierarchical neural sentence-level sequence tagging model for this task, which existing approaches to question generation have ignored.The approach is fully data-driven -with no sophisticated NLP pipelines or any hand-crafted rules/features -and compares favorably to a number of baselines when evaluated on the SQuAD data set.When incorporated into an existing neural question generation system, the resulting end-to-end system achieves stateof-the-art performance for paragraph-level question generation for reading comprehension. Xinya Du, Claire Cardie |
EMNLP | 2 |
| 2017 | Keeping Apace with Progress in Natural Language ProcessingabstractIncreasingly central to online search and information discovery are methods from Natural Language Processing (NLP). Research in the field, however, is progressing at what feels a frenetic pace, making it a daunting task to keep up with the latest results. Claire Cardie |
WSDM | 1 |
| 2017 | Using Argumentative Structure to Interpret Debates in Online Deliberative Democracy and eRulemakingabstractGovernments around the world are increasingly utilising online platforms and social media to engage with, and ascertain the opinions of, their citizens. Whilst policy makers could potentially benefit from such enormous feedback from society, they first face the challenge of making sense out of the large volumes of data produced. In this article, we show how the analysis of argumentative and dialogical structures allows for the principled identification of those issues that are central, controversial, or popular in an online corpus of debates. Although areas such as controversy mining work towards identifying issues that are a source of disagreement, by looking at the deeper argumentative structure, we show that a much richer understanding can be obtained. We provide results from using a pipeline of argument-mining techniques on the debate corpus, showing that the accuracy obtained is sufficient to automatically identify those issues that are key to the discussion, attracting proportionately more support than others, and those that are divisive, attracting proportionately more conflicting viewpoints. John Lawrence, Joonsuk Park, Katarzyna Budzynska, Claire Cardie, Barbara Konat, Chris Reed 0001 |
ACM Trans. Internet Techn. | 4 |
| 2016 | Investigating LSTMs for Joint Extraction of Opinion Entities and RelationsabstractWe investigate the use of deep bidirectional LSTMs for joint extraction of opinion entities and the IS-FROM and IS-ABOUT relations that connect them -the first such attempt using a deep learning approach.Perhaps surprisingly, we find that standard LSTMs are not competitive with a state-of-the-art CRF+ILP joint inference approach (Yang and Cardie, 2013) to opinion entities extraction, performing below even the standalone sequencetagging CRF.Incorporating sentence-level and a novel relation-level optimization, however, allows the LSTM to identify opinion relations and to perform within 1-3% of the state-of-the-art joint model for opinion entities and the IS-FROM relation; and to perform as well as the state-of-theart for the IS-ABOUT relation -all without access to opinion lexicons, parsers and other preprocessing components required for the feature-rich CRF+ILP approach. Arzoo Katiyar, Claire Cardie |
ACL (1) | 2 |
| 2015 | Toward machine-assisted participation in eRulemaking: an argumentation model of evaluabilityabstracteRulemaking is an ongoing effort to use online tools to foster broader and better public participation in rulemaking --- the multi-step process that federal agencies use to develop new health, safety, and economic regulations. The increasing participation of non-expert citizens, however, has led to a growth in the amount of arguments whose validity or strength are difficult to evaluate, both by the government agencies and fellow citizens. Such arguments typically neglect to provide the reasons for the conclusions and objective evidence for factual claims upon which the arguments are based. In this paper, we propose a novel argumentation model for capturing the evaluability of user comments in eRulemaking. This model is intended to be used for implementing automated systems to assist users in constructing evaluable arguments under online commenting environment for the benefit of quick feedback at a low cost. Joonsuk Park, Cheryl Blake, Claire Cardie |
ICAIL | 3 |
| 2015 | Socially-Informed Timeline Generation for Complex EventsabstractExisting timeline generation systems for complex events consider only information from traditional media, ignoring the rich social context provided by user-generated content that reveals representative public interests or insightful opinions. We instead aim to generate socially-informed timelines that contain both news article summaries and selected user comments. We present an optimization framework designed to balance topical cohesion between the article and comment summaries along with their informativeness and coverage of the event. Automatic evaluations on real-world datasets that cover four complex events show that our system produces more informative timelines than state-of-the-art systems. In human evaluation, the associated comment summaries are furthermore rated more insightful than editor's picks and comments ranked highly by users. Lu Wang 0008, Claire Cardie, Galen Marchetti |
HLT-NAACL | 2 |
| 2015 | A Hierarchical Distance-dependent Bayesian Model for Event Coreference ResolutionabstractWe present a novel hierarchical distance-dependent Bayesian model for event coreference resolution. While existing generative models for event coreference resolution are completely unsupervised, our model allows for the incorporation of pairwise distances between event mentions — information that is widely used in supervised coreference models to guide the generative clustering processing for better event clustering both within and across documents. We model the distances between event mentions using a feature-rich learnable distance function and encode them as Bayesian priors for nonparametric clustering. Experiments on the ECB+ corpus show that our model outperforms state-of-the-art methods for both within- and cross-document event coreference resolution. Bishan Yang, Claire Cardie, Peter I. Frazier |
Trans. Assoc. Comput. Linguistics | 2 |
| 2014 | Towards a General Rule for Identifying Deceptive Opinion SpamabstractConsumers' purchase decisions are increasingly influenced by user-generated online reviews.Accordingly, there has been growing concern about the potential for posting deceptive opinion spamfictitious reviews that have been deliberately written to sound authentic, to deceive the reader.In this paper, we explore generalized approaches for identifying online deceptive opinion spam based on a new gold standard dataset, which is comprised of data from three different domains (i.e.Hotel, Restaurant, Doctor), each of which contains three types of reviews, i.e. customer generated truthful reviews, Turker generated deceptive reviews and employee (domain-expert) generated deceptive reviews.Our approach tries to capture the general difference of language usage between deceptive and truthful reviews, which we hope will help customers when making purchase decisions and review portal operators, such as TripAdvisor or Yelp, investigate possible fraudulent activity on their sites.1 Jiwei Li 0001, Myle Ott, Claire Cardie, Eduard H. Hovy |
ACL (1) | 3 |
| 2014 | Context-aware Learning for Sentence-level Sentiment Analysis with Posterior RegularizationabstractThis paper proposes a novel context-aware method for analyzing sentiment at the level of individual sentences. Most existing machine learning approaches suffer from limitations in the modeling of complex linguistic structures across sentences and often fail to capture nonlocal contextual cues that are important for sentiment interpretation. In contrast, our approach allows structured modeling of sentiment while taking into account both local and global contextual information. Specifically, we encode intuitive lexical and discourse knowledge as expressive constraints and integrate them into the learning of conditional random field models via posterior regularization. The context-aware constraints provide additional power to the CRF model and can guide semi-supervised learning when labeled data is limited. Experiments on standard product review datasets show that our method outperforms the state-of-theart methods in both the supervised and semi-supervised settings. Bishan Yang, Claire Cardie |
ACL (1) | 2 |
| 2014 | Query-Focused Opinion Summarization for User-Generated Content
Lu Wang 0008, Hema Raghavan, Claire Cardie, Vittorio Castelli |
COLING | 3 |
| 2014 | AsseSS: A Tool for Assessing the Support Structures of Arguments in User CommentsabstractWe present AsseSS, a tool for identifying and assessing the support structures of arguments in user comments. Given a comment, the system first classifies elementary units of arguments comprising the comment based on the type of appropriate support. Then, it detects support relations among the elementary units. With this information, it is possible to decide whether the existing support relation is of suitable type. Also, in the case that no support has been provided for an elementary unit, an appropriate type of support can be determined. Joonsuk Park, Claire Cardie |
COMMA | 2 |
| 2014 | Opinion Mining with Deep Recurrent Neural NetworksabstractRecurrent neural networks (RNNs) are connectionist models of sequential data that are naturally applicable to the analysis of natural language. Recently, “depth in space” — as an orthogonal notion to “depth in time” — in RNNs has been investigated by stacking multiple layers of RNNs and shown empirically to bring a temporal hierarchy to the architecture. In this work we apply these deep RNNs to the task of opinion expression extraction formulated as a token-level sequence-labeling task. Experimental results show that deep, narrow RNNs outperform traditional shallow, wide RNNs with the same number of parameters. Furthermore, our approach outperforms previous CRF-based baselines, including the state-of-the-art semi-Markov CRF model, and does so without access to the powerful opinion lexicons and syntactic features relied upon by the semi-CRF, as well as without the standard layer-by-layer pre-training typically required of RNN architectures. Ozan Irsoy, Claire Cardie |
EMNLP | 2 |
| 2014 | Major Life Event Extraction from Twitter based on Congratulations/Condolences Speech ActsabstractSocial media websites provide a platform for anyone to describe significant events taking place in their lives in realtime.Currently, the majority of personal news and life events are published in a textual format, motivating information extraction systems that can provide a structured representations of major life events (weddings, graduation, etc. . .).This paper demonstrates the feasibility of accurately extracting major life events.Our system extracts a fine-grained description of users' life events based on their published tweets.We are optimistic that our system can help Twitter users more easily grasp information from users they take interest in following and also facilitate many downstream applications, for example realtime friend recommendation. Jiwei Li 0001, Alan Ritter, Claire Cardie, Eduard H. Hovy |
EMNLP | 3 |
| 2014 | Deep Recursive Neural Networks for Compositionality in Language
Ozan Irsoy, Claire Cardie |
NIPS | 2 |
| 2014 | Sentiment analysis on evolving social streams: how self-report imbalances can helpabstractReal-time sentiment analysis is a challenging machine learning task, due to scarcity of labeled data and sudden changes in sentiment caused by real-world events that need to be instantly interpreted. In this paper we propose solutions to acquire labels and cope with concept drift in this setting, by using findings from social psychology on how humans prefer to disclose some types of emotions. In particular, we use findings that humans are more motivated to report positive feelings rather than negative feelings and also prefer to report extreme feelings rather than average feelings. Pedro Henrique Calais Guerra, Wagner Meira Jr., Claire Cardie |
WSDM | 3 |
| 2014 | Timeline generation: tracking individuals on twitterabstractIn this paper, we preliminarily learn the problem of reconstructing users' life history based on the their Twitter stream and proposed an unsupervised framework that create a chronological list for personal important events (PIE) of individuals. By analyzing individ- ual tweet collections, we find that what are suitable for inclusion in the personal timeline should be tweets talking about personal (as opposed to public) and time-specific (as opposed to time-general) topics. To further extract these types of topics, we introduce a non-parametric multi-level Dirichlet Process model to recognize four types of tweets: personal time-specific (PersonTS), personal time-general (PersonTG), public time-specific (PublicTS) and pub- lic time-general (PublicTG) topics, which, in turn, are used for fur- ther personal event extraction and timeline generation. To the best of our knowledge, this is the first work focused on the generation of timeline for individuals from Twitter data. For evaluation, we have built gold standard timelines that contain PIE related events from 20 ordinary twitter users and 20 celebrities. Experimental results demonstrate that it is feasible to automatically extract chronologi- cal timelines for Twitter users from their tweet collection Jiwei Li 0001, Claire Cardie |
WWW | 2 |
| 2014 | Sentiment Analysis and Opinion Mining Bing Liu (University of Illinois at Chicago) Morgan & Claypool (Synthesis Lectures on Human Language Technologies, edited by Graeme Hirst, 5(1)), 2012, 167 pp; paperbound, ISBN 978-1-60845-884-4abstractThis 2012 book is written as a comprehensive introductory and survey text for sentiment analysis and opinion mining, a field of study that investigates computational techniques for analyzing text to uncover the opinions, sentiment, emotions, and evaluations expressed therein. As such, it aims to be accessible to a broad audience that includes students, researchers, and practitioners, as well as to cover all important topics in the field.With regard to the first aim, the book is very much a success: The writing is clear and concise, informative examples motivate each new topic, terminology is clearly defined, and descriptions of key algorithms are provided in the running text along with short (usually one-line) descriptions of each piece of relevant related work. The latter, in particular, makes the book an excellent platform from which to dive into the quickly expanding body of literature on sentiment and opinion analysis. In addition, I believe that the book should be easily accessible to anyone with a computer science background.With regard to Liu's second aim of covering all important topics in the field, the degree to which the book succeeds is a matter of, well, opinion. Let me explain. Liu's early research was in data mining and Web mining; not surprisingly then, the book is written from this perspective. It is very much centered around the analysis of user-generated opinions in social media. Liu's particular expertise is in the area of product reviews; hence, the bulk of the book's examples are from this domain. Furthermore, the book focuses on techniques that are first and foremost applicable to aspect-based sentiment analysis—fine-grained analysis of opinions regarding specific aspects of products and services. For the most part, investigations in this area have been restricted to reviews of electronics products (e.g., cameras), hotels, and restaurants with their associated entity-specific aspects (e.g., weight, photo quality, and ease of use for cameras; rooms, front desk service, and cleanliness for hotels; and food, service, ambience, and cost for restaurants).In contrast, the similarly named survey of Pang and Lee (2008)—Opinion Mining and Sentiment Analysis—is more even-handed in its selection of topics and techniques and is written from the point of view of natural language processing (NLP) and computational linguistics. Pang and Lee, for example, are aware of prior work in the field on fact and event-based text analysis and, within that context, focus consciously on the description of “new challenges1 raised by sentiment-aware applications” as well as the methods proposed to address them. As a result, the survey proves to be an easy, comfortable (and entertaining) read for those with an NLP-centric ancestry.Not so with the Liu survey, because the goals and assumptions that underlie aspect-based sentiment analysis can be at odds with many of those at the core of computational linguistics and NLP. But do not despair! This is a good thing! Precisely because of Liu's different tack on opinion and sentiment analysis, for many readers the book will be a wonderful source of ideas for new problems to work on in the field. In particular, the language of product reviews is quite different from that of other opinion-oriented text (e.g., editorials, blogs, position papers, political arguments, and even movie reviews). Product reviews tend to be quite short; they describe a single, known product; the opinion expressions themselves tend to be product-specific. Other genres of opinion-oriented text are generally longer; they can discuss virtually any topic or set of topics and, hence, are likely to exhibit a greater variety of sentence and discourse structure, including a virtually unlimited (and, out of context, ambiguous) opinion expression vocabulary, the presence of opinion holders that are not the author, and implicit opinion targets.With this contrast in writing genres in mind, reading the book becomes a thought experiment in determining whether the techniques it covers will perform well on opinion-oriented texts beyond product reviews; if not, how and when will they fail; and in what circumstances might more complex language understanding components like parsing, semantic interpretation, or discourse analysis be helpful in analyzing product reviews?The book contains eleven chapters (with a short summary of the book in a final twelfth chapter). Chapter 1 introduces the problem of sentiment analysis. It discusses the differences in terminology that exist in industry vs. academia and briefly describes approximately 20 recent2 applications-oriented sentiment analysis research efforts published largely at venues outside of NLP. The latter is a nice entree into the applied sentiment analysis literature beyond the standard NLP conferences.Chapter 2 provides an abstraction of the opinion mining problem and formally defines an opinion in terms of its components—the opinion holder, the entity and aspect of that entity that is the target of the opinion, the sentiment expressed, and the time that the opinion was expressed. The remaining chapters are organized around this definition and the sentiment-based applications that it enables. Thus, there are chapters on document-level sentiment classification and rating prediction (Chapter 3), sentence-level subjectivity and sentiment classification (Chapter 4), and aspect-based sentiment analysis (Chapter 5, roughly one-quarter of the book). Following these are shorter chapters on sentiment lexicon generation, opinion summarization, the analysis of comparative opinions, opinion retrieval (vs. Web search), and determining review quality (Chapters 6, 7, 8, 9, and 11, respectively).Chapter 10 is a somewhat longer chapter devoted to opinion spam detection. Here Liu describes techniques to identify fake product reviews—some that rely primarily on the review content and available meta-data, and others based on identifying atypical behaviors of the reviewer(s).The chapter on aspect-based sentiment analysis is, by far, the most detailed, but there are some nice surprises in other chapters including sections on handling sarcastic sentences, learning a priori objective terms that imply an opinion, and analyzing opinions in contexts, as well as multiple sections that address cross-language and cross-domain issues.Overall, the book is a very valuable resource for those interested in understanding the quickly expanding literature on sentiment and opinion analysis, especially techniques for aspect-based sentiment analysis. For NLP researchers, it can also serve as a source of new problems to tackle in the analysis of opinion-oriented text. Claire Cardie |
Comput. Linguistics | 1 |
| 2014 | Joint Modeling of Opinion Expression Extraction and Attribute ClassificationabstractIn this paper, we study the problems of opinion expression extraction and expression-level polarity and intensity classification. Traditional fine-grained opinion analysis systems address these problems in isolation and thus cannot capture interactions among the textual spans of opinion expressions and their opinion-related properties. We present two types of joint approaches that can account for such interactions during 1) both learning and inference or 2) only during inference. Extensive experiments on a standard dataset demonstrate that our approaches provide substantial improvements over previously published results. By analyzing the results, we gain some insight into the advantages of different joint models. Bishan Yang, Claire Cardie |
Trans. Assoc. Comput. Linguistics | 2 |
| 2013 | Domain-Independent Abstract Generation for Focused Meeting Summarization
Lu Wang 0008, Claire Cardie |
ACL (1) | 2 |
| 2013 | A Sentence Compression Based Framework to Query-Focused Multi-Document Summarization
Lu Wang 0008, Hema Raghavan, Vittorio Castelli, Radu Florian, Claire Cardie |
ACL (1) | 5 |
| 2013 | Joint Inference for Fine-grained Opinion Extraction
Bishan Yang, Claire Cardie |
ACL (1) | 2 |
| 2013 | Identifying Manipulated Offerings on Review PortalsabstractRecent work has developed supervised methods for detecting deceptive opinion spamfake reviews written to sound authentic and deliberately mislead readers.And whereas past work has focused on identifying individual fake reviews, this paper aims to identify offerings (e.g., hotels) that contain fake reviews.We introduce a semi-supervised manifold ranking algorithm for this task, which relies on a small set of labeled individual reviews for training.Then, in the absence of gold standard labels (at an offering level), we introduce a novel evaluation procedure that ranks artificial instances of real offerings, where each artificial offering contains a known number of injected deceptive reviews.Experiments on a novel dataset of hotel reviews show that the proposed method outperforms state-of-art learning baselines. Jiwei Li 0001, Myle Ott, Claire Cardie |
EMNLP | 3 |
| 2013 | A Measure of Polarization on Social Media Networks Based on Community Boundaries
Pedro Henrique Calais Guerra, Wagner Meira Jr., Claire Cardie, Robert D. Kleinberg |
ICWSM | 3 |
| 2013 | Properties, Prediction, and Prevalence of Useful User-Generated Comments for Descriptive Annotation of Social Media Objects
Elaheh Momeni, Claire Cardie, Myle Ott |
ICWSM | 2 |
| 2013 | Negative Deceptive Opinion Spam
Myle Ott, Claire Cardie, Jeffrey T. Hancock |
HLT-NAACL | 2 |
| 2012 | Extracting Opinion Expressions with semi-Markov Conditional Random Fields
Bishan Yang, Claire Cardie |
EMNLP-CoNLL | 2 |
| 2012 | Improving Implicit Discourse Relation Recognition Through Feature Set Optimization
Joonsuk Park, Claire Cardie |
SIGDIAL Conference | 2 |
| 2012 | Unsupervised Topic Modeling Approaches to Decision Summarization in Spoken Meetings
Lu Wang 0008, Claire Cardie |
SIGDIAL Conference | 2 |
| 2012 | Focused Meeting Summarization via Unsupervised Relation Extraction
Lu Wang 0008, Claire Cardie |
SIGDIAL Conference | 2 |
| 2012 | Estimating the prevalence of deception in online review communitiesabstractConsumers' purchase decisions are increasingly influenced by user-generated online reviews. Accordingly, there has been growing concern about the potential for posting deceptive opinion spam---fictitious reviews that have been deliberately written to sound authentic, to deceive the reader. But while this practice has received considerable public attention and concern, relatively little is known about the actual prevalence, or rate, of deception in online review communities, and less still about the factors that influence it. Myle Ott, Claire Cardie, Jeffrey T. Hancock |
WWW | 2 |
| 2011 | Joint Bilingual Sentiment Classification with Unlabeled Parallel Corpora
Bin Lu 0001, Chenhao Tan, Claire Cardie, Benjamin Ka-Yin T'sou |
ACL | 3 |
| 2011 | Finding Deceptive Opinion Spam by Any Stretch of the Imagination
Myle Ott, Yejin Choi 0001, Claire Cardie, Jeffrey T. Hancock |
ACL | 3 |
| 2011 | Compositional Matrix-Space Models for Sentiment Analysis
Ainur Yessenalina, Claire Cardie |
EMNLP | 2 |
| 2010 | Multi-Level Structured Models for Document-Level Sentiment Classification
Ainur Yessenalina, Yisong Yue, Claire Cardie |
EMNLP | 3 |
| 2009 | Conundrums in Noun Phrase Coreference Resolution: Making Sense of the State-of-the-Art
Veselin Stoyanov, Nathan Gilbert, Claire Cardie, Ellen Riloff |
ACL/IJCNLP | 3 |
| 2009 | Adapting a Polarity Lexicon using Integer Linear Programming for Domain-Specific Sentiment Classification
Yejin Choi 0001, Claire Cardie |
EMNLP | 2 |
| 2008 | Topic Identification for Fine-Grained Opinion Analysis
Veselin Stoyanov, Claire Cardie |
COLING | 2 |
| 2008 | Learning with Compositional Semantics as Structural Inference for Subsentential Sentiment Analysis
Yejin Choi 0001, Claire Cardie |
EMNLP | 2 |
| 2008 | An eRulemaking Corpus: Identifying Substantive Issues in Public Comments
Claire Cardie, Cynthia Farina, Matt Rawding, Adil Aijaz |
LREC | 1 |
| 2008 | Annotating Topics of Opinions
Veselin Stoyanov, Claire Cardie |
LREC | 2 |
| 2007 | Identifying Expressions of Opinion in Context
Eric Breck, Yejin Choi 0001, Claire Cardie |
IJCAI | 3 |
| 2007 | Structured Local Training and Biased Potential Functions for Conditional Random Fields with Application to Coreference Resolution
Yejin Choi 0001, Claire Cardie |
HLT-NAACL | 2 |
| 2006 | Joint Extraction of Entities and Relations for Opinion Recognition
Yejin Choi 0001, Eric Breck, Claire Cardie |
EMNLP | 3 |
| 2006 | Partially Supervised Coreference Resolution for Opinion Summarization through Structured Rule Learning
Veselin Stoyanov, Claire Cardie |
EMNLP | 2 |
| 2005 | Machine Learning for Natural Language Processing (and Vice Versa?)
Claire Cardie |
PKDD | 1 |
| 2004 | Playing the Telephone Game: Determining the Hierarchical Structure of Perspective and Speech Expressions
Eric Breck, Claire Cardie |
COLING | 2 |
| 2003 | Bootstrapping Coreference Classifiers with Multiple Machine Learning Algorithms
Vincent Ng 0001, Claire Cardie |
EMNLP | 2 |
| 2003 | Weakly Supervised Natural Language Learning Without Redundant Views
Vincent Ng 0001, Claire Cardie |
HLT-NAACL | 2 |
| 2002 | Identifying Anaphoric and Non-Anaphoric Noun Phrases to Improve Coreference Resolution
Vincent Ng 0001, Claire Cardie |
COLING | 2 |
| 2002 | Combining Sample Selection and Error-Driven Pruning for Machine Learning of Coreference RulesabstractMost machine learning solutions to noun phrase coreference resolution recast the problem as a classification task. We examine three potential problems with this reformulation, namely, skewed class distributions, the inclusion of "hard" training instances, and the loss of transitivity inherent in the original coreference relation. We show how these problems can be handled via intelligent sample selection and error-driven pruning of classification rule-sets. The resulting system achieves an F-measure of 69.5 and 63.4 on the MUC-6 and MUC-7 coreference resolution data sets, respectively, surpassing the performance of the best MUC-6 and MUC-7 coreference systems. In particular, the system outperforms the best-performing learning-based coreference system to date. Vincent Ng 0001, Claire Cardie |
EMNLP | 2 |
| 2001 | Limitations of Co-Training for Natural Language Learning from Large Datasets
David R. Pierce, Claire Cardie |
EMNLP | 2 |
| 2001 | Constrained K-means Clustering with Background Knowledge
Kiri Wagstaff, Claire Cardie, Seth Rogers, Stefan Schrödl |
ICML | 2 |
| 2000 | Clustering with Instance-level Constraints
Kiri Wagstaff, Claire Cardie |
ICML | 2 |
| 2000 | Using clustering and SuperConcepts within SMART: TREC 6
Chris Buckley, Mandar Mitra, Janet A. Walz, Claire Cardie |
Inf. Process. Manag. | 4 |
| 2000 | A Cognitive Bias Approach to Feature Selection and Weighting for Case-Based Learners
Claire Cardie |
Mach. Learn. | 1 |
| 1999 | Noun Phrase Coreference as Clustering
Claire Cardie, Kiri Wagstaff |
EMNLP | 1 |
| 1999 | Combining Error-Driven Pruning and Classification for Partial Parsing
Claire Cardie, Scott Anthony Mardis, David R. Pierce |
ICML | 1 |
| 1999 | Integrating case-based learning and cognitive biases for machine learning of natural language
Claire Cardie |
J. Exp. Theor. Artif. Intell. | 1 |
| 1999 | Guest Editors' Introduction: Machine Learning and Natural Language
Claire Cardie, Raymond J. Mooney |
Mach. Learn. | 1 |
| 1997 | Examining Locally Varying Weights for Nearest Neighbor Algorithms
Nicholas R. Howe, Claire Cardie |
ICCBR | 2 |
| 1997 | Improving Minority Class Prediction Using Case-Specific Feature Weights
Claire Cardie, Nicholas Nowe |
ICML | 1 |
| 1996 | Automating Feature Set Selection for Case-Based Learning of Linguistic Knowledge
Claire Cardie |
EMNLP | 1 |
| 1993 | A Case-Based Approach to Knowledge Acquisition for Domain-Specific Sentence Analysis
Claire Cardie |
AAAI | 1 |
| 1993 | Using Decision Trees to Improve Case-Based Learning
Claire Cardie |
ICML | 1 |
| 1992 | Learning to Disambiguate Relative Pronouns
Claire Cardie |
AAAI | 1 |
| 1992 | Corpus-Based Acquisition of Relative Pronoun Disambiguation HeuristicsabstractThis paper presents a corpus-based approach for deriving heuristics to locate the antecedents of relative pronouns. The technique duplicates the performance of hand-coded rules and requires human intervention only during the training phase. Because the training instances are built on parser output rather than word cooccurrences, the technique requires a small number of training examples and can be used on small to medium-sized corpora. Our initial results suggest that the approach may provide a general method for the automated acquisition of a variety of disambiguation heuristics for natural language systems, especially for problems that require the assimilation of syntactic and semantic knowledge. Claire Cardie |
ACL | 1 |
| 1991 | A Cognitively Plausible Approach to Understanding Complex Syntax
Claire Cardie, Wendy G. Lehnert |
AAAI | 1 |