EDBT 2026 Demo / reviewers in the wild / expert
Jordan L. Boyd-Graber
dblp:57/5950 · also Jordan Lee Boyd-Graber
· DBLP profile ↗
124ranked-venue papers
8as first author
41since 2021 · last 2026
0000-0002-7770-4431ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 113 · 7 first-author · 41 since 2021Human-computer interaction and ubiquitous computing · 9 · 1 first-authorDatabases, data management, data science and information retrieval · 4Applied, interdisciplinary, general and emerging computing · 4Graphics, computer vision, multimedia, augmented reality and games · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real UsersabstractNishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman, Amanpreet Singh, Pao Siangliulue, Joseph Chee Chang, Eunsol Choi, Jordan Lee Boyd-Graber, Aakanksha Naik. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Nishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman, Amanpreet Singh, Pao Siangliulue, Joseph Chee Chang, Eunsol Choi, Jordan L. Boyd-Graber, Aakanksha Naik |
ACL (1) | 9 |
| 2026 | BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice BenchmarksabstractNishant Balepur, Bhavya Rajasekaran, Hyunjin Jane Oh, Michael Xie, Atrey Desai, Vipul Gupta, Steven James Moore, Eunsol Choi, Rachel Rudinger, Jordan Lee Boyd-Graber. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Nishant Balepur, Bhavya Rajasekaran, Hyunjin Jane Oh, Michael Xie, Atrey Desai, Steven James Moore, Eunsol Choi, Rachel Rudinger, Jordan L. Boyd-Graber |
ACL (1) | 10 |
| 2026 | Measuring User's Mental Models of Speech Translation in Human-AI CollaborationabstractMillions of people use machine translation (MT) tools daily, yet little is known about their perception of what systems can and cannot do.This paper studies users' mental models of speech translation systems through a new framework based on cross-lingual question answering, where users either accept MT output or request professional re-translation to answer questions based on the information presented in a foreign language.By analyzing user behavior and accuracy trends across varying translation qualities, we examine to what extent they can predict where the system is likely to be wrong, and how this mental model evolves.Users develop stronger mental models with practice, especially when they have some knowledge of the source language, primarily by relying on surface-level error cues.Moreover, providing speech transcriptions can help users develop better mental models.Our results show the promise of cross-lingual question answering as a downstream task for studying MT mental models, and advancing our understanding of human-AI collaboration. HyoJung Han 0001, Nishant Balepur, Jordan L. Boyd-Graber, Marine Carpuat |
ACL (1) | 3 |
| 2025 | Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User PersonasabstractNishant Balepur, Vishakh Padmakumar, Fumeng Yang, Shi Feng, Rachel Rudinger, Jordan Lee Boyd-Graber. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Nishant Balepur, Vishakh Padmakumar, Fumeng Yang, Shi Feng 0005, Rachel Rudinger, Jordan L. Boyd-Graber |
ACL (1) | 6 |
| 2025 | Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the AboveabstractMultiple choice question answering (MCQA) is popular for LLM evaluation due to its simplicity and human-like testing, but we argue for its reform.We first reveal flaws in MCQA's format, as it struggles to: 1) test generation/subjectivity; 2) match LLM use cases; and 3) fully test knowledge.We instead advocate for generative formats based on human testing-where LLMs construct and explain answers-better capturing user needs and knowledge while remaining easy to score.We then show even when MCQA is a useful format, its datasets suffer from: leakage; unanswerability; shortcuts; and saturation.In each issue, we give fixes from education, like rubrics to guide MCQ writing; scoring methods to bridle guessing; and Item Response Theory to build harder MCQs.Lastly, we discuss LLM errors in MCQA-robustness, biases, and unfaithful explanations-showing how our prior solutions better measure or address these issues.While we do not need to desert MCQA, we encourage more efforts in refining the task based on educational testing, advancing evaluations. Q1. What's wrong with MCQA's format?A) It doesn't apply to many tasks ( §3.1) B) It's misaligned with LLM use cases ( §3.2) C) It doesn't fully test knowledge ( §3.3) Q2.What's wrong with MCQA datasets?A) Test sets are contaminated ( §5.1) B) They have unanswerable questions ( §5.2) C) They contain shortcuts ( §5.3) D) They're too easy for LLMs ( §5.4) Q3.How do LLMs struggle with MCQA?A) They lack robustness ( §6.1) B) They exhibit biases ( §6.2) C) They give unfaithful explanations ( §6.3) Q4.How can insights from education improve MCQA?A) Improve knowledge testing via generative formats ( §4) B) Combat test set leakage with fresh questions ( §5.1) C) Write MCQs informed by educational rubrics ( §5.2) D) Use calibration scoring to curb guessing ( §5.3.1)E) Find harder MCQs with item response theory ( §5.4.1) Nishant Balepur, Rachel Rudinger, Jordan L. Boyd-Graber |
ACL (1) | 3 |
| 2025 | ProxAnn: Use-Oriented Evaluations of Topic Models and Document ClusteringabstractAlexander Miserlis Hoyle, Lorena Calvo-Bartolomé, Jordan Lee Boyd-Graber, Philip Resnik. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Alexander Miserlis Hoyle, Lorena Calvo-Bartolomé, Jordan L. Boyd-Graber, Philip Resnik |
ACL (1) | 3 |
| 2025 | Large Language Models Struggle to Describe the Haystack without Human Help: A Social Science-Inspired Evaluation of Topic ModelsabstractA common use of NLP by social scientists is to understand large document collections. Recent data exploration and content analysis have shifted from probabilistic topic models to Large Language Models (LLMs). Yet their effectiveness in helping users understand content in real-world applications remains under explored. This study compares the knowledge users gain from unsupervised LLMs, supervised LLMs, and traditional topic models across two datasets. While unsupervised LLMs generate more human-readable topics, their topics are overly generic for domain-specific datasets and do not help users learn much about the documents. Adding human supervision to LLM generation improves data exploration by mitigating hallucination and over-genericity but requires greater human effort. Traditional topic models, such as Latent Dirichlet Allocation (LDA), remain effective for exploration but are less user-friendly. LLMs struggle to describe the haystack of large corpora without human help, particularly domain-specific data, and face scaling and hallucination limitations due to context length constraints. Zongxia Li, Lorena Calvo-Bartolomé, Alexander Miserlis Hoyle, Paiheng Xu, Daniel Kofi Stephens, Juan Francisco Fung, Alden Dima, Jordan L. Boyd-Graber |
ACL (1) | 8 |
| 2025 | No Questions are Stupid, but some are Poorly Posed: Understanding Poorly-Posed Information-Seeking QuestionsabstractQuestions help unlock information to satisfy users' information needs.However, when the question is poorly posed, answerers (whether human or computer) may struggle to answer the question in a way that satisfies the asker, despite possibly knowing everything necessary to address the asker's latent information need.Using Reddit question-answer interactions from r/NoStupidQuestions, we develop a computational framework grounded in linguistic theory to study poorly-posedness of questions by generating spaces of potential interpretations of questions and computing distributions over these spaces based on interpretations chosen by both human answerers in the Reddit question thread, as well as by a suite of large language models.Both humans and models struggle to converge on dominant interpretations when faced with poorly posed questions, but employ different strategies: humans focus on specific interpretations through question negotiation, while models attempt comprehensive coverage by addressing many interpretations simultaneously. Neha Srikanth, Rachel Rudinger, Jordan L. Boyd-Graber |
ACL (1) | 3 |
| 2025 | GRACE: A Granular Benchmark for Evaluating Model Calibration against Human CalibrationabstractLanguage models are often miscalibrated, leading to confidently incorrect answers. We introduce GRACE, a benchmark for language model calibration that incorporates comparison with human calibration. GRACE consists of question-answer pairs, in which each question contains a series of clues that gradually become easier, all leading to the same answer; models must answer correctly as early as possible as the clues are revealed. This setting permits granular measurement of model calibration based on how early, accurately, and confidently a model answers. After collecting these questions, we host live human vs. model competitions to gather 1,749 data points on human and model teams’ timing, accuracy, and confidence. We propose a metric, CalScore, that uses GRACE to analyze model calibration errors and identify types of model miscalibration that differ from human behavior. We find that although humans are less accurate than models, humans are generally better calibrated. Since state-of-the-art models struggle on GRACE, it effectively evaluates progress on improving model calibration. Yoo Yeon Sung, Eve Fleisig, Ishan Upadhyay, Jordan L. Boyd-Graber |
ACL (1) | 5 |
| 2025 | ADAPTIVE IE: Investigating the Complementarity of Human-AI Collaboration to Adaptively Extract Information on-the-flyabstractInformation extraction (IE) needs vary over time, where a flexible information extraction (IE) system can be useful. Despite this, existing IE systems are either fully supervised, requiring expensive human annotations, or fully unsupervised, extracting information that often do not cater to user’s needs. To address these issues, we formally introduce the task of “IE on-the-fly”, and address the problem using our proposed Adaptive IE framework that uses human-in-the-loop refinement to adapt to changing user questions. Through human experiments on three diverse datasets, we demonstrate that Adaptive IE is a domain-agnostic, responsive, efficient framework for helping users access useful information while quickly reorganizing information in response to evolving information needs. Ishani Mondal, Michelle Yuan, Anandhavelu Natarajan, Aparna Garimella, Francis Ferraro, Andrew Blair-Stanek, Benjamin Van Durme, Jordan L. Boyd-Graber |
COLING | 8 |
| 2025 | A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps UsersabstractNishant Balepur, Matthew Shu, Yoo Yeon Sung, Seraphina Goldfarb-Tarrant, Shi Feng, Fumeng Yang, Rachel Rudinger, Jordan Lee Boyd-Graber. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Nishant Balepur, Matthew Shu, Yoo Yeon Sung, Seraphina Goldfarb-Tarrant, Shi Feng 0005, Fumeng Yang, Rachel Rudinger, Jordan L. Boyd-Graber |
EMNLP | 8 |
| 2025 | Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question AnsweringabstractLorena Calvo-Bartolomé, Valérie Aldana, Karla Cantarero, Alonso Madroñal de Mesa, Jerónimo Arenas-García, Jordan Lee Boyd-Graber. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Lorena Calvo-Bartolomé, Valérie Aldana, Karla Cantarero, Alonso Madroñal de Mesa, Jerónimo Arenas-García, Jordan L. Boyd-Graber |
EMNLP | 6 |
| 2025 | MoDS: Moderating a Mixture of Document Speakers to Summarize Debatable Queries in Document CollectionsabstractNishant Balepur, Alexa Siu, Nedim Lipka, Franck Dernoncourt, Tong Sun, Jordan Lee Boyd-Graber, Puneet Mathur. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Nishant Balepur, Alexa F. Siu, Nedim Lipka, Franck Dernoncourt, Tong Sun 0005, Jordan L. Boyd-Graber, Puneet Mathur |
NAACL (Long Papers) | 6 |
| 2025 | Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded AdversarialnessabstractYoo Yeon Sung, Maharshi Gor, Eve Fleisig, Ishani Mondal, Jordan Lee Boyd-Graber. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yoo Yeon Sung, Maharshi Gor, Eve Fleisig, Ishani Mondal, Jordan L. Boyd-Graber |
NAACL (Long Papers) | 5 |
| 2025 | VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video UnderstandingabstractVision Language models (VLMs) have achieved remarkable success in video understanding tasks. Yet, a key question remains: Do they comprehend visual information or merely learn superficial mappings between visual and textual patterns?
Understanding visual cues, particularly those related to physics and common sense, is crucial for AI systems interacting with the physical world. However, existing VLM evaluations primarily rely on positive-control tests using real-world videos that resemble training distributions. While VLMs perform well on such benchmarks, it is unclear whether they grasp underlying visual and contextual signals or simply exploit visual-language correlations. To fill this gap, we propose incorporating negative-control tests, i.e., videos depicting physically impossible or logically inconsistent scenarios, and evaluating whether models can recognize these violations. True visual understanding should evince comparable performance across both positive and negative tests. Since such content is rare in the real world, we introduce VideoHallu, a synthetic video dataset featuring physics- and commonsense-violating scenes generated using state-of-the-art tools such as Veo2, Sora, and Kling. The dataset includes expert-annotated question-answer pairs spanning four categories of physical and commonsense violations, designed to be straightforward for human reasoning. We evaluate several leading VLMs, including Qwen-2.5-VL, Video-R1, and VideoChat-R1. Despite their strong performance on real-world benchmarks (e.g., MVBench, MMVU), these models hallucinate or fail to detect physical or logical violations, revealing fundamental weaknesses in visual understanding. Finally, we explore reinforcement learning-based post-training on our negative dataset: fine-tuning improves performance on VideoHallu without degrading results on standard benchmarks, indicating enhanced visual reasoning in VLMs. Our data is available at https://github.com/zli12321/VideoHallu.git. Zongxia Li, Xiyang Wu, Guangyao Shi, Yubin Qin, Hongyang Du 0002, Tianyi Zhou 0001, Dinesh Manocha, Jordan L. Boyd-Graber |
NeurIPS | 8 |
| 2024 | More Victories, Less Cooperation: Assessing Cicero's Diplomacy PlayabstractWichayaporn Wongkamjan, Feng Gu, Yanze Wang, Ulf Hermjakob, Jonathan May, Brandon M. Stewart, Jonathan K. Kummerfeld, Denis Peskoff, Jordan Lee Boyd-Graber. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Wichayaporn Wongkamjan, Ulf Hermjakob, Jonathan May, Brandon M. Stewart, Jonathan K. Kummerfeld, Denis Peskoff, Jordan L. Boyd-Graber |
ACL (1) | 9 |
| 2024 | Rapidly Piloting Real-time Linguistic Assistance for Simultaneous Interpreters with Untrained Bilingual SurrogatesabstractSimultaneous interpretation is a cognitively taxing task, and even seasoned professionals benefit from real-time assistance. However, both recruiting professional interpreters and evaluating new assistance techniques are difficult. We present a novel, realistic simultaneous interpretation task that mimics the cognitive load of interpretation with crowdworker surrogates. Our task tests different real-time assistance methods in a Wizard-of-Oz experiment with a large pool of proxy users and compares against professional interpreters. Both professional and proxy participants respond similarly to changes in interpreting conditions, including improvement with two assistance interventions—translation of specific terms and of numbers—compared to a no-assistance control. Alvin Grissom II, Jo Shoemaker, Benjamin Goldman, Ruikang Shi, Craig Stewart, C. Anton Rytting, Leah Findlater, Jordan L. Boyd-Graber |
LREC/COLING | 8 |
| 2024 | Improving the TENOR of Labeling: Re-evaluating Topic Models for Content AnalysisabstractZongxia Li, Andrew Mao, Daniel Stephens, Pranav Goel, Emily Walpole, Alden Dima, Juan Fung, Jordan Boyd-Graber. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zongxia Li, Andrew Mao, Daniel Kofi Stephens, Pranav Goel 0001, Emily Walpole, Alden Dima, Juan Fung, Jordan L. Boyd-Graber |
EACL (1) | 8 |
| 2024 | Presentations by the Humans and For the Humans: Harnessing LLMs for Generating Persona-Aware Slides from DocumentsabstractIshani Mondal, Shwetha S, Anandhavelu Natarajan, Aparna Garimella, Sambaran Bandyopadhyay, Jordan Boyd-Graber. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Ishani Mondal, Shwetha S, Anandhavelu Natarajan, Aparna Garimella, Sambaran Bandyopadhyay, Jordan L. Boyd-Graber |
EACL (1) | 6 |
| 2024 | A SMART Mnemonic Sounds like "Glue Tonic": Mixing LLMs with Student Feedback to Make Mnemonic Learning StickabstractNishant Balepur, Matthew Shu, Alexander Hoyle, Alison Robey, Shi Feng, Seraphina Goldfarb-Tarrant, Jordan Lee Boyd-Graber. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Nishant Balepur, Matthew Shu, Alexander Miserlis Hoyle, Alison Robey, Shi Feng 0005, Seraphina Goldfarb-Tarrant, Jordan L. Boyd-Graber |
EMNLP | 7 |
| 2024 | Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRAabstractRecent advancements of large language models (LLMs) have led to claims of AI surpassing humans in natural language processing (NLP) tasks such as textual understanding and reasoning.This work investigates these assertions by introducing CAIMIRA, a novel framework rooted in item response theory (IRT) that enables quantitative assessment and comparison of problem-solving abilities in questionanswering (QA) agents.Through analysis of over 300,000 responses from ~70 AI systems and 155 humans across thousands of quiz questions, CAIMIRA uncovers distinct proficiency patterns in knowledge domains and reasoning skills.Humans outperform AI systems in knowledge-grounded abductive and conceptual reasoning, while state-of-the-art LLMs like GPT-4-TURBO and LLAMA-3-70B demonstrate superior performance on targeted information retrieval and fact-based reasoning, particularly when information gaps are well-defined and addressable through pattern matching or data retrieval.These findings identify key areas for future QA tasks and model development, highlighting the critical need for questions that not only challenge higher-order reasoning and scientific thinking, but also demand nuanced linguistic and cross-contextual application. Maharshi Gor, Hal Daumé III, Tianyi Zhou 0001, Jordan L. Boyd-Graber |
EMNLP | 4 |
| 2024 | You Make me Feel like a Natural Question: Training QA Systems on Transformed Trivia QuestionsabstractTraining question answering (QA) and information retrieval systems for web queries require large, expensive datasets that are difficult to annotate and time-consuming to gather.Moreover, while natural datasets of informationseeking questions are often prone to ambiguity or ill-formed, there are troves of freely available, carefully crafted question datasets for many languages.Thus, we automatically generate shorter, information-seeking questions, resembling web queries in the style of the Natural Questions (NQ) dataset from longer trivia data.Training a QA system on these transformed questions is a viable strategy for alternating to more expensive training setups showing the F1 score difference of less than six points and contrasting the final systems. 1 We also combine NQ with QB-TRANS as training data in our supervised setting (Section 5), improving F1 (tested on NQ test set) by 10 points compared to training on only NQ.QB-TRANS lacks issues that plague NQ: presupposition and ambiguity (Section 6).Moreover, NATURALIZATION generalizes to other datasets (Section 6.4).Our contributions are naturalizing Manchester QB questions into Cranfield QB-TRANS while retaining the positive traits of QB samples, thereby improving QA with a more affordable process.The dataset generated from NATURALIZATION can be used to answer non-NQ data (Section 6.5) which proves the generalization of NATURALIZATION.Section 8 shows how this can ensure a cheaper and more up-to-date alternative to NQ data by generating large-scale Cranfield dataset that benefits training questionanswering models and generalizes to other non-NQ datasets. Artful but Arcane QB datasetThis section discusses why we use QB data and how different they are from NQ questions.The next section explains NATURALIZATION (Section 3).Elicitations from QB dataset Consider this QB example: Tasnim Kabir, Yoo Yeon Sung, Saptarashmi Bandyopadhyay, Abhranil Chandra, Jordan L. Boyd-Graber |
EMNLP | 6 |
| 2024 | KARL: Knowledge-Aware Retrieval and Representations aid Retention and Learning in StudentsabstractFlashcard schedulers rely on 1) student models to predict the flashcards a student knows; and 2) teaching policies to pick which cards to show next via these predictions.Prior student models, however, just use study data like the student's past responses, ignoring the text on cards.We propose content-aware scheduling, the first schedulers exploiting flashcard content.To give the first evidence that such schedulers enhance student learning, we build KAR 3 L, a simple but effective content-aware student model employing deep knowledge tracing (DKT), retrieval, and BERT to predict student recall.We train KAR 3 L by collecting a new dataset of 123,143 study logs on diverse trivia questions.KAR 3 L bests existing student models in AUC and calibration error.To ensure our improved predictions lead to better student learning, we create a novel delta-based teaching policy to deploy KAR 3 L online.Based on 32 study paths from 27 users, KAR 3 L improves learning efficiency over SOTA, showing KAR 3 L's strength and encouraging researchers to look beyond historical study data to fully capture student abilities.1 * Equal contribution. Matthew Shu, Nishant Balepur, Shi Feng 0005, Jordan L. Boyd-Graber |
EMNLP | 4 |
| 2024 | Large Language Models Help Humans Verify Truthfulness - Except When They Are Convincingly WrongabstractChenglei Si, Navita Goyal, Tongshuang Wu, Chen Zhao, Shi Feng, Hal Daumé Iii, Jordan Boyd-Graber. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Chenglei Si, Navita Goyal, Sherry Tongshuang Wu, Chen Zhao 0013, Shi Feng 0005, Hal Daumé III, Jordan L. Boyd-Graber |
NAACL-HLT | 7 |
| 2024 | Pregnant Questions: The Importance of Pragmatic Awareness in Maternal Health Question AnsweringabstractNeha Srikanth, Rupak Sarkar, Heran Mane, Elizabeth Aparicio, Quynh Nguyen, Rachel Rudinger, Jordan Boyd-Graber. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Neha Srikanth, Rupak Sarkar, Heran Mane, Elizabeth Aparicio, Quynh C. Nguyen, Rachel Rudinger, Jordan L. Boyd-Graber |
NAACL-HLT | 7 |
| 2023 | Bridging Background Knowledge Gaps in Translation with Automatic ExplicitationabstractTranslations help people understand content written in another language.However, even correct literal translations do not fulfill that goal when people lack the necessary background to understand them.Professional translators incorporate explicitations to explain the missing context by considering cultural differences between source and target audiences.Despite its potential to help users, NLP research on explicitation is limited because of the dearth of adequate evaluation methods.This work introduces techniques for automatically generating explicitations, motivated by WIKIEXPL 1 : a dataset that we collect from Wikipedia and annotate with human translators.The resulting explicitations are useful as they help answer questions more accurately in a multilingual question answering framework. HyoJung Han 0001, Jordan L. Boyd-Graber, Marine Carpuat |
EMNLP | 2 |
| 2023 | Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs Through a Global Prompt Hacking CompetitionabstractSander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-François Bouchard, Chenglei Si, Svetlina Anati, Valen Tagliabue, Anson Kost, Christopher Carnahan, Jordan Boyd-Graber. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Sander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-François Bouchard, Chenglei Si, Svetlina Anati, Valen Tagliabue, Anson Liu Kost, Christopher Carnahan, Jordan L. Boyd-Graber |
EMNLP | 10 |
| 2023 | Not all Fake News is Written: A Dataset and Analysis of Misleading Video HeadlinesabstractPolarization and the marketplace for impressions have conspired to make navigating information online difficult for users, and while there has been a significant effort to detect false or misleading text, multimodal datasets have received considerably less attention.To complement existing resources, we present multimodal Video Misleading Headline (VMH), a dataset that consists of videos and whether annotators believe the headline is representative of the video's contents.After collecting and annotating this dataset, we analyze multimodal baselines for detecting misleading headlines.Our annotation process also focuses on why annotators view a video as misleading, allowing us to better understand the interplay of annotators' background and the content of the videos. Yoo Yeon Sung, Jordan L. Boyd-Graber, Naeemul Hassan |
EMNLP | 2 |
| 2023 | Prompting GPT-3 To Be Reliable
Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jordan L. Boyd-Graber |
ICLR | 6 |
| 2022 | Match the Script, Adapt if Multilingual: Analyzing the Effect of Multilingual Pretraining on Cross-lingual TransferabilityabstractPretrained multilingual models enable zeroshot learning even for unseen languages, and that performance can be further improved via adaptation prior to finetuning.However, it is unclear how the number of pretraining languages influences a model's zero-shot learning for languages unseen during pretraining.To fill this gap, we ask the following research questions: (1) How does the number of pretraining languages influence zero-shot performance on unseen target languages?( 2) Does the answer to that question change with model adaptation?(3) Do the findings for our first question change if the languages used for pretraining are all related?Our experiments on pretraining with related languages indicate that choosing a diverse set of languages is crucial.Without model adaptation, surprisingly, increasing the number of pretraining languages yields better results up to adding related languages, after which performance plateaus.In contrast, with model adaptation via continued pretraining, pretraining on a larger number of languages often gives further improvement, suggesting that model adaptation is crucial to exploit additional pretraining languages.1 Yoshinari Fujinuma, Jordan L. Boyd-Graber, Katharina Kann |
ACL (1) | 2 |
| 2022 | Adapting Coreference Resolution Models through Active LearningabstractNeural coreference resolution models trained on one dataset may not transfer to new, lowresource domains.Active learning mitigates this problem by sampling a small subset of data for annotators to label.While active learning is well-defined for classification tasks, its application to coreference resolution is neither well-defined nor fully understood.This paper explores how to actively label coreference, examining sources of model uncertainty and document reading costs.We compare uncertainty sampling strategies and their advantages through thorough error analysis.In both synthetic and human experiments, labeling spans within the same document is more effective than annotating spans across documents.The findings contribute to a more realistic development of coreference resolution models. Michelle Yuan, Patrick Xia 0002, Chandler May, Benjamin Van Durme, Jordan L. Boyd-Graber |
ACL (1) | 5 |
| 2022 | Learning to Explain Selectively: A Case Study on Question AnsweringabstractExplanations promise to bridge the gap between humans and AI, yet it remains difficult to achieve consistent improvement in AIaugmented human decision making.The usefulness of AI explanations depends on many factors, and always showing the same type of explanation in all cases is suboptimal-so is relying on heuristics to adapt explanations for each scenario.We propose learning to explain "selectively": for each decision that the user makes, we use a model to choose the best explanation from a set of candidates and update this model with feedback to optimize human performance.We experiment on a question answering task, Quizbowl, and show that selective explanations improve human performance for both experts and crowdworkers. Shi Feng 0005, Jordan L. Boyd-Graber |
EMNLP | 2 |
| 2022 | SimQA: Detecting Simultaneous MT Errors through Word-by-Word Question AnsweringabstractDetractors of neural machine translation admit that while its translations are fluent, it sometimes gets key facts wrong.This is particularly important in simultaneous interpretation where translations have to be provided as fast as possible: before a sentence is complete.Yet, evaluations of simultaneous machine translation (SIMULMT) fail to capture if systems correctly translate the most salient elements of a question: people, places, and dates.To address this problem, we introduce a downstream word-by-word question answering evaluation task (SIMQA): given a source language question, translate the question word by word into the target language, and answer as soon as possible.SIMQA jointly measures whether the SIMULMT models translate the question quickly and accurately, and can reveal shortcomings in existing neural systemshallucinating or omitting facts. HyoJung Han 0001, Marine Carpuat, Jordan L. Boyd-Graber |
EMNLP | 3 |
| 2021 | Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?abstractPedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor, Robin Jia, Jordan Boyd-Graber. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Pedro Rodríguez 0001, Joe Barrow, Alexander Miserlis Hoyle, John Lalor, Robin Jia, Jordan L. Boyd-Graber |
ACL/IJCNLP (1) | 6 |
| 2021 | Toward Deconfounding the Effect of Entity Demographics for Question Answering AccuracyabstractThe goal of question answering (QA) is to answer any question.However, major QA datasets have skewed distributions over gender, profession, and nationality.Despite that skew, model accuracy analysis reveals little evidence that accuracy is lower for people based on gender or nationality; instead, there is more variation on professions (question topic).But QA's lack of representation could itself hide evidence of bias, necessitating QA datasets that better represent global diversity.1 For NQ, we only consider questions with short answers. 2 https://cloud.google.com/natural-language/docs/ analyzing-entities 3 We analyze the dev fold, which is consistent with the training fold (Table 1 and 2), as we examine accuracy. Maharshi Gor, Kellie Webster, Jordan L. Boyd-Graber |
EMNLP (1) | 3 |
| 2021 | Evaluation Paradigms in Question AnsweringabstractQuestion answering (QA) primarily descends from two branches of research: (1) Alan Turing's investigation of machine intelligence at Manchester University and (2) Cyril Cleverdon's comparison of library card catalog indices at Cranfield University.This position paper names and distinguishes these paradigms.Despite substantial overlap, subtle but significant distinctions exert an outsize influence on research.While one evaluation paradigm values creating more intelligent QA systems, the other paradigm values building QA systems that appeal to users.By better understanding the epistemic heritage of QA, researchers, academia, and industry can more effectively accelerate QA research. Pedro Rodríguez 0001, Jordan L. Boyd-Graber |
EMNLP (1) | 2 |
| 2021 | What's in a Name? Answer Equivalence For Open-Domain Question AnsweringabstractA flaw in QA evaluation is that annotations often only provide one gold answer.Thus, model predictions semantically equivalent to the answer but superficially different are considered incorrect.This work explores mining alias entities from knowledge bases and using them as additional gold answers (i.e., equivalent answers).We incorporate answers for two settings: evaluation with additional answers and model training with equivalent answers.We analyse three QA benchmarks: Natural Questions, TriviaQA and SQuAD.Answer expansion increases the exact match score on all datasets for evaluation, while incorporating it helps model training over real-world datasets.We ensure the additional answers are valid through a human post hoc evaluation. 1 Chenglei Si, Chen Zhao 0013, Jordan L. Boyd-Graber |
EMNLP (1) | 3 |
| 2021 | Distantly-Supervised Dense Retrieval Enables Open-Domain Question Answering without Evidence AnnotationabstractOpen-domain question answering answers a question based on evidence retrieved from a large corpus.State-of-the-art neural approaches require intermediate evidence annotations for training.However, such intermediate annotations are expensive, and methods that rely on them cannot transfer to the more common setting, where only questionanswer pairs are available.This paper investigates whether models can learn to find evidence from a large corpus, with only distant supervision from answer labels for model training, thereby generating no additional annotation cost.We introduce a novel approach (DISTDR) that iteratively improves over a weak retriever by alternately finding evidence from the up-to-date model and encouraging the model to learn the most likely evidence.Without using any evidence labels, DISTDR is on par with fully-supervised state-of-theart methods on both multi-hop and singlehop QA benchmarks.Our analysis confirms that DISTDR finds more accurate evidence over iterations, which leads to model improvements.The code is available at https:// github.com/henryzhao5852/DistDR. Chen Zhao 0013, Chenyan Xiong, Jordan L. Boyd-Graber, Hal Daumé III |
EMNLP (1) | 3 |
| 2021 | Fool Me Twice: Entailment from Wikipedia GamificationabstractJulian Eisenschlos, Bhuwan Dhingra, Jannis Bulian, Benjamin Börschinger, Jordan Boyd-Graber. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Julian Martin Eisenschlos, Bhuwan Dhingra, Jannis Bulian, Benjamin Börschinger, Jordan L. Boyd-Graber |
NAACL-HLT | 5 |
| 2021 | Multi-Step Reasoning Over Unstructured Text with Beam Dense RetrievalabstractChen Zhao, Chenyan Xiong, Jordan Boyd-Graber, Hal Daumé III. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Chen Zhao 0013, Chenyan Xiong, Jordan L. Boyd-Graber, Hal Daumé III |
NAACL-HLT | 3 |
| 2021 | Is Automated Topic Model Evaluation Broken? The Incoherence of CoherenceabstractTopic model evaluation, like evaluation of other unsupervised methods, can be contentious. However, the field has coalesced around automated estimates of topic coherence, which rely on the frequency of word co-occurrences in a reference corpus. Contemporary neural topic models surpass classical ones according to these metrics. At the same time, topic model evaluation suffers from a validation gap: automated coherence, developed for classical models, has not been validated using human experimentation for neural models. In addition, a meta-analysis of topic modeling literature reveals a substantial standardization gap in automated topic modeling benchmarks. To address the validation gap, we compare automated coherence with the two most widely accepted human judgment tasks: topic rating and word intrusion. To address the standardization gap, we systematically evaluate a dominant classical model and two state-of-the-art neural models on two commonly used datasets. Automated evaluations declare a winning model when corresponding human evaluations do not, calling into question the validity of fully automatic evaluations independent of human judgments. Alexander Miserlis Hoyle, Pranav Goel 0001, Andrew Hian-Cheong, Denis Peskov, Jordan L. Boyd-Graber, Philip Resnik |
NeurIPS | 5 |
| 2020 | Exploiting Cross-Lingual Subword Similarities in Low-Resource Document ClassificationabstractText classification must sometimes be applied in a low-resource language with no labeled training data. However, training data may be available in a related language. We investigate whether character-level knowledge transfer from a related language helps text classification. We present a cross-lingual document classification framework (caco) that exploits cross-lingual subword similarity by jointly training a character-based embedder and a word-based classifier. The embedder derives vector representations for input words from their written forms, and the classifier makes predictions based on the word vectors. We use a joint character representation for both the source language and the target language, which allows the embedder to generalize knowledge about source language words to target language words with similar forms. We propose a multi-task objective that can further improve the model if additional cross-lingual or monolingual resources are available. Experiments confirm that character-level knowledge transfer is more data-efficient than word-level transfer between related languages. Mozhi Zhang, Yoshinari Fujinuma, Jordan L. Boyd-Graber |
AAAI | 3 |
| 2020 | What Question Answering can Learn from Trivia NerdsabstractQuestion answering (QA)is not just building systems; this NLP subfield also creates and curates challenging question datasets that reveal the best systems.We argue that QA datasets-and QA leaderboards-closely resemble trivia tournaments: the questions agents-humans or machines-answer reveals a "winner".However, the research community has ignored the lessons from decades of the trivia community creating vibrant, fair, and effective QA competitions.After detailing problems with existing QA datasets, we outline several lessons that transfer to QA research: removing ambiguity, identifying better QA agents, and adjudicating disputes. Jordan L. Boyd-Graber, Benjamin Börschinger |
ACL | 1 |
| 2020 | It Takes Two to Lie: One to Lie, and One to ListenabstractDenis Peskov, Benny Cheng, Ahmed Elgohary, Joe Barrow, Cristian Danescu-Niculescu-Mizil, Jordan Boyd-Graber. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Denis Peskov, Benny Cheng, Ahmed Elgohary, Joe Barrow, Cristian Danescu-Niculescu-Mizil, Jordan L. Boyd-Graber |
ACL | 6 |
| 2020 | Why Overfitting Isn't Always Bad: Retrofitting Cross-Lingual Word Embeddings to DictionariesabstractCross-lingual word embeddings (CLWE) are often evaluated on bilingual lexicon induction (BLI).Recent CLWE methods use linear projections, which underfit the training dictionary, to generalize on BLI.However, underfitting can hinder generalization to other downstream tasks that rely on words from the training dictionary.We address this limitation by retrofitting CLWE to the training dictionary, which pulls training translation pairs closer in the embedding space and overfits the training dictionary.This simple post-processing step often improves accuracy on two downstream tasks, despite lowering BLI test accuracy.We also retrofit to both the training dictionary and a synthetic dictionary induced from CLWE, which sometimes generalizes even better on downstream tasks.Our results confirm the importance of fully exploiting the training dictionary in downstream tasks and explains why BLI is a flawed CLWE evaluation. Mozhi Zhang, Yoshinari Fujinuma, Michael J. Paul, Jordan L. Boyd-Graber |
ACL | 4 |
| 2020 | No Explainability without Accountability: An Empirical Study of Explanations and Feedback in Interactive MLabstractAutomatically generated explanations of how machine learning (ML) models reason can help users understand and accept them. However, explanations can have unintended consequences: promoting over-reliance or undermining trust. This paper investigates how explanations shape users' perceptions of ML models with or without the ability to provide feedback to them: (1) does revealing model flaws increase users' desire to "fix" them; (2) does providing explanations cause users to believe - wrongly - that models are introspective, and will thus improve over time. Through two controlled experiments - varying model quality - we show how the combination of explanations and user feedback impacted perceptions, such as frustration and expectations of model improvement. Explanations without opportunity for feedback were frustrating with a lower quality model, while interactions between explanation and feedback for the higher quality model suggest that detailed feedback should not be requested without explanation. Users expected model correction, regardless of whether they provided feedback or received explanations. Alison Smith-Renner, Ron Fan, Melissa Birchfield, Sherry Tongshuang Wu, Jordan L. Boyd-Graber, Daniel S. Weld, Leah Findlater |
CHI | 5 |
| 2020 | Cold-start Active Learning through Self-supervised Language ModelingabstractActive learning strives to reduce annotation costs by choosing the most critical examples to label.Typically, the active learning strategy is contingent on the classification model.For instance, uncertainty sampling depends on poorly calibrated model confidence scores.In the cold-start setting, active learning is impractical because of model instability and data scarcity.Fortunately, modern NLP provides an additional source of information: pretrained language models.The pre-training loss can find examples that surprise the model and should be labeled for efficient fine-tuning.Therefore, we treat the language modeling loss as a proxy for classification uncertainty.With BERT, we develop a simple strategy based on the masked language modeling loss that minimizes labeling costs for text classification.Compared to other baselines, our approach reaches higher accuracy within less sampling iterations and computation time. Michelle Yuan, Hsuan-Tien Lin, Jordan L. Boyd-Graber |
EMNLP (1) | 3 |
| 2020 | Interactive Refinement of Cross-Lingual Word EmbeddingsabstractCross-lingual word embeddings transfer knowledge between languages: models trained on high-resource languages can predict in low-resource languages.We introduce CLIME, an interactive system to quickly refine cross-lingual word embeddings for a given classification problem.First, CLIME ranks words by their salience to the downstream task.Then, users mark similarity between keywords and their nearest neighbors in the embedding space.Finally, CLIME updates the embeddings using the annotations.We evaluate CLIME on identifying health-related text in four low-resource languages: Ilocano, Sinhalese, Tigrinya, and Uyghur.Embeddings refined by CLIME capture more nuanced word semantics and have higher test accuracy than the original embeddings.CLIME often improves accuracy faster than an active learning baseline and can be easily combined with active learning to improve results. Michelle Yuan, Mozhi Zhang, Benjamin Van Durme, Leah Findlater, Jordan L. Boyd-Graber |
EMNLP (1) | 5 |
| 2020 | Digging into user control: perceptions of adherence and instability in transparent modelsabstractWe explore predictability and control in interactive systems where controls are easy to validate. Human-in-the-loop techniques allow users to guide unsupervised algorithms by exposing and supporting interaction with underlying model representations, increasing transparency and promising fine-grained control. However, these models must balance user input and the underlying data, meaning they sometimes update slowly, poorly, or unpredictably---either by not incorporating user input as expected (adherence) or by making other unexpected changes (instability). While prior work exposes model internals and supports user feedback, less attention has been paid to users' reactions when transparent models limit control. Focusing on interactive topic models, we explore user perceptions of control using a study where 100 participants organize documents with one of three distinct topic modeling approaches. These approaches incorporate input differently, resulting in varied adherence, stability, update speeds, and model quality. Participants disliked slow updates most, followed by lack of adherence. Instability was polarizing: some participants liked it when it surfaced interesting information, while others did not. Across modeling approaches, participants differed only in whether they noticed adherence. Alison Smith-Renner, Jordan L. Boyd-Graber, Kevin D. Seppi, Leah Findlater |
IUI | 3 |
| 2020 | Which Evaluations Uncover Sense Representations that Actually Make Sense?abstractText representations are critical for modern natural language processing. One form of text representation, sense-specific embeddings, reflect a word’s sense in a sentence better than single-prototype word embeddings tied to each type. However, existing sense representations are not uniformly better: although they work well for computer-centric evaluations, they fail for human-centric tasks like inspecting a language’s sense inventory. To expose this discrepancy, we propose a new coherence evaluation for sense embeddings. We also describe a minimal model (Gumbel Attention for Sense Induction) optimized for discovering interpretable sense representations that are more coherent than existing sense embeddings. Jordan L. Boyd-Graber, Fenfei Guo, Leah Findlater, Mohit Iyyer |
LREC | 1 |
| 2020 | Complex Factoid Question Answering with a Free-Text Knowledge GraphabstractWe introduce delft, a factoid question answering system which combines the nuance and depth of knowledge graph question answering approaches with the broader coverage of free-text. delft builds a free-text knowledge graph from Wikipedia, with entities as nodes and sentences in which entities co-occur as edges. For each question, delft finds the subgraph linking question entity nodes to candidates using text sentences as edges, creating a dense and high coverage semantic graph. A novel graph neural network reasons over the free-text graph—combining evidence on the nodes via information along edge sentences—to select a final answer. Experiments on three question answering datasets show delft can answer entity-rich questions better than machine reading based models, bert-based answer ranking and memory networks. delft’s advantage comes from both the high coverage of its free-text knowledge graph—more than double that of dbpedia relations—and the novel graph neural network which reasons on the rich but noisy free-text evidence. Chen Zhao 0013, Chenyan Xiong, Jordan L. Boyd-Graber |
WWW | 4 |
| 2019 | Misleading Failures of Partial-input BaselinesabstractRecent work establishes dataset difficulty and removes annotation artifacts via partial-input baselines (e.g., hypothesis-only model for SNLI or question-only model for VQA). A successful partial-input baseline indicates that the dataset is cheatable. But the converse is not necessarily true: failures of partial-input baselines do not mean the dataset is free of artifacts. We first design artificial datasets to illustrate how the trivial patterns that are only visible in the full input can evade any partial-input baseline. Next, we identify such artifacts in the SNLI dataset—a hypothesis-only model augmented with trivial patterns in the premise can solve 15% of previously-thought “hard” examples. Our work provides a caveat for the use and creation of partial-input baselines for datasets. Shi Feng 0005, Eric Wallace, Jordan L. Boyd-Graber |
ACL (1) | 3 |
| 2019 | A Resource-Free Evaluation Metric for Cross-Lingual Word Embeddings Based on Graph ModularityabstractCross-lingual word embeddings encode the meaning of words from different languages into a shared low-dimensional space. An important requirement for many downstream tasks is that word similarity should be independent of language - i.e., word vectors within one language should not be more similar to each other than to words in another language. We measure this characteristic using modularity, a network measurement that measures the strength of clusters in a graph. Modularity has a moderate to strong correlation with three downstream tasks, even though modularity is based only on the structure of embeddings and does not require any external resources. We show through experiments that modularity can serve as an intrinsic validation metric to improve unsupervised cross-lingual word embeddings, particularly on distant language pairs in low-resource settings. Yoshinari Fujinuma, Jordan L. Boyd-Graber, Michael J. Paul |
ACL (1) | 2 |
| 2019 | Why Didn't You Listen to Me? Comparing User Control of Human-in-the-Loop Topic ModelsabstractTo address the lack of comparative evaluation of Human-in-the-Loop Topic Modeling (HLTM) systems, we implement and evaluate three contrasting HLTM modeling approaches using simulation experiments.These approaches extend previously proposed frameworks, including constraints and informed prior-based methods.Users should have a sense of control in HLTM systems, so we propose a control metric to measure whether refinement operations' results match users' expectations.Informed prior-based methods provide better control than constraints, but constraints yield higher quality topics. Alison Smith-Renner, Leah Findlater, Kevin D. Seppi, Jordan L. Boyd-Graber |
ACL (1) | 5 |
| 2019 | Automatic Evaluation of Local Topic QualityabstractJeffrey Lund, Piper Armstrong, Wilson Fearn, Stephen Cowley, Courtni Byun, Jordan Boyd-Graber, Kevin Seppi. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Jeffrey Lund, Piper Armstrong, Wilson Fearn, Stephen Cowley, Courtni Byun, Jordan L. Boyd-Graber, Kevin D. Seppi |
ACL (1) | 6 |
| 2019 | Are Girls Neko or Shōjo? Cross-Lingual Alignment of Non-Isomorphic Embeddings with Iterative NormalizationabstractCross-lingual word embeddings (CLWE) underlie many multilingual natural language processing systems, often through orthogonal transformations of pre-trained monolingual embeddings.However, orthogonal mapping only works on language pairs whose embeddings are naturally isomorphic.For nonisomorphic pairs, our method (Iterative Normalization) transforms monolingual embeddings to make orthogonal alignment easier by simultaneously enforcing that (1) individual word vectors are unit length, and (2) each language's average vector is zero.Iterative Normalization consistently improves word translation accuracy of three CLWE methods, with the largest improvement observed on English-Japanese (from 2% to 44% test accuracy). Mozhi Zhang, Keyulu Xu, Ken-ichi Kawarabayashi, Stefanie Jegelka, Jordan L. Boyd-Graber |
ACL (1) | 5 |
| 2019 | Can You Unpack That? Learning to Rewrite Questions-in-ContextabstractAhmed Elgohary, Denis Peskov, Jordan Boyd-Graber. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ahmed Elgohary, Denis Peskov, Jordan L. Boyd-Graber |
EMNLP/IJCNLP (1) | 3 |
| 2019 | A Multilingual Topic Model for Learning Weighted Topic Links Across Corpora with Low ComparabilityabstractWeiwei Yang, Jordan Boyd-Graber, Philip Resnik. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jordan L. Boyd-Graber, Philip Resnik |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Mitigating Noisy Inputs for Question AnsweringabstractNatural language processing systems are often downstream of unreliable inputs: machine translation, optical character recognition, or speech recognition. For instance, virtual assistants can only answer your questions after understanding your speech. We investigate and mitigate the effects of noise from Automatic Speech Recognition systems on two factoid Question Answering (QA) tasks. Integrating confidences into the model and forced decoding of unknown words are empirically shown to improve the accuracy of downstream neural QA systems. We create and train models on a synthetic corpus of over 500,000 noisy sentences and evaluate on two human corpora from Quizbowl and Jeopardy! competitions. Denis Peskov, Joe Barrow, Pedro Rodríguez 0001, Graham Neubig, Jordan L. Boyd-Graber |
INTERSPEECH | 5 |
| 2019 | What can AI do for me?: evaluating machine learning interpretations in cooperative playabstractMachine learning is an important tool for decision making, but its ethical and responsible application requires rigorous vetting of its interpretability and utility: an understudied problem, particularly for natural language processing models. We propose an evaluation of interpretation on a real task with real human users, where the effectiveness of interpretation is measured by how much it improves human performance. We design a grounded, realistic human-computer cooperative setting using a question answering task, Quizbowl. We recruit both trivia experts and novices to play this game with computer as their teammate, who communicates its prediction via three different interpretations. We also provide design guidance for natural language processing human-in-the-loop settings. Shi Feng 0005, Jordan L. Boyd-Graber |
IUI | 2 |
| 2019 | Trick Me If You Can: Human-in-the-loop Generation of Adversarial Question Answering ExamplesabstractAdversarial evaluation stress-tests a model’s understanding of natural language. Because past approaches expose superficial patterns, the resulting adversarial examples are limited in complexity and diversity. We propose human- in-the-loop adversarial generation, where human authors are guided to break models. We aid the authors with interpretations of model predictions through an interactive user interface. We apply this generation framework to a question answering task called Quizbowl, where trivia enthusiasts craft adversarial questions. The resulting questions are validated via live human–computer matches: Although the questions appear ordinary to humans, they systematically stump neural and information retrieval models. The adversarial questions cover diverse phenomena from multi-hop reasoning to entity type distractors, exposing open challenges in robust question answering. Eric Wallace, Pedro Rodríguez 0001, Shi Feng 0005, Ikuya Yamada, Jordan L. Boyd-Graber |
Trans. Assoc. Comput. Linguistics | 5 |
| 2018 | Learning from Measurements in Crowdsourcing Models: Inferring Ground Truth from Diverse Annotation TypesabstractAnnotated corpora enable supervised machine learning and data analysis. To reduce the cost of manual annotation, tasks are often assigned to internet workers whose judgments are reconciled by crowdsourcing models. We approach the problem of crowdsourcing using a framework for learning from rich prior knowledge, and we identify a family of crowdsourcing models with the novel ability to combine annotations with differing structures: e.g., document labels and word labels. Annotator judgments are given in the form of the predicted expected value of measurement functions computed over annotations and the data, unifying annotation models. Our model, a specific instance of this framework, compares favorably with previous work. Furthermore, it enables active sample selection, jointly selecting annotator, data item, and annotation structure to reduce annotation effort. Paul Felt, Eric K. Ringger, Jordan L. Boyd-Graber, Kevin D. Seppi |
COLING | 3 |
| 2018 | A dataset and baselines for sequential open-domain question answeringabstractPrevious work on question-answering systems mainly focuses on answering individual questions, assuming they are independent and devoid of context.Instead, we investigate sequential question answering, asking multiple related questions.We present QBLink, a new dataset of fully human-authored questions.We extend existing strong question answering frameworks to include previous questions to improve the overall question-answering accuracy in open-domain question answering.The dataset is publicly available at http:// sequential.qanta.org. Ahmed Elgohary, Chen Zhao 0013, Jordan L. Boyd-Graber |
EMNLP | 3 |
| 2018 | Pathologies of Neural Models Make Interpretation DifficultabstractOne way to interpret neural model predictions is to highlight the most important input features-for example, a heatmap visualization over the words in an input sentence.In existing interpretation methods for NLP, a word's importance is determined by either input perturbation-measuring the decrease in model confidence when that word is removed-or by the gradient with respect to that word.To understand the limitations of these methods, we use input reduction, which iteratively removes the least important word from the input.This exposes pathological behaviors of neural models: the remaining words appear nonsensical to humans and are not the ones determined as important by interpretation methods.As we confirm with human experiments, the reduced examples lack information to support the prediction of any label, but models still make the same predictions with high confidence.To explain these counterintuitive results, we draw connections to adversarial examples and confidence calibration: pathological behaviors reveal difficulties in interpreting neural models trained with maximum likelihood.To mitigate their deficiencies, we fine-tune the models by encouraging high entropy outputs on reduced examples.Fine-tuned models become more interpretable under input reduction without accuracy loss on regular examples. Shi Feng 0005, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodríguez 0001, Jordan L. Boyd-Graber |
EMNLP | 6 |
| 2018 | Closing the Loop: User-Centered Design and Evaluation of a Human-in-the-Loop Topic Modeling SystemabstractHuman-in-the-loop topic modeling allows users to guide the creation of topic models and to improve model quality without having to be experts in topic modeling algorithms. Prior work in this area has focused either on algorithmic implementation without understanding how users actually wish to improve the model or on user needs but without the context of a fully interactive system. To address this disconnect, we implemented a set of model refinements requested by users in prior work and conducted a study with twelve non-expert participants to examine how end users are affected by issues that arise with a fully interactive, user-centered system. As these issues mirror those identified in interactive machine learning more broadly, such as unpredictability, latency, and trust, we also examined interactive machine learning challenges with non-expert end users through the lens of human-in-the-loop topic modeling. We found that although users experience unpredictability, their reactions vary from positive to negative, and, surprisingly, we did not find any cases of distrust, but instead noted instances where users perhaps trusted the system too much or had too little confidence in themselves. Alison Smith-Renner, Jordan L. Boyd-Graber, Kevin D. Seppi, Leah Findlater |
IUI | 3 |
| 2018 | Lessons from the Bible on Modern Topics: Low-Resource Multilingual Topic Model EvaluationabstractShudong Hao, Jordan Boyd-Graber, Michael J. Paul. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Shudong Hao, Jordan L. Boyd-Graber, Michael J. Paul |
NAACL-HLT | 2 |
| 2017 | Tandem Anchoring: a Multiword Anchor Approach for Interactive Topic ModelingabstractInteractive topic models are powerful tools for understanding large collections of text.However, existing sampling-based interactive topic modeling approaches scale poorly to large data sets.Anchor methods, which use a single word to uniquely identify a topic, offer the speed needed for interactive work but lack both a mechanism to inject prior knowledge and lack the intuitive semantics needed for userfacing applications.We propose combinations of words as anchors, going beyond existing single word anchor algorithmsan approach we call "Tandem Anchors".We begin with a synthetic investigation of this approach then apply the approach to interactive topic modeling in a user study and compare it to interactive and noninteractive approaches.Tandem anchors are faster and more intuitive than existing interactive approaches. Jeffrey Lund, Connor Cook, Kevin D. Seppi, Jordan L. Boyd-Graber |
ACL (1) | 4 |
| 2017 | The Amazing Mysteries of the Gutter: Drawing Inferences Between Panels in Comic Book NarrativesabstractVisual narrative is often a combination of explicit information and judicious omissions, relying on the viewer to supply missing details. In comics, most movements in time and space are hidden in the gutters between panels. To follow the story, readers logically connect panels together by inferring unseen actions through a process called closure. While computers can now describe the content of natural images, in this paper we examine whether they can understand the closure-driven narratives conveyed by stylized artwork and dialogue in comic book panels. We collect a dataset, COMICS, that consists of over 1.2 million panels (120 GB) paired with automatic textbox transcriptions. An in-depth analysis of COMICS demonstrates that neither text nor image alone can tell a comic book story, so a computer must understand both modalities to keep up with the plot. We introduce three cloze-style tasks that ask models to predict narrative and character-centric aspects of a panel given n preceding panels as context. Various deep neural architectures underperform human baselines on these tasks, suggesting that COMICS contains fundamental challenges for both vision and language. Mohit Iyyer, Varun Manjunatha, Anupam Guha, Yogarshi Vyas, Jordan L. Boyd-Graber, Hal Daumé III, Larry Davis 0001 |
CVPR | 5 |
| 2017 | Why ADAGRAD Fails for Online Topic ModelingabstractOnline topic modeling, i.e., topic modeling with stochastic variational inference, is a powerful and efficient technique for analyzing large datasets, and ADAGRAD is a widely-used technique for tuning learning rates during online gradient optimization.However, these two techniques do not work well together.We show that this is because ADAGRAD uses accumulation of previous gradients as the learning rates' denominators.For online topic modeling, the magnitude of gradients is very large.It causes learning rates to shrink very quickly, so the parameters cannot fully converge until the training ends. You Lu 0003, Jeffrey Lund, Jordan L. Boyd-Graber |
EMNLP | 3 |
| 2017 | Reinforcement Learning for Bandit Neural Machine Translation with Simulated Human FeedbackabstractMachine translation is a natural candidate problem for reinforcement learning from human feedback: users provide quick, dirty ratings on candidate translations to guide a system to improve.Yet, current neural machine translation training focuses on expensive human-generated reference translations.We describe a reinforcement learning algorithm that improves neural machine translation systems from simulated human feedback.Our algorithm combines the advantage actor-critic algorithm (Mnih et al., 2016) with the attention-based neural encoderdecoder architecture (Luong et al., 2015).This algorithm (a) is well-designed for problems with a large action space and delayed rewards, (b) effectively optimizes traditional corpus-level machine translation metrics, and (c) is robust to skewed, high-variance, granular feedback modeled after actual human behaviors. Hal Daumé III, Jordan L. Boyd-Graber |
EMNLP | 3 |
| 2017 | Adapting Topic Models using Lexical Associations with Tree PriorsabstractModels work best when they are optimized taking into account the evaluation criteria that people care about.For topic models, people often care about interpretability, which can be approximated using measures of lexical association.We integrate lexical association into topic optimization using tree priors, which provide a flexible framework that can take advantage of both first order word associations and the higher-order associations captured by word embeddings.Tree priors improve topic interpretability without hurting extrinsic performance. Jordan L. Boyd-Graber, Philip Resnik |
EMNLP | 2 |
| 2017 | The human touch: How non-expert users perceive, interpret, and fix topic modelsabstractTopic modeling is a common tool for understanding large bodies of text, but is typically provided as a “take it or leave it” proposition. Incorporating human knowledge in unsupervised learning is a promising approach to create high-quality topic models. Existing interactive systems and modeling algorithms support a wide range of refinement operations to express feedback. However, these systems’ interactions are primarily driven by algorithmic convenience, ignoring users who may lack expertise in topic modeling. To better understand how non-expert users understand, assess, and refine topics, we conducted two user studies—an in-person interview study and an online crowdsourced study. These studies demonstrate a disconnect between what non-expert users want and the complex, low-level operations that current interactive systems support . In particular, our findings include: (1) analysis of how non-expert users perceive topic models; (2) characterization of primary refinement operations expected by non-expert users and ordered by relative preference; (3) further evidence of the benefits of supporting users in directly refining a topic model; (4) design implications for future human-in-the-loop topic modeling interfaces. Tak Yeon Lee, Alison Smith-Renner, Kevin D. Seppi, Niklas Elmqvist, Jordan L. Boyd-Graber, Leah Findlater |
Int. J. Hum. Comput. Stud. | 5 |
| 2017 | Evaluating Visual Representations for Topic Understanding and Their Effects on Manually Generated LabelsabstractProbabilistic topic models are important tools for indexing, summarizing, and analyzing large document collections by their themes. However, promoting end-user understanding of topics remains an open research problem. We compare labels generated by users given four topic visualization techniques—word lists, word lists with bars, word clouds, and network graphs—against each other and against automatically generated labels. Our basis of comparison is participant ratings of how well labels describe documents from the topic. Our study has two phases: a labeling phase where participants label visualized topics and a validation phase where different participants select which labels best describe the topics’ documents. Although all visualizations produce similar quality labels, simple visualizations such as word lists allow participants to quickly understand topics, while complex visualizations take longer but expose multi-word expressions that simpler visualizations obscure. Automatic labels lag behind user-created labels, but our dataset of manually labeled topics highlights linguistic patterns (e.g., hypernyms, phrases) that can be used to improve automatic topic labeling algorithms. Alison Smith-Renner, Tak Yeon Lee, Forough Poursabzi-Sangdeh, Jordan L. Boyd-Graber, Niklas Elmqvist, Leah Findlater |
Trans. Assoc. Comput. Linguistics | 4 |
| 2016 | Learning Text Pair Similarity with Context-sensitive Autoencoders
Hadi Amiri, Philip Resnik, Jordan L. Boyd-Graber, Hal Daumé III |
ACL (1) | 3 |
| 2016 | ALTO: Active Learning with Topic Overviews for Speeding Label Induction and Document LabelingabstractEffective text classification requires experts to annotate data with labels; these training data are time-consuming and expensive to obtain.If you know what labels you want, active learning can reduce the number of labeled documents needed.However, establishing the label set remains difficult.Annotators often lack the global knowledge needed to induce a label set.We introduce ALTO: Active Learning with Topic Overviews, an interactive system to help humans annotate documents: topic models provide a global overview of what labels to create and active learning directs them to the right documents to label.Our forty-annotator user study shows that while active learning alone is best in extremely resource limited conditions, topic models (even by themselves) lead to better label sets, and ALTO's combination is best overall. Forough Poursabzi-Sangdeh, Jordan L. Boyd-Graber, Leah Findlater, Kevin D. Seppi |
ACL (1) | 2 |
| 2016 | A Discriminative Topic Model using Document Network StructureabstractDocument collections often have links between documents—citations, hyperlinks, or revisions—and which links are added is often based on topical similarity. To model these intuitions, we introduce a new topic model for documents situated within a network structure, integrating latent blocks of documents with a max-margin learning criterion for link prediction using topicand word-level features. Experiments on a scientific paper dataset and collection of webpages show that, by more robustly exploiting the rich link structure within a document network, our model improves link prediction, topic quality, and block distributions. Jordan L. Boyd-Graber, Philip Resnik |
ACL (1) | 2 |
| 2016 | Incremental Prediction of Sentence-final Verbs: Humans versus MachinesabstractVerb prediction is important in human sentence processing and, practically, in simultaneous machine translation.In verb-final languages, speakers select the final verb before it is uttered, and listeners predict it before it is uttered.Simultaneous interpreters must do the same to translate in real-time.Motivated by the problem of SOV-SVO simultaneous machine translation, we provide a study of incremental verb prediction in verb-final languages.As a basis of comparison, we examine incremental verb prediction with human participants in a multiple choice setting using crowdsourcing to gain insight into incremental human performance in a constrained setting.We then examine a computational approach to incremental verb prediction using discriminative classification with shallow features.Both humans and machines predict verbs more accurately as more of a sentence becomes available, and case markers-when available-help humans and sometimes machines predict final verbs. Alvin Grissom II, Naho Orita, Jordan L. Boyd-Graber |
CoNLL | 3 |
| 2016 | Opponent Modeling in Deep Reinforcement LearningabstractOpponent modeling is necessary in multi-agent settings where secondary agents with competing goals also adapt their strategies, yet it remains challenging because of strategies’ complex interaction and the non-stationary nature. Most previous work focuses on developing probabilistic models or parameterized strategies for specific applications. Inspired by the recent success of deep reinforcement learning, we present neural-based models that jointly learn a policy and the behavior of opponents. Instead of explicitly predicting the opponent’s action, we encode observation of the opponents into a deep Q-Network (DQN), while retaining explicit modeling under multitasking. By using a Mixture-of-Experts architecture, our model automatically discovers different strategy patterns of opponents even without extra supervision. We evaluate our models on a simulated soccer game and a popular trivia game, showing superior performance over DQN and its variants. He He 0001, Jordan L. Boyd-Graber |
ICML | 2 |
| 2016 | Interpretese vs. Translationese: The Uniqueness of Human Strategies in Simultaneous InterpretationabstractComputational approaches to simultaneous interpretation are stymied by how little we know about the tactics human interpreters use.We produce a parallel corpus of translated and simultaneously interpreted text and study differences between them through a computational approach.Our analysis reveals that human interpreters regularly apply several effective tactics to reduce translation latency, including sentence segmentation and passivization.In addition to these unique, clever strategies, we show that limited human memory also causes other idiosyncratic properties of human interpretation such as generalization and omission of source content. He He 0001, Jordan L. Boyd-Graber, Hal Daumé III |
HLT-NAACL | 2 |
| 2016 | Feuding Families and Former Friends: Unsupervised Learning for Dynamic Fictional RelationshipsabstractMohit Iyyer, Anupam Guha, Snigdha Chaturvedi, Jordan Boyd-Graber, Hal Daumé III. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Mohit Iyyer, Anupam Guha, Snigdha Chaturvedi, Jordan L. Boyd-Graber, Hal Daumé III |
HLT-NAACL | 4 |
| 2016 | Bayesian Supervised Domain Adaptation for Short Text SimilarityabstractMd Arafat Sultan, Jordan Boyd-Graber, Tamara Sumner. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Md. Arafat Sultan, Jordan L. Boyd-Graber, Tamara Sumner |
HLT-NAACL | 2 |
| 2015 | Deep Unordered Composition Rivals Syntactic Methods for Text ClassificationabstractMohit Iyyer, Varun Manjunatha, Jordan Boyd-Graber, Hal Daumé III. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Mohit Iyyer, Varun Manjunatha, Jordan L. Boyd-Graber, Hal Daumé III |
ACL (1) | 3 |
| 2015 | Tea Party in the House: A Hierarchical Ideal Point Topic Model and Its Application to Republican Legislators in the 112th CongressabstractViet-An Nguyen, Jordan Boyd-Graber, Philip Resnik, Kristina Miler. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Viet-An Nguyen, Jordan L. Boyd-Graber, Philip Resnik, Kristina Miler |
ACL (1) | 2 |
| 2015 | Linguistic Harbingers of Betrayal: A Case Study on an Online Strategy GameabstractVlad Niculae, Srijan Kumar, Jordan Boyd-Graber, Cristian Danescu-Niculescu-Mizil. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Vlad Niculae, Srijan Kumar, Jordan L. Boyd-Graber, Cristian Danescu-Niculescu-Mizil |
ACL (1) | 3 |
| 2015 | Making the Most of Crowdsourced Document Annotations: Confused Supervised LDAabstractCorpus labeling projects frequently use low-cost workers from microtask marketplaces; however, these workers are often inexperienced or have misaligned incentives.Crowdsourcing models must be robust to the resulting systematic and nonsystematic inaccuracies.We introduce a novel crowdsourcing model that adapts the discrete supervised topic model sLDA to handle multiple corrupt, usually conflicting (hence "confused") supervision signals.Our model achieves significant gains over previous work in the accuracy of deduced ground truth. Paul Felt, Eric K. Ringger, Jordan L. Boyd-Graber, Kevin D. Seppi |
CoNLL | 3 |
| 2015 | Syntax-based Rewriting for Simultaneous Machine TranslationabstractDivergent word order between languages causes delay in simultaneous machine translation.We present a sentence rewriting method that generates more monotonic translations to improve the speedaccuracy tradeoff.We design grammaticality and meaning-preserving syntactic transformation rules that operate on constituent parse trees.We apply the rules to reference translations to make their word order closer to the source language word order.On Japanese-English translation (two languages with substantially different structure), incorporating the rewritten, more monotonic reference translation into a phrase-based machine translation system enables better translations faster than a baseline system that only uses gold reference translations. He He 0001, Alvin Grissom II, John Morgan, Jordan L. Boyd-Graber, Hal Daumé III |
EMNLP | 4 |
| 2015 | Birds of a Feather Linked Together: A Discriminative Topic Model using Link-based PriorsabstractA wide range of applications, from social media to scientific literature analysis, involve graphs in which documents are connected by links. We introduce a topic model for link prediction based on the intuition that linked documents will tend to have similar topic distributions, integrating a max-margin learning criterion and lexical term weights in the loss function. We validate our approach on the tweets from 2,000 Sina Weibo users and evaluate our model’s reconstruction of the social network. Jordan L. Boyd-Graber, Philip Resnik |
EMNLP | 2 |
| 2015 | Efficient Methods for Incorporating Knowledge into Topic ModelsabstractLatent Dirichlet allocation (LDA) is a popular topic modeling technique for exploring hidden topics in text corpora.Increasingly, topic modeling needs to scale to larger topic spaces and use richer forms of prior knowledge, such as word correlations or document labels.However, inference is cumbersome for LDA models with prior knowledge.As a result, LDA models that use prior knowledge only work in small-scale scenarios.In this work, we propose a factor graph framework, Sparse Constrained LDA (SC-LDA), for efficiently incorporating prior knowledge into LDA.We evaluate SC-LDA's ability to incorporate word correlation knowledge and document label knowledge on three benchmark datasets.Compared to several baseline methods, SC-LDA achieves comparable performance but is significantly faster. Yi Yang 0042, Doug Downey, Jordan L. Boyd-Graber |
EMNLP | 3 |
| 2015 | Paired-Dual Learning for Fast Training of Latent Variable Hinge-Loss MRFsabstractLatent variables allow probabilistic graphical models to capture nuance and structure in important domains such as network science, natural language processing, and computer vision. Naive approaches to learning such complex models can be prohibitively expensive—because they require repeated inferences to update beliefs about latent variables—so lifting this restriction for useful classes of models is an important problem. Hinge-loss Markov random fields (HL-MRFs) are graphical models that allow highly scalable inference and learning in structured domains, in part by representing structured problems with continuous variables. However, this representation leads to challenges when learning with latent variables. We introduce paired-dual learning, a framework that greatly speeds up training by using tractable entropy surrogates and avoiding repeated inferences. Paired-dual learning optimizes an objective with a pair of dual inference problems. This allows fast, joint optimization of parameters and dual variables. We evaluate on social-group detection, trust prediction in social networks, and image reconstruction, finding that paired-dual learning trains models as accurate as those trained by traditional methods in much less time, often before traditional methods make even a single parameter update. Stephen H. Bach, Bert Huang, Jordan L. Boyd-Graber, Lise Getoor |
ICML | 3 |
| 2015 | Removing the Training Wheels: A Coreference Dataset that Entertains Humans and Challenges ComputersabstractAnupam Guha, Mohit Iyyer, Danny Bouman, Jordan Boyd-Graber. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Anupam Guha, Mohit Iyyer, Danny Bouman, Jordan L. Boyd-Graber |
HLT-NAACL | 4 |
| 2015 | Is Your Anchor Going Up or Down? Fast and Accurate Supervised Topic ModelsabstractThang Nguyen, Jordan Boyd-Graber, Jeffrey Lund, Kevin Seppi, Eric Ringger. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Thang Nguyen 0001, Jordan L. Boyd-Graber, Jeffrey Lund, Kevin D. Seppi, Eric K. Ringger |
HLT-NAACL | 2 |
| 2015 | Speeding Document Annotation with Topic ModelsabstractDocument classification and topic models are useful tools for managing and understanding large corpora.Topic models are used to uncover underlying semantic and structure of document collections.Categorizing large collection of documents requires hand-labeled training data, which is time consuming and needs human expertise.We believe engaging user in the process of document labeling helps reduce annotation time and address user needs.We present an interactive tool for document labeling.We use topic models to help users in this procedure.Our preliminary results show that users can more effectively and efficiently apply labels to documents using topic model information. Forough Poursabzi-Sangdeh, Jordan L. Boyd-Graber |
HLT-NAACL | 2 |
| 2014 | Polylingual Tree-Based Topic Models for Translation Domain AdaptationabstractTopic models, an unsupervised technique for inferring translation domains improve machine translation quality.However, previous work uses only the source language and completely ignores the target language, which can disambiguate domains.We propose new polylingual tree-based topic models to extract domain knowledge that considers both source and target languages and derive three different inference schemes.We evaluate our model on a Chinese to English translation task and obtain up to 1.2 BLEU improvement over strong baselines. Yuening Hu, Ke Zhai 0001, Vladimir Eidelman, Jordan L. Boyd-Graber |
ACL (1) | 4 |
| 2014 | Political Ideology Detection Using Recursive Neural NetworksabstractAn individual's words often reveal their political ideology.Existing automated techniques to identify ideology from text focus on bags of words or wordlists, ignoring syntax.Taking inspiration from recent work in sentiment analysis that successfully models the compositional aspect of language, we apply a recursive neural network (RNN) framework to the task of identifying the political position evinced by a sentence.To show the importance of modeling subsentential elements, we crowdsource political annotations at a phrase and sentence level.Our model outperforms existing models on our newly annotated dataset and an existing dataset. Mohit Iyyer, Peter Enns, Jordan L. Boyd-Graber, Philip Resnik |
ACL (1) | 3 |
| 2014 | Anchors Regularized: Adding Robustness and Extensibility to Scalable Topic-Modeling AlgorithmsabstractSpectral methods offer scalable alternatives to Markov chain Monte Carlo and expectation maximization.However, these new methods lack the rich priors associated with probabilistic models.We examine Arora et al.'s anchor words algorithm for topic modeling and develop new, regularized algorithms that not only mathematically resemble Gaussian and Dirichlet priors but also improve the interpretability of topic models.Our new regularization approaches make these efficient algorithms more flexible; we also show that these methods can be combined with informed priors. Thang Nguyen 0001, Yuening Hu, Jordan L. Boyd-Graber |
ACL (1) | 3 |
| 2014 | Don't Until the Final Verb Wait: Reinforcement Learning for Simultaneous Machine TranslationabstractWe introduce a reinforcement learningbased approach to simultaneous machine translation-producing a translation while receiving input wordsbetween languages with drastically different word orders: from verb-final languages (e.g., German) to verb-medial languages (English).In traditional machine translation, a translator must "wait" for source material to appear before translation begins.We remove this bottleneck by predicting the final verb in advance.We use reinforcement learning to learn when to trust predictions about unseen, future portions of the sentence.We also introduce an evaluation metric to measure expeditiousness and quality.We show that our new translation model outperforms batch and monotone translation strategies. Alvin Grissom II, He He 0001, Jordan L. Boyd-Graber, John Morgan, Hal Daumé III |
EMNLP | 3 |
| 2014 | A Neural Network for Factoid Question Answering over ParagraphsabstractText classification methods for tasks like factoid question answering typi-cally use manually defined string match-ing rules or bag of words representa-tions. These methods are ineffective when question text contains very few individual words (e.g., named entities) that are indicative of the answer. We introduce a recursive neural network (rnn) model that can reason over such input by modeling textual composition-ality. We apply our model, qanta, to a dataset of questions from a trivia competition called quiz bowl. Unlike previous rnn models, qanta learns word and phrase-level representations that combine across sentences to reason about entities. The model outperforms multiple baselines and, when combined with information retrieval methods, ri-vals the best human players. 1 Mohit Iyyer, Jordan L. Boyd-Graber, Leonardo Max Batista Claudino, Richard Socher, Hal Daumé III |
EMNLP | 2 |
| 2014 | Sometimes Average is Best: The Importance of Averaging for Prediction using MCMC Inference in Topic ModelingabstractMarkov chain Monte Carlo (MCMC) approximates the posterior distribution of latent variable models by generating many samples and averaging over them.In practice, however, it is often more convenient to cut corners, using only a single sample or following a suboptimal averaging strategy.We systematically study different strategies for averaging MCMC samples and show empirically that averaging properly leads to significant improvements in prediction. Viet-An Nguyen, Jordan L. Boyd-Graber, Philip Resnik |
EMNLP | 2 |
| 2014 | "Our Grief is Unspeakable": Automatically Measuring the Community Impact of a Tragedy
Kimberly Glasgow, Clayton Fink, Jordan L. Boyd-Graber |
ICWSM | 3 |
| 2014 | Learning a Concept Hierarchy from Multi-labeled Documents
Viet-An Nguyen, Jordan L. Boyd-Graber, Philip Resnik, Jonathan D. Chang |
NIPS | 2 |
| 2014 | Interactive topic modeling
Yuening Hu, Jordan L. Boyd-Graber, Brianna Satinoff, Alison Smith-Renner |
Mach. Learn. | 2 |
| 2014 | Modeling topic control to detect influence in conversations using nonparametric topic models
Viet-An Nguyen, Jordan L. Boyd-Graber, Philip Resnik, Deborah A. Cai, Jennifer E. Midberry |
Mach. Learn. | 2 |
| 2014 | Online Adaptor Grammars with Hybrid InferenceabstractAdaptor grammars are a flexible, powerful formalism for defining nonparametric, unsupervised models of grammar productions. This flexibility comes at the cost of expensive inference. We address the difficulty of inference through an online algorithm which uses a hybrid of Markov chain Monte Carlo and variational inference. We show that this inference strategy improves scalability without sacrificing performance on unsupervised word segmentation and topic modeling tasks. Ke Zhai 0001, Jordan L. Boyd-Graber, Shay B. Cohen |
Trans. Assoc. Comput. Linguistics | 2 |
| 2013 | Discovering Pronoun Categories using Discourse Information
Naho Orita, Rebecca McKeown, Naomi Feldman, Jeffrey Lidz, Jordan L. Boyd-Graber |
CogSci | 5 |
| 2013 | Online Latent Dirichlet Allocation with Infinite VocabularyabstractTopic models based on latent Dirichlet allocation (LDA) assume a predefined vocabulary a priori. This is reasonable in batch settings, but it is not reasonable when data are revealed over time, as is the case with streaming / online algorithms. To address this lacuna, we extend LDA by drawing topics from a Dirichlet process whose base distribution is a distribution over all strings rather than from a finite Dirichlet. We develop inference using online variational inference and because we only can consider a finite number of words for each truncated topic propose heuristics to dynamically organize, expand, and contract the set of words we consider in our vocabulary truncation. We show our model can successfully incorporate new words as it encounters new terms and that it performs better than online LDA in evaluations of topic quality and classification performance. Ke Zhai 0001, Jordan L. Boyd-Graber |
ICML (1) | 2 |
| 2013 | Argviz: Interactive Visualization of Topic Dynamics in Multi-party Conversations
Viet-An Nguyen, Yuening Hu, Jordan L. Boyd-Graber, Philip Resnik |
HLT-NAACL | 3 |
| 2013 | Binary to Bushy: Bayesian Hierarchical Clustering with the Beta CoalescentabstractDiscovering hierarchical regularities in data is a key problem in interacting with large datasets, modeling cognition, and encoding knowledge. A previous Bayesian solution---Kingman's coalescent---provides a convenient probabilistic model for data represented as a binary tree. Unfortunately, this is inappropriate for data better described by bushier trees. We generalize an existing belief propagation framework of Kingman's coalescent to the beta coalescent, which models a wider range of tree structures. Because of the complex combinatorial search over possible structures, we develop new sampling schemes using sequential Monte Carlo and Dirichlet process mixture models, which render inference efficient and tractable. We present results on both synthetic and real data that show the beta coalescent outperforms Kingman's coalescent on real datasets and is qualitatively better at capturing data in bushy hierarchies. Yuening Hu, Jordan L. Boyd-Graber, Hal Daumé III, Z. Irene Ying |
NIPS | 2 |
| 2013 | Lexical and Hierarchical Topic RegressionabstractInspired by a two-level theory that unifies agenda setting and ideological framing, we propose supervised hierarchical latent Dirichlet allocation (SHLDA) which jointly captures documents' multi-level topic structure and their polar response variables. Our model extends the nested Chinese restaurant process to discover a tree-structured topic hierarchy and uses both per-topic hierarchical and per-word lexical regression parameters to model the response variables. Experiments in a political domain and on sentiment analysis tasks show that SHLDA improves predictive accuracy while adding a new dimension of insight into how topics under discussion are framed. Viet-An Nguyen, Jordan L. Boyd-Graber, Philip Resnik |
NIPS | 2 |
| 2012 | SITS: A Hierarchical Nonparametric Model using Speaker Identity for Topic Segmentation in Multiparty Conversations
Viet-An Nguyen, Jordan L. Boyd-Graber, Philip Resnik |
ACL (1) | 2 |
| 2012 | Besting the Quiz Master: Crowdsourcing Incremental Classification Games
Jordan L. Boyd-Graber, Brianna Satinoff, He He 0001, Hal Daumé III |
EMNLP-CoNLL | 1 |
| 2012 | Modeling Images using Transformed Indian Buffet Processes
Ke Zhai 0001, Yuening Hu, Jordan L. Boyd-Graber, Sinead Williamson |
ICML | 3 |
| 2012 | Grammatical structures for word-level sentiment detection
Asad B. Sayeed, Jordan L. Boyd-Graber, Bryan Rusk, Amy Weinberg |
HLT-NAACL | 2 |
| 2012 | Mr. LDA: a flexible large scale topic modeling package using variational inference in MapReduceabstractLatent Dirichlet Allocation (LDA) is a popular topic modeling technique for exploring document collections. Because of the increasing prevalence of large datasets, there is a need to improve the scalability of inference for LDA. In this paper, we introduce a novel and flexible large scale topic modeling package in MapReduce (Mr. LDA). As opposed to other techniques which use Gibbs sampling, our proposed framework uses variational inference, which easily fits into a distributed environment. More importantly, this variational implementation, unlike highly tuned and specialized implementations based on Gibbs sampling, is easily extensible. We demonstrate two extensions of the models possible with this scalable framework: informed priors to guide topic discovery and extracting topics from a multilingual corpus. We compare the scalability of Mr. LDA against Mahout, an existing large scale topic modeling package. Mr. LDA out-performs Mahout both in execution speed and held-out likelihood. Ke Zhai 0001, Jordan L. Boyd-Graber, Sebastian Bruch 0001, Mohamad L. Alkhouja |
WWW | 2 |
| 2011 | Interactive Topic Modeling
Yuening Hu, Jordan L. Boyd-Graber, Brianna Satinoff |
ACL | 2 |
| 2010 | Holistic Sentiment Analysis Across Languages: Multilingual Supervised Latent Dirichlet Allocation
Jordan L. Boyd-Graber, Philip Resnik |
EMNLP | 1 |
| 2010 | Modeling Perspective Using Adaptor Grammars
Eric Hardisty, Jordan L. Boyd-Graber, Philip Resnik |
EMNLP | 2 |
| 2009 | Speaking through pictures: images vs. iconsabstractPeople with aphasia, a condition that impairs the ability to understand or generate written or spoken language, are aided by assistive technology that helps them communicate through a vocabulary of icons. These systems are akin to language translation systems, translating icon arrangements into spoken or written language and vice versa. However, these icon-based systems have little vocabulary breadth or depth, making it difficult for people with aphasia to apply their usage to multiple real world situations. Pictures from the web are numerous, varied, and easily accessible and thus, could potentially address the small size issues of icon-based systems. We present results from two studies that investigate this potential and demonstrate that images can be as effective as icons when used as a replacement for English language communication. The first study uses elderly subjects to investigate the efficacy of images vs. icons in conveying word meaning; the second study examines the retention of word-level meaning by both images and icons with a population of aphasics. We conclude that images collected from the web are as functional as icons in conveying information and thus, are feasible to use in assistive technology that supports people with aphasia. Xiaojuan Ma, Jordan L. Boyd-Graber, Sonya S. Nikolova, Perry R. Cook |
ASSETS | 2 |
| 2009 | Better vocabularies for assistive communication aids: connecting terms using semantic networks and untrained annotatorsabstractThe difficulties of navigating vocabulary in an assistive communication device are exacerbated for individuals with lexical access disorders like those due to aphasia. We present the design and implementation of a vocabulary network based on WordNet, a resource that attempts to model human semantic memory, that enables users to find words easily. To correct for the sparsity of links among words, we augment WordNet with additional connections derived from human judgments of semantic similarity collected in an online experiment. We evaluate the resulting system, the visual vocabulary for aphasia (ViVA), and describe its potential to adapt to a user's profile and enable faster search and improved navigation. Sonya S. Nikolova, Jordan L. Boyd-Graber, Christiane Fellbaum, Perry R. Cook |
ASSETS | 2 |
| 2009 | Connections between the lines: augmenting social networks with textabstractNetwork data is ubiquitous, encoding collections of relationships between entities such as people, places, genes, or corporations. While many resources for networks of interesting entities are emerging, most of these can only annotate connections in a limited fashion. Although relationships between entities are rich, it is impractical to manually devise complete characterizations of these relationships for every pair of entities on large, real-world corpora. Jonathan D. Chang, Jordan L. Boyd-Graber, David M. Blei |
KDD | 2 |
| 2009 | Reading Tea Leaves: How Humans Interpret Topic ModelsabstractProbabilistic topic models are a popular tool for the unsupervised analysis of text, providing both a predictive model of future text and a latent topic representation of the corpus. Practitioners typically assume that the latent space is semantically meaningful. It is used to check models, summarize the corpus, and guide exploration of its contents. However, whether the latent space is interpretable is in need of quantitative evaluation. In this paper, we present new quantitative methods for measuring semantic meaning in inferred topics. We back these measures with large-scale user studies, showing that they capture aspects of the model that are undetected by previous measures of model quality based on held-out likelihood. Surprisingly, topic models which perform better on held-out likelihood may infer less semantically meaningful topics. Jonathan D. Chang, Jordan L. Boyd-Graber, Sean Gerrish, Chong Wang 0002, David M. Blei |
NIPS | 2 |
| 2009 | Multilingual Topic Models for Unaligned Text
Jordan L. Boyd-Graber, David M. Blei |
UAI | 1 |
| 2008 | Syntactic Topic ModelsabstractWe develop \name\ (STM), a nonparametric Bayesian model of parsed documents. \Shortname\ generates words that are both thematically and syntactically constrained, which combines the semantic insights of topic models with the syntactic information available from parse trees. Each word of a sentence is generated by a distribution that combines document-specific topic weights and parse-tree specific syntactic transitions. Words are assumed generated in an order that respects the parse tree. We derive an approximate posterior inference method based on variational methods for hierarchical Dirichlet processes, and we report qualitative and quantitative results on both synthetic data and hand-parsed documents. Jordan L. Boyd-Graber, David M. Blei |
NIPS | 1 |
| 2007 | A Topic Model for Word Sense Disambiguation
Jordan L. Boyd-Graber, David M. Blei, Xiaojin Zhu 0001 |
EMNLP-CoNLL | 1 |
| 2006 | Participatory design with proxies: developing a desktop-PDA system to support people with aphasiaabstractIn this paper, we describe the design and preliminary evaluation of a hybrid desktop-handheld system developed to support individuals with aphasia, a disorder which impairs the ability to speak, read, write, or understand language. The system allows its users to develop speech communication through images and sound on a desktop computer and download this speech to a mobile device that can then support communication outside the home. Using a desktop computer for input addresses some of this population's difficulties interacting with handheld devices, while the mobile device addresses stigma and portability issues. A modified participatory design approach was used in which proxies, that is, speech-language pathologists who work with aphasic individuals, assumed the role normally filled by users. This was done because of the difficulties in communicating with the target population and the high variability in aphasic disorders. In addition, the paper presents a case study of the proxy-use participatory design process that illustrates how different interview techniques resulted in different user feedback. Jordan L. Boyd-Graber, Sonya S. Nikolova, Karyn Moffatt, Kenrick C. Kin, Joshua Y. Lee, Lester Mackey, Marilyn Tremaine, Maria M. Klawe |
CHI | 1 |