EDBT 2026 Demo / reviewers in the wild / expert
Mohit Iyyer
dblp:148/9178
· DBLP profile ↗
71ranked-venue papers
7as first author
42since 2021 · last 2026
0000-0001-7340-0804ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 66 · 7 first-author · 39 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Frankentext: Stitching random text fragments into long-form narrativesabstractAs AI text detectors are increasingly used to flag LLM-generated writing, a natural question arises: are there forms of high-quality generated narrative that can evade such detection?We introduce Frankentexts, a longform narrative generation paradigm that treats an LLM as a composer of existing texts rather than as an author.Given a writing prompt and thousands of randomly sampled human-written snippets, the model assembles a coherent narrative where most tokens (e.g., 90%) are copied verbatim from the source passages.Despite the extreme challenge of the task, we observe through extensive automatic and human evaluation that Frankentexts improve over vanilla LLM generations in key writing quality metrics such as diversity and novelty while remaining mostly coherent and relevant to the prompt.Furthermore, Frankentexts pose a fundamental challenge to current AI text detectors: 72% of Frankentexts produced by our best configuration (Gemini-2.5-Prowith 5K input snippets) are misclassified as human-written by Pangram, a state-of-the-art detector.Human annotators praise Frankentexts for their inventive premises, vivid descriptions, and dry humor; however, they still identify issues with abrupt tonal shifts and uneven grammar across segments.Overall, the emergence of highquality yet low-detectability Frankentexts challenges established authorship norms while raising concerns about the publishing economy. Chau Minh Pham, Jenna Russell, Dzung Pham 0002, Mohit Iyyer |
ACL (1) | 4 |
| 2026 | AI use in American newspapers is widespread, uneven, and rarely disclosedabstractJenna Russell, Marzena Karpinska, Destiny Akinode, James Zhou, Katherine Thai, Bradley Emi, Max Spero, Mohit Iyyer. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jenna Russell, Marzena Karpinska, Destiny Akinode, James Zhou, Katherine Thai, Bradley Emi, Max Spero, Mohit Iyyer |
ACL (1) | 8 |
| 2025 | CaLMQA: Exploring culturally specific long-form question answering across 23 languagesabstractShane Arora, Marzena Karpinska, Hung-Ting Chen, Ipsita Bhattacharjee, Mohit Iyyer, Eunsol Choi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shane Arora, Marzena Karpinska, Hung-Ting Chen, Ipsita Bhattacharjee, Mohit Iyyer, Eunsol Choi |
ACL (1) | 5 |
| 2025 | People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated textabstractIn this paper, we study how well humans can detect text generated by commercial LLMs (GPT-4o, Claude, o1). We hire annotators to read 300 non-fiction English articles, label them as either human-written or AI-generated, and provide paragraph-length explanations for their decisions. Our experiments show that annotators who frequently use LLMs for writing tasks excel at detecting AI-generated text, even without any specialized training or feedback. In fact, the majority vote among five such “expert” annotators misclassifies only 1 of 300 articles, significantly outperforming most commercial and open-source detectors we evaluated even in the presence of evasion tactics like paraphrasing and humanization. Qualitative analysis of the experts’ free-form explanations shows that while they rely heavily on specific lexical clues (‘AI vocabulary’), they also pick up on more complex phenomena within the text (e.g., formality, originality, clarity) that are challenging to assess for automatic detectors. We release our annotated dataset and code to spur future research into both human and automated detection of AI-generated text. Jenna Russell, Marzena Karpinska, Mohit Iyyer |
ACL (1) | 3 |
| 2025 | Does quantization affect models' performance on long-context tasks?abstractLarge language models (LLMs) now support context windows exceeding 128K tokens, but this comes with significant memory requirements and high inference latency.Quantization can mitigate these costs, but may degrade performance.In this work, we present the first systematic evaluation of quantized LLMs on tasks with long inputs (≥64K tokens) and long-form outputs.Our evaluation spans 9.7K test examples, five quantization methods (FP8, GPTQ-int8, AWQ-int4, GPTQ-int4, BNB-nf4), and five models (Llama-3.1 8B and 70B; Qwen-2.5 7B, 32B, and 72B).We find that, on average, 8-bit quantization preserves accuracy (~0.8% drop), whereas 4-bit methods lead to substantial losses, especially for tasks involving longcontext inputs (drops of up to 59%).This degradation tends to worsen when the input is in a language other than English.Crucially, the effects of quantization depend heavily on the quantization method, model, and task.For instance, while Qwen-2.5 72B remains robust under BNB-nf4, Llama-3.1 70B experiences a 32% performance drop on the same task.These findings highlight the importance of a careful, task-specific evaluation before deploying quantized LLMs, particularly in long-context scenarios and for languages other than English. github.com/molereddy/long-context-quantization FP8 GPTQ-int8 AWQ-int4 GPTQ-int4 BNB-nf4Ruler Anmol Mekala, Anirudh Atmakuru, Yixiao Song, Marzena Karpinska, Mohit Iyyer |
EMNLP | 5 |
| 2025 | OWL: Probing Cross-Lingual Recall of Memorized Texts via World LiteratureabstractAlisha Srivastava, Emir Kaan Korukluoglu, Minh Nhat Le, Duyen Tran, Chau Minh Pham, Marzena Karpinska, Mohit Iyyer. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Alisha Srivastava, Emir Korukluoglu, Minh Nhat Le, Duyen Tran, Chau Minh Pham, Marzena Karpinska, Mohit Iyyer |
EMNLP | 7 |
| 2025 | BLEUBERI: BLEU is a surprisingly effective reward for instruction followingabstractReward models are central to aligning LLMs with human preferences, but they are costly to train, requiring large-scale human-labeled preference data and powerful pretrained LLM backbones. Meanwhile, the increasing availability of high-quality synthetic instruction-following datasets raises the question: can simpler, reference-based metrics serve as viable alternatives to reward models during RL-based alignment? In this paper, we show first that BLEU, a basic string-matching metric, surprisingly matches strong reward models in agreement with human preferences on general instruction-following datasets. Based on this insight, we develop BLEUBERI, a method that first identifies challenging instructions and then applies Group Relative Policy Optimization (GRPO) using BLEU directly as the reward function. We demonstrate that BLEUBERI-trained models are competitive with models trained via reward model-guided RL across four challenging instruction-following benchmarks and three different base language models. A human evaluation further supports that the quality of BLEUBERI model outputs is on par with those from reward model-aligned models. Moreover, BLEUBERI models generate outputs that are more factually grounded than competing methods. Overall, we show that given access to high-quality reference outputs (easily obtained via existing instruction-following datasets or synthetic data generation), string matching-based metrics are cheap yet effective proxies for reward models during alignment. We release our code and data at https://github.com/lilakk/BLEUBERI. Yapei Chang, Yekyung Kim, Michael Krumdick, Chris Tanner, Mohit Iyyer |
NeurIPS | 7 |
| 2025 | Contextualized Evaluations: Judging Language Model Responses to Underspecified QueriesabstractAbstract Language model users often issue queries that lack specification, where the context under which a query was issued—such as the user’s identity, the query’s intent, and the criteria for a response to be useful—is not explicit. For instance, a good response to a subjective query like “What book should I read next?” would depend on the user’s preferences, and a good response to an open-ended query like “How do antibiotics work against bacteria?” would depend on the user’s expertise. This makes evaluation of responses to such queries an ill-posed task, as evaluators may make arbitrary judgments about the response quality. To remedy this, we present contextualized evaluations, a protocol that synthetically constructs context surrounding an underspecified query and provides it during evaluation. We find that the presence of context can 1) alter conclusions drawn from evaluation, even flipping benchmark rankings between model pairs, 2) nudge evaluators to make fewer judgments based on surface-level criteria, like style, and 3) provide new insights about model behavior across diverse contexts. Specifically, our procedure suggests a potential bias towards WEIRD (Western, Educated, Industrialized, Rich and Democratic) contexts in models’ “default” responses and we find that models are not equally sensitive to following different contexts, even when they are provided in prompts.1 Chaitanya Malaviya, Joseph Chee Chang, Dan Roth 0001, Mohit Iyyer, Mark Yatskar, Kyle Lo |
Trans. Assoc. Comput. Linguistics | 4 |
| 2024 | Multistage Collaborative Knowledge Distillation from a Large Language Model for Semi-Supervised Sequence GenerationabstractJiachen Zhao, Wenlong Zhao, Andrew Drozdov, Benjamin Rozonoyer, Md Arafat Sultan, Jay-Yoon Lee, Mohit Iyyer, Andrew McCallum. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Wenlong Zhao 0001, Andrew Drozdov, Benjamin Rozonoyer, Md. Arafat Sultan, Jay-Yoon Lee, Mohit Iyyer, Andrew McCallum |
ACL (1) | 7 |
| 2024 | PEARL: Prompting Large Language Models to Plan and Execute Actions Over Long DocumentsabstractSimeng Sun, Yang Liu, Shuohang Wang, Dan Iter, Chenguang Zhu, Mohit Iyyer. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Simeng Sun, Yang Liu 0124, Shuohang Wang, Dan Iter, Chenguang Zhu 0001, Mohit Iyyer |
EACL (1) | 6 |
| 2024 | PostMark: A Robust Blackbox Watermark for Large Language ModelsabstractThe most effective techniques to detect LLMgenerated text rely on inserting a detectable signature-or watermark-during the model's decoding process.Most existing watermarking methods require access to the underlying LLM's logits, which LLM API providers are loath to share due to fears of model distillation.As such, these watermarks must be implemented independently by each LLM provider.In this paper, we develop POSTMARK, a modular post-hoc watermarking procedure in which an input-dependent set of words (determined via a semantic embedding) is inserted into the text after the decoding process has completed.Critically, POSTMARK does not require logit access, which means it can be implemented by a third party.We also show that POST-MARK is more robust to paraphrasing attacks than existing watermarking methods: our experiments cover eight baseline algorithms, five base LLMs, and three datasets.Finally, we evaluate the impact of POSTMARK on text quality using both automated and human assessments, highlighting the trade-off between quality and robustness to paraphrasing.We release our code, outputs, and annotations at https://github.com/lilakk/PostMark. Yapei Chang, Kalpesh Krishna, Amir Houmansadr, John Wieting, Mohit Iyyer |
EMNLP | 5 |
| 2024 | One Thousand and One Pairs: A "novel" challenge for long-context language modelsabstractSynthetic long-context LLM benchmarks (e.g., "needle-in-the-haystack") test only surfacelevel retrieval capabilities, but how well can long-context LLMs retrieve, synthesize, and reason over information across book-length inputs?We address this question by creating NOCHA, a dataset of 1,001 minimally different pairs of true and false claims about 67 recentlypublished English fictional books, written by human readers of those books.In contrast to existing long-context benchmarks, our annotators confirm that the largest share of pairs in NOCHA require global reasoning over the entire book to verify.Our experiments show that while human readers easily perform this task, it is enormously challenging for all ten long-context LLMs that we evaluate: no open-weight model performs above random chance (despite their strong performance on synthetic benchmarks), while GPT-4O achieves the highest accuracy at 55.8%.Further analysis reveals that (1) on average, models perform much better on pairs that require only sentence-level retrieval vs. global reasoning; (2) model-generated explanations for their decisions are often inaccurate even for correctly-labeled claims; and (3) models perform substantially worse on speculative fiction books that contain extensive world-building.The methodology proposed in NOCHA allows for the evolution of the benchmark dataset and the easy analysis of future models.TRUE TRUE Niema takes her students out to world's end and tells them that the fog kills anything it touches, a statement backed by her extensive research into the fog.When Niema takes her students out to world's end and tells them that the fog kills anything it touches, she is intentionally lying to them. Marzena Karpinska, Katherine Thai, Kyle Lo, Tanya Goyal, Mohit Iyyer |
EMNLP | 5 |
| 2024 | BooookScore: A systematic exploration of book-length summarization in the era of LLMsabstractSummarizing book-length documents ($>$100K tokens) that exceed the context window size of large language models (LLMs) requires first breaking the input document into smaller chunks and then prompting an LLM to merge, update, and compress chunk-level summaries. Despite the complexity and importance of this task, it has yet to be meaningfully studied due to the challenges of evaluation: existing book-length summarization datasets (e.g., BookSum) are in the pretraining data of most public LLMs, and existing evaluation methods struggle to capture errors made by modern LLM summarizers. In this paper, we present the first study of the coherence of LLM-based book-length summarizers implemented via two prompting workflows: (1) hierarchically merging chunk-level summaries, and (2) incrementally updating a running summary. We obtain 1193 fine-grained human annotations on GPT-4 generated summaries of 100 recently-published books and identify eight common types of coherence errors made by LLMs. Because human evaluation is expensive and time-consuming, we develop an automatic metric, BooookScore, that measures the proportion of sentences in a summary that do not contain any of the identified error types. BooookScore has high agreement with human annotations and allows us to systematically evaluate the impact of many other critical parameters (e.g., chunk size, base LLM) while saving \$15K USD and 500 hours in human evaluation costs. We find that closed-source LLMs such as GPT-4 and Claude 2 produce summaries with higher BooookScore than those generated by open-source models. While LLaMA 2 falls behind other models, Mixtral achieves performance on par with GPT-3.5-Turbo. Incremental updating yields lower BooookScore but higher level of detail than hierarchical merging, a trade-off sometimes preferred by annotators. We release code and annotations to spur more principled research on book-length summarization. Yapei Chang, Kyle Lo, Tanya Goyal, Mohit Iyyer |
ICLR | 4 |
| 2024 | TopicGPT: A Prompt-based Topic Modeling FrameworkabstractChau Minh Pham, Alexander Hoyle, Simeng Sun, Philip Resnik, Mohit Iyyer. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Alexander Miserlis Hoyle, Simeng Sun, Philip Resnik, Mohit Iyyer |
NAACL-HLT | 5 |
| 2024 | Triage of Messages and Conversations in a Large-Scale Child Victimization CorpusabstractChildren are among the most vulnerable online populations. Reports of child sexual exploitation on social media and apps have grown annually at an alarming rate and are overwhelming investigators. Even a single case can require examining millions of messages involving hundreds of victims. Triage and prioritization based on victims' experiences is an unfortunate necessity. Using a chat dataset of more than 3 million messages between victims and perpetrators, we evaluate and contribute tools for analyzing the experiences of victims of sexual exploitation. We develop both supervised and unsupervised methods to classify messages into categories of interest to law enforcement, such as age requests, persuasion, and sexual messages. We also introduce a conversation clustering technique to illuminate differences among victims' experiences based on their chat history. Through a qualitative analysis, we demonstrate that the learned clusters are coherent and represent distinct conversation patterns. For example, we can distinguish groups of users who never comply with sexual requests, comply after a few conversations, or comply immediately after being targeted. We expect this approach and associated visualizations will aid law enforcement, industry moderators, and sociologists who need to analyze massive corpora in this domain. Finally, we validate prior models derived from conversations involving adults pretending to be minors and provide statistics that could help undercover adults more accurately portray minor victims. Prasanna Lakkur Subramanyam, Mohit Iyyer, Brian Neil Levine |
WWW | 2 |
| 2023 | A Critical Evaluation of Evaluations for Long-form Question AnsweringabstractLong-form question answering (LFQA) enables answering a wide range of questions, but its flexibility poses enormous challenges for evaluation.We perform the first targeted study of the evaluation of long-form answers, covering both human and automatic evaluation practices.We hire domain experts in seven areas to provide preference judgments over pairs of answers, along with free-form justifications for their choices.We present a careful analysis of experts' evaluation, which focuses on new aspects such as the comprehensiveness of the answer.Next, we examine automatic text generation metrics, finding that no existing metrics are predictive of human preference judgments.However, some metrics correlate with fine-grained aspects of answers (e.g., coherence).We encourage future work to move away from a single "overall score" of the answer and adopt a multi-faceted evaluation, targeting aspects such as factuality and completeness.We publicly release all of our annotations and code to spur future work into LFQA evaluation. 1 Yixiao Song, Mohit Iyyer, Eunsol Choi |
ACL (1) | 3 |
| 2023 | Stealing the Decoding Algorithms of Language ModelsabstractA key component of generating text from modern language models (LM) is the selection and tuning of decoding algorithms. These algorithms determine how to generate text from the internal probability distribution generated by the LM. The process of choosing a decoding algorithm and tuning its hyperparameters takes significant time, manual effort, and computation, and it also requires extensive human evaluation. Therefore, the identity and hyperparameters of such decoding algorithms are considered to be extremely valuable to their owners. In this work, we show, for the first time, that an adversary with typical API access to an LM can steal the type and hyperparameters of its decoding algorithms at very low monetary costs. Our attack is effective against popular LMs used in text generation APIs, including GPT-2, GPT-3 and GPT-Neo. We demonstrate the feasibility of stealing such information with only a few dollars, e.g., 0.8, 1, 4, and 40 for the four versions of GPT-3. Ali Naseh, Kalpesh Krishna, Mohit Iyyer, Amir Houmansadr |
CCS | 3 |
| 2023 | LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form SummarizationabstractKalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, Kyle Lo. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Kalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, Kyle Lo |
EACL | 4 |
| 2023 | FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationabstractSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, Hannaneh Hajishirzi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Scott Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, Hannaneh Hajishirzi |
EMNLP | 7 |
| 2023 | kNN-LM Does Not Improve Open-ended Text GenerationabstractIn this paper, we study the generation quality of interpolation-based retrieval-augmented language models (LMs).These methods, best exemplified by the kNN-LM (Khandelwal et al., 2020), interpolate the LM's predicted distribution of the next word with a distribution formed from the most relevant retrievals for a given prefix.While the kNN-LM and related methods yield impressive decreases in perplexity, we discover that they do not exhibit corresponding improvements in open-ended generation quality, as measured by both automatic evaluation metrics (e.g., MAUVE) and human evaluations.Digging deeper, we find that interpolating with a retrieval distribution actually increases perplexity compared to the baseline LM for the majority of tokens in the WikiText-103 test set, even though the overall perplexity is lower due to a smaller number of tokens for which perplexity dramatically decreases after interpolation.However, when decoding a long sequence at inference time, significant improvements on this smaller subset of tokens are washed out by slightly worse predictions on most tokens.Furthermore, we discover that the entropy of the retrieval distribution increases faster than that of the base LM as the generated sequence becomes longer, which indicates that retrieval is less reliable when using model-generated text as queries (i.e., is subject to exposure bias).We hope that our analysis spurs future work on improved decoding algorithms and interpolation strategies for retrieval-augmented language models. Shufan Wang, Yixiao Song, Andrew Drozdov, Aparna Garimella, Varun Manjunatha, Mohit Iyyer |
EMNLP | 6 |
| 2023 | Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseabstractThe rise in malicious usage of large language models, such as fake content creation and academic plagiarism, has motivated the development of approaches that identify AI-generated text, including those based on watermarking or outlier detection. However, the robustness of these detection algorithms to paraphrases of AI-generated text remains unclear. To stress test these detectors, we build a 11B parameter paraphrase generation model (DIPPER) that can paraphrase paragraphs, condition on surrounding context, and control lexical diversity and content reordering. Paraphrasing text generated by three large language models (including GPT3.5-davinci-003) with DIPPER successfully evades several detectors, including watermarking, GPTZero, DetectGPT, and OpenAI's text classifier. For example, DIPPER drops detection accuracy of DetectGPT from 70.3% to 4.6% (at a constant false positive rate of 1%), without appreciably modifying the input semantics.
To increase the robustness of AI-generated text detection to paraphrase attacks, we introduce a simple defense that relies on retrieving semantically-similar generations and must be maintained by a language model API provider. Given a candidate text, our algorithm searches a database of sequences previously generated by the API, looking for sequences that match the candidate text within a certain threshold. We empirically verify our defense using a database of 15M generations from a fine-tuned T5-XXL model and find that it can detect 80% to 97% of paraphrased generations across different settings while only classifying 1% of human-written sequences as AI-generated. We open-source our models, code and data. Kalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting, Mohit Iyyer |
NeurIPS | 5 |
| 2022 | RELiC: Retrieving Evidence for Literary ClaimsabstractHumanities scholars commonly provide evidence for claims that they make about a work of literature (e.g., a novel) in the form of quotations from the work.We collect a large-scale dataset (RELiC) of 78K literary quotations and surrounding critical analysis and use it to formulate the novel task of literary evidence retrieval, in which models are given an excerpt of literary analysis surrounding a masked quotation and asked to retrieve the quoted passage from the set of all passages in the work.Solving this retrieval task requires a deep understanding of complex literary and linguistic phenomena, which proves challenging to methods that overwhelmingly rely on lexical and semantic similarity matching.We implement a RoBERTa-based dense passage retriever for this task that outperforms existing pretrained information retrieval baselines; however, experiments and analysis by human domain experts indicate that there is substantial room for improvement over our dense retriever. Katherine Thai, Yapei Chang, Kalpesh Krishna, Mohit Iyyer |
ACL (1) | 4 |
| 2022 | DEMETR: Diagnosing Evaluation Metrics for TranslationabstractWhile machine translation evaluation metrics based on string overlap (e.g., BLEU) have their limitations, their computations are transparent: the BLEU score assigned to a particular candidate translation can be traced back to the presence or absence of certain words.The operations of newer learned metrics (e.g., BLEURT, COMET), which leverage pretrained language models to achieve higher correlations with human quality judgments than BLEU, are opaque in comparison.In this paper, we shed light on the behavior of these learned metrics by creating DEMETR, a diagnostic dataset with 31K English examples (translated from 10 source languages) for evaluating the sensitivity of MT evaluation metrics to 35 different linguistic perturbations spanning semantic, syntactic, and morphological error categories.All perturbations were carefully designed to form minimal pairs with the actual translation (i.e., differ in only one aspect).We find that learned metrics perform substantially better than string-based metrics on DEMETR.Additionally, learned metrics differ in their sensitivity to various phenomena (e.g., BERTSCORE is sensitive to untranslated words but relatively insensitive to gender manipulation, while COMET is much more sensitive to word repetition than to aspectual changes).We publicly release DEMETR to spur more informed future development of machine translation evaluation metrics 1 . Marzena Karpinska, Nishant Raj, Katherine Thai, Yixiao Song, Mohit Iyyer |
EMNLP | 6 |
| 2022 | RankGen: Improving Text Generation with Large Ranking ModelsabstractGiven an input sequence (or prefix), modern language models often assign high probabilities to output sequences that are repetitive, incoherent, or irrelevant to the prefix; as such, model-generated text also contains such artifacts.To address these issues we present RANKGEN, a 1.2B parameter encoder model for English that scores model generations given a prefix.RANKGEN can be flexibly incorporated as a scoring function in beam search and used to decode from any pretrained language model.We train RANKGEN using large-scale contrastive learning to map a prefix close to the ground-truth sequence that follows it and far away from two types of negatives:(1) random sequences from the same document as the prefix, and (2) sequences generated from a large language model conditioned on the prefix.Experiments across four different language models (345M-11B parameters) and two domains show that RANKGEN significantly outperforms decoding algorithms like nucleus, top-k, and typical sampling on both automatic metrics (85.0 vs 77.3 MAUVE) as well as human evaluations with English writers (74.5% human preference over nucleus sampling).Analysis reveals that RANKGEN outputs are more relevant to the prefix and improve continuity and coherence compared to baselines.We release our model checkpoints, code, and human preference data with explanations to facilitate future research.1 Kalpesh Krishna, Yapei Chang, John Wieting, Mohit Iyyer |
EMNLP | 4 |
| 2022 | SLING: Sino Linguistic Evaluation of Large Language ModelsabstractTo understand what kinds of linguistic knowledge are encoded by pretrained Chinese language models (LMs), we introduce the benchmark of Sino LINGuistics (SLING), which consists of 38K minimal sentence pairs in Mandarin Chinese grouped into 9 high-level linguistic phenomena.Each pair demonstrates the acceptability contrast of a specific syntactic or semantic phenomenon (e.g., The keys are lost vs.The keys is lost), and an LM should assign lower perplexity to the acceptable sentence.In contrast to the CLiMP dataset (Xiang et al., 2021), which also contains Chinese minimal pairs and was created by translating the vocabulary of the English BLiMP dataset, the minimal pairs in SLING are derived primarily by applying syntactic and lexical transformations to naturally-occurring, linguist-annotated sentences from the Chinese Treebank 9.0, thus addressing severe issues in CLiMP's data generation process.We test 18 publicly available pretrained monolingual (e.g., BERT-base-zh, CPM) and multi-lingual (e.g., mT5, XLM) language models on SLING.Our experiments show that the average accuracy for LMs is far below human performance (69.7% vs. 97.1%),while BERT-base-zh achieves the highest accuracy (84.8%) of all tested LMs, even much larger ones.Additionally, we find that most LMs have a strong gender and number (singular/plural) bias, and they perform better on local phenomena than hierarchical ones. 1 Yixiao Song, Kalpesh Krishna, Rajesh Bhatt, Mohit Iyyer |
EMNLP | 4 |
| 2022 | Exploring Document-Level Literary Machine Translation with Parallel Paragraphs from World LiteratureabstractLiterary translation is a culturally significant task, but it is bottlenecked by the small number of qualified literary translators relative to the many untranslated works published around the world.Machine translation (MT) holds potential to complement the work of human translators by improving both training procedures and their overall efficiency.Literary translation is less constrained than more traditional MT settings since translators must balance meaning equivalence, readability, and critical interpretability in the target language.This property, along with the complex discourse-level context present in literary texts, also makes literary MT more challenging to computationally model and evaluate.To explore this task, we collect a dataset (PAR3) of non-English language novels in the public domain, each aligned at the paragraph level to both human and automatic English translations.Using PAR3, we discover that expert literary translators prefer reference human translations over machinetranslated paragraphs at a rate of 84%, while state-of-the-art automatic MT metrics do not correlate with those preferences.The experts note that MT outputs contain not only mistranslations, but also discourse-disrupting errors and stylistic inconsistencies.To address these problems, we train a post-editing model whose output is preferred over normal MT output at a rate of 69% by experts.We publicly release PAR3 to spur future research into literary MT. 1 Katherine Thai, Marzena Karpinska, Kalpesh Krishna, Bill Ray, Moira Inghilleri, John Wieting, Mohit Iyyer |
EMNLP | 7 |
| 2022 | Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual GenerationabstractIn this paper, we explore the challenging problem of performing a generative task in a target language when labeled data is only available in English, using summarization as a case study.We assume a strict setting with no access to parallel data or machine translation and find that common transfer learning approaches struggle in this setting, as a generative multilingual model fine-tuned purely on English catastrophically forgets how to generate non-English.Given the recent rise of parameter-efficient adaptation techniques, we conduct the first investigation into how one such method, prompt tuning (Lester et al., 2021), can overcome catastrophic forgetting to enable zero-shot cross-lingual generation.Our experiments show that parameter-efficient prompt tuning provides gains over standard fine-tuning when transferring between lessrelated languages, e.g., from English to Thai.However, a significant gap still remains between these methods and fully-supervised baselines.To improve cross-lingual transfer further, we explore several approaches, including: (1) mixing in unlabeled multilingual data, and (2) explicitly factoring prompts into recombinable language and task components.Our approaches can provide further quality gains, suggesting that robust zero-shot crosslingual generation is within reach.ING and standard MODELTUNING for zero-shot cross-lingual generation (XGEN).We show that increasing model scale and decreasing tunable parameter capacity are key for overcoming catastrophic forgetting on XGEN.• We propose WIKILINGUA-0, a challenging XGEN benchmark and an associated SP-ROUGE evaluation metric, which we hope will facilitate future work evaluating multilingual summarization.• We show that mixing in unsupervised multilingual data can boost XGEN performance, and are the first to combine this approach with PROMPTTUNING.• We propose "factorized prompts", a novel approach that can also help PROMPTTUNING overcome severe catastrophic forgetting.• To facilitate future work, we release our data, pretrained models Tu Vu, Aditya Barua, Brian Lester, Daniel M. Cer, Mohit Iyyer, Noah Constant |
EMNLP | 5 |
| 2022 | ChapterBreak: A Challenge Dataset for Long-Range Language ModelsabstractWhile numerous architectures for long-range language models (LRLMs) have recently been proposed, a meaningful evaluation of their discourse-level language understanding capabilities has not yet followed.To this end, we introduce CHAPTERBREAK, a challenge dataset that provides an LRLM with a long segment from a narrative that ends at a chapter boundary and asks it to distinguish the beginning of the ground-truth next chapter from a set of negative segments sampled from the same narrative.A fine-grained human annotation reveals that our dataset contains many complex types of chapter transitions (e.g., parallel narratives, cliffhanger endings) that require processing global context to comprehend.Experiments on CHAPTERBREAK show that existing LRLMs fail to effectively leverage long-range context, substantially underperforming a segment-level model trained directly for this task.We publicly release our CHAPTERBREAK dataset to spur more principled future research into LRLMs. 1 Simeng Sun, Katherine Thai, Mohit Iyyer |
NAACL-HLT | 3 |
| 2022 | Modeling Exemplification in Long-form Question Answering via RetrievalabstractShufan Wang, Fangyuan Xu, Laure Thompson, Eunsol Choi, Mohit Iyyer. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Shufan Wang, Laure Thompson, Eunsol Choi, Mohit Iyyer |
NAACL-HLT | 5 |
| 2021 | Energy-Based Reranking: Improving Neural Machine Translation Using Energy-Based ModelsabstractSumanta Bhattacharyya, Amirmohammad Rooshenas, Subhajit Naskar, Simeng Sun, Mohit Iyyer, Andrew McCallum. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Sumanta Bhattacharyya, Amirmohammad Rooshenas, Subhajit Naskar, Simeng Sun, Mohit Iyyer, Andrew McCallum |
ACL/IJCNLP (1) | 5 |
| 2021 | WiFiMod: Transformer-based Indoor Human Mobility Modeling using Passive SensingabstractModeling human mobility has a wide range of applications from urban planning to simulations of disease spread. It is well known that humans spend 80% of their time indoors but modeling indoor human mobility is challenging due to three main reasons: (i) the absence of easily acquirable, reliable, low-cost indoor mobility datasets, (ii) high prediction space in modeling the frequent indoor mobility, and (iii) multi-scalar periodicity and correlations in mobility. To deal with all these challenges, we propose WiFiMod, a Transformer-based, data-driven approach that models indoor human mobility at multiple spatial scales using WiFi system logs. WiFiMod takes as input enterprise WiFi system logs to extract human mobility trajectories from smartphone digital traces. Next, for each extracted trajectory, we identify the mobility features at multiple spatial scales, macro and micro, to design a multi-modal embedding Transformer that predicts user mobility for several hours to an entire day across multiple spatial granularities. Multi-modal embedding captures the mobility periodicity and correlations across various scales while Transformers capture long term mobility dependencies boosting model prediction performance. This approach significantly reduces the prediction space by first predicting macro mobility, then modeling indoor scale mobility, micro mobility, conditioned on the estimated macro mobility distribution, thereby using the topological constraint of the macro-scale. Experimental results show that WiFiMod achieves a prediction accuracy of at least 10% points higher than the current state-of-art models. Additionally, we present 3 real-world applications of WiFiMod - (i) predict high density hot pockets and space utilization for policy making decisions for COVID19 or ILI, (ii) generate a realistic simulation of indoor mobility data to simulate spread of diseases, (iii) design personal assistants. Amee Trivedi, Kate Silverstein, Emma Strubell, Prashant J. Shenoy, Mohit Iyyer |
COMPASS | 5 |
| 2021 | Changing the Mind of Transformers for Topically-Controllable Language GenerationabstractLarge Transformer-based language models can aid human authors by suggesting plausible continuations of text written so far.However, current interactive writing assistants do not allow authors to guide text generation in desired topical directions.To address this limitation, we design a framework that displays multiple candidate upcoming topics, of which a user can select a subset to guide the generation.Our framework consists of two components: (1) a method that produces a set of candidate topics by predicting the centers of word clusters in the possible continuations, and ( 2) a text generation model whose output adheres to the chosen topics.The training of both components is self-supervised, using only unlabeled text.Our experiments demonstrate that our topic options are better than those of standard clustering approaches, and our framework often generates fluent sentences related to the chosen topics, as judged by automated metrics and crowdsourced workers. Haw-Shiuan Chang, Jiaming Yuan, Mohit Iyyer, Andrew McCallum |
EACL | 3 |
| 2021 | Weakly-Supervised Open-Retrieval Conversational Question Answering
Chen Qu 0001, Liu Yang 0005, Cen Chen 0001, W. Bruce Croft, Kalpesh Krishna, Mohit Iyyer |
ECIR (1) | 6 |
| 2021 | The Perils of Using Mechanical Turk to Evaluate Open-Ended Text GenerationabstractRecent text generation research has increasingly focused on open-ended domains such as story and poetry generation.Because models built for such tasks are difficult to evaluate automatically, most researchers in the space justify their modeling choices by collecting crowdsourced human judgments of text quality (e.g., Likert scores of coherence or grammaticality) from Amazon Mechanical Turk (AMT).In this paper, we first conduct a survey of 45 open-ended text generation papers and find that the vast majority of them fail to report crucial details about their AMT tasks, hindering reproducibility.We then run a series of story evaluation experiments with both AMT workers and English teachers and discover that even with strict qualification filters, AMT workers (unlike teachers) fail to distinguish between model-generated text and human-generated references.We show that AMT worker judgments improve when they are shown model-generated output alongside human-generated references, which enables the workers to better calibrate their ratings.Finally, interviews with the English teachers provide deeper insights into the challenges of the evaluation process, particularly when rating model-generated text. Marzena Karpinska, Nader Akoury, Mohit Iyyer |
EMNLP (1) | 3 |
| 2021 | Do Long-Range Language Models Actually Use Long-Range Context?abstractLanguage models are generally trained on short, truncated input sequences, which limits their ability to use discourse-level information present in long-range context to improve their predictions.Recent efforts to improve the efficiency of self-attention have led to a proliferation of long-range Transformer language models, which can process much longer sequences than models of the past.However, the ways in which such models take advantage of the longrange context remain unclear.In this paper, we perform a fine-grained analysis of two longrange Transformer language models (including the Routing Transformer, which achieves state-of-the-art perplexity on the PG-19 longsequence LM benchmark dataset) that accept input sequences of up to 8K tokens.Our results reveal that providing long-range context (i.e., beyond the previous 2K tokens) to these models only improves their predictions on a small set of tokens (e.g., those that can be copied from the distant context) and does not help at all for sentence-level prediction tasks.Finally, we discover that PG-19 contains a variety of different document types and domains, and that long-range context helps most for literary novels (as opposed to textbooks or magazines). Simeng Sun, Kalpesh Krishna, Andrew Mattarella-Micke, Mohit Iyyer |
EMNLP (1) | 4 |
| 2021 | IGA: An Intent-Guided Authoring AssistantabstractSimeng Sun, Wenlong Zhao, Varun Manjunatha, Rajiv Jain, Vlad Morariu, Franck Dernoncourt, Balaji Vasan Srinivasan, Mohit Iyyer. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Simeng Sun, Wenlong Zhao 0001, Varun Manjunatha, Rajiv Jain, Vlad I. Morariu, Franck Dernoncourt, Balaji Vasan Srinivasan, Mohit Iyyer |
EMNLP (1) | 8 |
| 2021 | STraTA: Self-Training with Task Augmentation for Better Few-shot LearningabstractDespite their recent successes in tackling many NLP tasks, large-scale pre-trained language models do not perform as well in few-shot settings where only a handful of training examples are available.To address this shortcoming, we propose STraTA, which stands for Self-Training with Task Augmentation, an approach that builds on two key ideas for effective leverage of unlabeled data.First, STraTA uses task augmentation, a novel technique that synthesizes a large amount of data for auxiliary-task fine-tuning from target-task unlabeled texts.Second, STraTA performs selftraining by further fine-tuning the strong base model created by task augmentation on a broad distribution of pseudo-labeled data.Our experiments demonstrate that STraTA can substantially improve sample efficiency across 12 fewshot benchmarks.Remarkably, on the SST-2 sentiment dataset, STraTA, with only 8 training examples per class, achieves comparable results to standard fine-tuning with 67K training examples.Our analyses reveal that task augmentation and self-training are both complementary and independently effective. Tu Vu, Minh-Thang Luong, Quoc V. Le, Grady Simon, Mohit Iyyer |
EMNLP (1) | 5 |
| 2021 | Phrase-BERT: Improved Phrase Embeddings from BERT with an Application to Corpus ExplorationabstractPhrase representations derived from BERT often do not exhibit complex phrasal compositionality, as the model relies instead on lexical similarity to determine semantic relatedness.In this paper, we propose a contrastive fine-tuning objective that enables BERT to produce more powerful phrase embeddings.Our approach (Phrase-BERT) relies on a dataset of diverse phrasal paraphrases, which is automatically generated using a paraphrase generation model, as well as a large-scale dataset of phrases in context mined from the Books3 corpus.Phrase-BERT outperforms baselines across a variety of phrase-level similarity tasks, while also demonstrating increased lexical diversity between nearest neighbors in the vector space.Finally, as a case study, we show that Phrase-BERT embeddings can be easily integrated with a simple autoencoder to build a phrase-based neural topic model that interprets topics as mixtures of words and phrases by performing a nearest neighbor search in the embedding space.Crowdsourced evaluations demonstrate that this phrase-based topic model produces more coherent and meaningful topics than baseline word and phrase-level topic models, further validating the utility of Phrase-BERT. Shufan Wang, Laure Thompson, Mohit Iyyer |
EMNLP (1) | 3 |
| 2021 | Improved Latent Tree Induction with Distant Supervision via Span ConstraintsabstractZhiyang Xu, Andrew Drozdov, Jay Yoon Lee, Tim O’Gorman, Subendhu Rongali, Dylan Finkbeiner, Shilpa Suresh, Mohit Iyyer, Andrew McCallum. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Zhiyang Xu, Andrew Drozdov, Jay-Yoon Lee, Tim O'Gorman, Subendhu Rongali, Dylan Finkbeiner, Shilpa Suresh, Mohit Iyyer, Andrew McCallum |
EMNLP (1) | 8 |
| 2021 | TABBIE: Pretrained Representations of Tabular DataabstractHiroshi Iida, Dung Thai, Varun Manjunatha, Mohit Iyyer. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Hiroshi Iida, Dung Thai, Varun Manjunatha, Mohit Iyyer |
NAACL-HLT | 4 |
| 2021 | Hurdles to Progress in Long-form Question AnsweringabstractThe task of long-form question answering (LFQA) involves retrieving documents relevant to a given question and using them to generate a paragraph-length answer.While many models have recently been proposed for LFQA, we show in this paper that the task formulation raises fundamental challenges regarding evaluation and dataset creation that currently preclude meaningful modeling progress.To demonstrate these challenges, we first design a new system that relies on sparse attention and contrastive retriever learning to achieve state-of-the-art performance on the ELI5 LFQA dataset.While our system tops the public leaderboard, a detailed analysis reveals several troubling trends: (1) our system's generated answers are not actually grounded in the documents that it retrieves; (2) ELI5 contains significant train / validation overlap, as at least 81% of ELI5 validation questions occur in paraphrased form in the training set; (3) ROUGE-L is not an informative metric of generated answer quality and can be easily gamed; and (4) human evaluations used for other text generation tasks are unreliable for LFQA.We offer suggestions to mitigate each of these issues, which we hope will lead to more rigorous LFQA research and meaningful progress in the future. 1 * Work done during an internship at Google Research. 1 Resources accompanying our paper can be found in https://github.com/martiansideofthemoon/ hurdles-longform-qa Val Q: Can you protect electronics from EMPs/solar flares?If so, how?Generation using random retrievals, 24.8 ROUGE-L Yes, you absolutely can, in fact you can build a Faraday cage around your electronics, and protect them from solar flares...This is what is done with the Faraday cage around your electronics, which is the problem.The reason it is expensive is because it requires a huge amount of power and is expensive to replace... designed to shield your electronics from solar flares, you will have to pay for the protection.This is because you have to buy a piece of equipment that is designed to shield your electronics from solar flares, and that is expensive.... This is also expensive, but not as expensive as the protection you need to shield your electronics from solar flares... designed to be as cheap as possible... Generation using predicted retrievals, 19.0 ROUGE-L Yes, you can shield them.But it's a slow process... Also, the equipment that's powered by them is a lot more expensive than you'd think, so it's hard to make sure that you're not just shielding them from your remote control.Train Q1: How does an EMP ruin electronics?What does it do?How would they be fixed?Can It be protected against?How?Train Q2: If Earth were hit with a massive EMP, would all of our currently technology be completely unusable permanently?Train Q3: Whenever a electromagnetic pulse (EMP) is released what does it do to electronics to disable them?Train Q4: If earth was hit with an EMP, could we ever restore electricity?If not, why?Train Q5: What are solar flares and why does it impact our electronics?Train Q6.When an EMP goes off, can the electronics affected be replaced?Gold Answer, 18.6 ROUGE-L I'll start with the grounding question, because that's the easiest to answer: Doesn't help a bit.All that matters is that the metal container is conductive and doesn't have gaps...completely seal your Faraday cage.Consider soldering the lid on to that paint can... look at little baggie it comes in.Sealed mylar.That protected that chip from air travel at 35,000 feet, land travel through rural, urban, and suburban areas, and all the electromagnetic radiation that the trip entails...No lead shielding.No safes.... Random Train Ans, 19.4 ROUGE-LThe fast lane/slow lane is a bit of a misnomer.It gives the impression that new, faster lanes are being built.In reality, normal speed will be... Kalpesh Krishna, Aurko Roy, Mohit Iyyer |
NAACL-HLT | 3 |
| 2021 | Revisiting Simple Neural Probabilistic Language ModelsabstractRecent progress in language modeling has been driven not only by advances in neural architectures, but also through hardware and optimization improvements.In this paper, we revisit the neural probabilistic language model (NPLM) of Bengio et al. (2003), which simply concatenates word embeddings within a fixed window and passes the result through a feed-forward network to predict the next word.When scaled up to modern hardware, this model (despite its many limitations) performs much better than expected on word-level language model benchmarks.Our analysis reveals that the NPLM achieves lower perplexity than a baseline Transformer with short input contexts but struggles to handle long-term dependencies.Inspired by this result, we modify the Transformer by replacing its first selfattention layer with the NPLM's local concatenation layer, which results in small but consistent perplexity decreases across three wordlevel language modeling datasets. Simeng Sun, Mohit Iyyer |
NAACL-HLT | 2 |
| 2020 | Hard-Coded Gaussian Attention for Neural Machine TranslationabstractRecent work has questioned the importance of the Transformer's multi-headed attention for achieving high translation quality.We push further in this direction by developing a "hardcoded" attention variant without any learned parameters.Surprisingly, replacing all learned self-attention heads in the encoder and decoder with fixed, input-agnostic Gaussian distributions minimally impacts BLEU scores across four different language pairs.However, additionally hard-coding cross attention (which connects the decoder to the encoder) significantly lowers BLEU, suggesting that it is more important than self-attention.Much of this BLEU drop can be recovered by adding just a single learned cross attention head to an otherwise hard-coded Transformer.Taken as a whole, our results offer insight into which components of the Transformer are actually important, which we hope will guide future work into the development of simpler and more efficient attention-based models. Weiqiu You, Simeng Sun, Mohit Iyyer |
ACL | 3 |
| 2020 | STORIUM: A Dataset and Evaluation Platform for Machine-in-the-Loop Story GenerationabstractSystems for story generation are asked to produce plausible and enjoyable stories given an input context.This task is underspecified, as a vast number of diverse stories can originate from a single input.The large output space makes it difficult to build and evaluate story generation models, as (1) existing datasets lack rich enough contexts to meaningfully guide models, and (2) existing evaluations (both crowdsourced and automatic) are unreliable for assessing long-form creative text.To address these issues, we introduce a dataset and evaluation platform built from STORIUM, an online collaborative storytelling community.Our author-generated dataset contains 6K lengthy stories (125M tokens) with fine-grained natural language annotations (e.g., character goals and attributes) interspersed throughout each narrative, forming a robust source for guiding models.We evaluate language models fine-tuned on our dataset by integrating them onto STORIUM, where real authors can query a model for suggested story continuations and then edit them.Automatic metrics computed over these edits correlate well with both user ratings of generated stories and qualitative feedback from semi-structured user interviews.We release both the STORIUM dataset and evaluation platform to spur more principled research into story generation. Nader Akoury, Shufan Wang, Josh Whiting, Stephen Hood, Nanyun Peng 0001, Mohit Iyyer |
EMNLP (1) | 6 |
| 2020 | Unsupervised Parsing with S-DIORA: Single Tree Encoding for Deep Inside-Outside Recursive AutoencodersabstractThe deep inside-outside recursive autoencoder (DIORA; Drozdov et al. 2019a) is a selfsupervised neural model that learns to induce syntactic tree structures for input sentences without access to labeled training data.In this paper, we discover that while DIORA exhaustively encodes all possible binary trees of a sentence with a soft dynamic program, its vector averaging approach is locally greedy and cannot recover from errors when computing the highest scoring parse tree in bottom-up chart parsing.To fix this issue, we introduce S-DIORA, an improved variant of DIORA that encodes a single tree rather than a softlyweighted mixture of trees by employing a hard argmax operation and a beam at each cell in the chart.Our experiments show that through fine-tuning a pre-trained DIORA with our new algorithm, we improve the state of the art in unsupervised constituency parsing on the English WSJ Penn Treebank by 2.2 6% F1, depending on the data used for fine-tuning. Andrew Drozdov, Subendhu Rongali, Yi-Pei Chen 0001, Tim O'Gorman, Mohit Iyyer, Andrew McCallum |
EMNLP (1) | 5 |
| 2020 | Reformulating Unsupervised Style Transfer as Paraphrase GenerationabstractModern NLP defines the task of style transfer as modifying the style of a given sentence without appreciably changing its semantics, which implies that the outputs of style transfer systems should be paraphrases of their inputs.However, many existing systems purportedly designed for style transfer inherently warp the input's meaning through attribute transfer, which changes semantic properties such as sentiment.In this paper, we reformulate unsupervised style transfer as a paraphrase generation problem, and present a simple methodology based on fine-tuning pretrained language models on automatically generated paraphrase data.Despite its simplicity, our method significantly outperforms state-of-the-art style transfer systems on both human and automatic evaluations.We also survey 23 style transfer papers and discover that existing automatic metrics can be easily gamed and propose fixed variants.Finally, we pivot to a more real-world style transfer setting by collecting a large dataset of 15M sentences in 11 diverse styles, which we use for an in-depth analysis of our system. Kalpesh Krishna, John Wieting, Mohit Iyyer |
EMNLP (1) | 3 |
| 2020 | Exploring and Predicting Transferability across NLP TasksabstractTu Vu, Tong Wang, Tsendsuren Munkhdalai, Alessandro Sordoni, Adam Trischler, Andrew Mattarella-Micke, Subhransu Maji, Mohit Iyyer. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Tu Vu, Tong Wang 0012, Tsendsuren Munkhdalai, Alessandro Sordoni, Adam Trischler, Andrew Mattarella-Micke, Subhransu Maji, Mohit Iyyer |
EMNLP (1) | 8 |
| 2020 | Thieves on Sesame Street! Model Extraction of BERT-based APIs
Kalpesh Krishna, Gaurav Tomar, Ankur P. Parikh, Nicolas Papernot, Mohit Iyyer |
ICLR | 5 |
| 2020 | Which Evaluations Uncover Sense Representations that Actually Make Sense?abstractText representations are critical for modern natural language processing. One form of text representation, sense-specific embeddings, reflect a word’s sense in a sentence better than single-prototype word embeddings tied to each type. However, existing sense representations are not uniformly better: although they work well for computer-centric evaluations, they fail for human-centric tasks like inspecting a language’s sense inventory. To expose this discrepancy, we propose a new coherence evaluation for sense embeddings. We also describe a minimal model (Gumbel Attention for Sense Induction) optimized for discovering interpretable sense representations that are more coherent than existing sense embeddings. Jordan L. Boyd-Graber, Fenfei Guo, Leah Findlater, Mohit Iyyer |
LREC | 4 |
| 2020 | Open-Retrieval Conversational Question AnsweringabstractConversational search is one of the ultimate goals of information retrieval. Recent research approaches conversational search by simplified settings of response ranking and conversational question answering, where an answer is either selected from a given candidate set or extracted from a given passage. These simplifications neglect the fundamental role of retrieval in conversational search. To address this limitation, we introduce an open-retrieval conversational question answering (ORConvQA) setting, where we learn to retrieve evidence from a large collection before extracting answers, as a further step towards building functional conversational search systems. We create a dataset, OR-QuAC, to facilitate research on ORConvQA. We build an end-to-end system for ORConvQA, featuring a retriever, a reranker, and a reader that are all based on Transformers. Our extensive experiments on OR-QuAC demonstrate that a learnable retriever is crucial for ORConvQA. We further show that our system can make a substantial improvement when we enable history modeling in all system components. Moreover, we show that the reranker component contributes to the model performance by providing a regularization effect. Finally, further in-depth analyses are performed to provide new insights into ORConvQA. Chen Qu 0001, Liu Yang 0005, Cen Chen 0001, Minghui Qiu, W. Bruce Croft, Mohit Iyyer |
SIGIR | 6 |
| 2019 | Syntactically Supervised Transformers for Faster Neural Machine TranslationabstractStandard decoders for neural machine translation autoregressively generate a single target token per time step, which slows inference especially for long outputs.While architectural advances such as the Transformer fully parallelize the decoder computations at training time, inference still proceeds sequentially.Recent developments in nonand semiautoregressive decoding produce multiple tokens per time step independently of the others, which improves inference speed but deteriorates translation quality.In this work, we propose the syntactically supervised Transformer (SynST), which first autoregressively predicts a chunked parse tree before generating all of the target tokens in one shot conditioned on the predicted parse.A series of controlled experiments demonstrates that SynST decodes sentences ∼ 5× faster than the baseline autoregressive Transformer while achieving higher BLEU scores than most competing methods on En-De and En-Fr datasets. Nader Akoury, Kalpesh Krishna, Mohit Iyyer |
ACL (1) | 3 |
| 2019 | Generating Question-Answer HierarchiesabstractThe process of knowledge acquisition can be viewed as a question-answer game between a student and a teacher in which the student typically starts by asking broad, open-ended questions before drilling down into specifics (Hintikka, 1981;Hakkarainen and Sintonen, 2002).This pedagogical perspective motivates a new way of representing documents.In this paper, we present SQUASH (Specificity-controlled Question-Answer Hierarchies), a novel and challenging text generation task that converts an input document into a hierarchy of question-answer pairs.Users can click on high-level questions (e.g., "Why did Frodo leave the Fellowship?") to reveal related but more specific questions (e.g., "Who did Frodo leave with?").Using a question taxonomy loosely based on Lehnert (1978), we classify questions in existing reading comprehension datasets as either GENERAL or SPECIFIC.We then use these labels as input to a pipelined system centered around a conditional neural language model.We extensively evaluate the quality of the generated QA hierarchies through crowdsourced experiments and report strong empirical results. Kalpesh Krishna, Mohit Iyyer |
ACL (1) | 2 |
| 2019 | Encouraging Paragraph Embeddings to Remember Sentence Identity Improves ClassificationabstractWhile paragraph embedding models are remarkably effective for downstream classification tasks, what they learn and encode into a single vector remains opaque.In this paper, we investigate a state-of-the-art paragraph embedding method proposed by Zhang et al. (2017) and discover that it cannot reliably tell whether a given sentence occurs in the input paragraph or not.We formulate a sentence content task to probe for this basic linguistic property and find that even a much simpler bag-of-words method has no trouble solving it.This result motivates us to replace the reconstructionbased objective of Zhang et al. (2017) with our sentence content probe objective in a semisupervised setting.Despite its simplicity, our objective improves over paragraph reconstruction in terms of (1) downstream classification accuracies on benchmark datasets, (2) faster training, and (3) better generalization ability. 1 Tu Vu, Mohit Iyyer |
ACL (1) | 2 |
| 2019 | Attentive History Selection for Conversational Question AnsweringabstractConversational question answering (ConvQA) is a simplified but concrete setting of conversational search. One of its major challenges is to leverage the conversation history to understand and answer the current question. In this work, we propose a novel solution for ConvQA that involves three aspects. First, we propose a positional history answer embedding method to encode conversation history with position information using BERT in a natural way. BERT is a powerful technique for text representation. Second, we design a history attention mechanism (HAM) to conduct a "soft selection" for conversation histories. This method attends to history turns with different weights based on how helpful they are on answering the current question. Third, in addition to handling conversation history, we take advantage of multi-task learning (MTL) to do answer prediction along with another essential conversation task (dialog act prediction) using a uniform model architecture. MTL is able to learn more expressive and generic representations to improve the performance of ConvQA. We demonstrate the effectiveness of our model with extensive experimental evaluations on QuAC, a large-scale ConvQA dataset. We show that position information plays an important role in conversation history modeling. We also visualize the history attention and provide new insights into conversation history understanding. Chen Qu 0001, Liu Yang 0005, Minghui Qiu, Yongfeng Zhang 0003, Cen Chen 0001, W. Bruce Croft, Mohit Iyyer |
CIKM | 7 |
| 2019 | Unsupervised Labeled Parsing with Deep Inside-Outside Recursive AutoencodersabstractAndrew Drozdov, Patrick Verga, Yi-Pei Chen, Mohit Iyyer, Andrew McCallum. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Andrew Drozdov, Patrick Verga, Yi-Pei Chen 0001, Mohit Iyyer, Andrew McCallum |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Investigating Sports Commentator Bias within a Large Corpus of American Football BroadcastsabstractJack Merullo, Luke Yeh, Abram Handler, Alvin Grissom II, Brendan O’Connor, Mohit Iyyer. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jack Merullo, Luke Yeh, Abram Handler, Alvin Grissom II, Brendan T. O'Connor 0001, Mohit Iyyer |
EMNLP/IJCNLP (1) | 6 |
| 2019 | BERT with History Answer Embedding for Conversational Question AnsweringabstractConversational search is an emerging topic in the information retrieval community. One of the major challenges to multi-turn conversational search is to model the conversation history to answer the current question. Existing methods either prepend history turns to the current question or use complicated attention mechanisms to model the history. We propose a conceptually simple yet highly effective approach referred to as history answer embedding. It enables seamless integration of conversation history into a conversational question answering (ConvQA) model built on BERT (Bidirectional Encoder Representations from Transformers). We first explain our view that ConvQA is a simplified but concrete setting of conversational search, and then we provide a general framework to solve ConvQA. We further demonstrate the effectiveness of our approach under this framework. Finally, we analyze the impact of different numbers of history turns under different settings to provide new insights into conversation history modeling in ConvQA. Chen Qu 0001, Liu Yang 0005, Minghui Qiu, W. Bruce Croft, Yongfeng Zhang 0003, Mohit Iyyer |
SIGIR | 6 |
| 2018 | QuAC: Question Answering in ContextabstractWe present QuAC, a dataset for Question Answering in Context that contains 14K information-seeking QA dialogs (100K questions in total).The dialogs involve two crowd workers: (1) a student who poses a sequence of freeform questions to learn as much as possible about a hidden Wikipedia text, and (2) a teacher who answers the questions by providing short excerpts from the text.QuAC introduces challenges not found in existing machine comprehension datasets: its questions are often more open-ended, unanswerable, or only meaningful within the dialog context, as we show in a detailed qualitative evaluation.We also report results for a number of reference models, including a recently state-ofthe-art reading comprehension architecture extended to model dialog context.Our best model underperforms humans by 20 F1, suggesting that there is significant room for future work on this data.Dataset, baseline, and leaderboard available at http://quac.ai.How was perversion handled?How long was he there?How popular did she become?How did Mark Felt contact Woodword?How did the meeting go?How did it do on the charts?When was she born?When was it founded?When was the breakup? Eunsol Choi, He He 0001, Mohit Iyyer, Mark Yatskar, Scott Yih, Yejin Choi 0001, Percy Liang, Luke Zettlemoyer |
EMNLP | 3 |
| 2018 | Pathologies of Neural Models Make Interpretation DifficultabstractOne way to interpret neural model predictions is to highlight the most important input features-for example, a heatmap visualization over the words in an input sentence.In existing interpretation methods for NLP, a word's importance is determined by either input perturbation-measuring the decrease in model confidence when that word is removed-or by the gradient with respect to that word.To understand the limitations of these methods, we use input reduction, which iteratively removes the least important word from the input.This exposes pathological behaviors of neural models: the remaining words appear nonsensical to humans and are not the ones determined as important by interpretation methods.As we confirm with human experiments, the reduced examples lack information to support the prediction of any label, but models still make the same predictions with high confidence.To explain these counterintuitive results, we draw connections to adversarial examples and confidence calibration: pathological behaviors reveal difficulties in interpreting neural models trained with maximum likelihood.To mitigate their deficiencies, we fine-tune the models by encouraging high entropy outputs on reduced examples.Fine-tuned models become more interpretable under input reduction without accuracy loss on regular examples. Shi Feng 0005, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodríguez 0001, Jordan L. Boyd-Graber |
EMNLP | 4 |
| 2018 | Revisiting the Importance of Encoding Logic Rules in Sentiment ClassificationabstractWe analyze the performance of different sentiment classification models on syntacticallycomplex inputs like A-but-B sentences.The first contribution of this analysis addresses reproducible research: to meaningfully compare different models, their accuracies must be averaged over far more random seeds than what has traditionally been reported.With proper averaging in place, we notice that the distillation model described in Hu et al. (2016), which incorporates explicit logic rules for sentiment classification, is ineffective.In contrast, using contextualized ELMo embeddings (Peters et al., 2018a) instead of logic rules yields significantly better performance.Additionally, we provide analysis and visualizations that demonstrate ELMo's ability to implicitly learn logic rules.Finally, a crowdsourced analysis reveals how ELMo outperforms baseline models even on sentences with ambiguous sentiment labels. Kalpesh Krishna, Preethi Jyothi, Mohit Iyyer |
EMNLP | 3 |
| 2018 | Adversarial Example Generation with Syntactically Controlled Paraphrase NetworksabstractMohit Iyyer, John Wieting, Kevin Gimpel, Luke Zettlemoyer. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Mohit Iyyer, John Wieting, Kevin Gimpel, Luke Zettlemoyer |
NAACL-HLT | 1 |
| 2018 | Deep Contextualized Word RepresentationsabstractMatthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, Luke Zettlemoyer. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner 0001, Kenton Lee, Luke Zettlemoyer |
NAACL-HLT | 3 |
| 2017 | Unsupervised Learning of Evolving Relationships Between Literary CharactersabstractUnderstanding inter-character relationships is fundamental for understanding character intentions and goals in a narrative. This paper addresses unsupervised modeling of relationships between characters. We model relationships as dynamic phenomenon, represented as evolving sequences of latent states empirically learned from data. Unlike most previous work our approach is completely unsupervised. This enables data-driven inference of inter-character relationship types beyond simple sentiment polarities, by incorporating lexical and semantic representations, and leveraging large quantities of raw text. We present three models based on rich sets of linguistic features that capture various cues about relationships. We compare these models with existing techniques and also demonstrate that relationship categories learned by our model are semantically coherent. Snigdha Chaturvedi, Mohit Iyyer, Hal Daumé III |
AAAI | 2 |
| 2017 | Search-based Neural Structured Learning for Sequential Question AnsweringabstractRecent work in semantic parsing for question answering has focused on long and complicated questions, many of which would seem unnatural if asked in a normal conversation between two humans.In an effort to explore a conversational QA setting, we present a more realistic task: answering sequences of simple but inter-related questions.We collect a dataset of 6,066 question sequences that inquire about semistructured tables from Wikipedia, with 17,553 question-answer pairs in total.To solve this sequential question answering task, we propose a novel dynamic neural semantic parsing framework trained using a weakly supervised reward-guided search.Our model effectively leverages the sequential context to outperform state-of-the-art QA systems that are designed to answer highly complex questions. Mohit Iyyer, Scott Yih, Ming-Wei Chang |
ACL (1) | 1 |
| 2017 | The Amazing Mysteries of the Gutter: Drawing Inferences Between Panels in Comic Book NarrativesabstractVisual narrative is often a combination of explicit information and judicious omissions, relying on the viewer to supply missing details. In comics, most movements in time and space are hidden in the gutters between panels. To follow the story, readers logically connect panels together by inferring unseen actions through a process called closure. While computers can now describe the content of natural images, in this paper we examine whether they can understand the closure-driven narratives conveyed by stylized artwork and dialogue in comic book panels. We collect a dataset, COMICS, that consists of over 1.2 million panels (120 GB) paired with automatic textbox transcriptions. An in-depth analysis of COMICS demonstrates that neither text nor image alone can tell a comic book story, so a computer must understand both modalities to keep up with the plot. We introduce three cloze-style tasks that ask models to predict narrative and character-centric aspects of a panel given n preceding panels as context. Various deep neural architectures underperform human baselines on these tasks, suggesting that COMICS contains fundamental challenges for both vision and language. Mohit Iyyer, Varun Manjunatha, Anupam Guha, Yogarshi Vyas, Jordan L. Boyd-Graber, Hal Daumé III, Larry Davis 0001 |
CVPR | 1 |
| 2016 | Ask Me Anything: Dynamic Memory Networks for Natural Language ProcessingabstractMost tasks in natural language processing can be cast into question answering (QA) problems over language input. We introduce the dynamic memory network (DMN), a neural network architecture which processes input sequences and questions, forms episodic memories, and generates relevant answers. Questions trigger an iterative attention process which allows the model to condition its attention on the inputs and the result of previous iterations. These results are then reasoned over in a hierarchical recurrent sequence model to generate answers. The DMN can be trained end-to-end and obtains state-of-the-art results on several types of tasks and datasets: question answering (Facebook’s bAbI dataset), text classification for sentiment analysis (Stanford Sentiment Treebank) and sequence modeling for part-of-speech tagging (WSJ-PTB). The training for these different tasks relies exclusively on trained word vector representations and input-question-answer triplets. Ozan Irsoy, Peter Ondruska, Mohit Iyyer, James Bradbury 0002, Ishaan Gulrajani, Victor Zhong, Romain Paulus, Richard Socher |
ICML | 4 |
| 2016 | Feuding Families and Former Friends: Unsupervised Learning for Dynamic Fictional RelationshipsabstractMohit Iyyer, Anupam Guha, Snigdha Chaturvedi, Jordan Boyd-Graber, Hal Daumé III. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Mohit Iyyer, Anupam Guha, Snigdha Chaturvedi, Jordan L. Boyd-Graber, Hal Daumé III |
HLT-NAACL | 1 |
| 2015 | Deep Unordered Composition Rivals Syntactic Methods for Text ClassificationabstractMohit Iyyer, Varun Manjunatha, Jordan Boyd-Graber, Hal Daumé III. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Mohit Iyyer, Varun Manjunatha, Jordan L. Boyd-Graber, Hal Daumé III |
ACL (1) | 1 |
| 2015 | Removing the Training Wheels: A Coreference Dataset that Entertains Humans and Challenges ComputersabstractAnupam Guha, Mohit Iyyer, Danny Bouman, Jordan Boyd-Graber. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Anupam Guha, Mohit Iyyer, Danny Bouman, Jordan L. Boyd-Graber |
HLT-NAACL | 2 |
| 2014 | Political Ideology Detection Using Recursive Neural NetworksabstractAn individual's words often reveal their political ideology.Existing automated techniques to identify ideology from text focus on bags of words or wordlists, ignoring syntax.Taking inspiration from recent work in sentiment analysis that successfully models the compositional aspect of language, we apply a recursive neural network (RNN) framework to the task of identifying the political position evinced by a sentence.To show the importance of modeling subsentential elements, we crowdsource political annotations at a phrase and sentence level.Our model outperforms existing models on our newly annotated dataset and an existing dataset. Mohit Iyyer, Peter Enns, Jordan L. Boyd-Graber, Philip Resnik |
ACL (1) | 1 |
| 2014 | A Neural Network for Factoid Question Answering over ParagraphsabstractText classification methods for tasks like factoid question answering typi-cally use manually defined string match-ing rules or bag of words representa-tions. These methods are ineffective when question text contains very few individual words (e.g., named entities) that are indicative of the answer. We introduce a recursive neural network (rnn) model that can reason over such input by modeling textual composition-ality. We apply our model, qanta, to a dataset of questions from a trivia competition called quiz bowl. Unlike previous rnn models, qanta learns word and phrase-level representations that combine across sentences to reason about entities. The model outperforms multiple baselines and, when combined with information retrieval methods, ri-vals the best human players. 1 Mohit Iyyer, Jordan L. Boyd-Graber, Leonardo Max Batista Claudino, Richard Socher, Hal Daumé III |
EMNLP | 1 |