Tanya Goyal

dblp:176/9145 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
11since 2021 · last 2025
0009-0009-6429-278XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 RefreshKV: Updating Small KV Cache During Long-form Generation
abstract
Generating long sequences of tokens given a long-context input is a very compute-intensive inference scenario for large language models (LLMs).One prominent inference speed-up approach is to construct a smaller key-value (KV) cache, relieving LLMs from computing attention over a long sequence of tokens.While such methods work well to generate short sequences, their performance degrades rapidly for longform generation.Most KV compression happens once, prematurely removing tokens that can be useful later in the generation.We propose a new inference method, RefreshKV, that flexibly alternates between full context attention and attention over a subset of input tokens during generation.After each full attention step, we update the smaller KV cache based on the attention pattern over the entire input.Applying our method to off-the-shelf LLMs achieves comparable speedup to eviction-based methods while improving performance for various long-form generation tasks.Lastly, we show that continued pretraining with our inference setting brings further gains in performance.
Tanya Goyal, Eunsol Choi
ACL (1)2
2024 LitSearch: A Retrieval Benchmark for Scientific Literature Search
abstract
Literature search questions, such as "Where can I find research on the evaluation of consistency in generated summaries?"pose significant challenges for modern search engines and retrieval systems.These questions often require a deep understanding of research concepts and the ability to reason across entire articles.In this work, we introduce LitSearch, a retrieval benchmark comprising 597 realistic literature search queries about recent ML and NLP papers.Lit-Search is constructed using a combination of (1) questions generated by GPT-4 based on paragraphs containing inline citations from research papers and (2) questions manually written by authors about their recently published papers.All LitSearch questions were manually examined or edited by experts to ensure high quality.We extensively benchmark state-ofthe-art retrieval models and also evaluate two LLM-based reranking pipelines.We find a significant performance gap between BM25 and state-of-the-art dense retrievers, with a 24.8% absolute difference in [email protected] LLMbased reranking strategies further improve the best-performing dense retriever by 4.4%.Additionally, commercial search engines and research tools like Google Search perform poorly on LitSearch, lagging behind the best dense retriever by up to 32 recall points.Taken together, these results show that LitSearch is an informative new testbed for retrieval systems while catering to a real-world use case.
Anirudh Ajith, Mengzhou Xia, Alexis Chevalier, Tanya Goyal, Danqi Chen 0001, Tianyu Gao 0001
EMNLP4
2024 One Thousand and One Pairs: A "novel" challenge for long-context language models
abstract
Synthetic long-context LLM benchmarks (e.g., "needle-in-the-haystack") test only surfacelevel retrieval capabilities, but how well can long-context LLMs retrieve, synthesize, and reason over information across book-length inputs?We address this question by creating NOCHA, a dataset of 1,001 minimally different pairs of true and false claims about 67 recentlypublished English fictional books, written by human readers of those books.In contrast to existing long-context benchmarks, our annotators confirm that the largest share of pairs in NOCHA require global reasoning over the entire book to verify.Our experiments show that while human readers easily perform this task, it is enormously challenging for all ten long-context LLMs that we evaluate: no open-weight model performs above random chance (despite their strong performance on synthetic benchmarks), while GPT-4O achieves the highest accuracy at 55.8%.Further analysis reveals that (1) on average, models perform much better on pairs that require only sentence-level retrieval vs. global reasoning; (2) model-generated explanations for their decisions are often inaccurate even for correctly-labeled claims; and (3) models perform substantially worse on speculative fiction books that contain extensive world-building.The methodology proposed in NOCHA allows for the evolution of the benchmark dataset and the easy analysis of future models.TRUE TRUE Niema takes her students out to world's end and tells them that the fog kills anything it touches, a statement backed by her extensive research into the fog.When Niema takes her students out to world's end and tells them that the fog kills anything it touches, she is intentionally lying to them.
Marzena Karpinska, Katherine Thai, Kyle Lo, Tanya Goyal, Mohit Iyyer
EMNLP4
2024 BooookScore: A systematic exploration of book-length summarization in the era of LLMs
abstract
Summarizing book-length documents ($>$100K tokens) that exceed the context window size of large language models (LLMs) requires first breaking the input document into smaller chunks and then prompting an LLM to merge, update, and compress chunk-level summaries. Despite the complexity and importance of this task, it has yet to be meaningfully studied due to the challenges of evaluation: existing book-length summarization datasets (e.g., BookSum) are in the pretraining data of most public LLMs, and existing evaluation methods struggle to capture errors made by modern LLM summarizers. In this paper, we present the first study of the coherence of LLM-based book-length summarizers implemented via two prompting workflows: (1) hierarchically merging chunk-level summaries, and (2) incrementally updating a running summary. We obtain 1193 fine-grained human annotations on GPT-4 generated summaries of 100 recently-published books and identify eight common types of coherence errors made by LLMs. Because human evaluation is expensive and time-consuming, we develop an automatic metric, BooookScore, that measures the proportion of sentences in a summary that do not contain any of the identified error types. BooookScore has high agreement with human annotations and allows us to systematically evaluate the impact of many other critical parameters (e.g., chunk size, base LLM) while saving \$15K USD and 500 hours in human evaluation costs. We find that closed-source LLMs such as GPT-4 and Claude 2 produce summaries with higher BooookScore than those generated by open-source models. While LLaMA 2 falls behind other models, Mixtral achieves performance on par with GPT-3.5-Turbo. Incremental updating yields lower BooookScore but higher level of detail than hierarchical merging, a trade-off sometimes preferred by annotators. We release code and annotations to spur more principled research on book-length summarization.
Yapei Chang, Kyle Lo, Tanya Goyal, Mohit Iyyer
ICLR3
2024 Evaluating Large Language Models at Evaluating Instruction Following
abstract
As research in large language models (LLMs) continues to accelerate, LLM-based evaluation has emerged as a scalable and cost-effective alternative to human evaluations for comparing the ever increasing list of models. This paper investigates the efficacy of these “LLM evaluators”, particularly in using them to assess instruction following, a metric that gauges how closely generated text adheres to the given instruction. We introduce a challenging meta-evaluation benchmark, LLMBar, designed to test the ability of an LLM evaluator in discerning instruction-following outputs. The authors manually curated 419 pairs of outputs, one adhering to instructions while the other diverging, yet may possess deceptive qualities that mislead an LLM evaluator, e.g., a more engaging tone. Contrary to existing meta-evaluation, we discover that different evaluators (i.e., combinations of LLMs and prompts) exhibit distinct performance on LLMBar and even the highest-scoring ones have substantial room for improvement. We also present a novel suite of prompting strategies that further close the gap between LLM and human evaluators. With LLMBar, we hope to offer more insight into LLM evaluators and foster future research in developing better instruction-following models.
Jiatong Yu, Tianyu Gao 0001, Yu Meng 0001, Tanya Goyal, Danqi Chen 0001
ICLR5
2023 Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error Detectors
abstract
Liyan Tang, Tanya Goyal, Alex Fabbri, Philippe Laban, Jiacheng Xu, Semih Yavuz, Wojciech Kryscinski, Justin Rousseau, Greg Durrett. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Liyan Tang, Tanya Goyal, Alexander R. Fabbri, Philippe Laban, Jiacheng Xu 0001, Semih Yavuz, Wojciech Kryscinski, Justin F. Rousseau, Greg Durrett
ACL (1)2
2023 Shortcomings of Question Answering Based Factuality Frameworks for Error Localization
abstract
Despite recent progress in abstractive summarization, models often generate summaries with factual errors.Numerous approaches to detect these errors have been proposed, the most popular of which are question answering (QA)based factuality metrics.These have been shown to work well at predicting summarylevel factuality and have potential to localize errors within summaries, but this latter capability has not been systematically evaluated in past research.In this paper, we conduct the first such analysis and find that, contrary to our expectations, QA-based frameworks fail to correctly identify error spans in generated summaries and are outperformed by trivial exact match baselines.Our analysis reveals a major reason for such poor localization: questions generated by the QG module often inherit errors from non-factual summaries which are then propagated further into downstream modules.Moreover, even human-in-the-loop question generation cannot easily offset these problems.Our experiments conclusively show that there exist fundamental issues with localization using the QA framework which cannot be fixed solely by stronger QA and QG models.
Ryo Kamoi, Tanya Goyal, Greg Durrett
EACL2
2023 WiCE: Real-World Entailment for Claims in Wikipedia
abstract
Textual entailment models are increasingly applied in settings like fact-checking, presupposition verification in question answering, or summary evaluation.However, these represent a significant domain shift from existing entailment datasets, and models underperform as a result.We propose WICE, a new fine-grained textual entailment dataset built on natural claim and evidence pairs extracted from Wikipedia.In addition to standard claim-level entailment, WICE provides entailment judgments over subsentence units of the claim, and a minimal subset of evidence sentences that support each subclaim.To support this, we propose an automatic claim decomposition strategy using GPT-3.5 which we show is also effective at improving entailment models' performance on multiple datasets at test time.Finally, we show that real claims in our dataset involve challenging verification and retrieval problems that existing models fail to address. 1
Ryo Kamoi, Tanya Goyal, Juan Diego Rodriguez, Greg Durrett
EMNLP2
2022 SNaC: Coherence Error Detection for Narrative Summarization
abstract
Progress in summarizing long texts is inhibited by the lack of appropriate evaluation frameworks.A long summary that appropriately covers the facets of that text must also present a coherent narrative, but current automatic and human evaluation methods fail to identify gaps in coherence.In this work, we introduce SNAC, a narrative coherence evaluation framework for fine-grained annotations of long summaries.We develop a taxonomy of coherence errors in generated narrative summaries and collect spanlevel annotations for 6.6k sentences across 150 book and movie summaries.Our work provides the first characterization of coherence errors generated by state-of-the-art summarization models and a protocol for eliciting coherence judgments from crowdworkers.Furthermore, we show that the collected annotations allow us to benchmark past work in coherence modeling and train a strong classifier for automatically localizing coherence errors in generated summaries.Finally, our SNAC framework can support future work in long document summarization and coherence evaluation, including improved summarization modeling and posthoc summary correction.
Tanya Goyal, Junyi Jessy Li, Greg Durrett
EMNLP1
2022 HydraSum: Disentangling Style Features in Text Summarization with Multi-Decoder Models
abstract
Summarization systems make numerous "decisions" about summary properties during inference, e.g.degree of copying, specificity and length of outputs, etc.However, these are implicitly encoded within model parameters and specific styles cannot be enforced.To address this, we introduce HYDRASUM, a new summarization architecture that extends the single decoder framework of current models to a mixture-of-experts version with multiple decoders.We show that HYDRASUM's multiple decoders automatically learn contrasting summary styles when trained under the standard training objective without any extra supervision.Through experiments on three summarization datasets (CNN, NEWSROOM and XSUM), we show that HYDRASUM provides a simple mechanism to obtain stylistically-diverse summaries by sampling from either individual decoders or their mixtures, outperforming baseline models.Finally, we demonstrate that a small modification to the gating strategy during training can enforce an even stricter style partitioning, e.g.high-vs low-abstractiveness or high-vs low-specificity, allowing users to sample from a larger area in the generation space and vary summary styles along multiple dimensions. 1Input Article: Insights into the workings of the human body that Leonardo da Vinci could only obtain by dissecting scores of corpses and recording the results in exquisite drawings will be displayed for the first time beside modern 3D films, CT and MRI scans, which show how close the Renaissance genius got to the truth of what lies under the skin.[…] the Edinburgh show will be the first to compare Leonardo's results with scalpel and pen with the best results of modern technology.[…] The exhibition will show how close Leonardo got in some of his last medical experiments to discovering the role of the beating heart in the circulation of the blood, a century before William Harvey worked it out.[…] Edinburgh show will be first to compare Renaissance genius's results with best results of modern technology.Edinburgh show will be first to compare Renaissance genius's results with the best results of modern technology.Edinburgh show will be first to compare Leonardo's results with the best results of modern technology. Low diversity Baseline BARTEdinburgh show will be first to compare Leonardo's results with best results of modern technology.Modern imaging techniques will be displayed alongside Leonardo da Vinci's anatomical drawings in Edinburgh exhibition.
Tanya Goyal, Nazneen Fatema Rajani, Wenhao Liu 0003, Wojciech Kryscinski
EMNLP1
2021 Annotating and Modeling Fine-grained Factuality in Summarization
abstract
Recent pre-trained abstractive summarization systems have started to achieve credible performance, but a major barrier to their use in practice is their propensity to output summaries that are not faithful to the input and that contain factual errors.While a number of annotated datasets and statistical models for assessing factuality have been explored, there is no clear picture of what errors are most important to target or where current techniques are succeeding and failing.We explore both synthetic and human-labeled data sources for training models to identify factual errors in summarization, and study factuality at the word-, dependency-, and sentence-level.Our observations are threefold.First, exhibited factual errors differ significantly across datasets, and commonly-used training sets of simple synthetic errors do not reflect errors made on abstractive datasets like XSUM.Second, human-labeled data with fine-grained annotations provides a more effective training signal than sentence-level annotations or synthetic data.Finally, we show that our best factuality detection model enables training of more factual XSUM summarization models by allowing us to identify non-factual tokens in the training data. 1Reference Summary: An early-medieval gold pendant created from an imitation of a Byzantine coin that was found in a Norfolk field is a "rare find", a museum expert has said. Source Article Fragment:Discovered on land at North Elmham, near Dereham, the circa 600 AD coin was created by French rulers of the time to increase their available currency.[…] The pendant was declared treasure by the Norfolk coroner on Wednesday.An 18th century coin believed to be worth more than #1m has been discovered.A gold pendant created from a necklace was found in a field Entitycentric (Ent-C)The pendant was declared a treasure by the Norfolk coroner on Wednesday.The pendant was declared a treasure by the Ohio coroner on March.
Tanya Goyal, Greg Durrett
NAACL-HLT1
2020 Neural Syntactic Preordering for Controlled Paraphrase Generation
abstract
Paraphrasing natural language sentences is a multifaceted process: it might involve replacing individual words or short phrases, local rearrangement of content, or high-level restructuring like topicalization or passivization.Past approaches struggle to cover this space of paraphrase possibilities in an interpretable manner.Our work, inspired by pre-ordering literature in machine translation, uses syntactic transformations to softly "reorder" the source sentence and guide our neural paraphrasing model.First, given an input sentence, we derive a set of feasible syntactic rearrangements using an encoder-decoder model.This model operates over a partially lexical, partially syntactic view of the sentence and can reorder big chunks.Next, we use each proposed rearrangement to produce a sequence of position embeddings, which encourages our final encoder-decoder paraphrase model to attend to the source words in a particular order.Our evaluation, both automatic and human, shows that the proposed system retains the quality of the baseline approaches while giving a substantial increase in the diversity of the generated paraphrases.
Tanya Goyal, Greg Durrett
ACL1
2019 Embedding Time Expressions for Deep Temporal Ordering Models
abstract
Data-driven models have demonstrated stateof-the-art performance in inferring the temporal ordering of events in text.However, these models often overlook explicit temporal signals, such as dates and time windows.Rule-based methods can be used to identify the temporal links between these time expressions (timexes), but they fail to capture timexes' interactions with events and are hard to integrate with the distributed representations of neural net models.In this paper, we introduce a framework to infuse temporal awareness into such models by learning a pre-trained model to embed timexes.We generate synthetic data consisting of pairs of timexes, then train a character LSTM to learn embeddings and classify the timexes' temporal relation.We evaluate the utility of these embeddings in the context of a strong neural model for event temporal ordering, and show a small increase in performance on the MATRES dataset and more substantial gains on an automatically collected dataset with more frequent event-timex interactions.1
Tanya Goyal, Greg Durrett
ACL (1)1
2018 Your Behavior Signals Your Reliability: Modeling Crowd Behavioral Traces to Ensure Quality Relevance Annotations
abstract
While peer-agreement and gold checks are well-established methods for ensuring quality in crowdsourced data collection, we explore a relatively new direction for quality control: estimating work quality directly from workers’ behavioral traces collected during annotation. We propose three behavior-based models to predict label correctness and worker accuracy, then further apply model predictions to label aggregation and optimization of label collection. As part of this work, we collect and share a new Mechanical Turk dataset of behavioral signals judging the relevance of search results. Results show that behavioral data can be effectively used to predict work quality, which could be especially useful with single labeling or in a cold start scenario in which individuals’ prior work history is unavailable. We further show improvement in label aggregation and reducing labeling cost while ensuring data quality.
Tanya Goyal, Tyler McDonnell, Mucahid Kutlu, Tamer Elsayed, Matthew Lease
HCOMP1
2018 Harvesting Knowledge from Cultural Heritage Artifacts in Museums of India
Abhilasha Sancheti, Paridhi Maheshwari, Rajat Chaturvedi, Anish V. Monsy, Tanya Goyal, Balaji Vasan Srinivasan
PAKDD (2)5
2017 An Empirical Analysis of Edit Importance between Document Versions
abstract
In this paper, we present a novel approach to infer significance of various textual edits to documents.An author may make several edits to a document; each edit varies in its impact to the content of the document.While some edits are surface changes and introduce negligible change, other edits may change the content/tone of the document significantly.In this paper, we perform an analysis of the human perceptions of edit importance while reviewing documents from one version to the next.We identify linguistic features that influence edit importance and model it in a regression based setting.We show that the predicted importance by our approach is highly correlated with the human perceived importance, established by a Mechanical Turk study.
Tanya Goyal, Sachin Kelkar, Manas Agarwal, Jeenu Grover
EMNLP1
2017 Preventing Inadvertent Information Disclosures via Automatic Security Policies
Tanya Goyal, Sanket Mehta, Balaji Vasan Srinivasan
PAKDD (1)1