Shashi Narayan

dblp:74/8458 · DBLP profile ↗
← Back
36ranked-venue papers
16as first author
11since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 16 first-author · 11 since 2021
YearPublicationVenuePosition
2024 Learning to Plan and Generate Text with Citations
abstract
Constanza Fierro, Reinald Kim Amplayo, Fantine Huot, Nicola De Cao, Joshua Maynez, Shashi Narayan, Mirella Lapata. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Constanza Fierro, Reinald Kim Amplayo, Fantine Huot, Nicola De Cao, Joshua Maynez, Shashi Narayan, Mirella Lapata
ACL (1)6
2024 Little Red Riding Hood Goes around the Globe: Crosslingual Story Planning and Generation with Large Language Models
abstract
Previous work has demonstrated the effectiveness of planning for story generation exclusively in a monolingual setting focusing primarily on English. We consider whether planning brings advantages to automatic story generation across languages. We propose a new task of crosslingual story generation with planning and present a new dataset for this task. We conduct a comprehensive study of different plans and generate stories in several languages, by leveraging the creative and reasoning capabilities of large pretrained language models. Our results demonstrate that plans which structure stories into three acts lead to more coherent and interesting narratives, while allowing to explicitly control their content and structure.
Evgeniia Razumovskaia, Joshua Maynez, Annie Louis, Mirella Lapata, Shashi Narayan
LREC/COLING5
2024 μPLAN: Summarizing using a Content Plan as Cross-Lingual Bridge
abstract
Fantine Huot, Joshua Maynez, Chris Alberti, Reinald Kim Amplayo, Priyanka Agrawal, Constanza Fierro, Shashi Narayan, Mirella Lapata. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Fantine Huot, Joshua Maynez, Christopher Alberti, Reinald Kim Amplayo, Priyanka Agrawal, Constanza Fierro, Shashi Narayan, Mirella Lapata
EACL (1)7
2023 Query Refinement Prompts for Closed-Book Long-Form QA
abstract
Reinald Kim Amplayo, Kellie Webster, Michael Collins, Dipanjan Das, Shashi Narayan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Reinald Kim Amplayo, Kellie Webster, Michael Collins 0001, Dipanjan Das 0001, Shashi Narayan
ACL (1)5
2023 SMART: Sentences as Basic Units for Text Evaluation
Reinald Kim Amplayo, Peter J. Liu, Shashi Narayan
ICLR4
2023 Calibrating Sequence likelihood Improves Conditional Language Generation
Misha Khalman, Rishabh Joshi, Shashi Narayan, Mohammad Saleh, Peter J. Liu
ICLR4
2023 Conditional Generation with a Question-Answering Blueprint
abstract
Abstract The ability to convey relevant and faithful information is critical for many tasks in conditional generation and yet remains elusive for neural seq-to-seq models whose outputs often reveal hallucinations and fail to correctly cover important details. In this work, we advocate planning as a useful intermediate representation for rendering conditional generation less opaque and more grounded. We propose a new conceptualization of text plans as a sequence of question-answer (QA) pairs and enhance existing datasets (e.g., for summarization) with a QA blueprint operating as a proxy for content selection (i.e., what to say) and planning (i.e., in what order). We obtain blueprints automatically by exploiting state-of-the-art question generation technology and convert input-output pairs into input-blueprint-output tuples. We develop Transformer-based models, each varying in how they incorporate the blueprint in the generated output (e.g., as a global plan or iteratively). Evaluation across metrics and datasets demonstrates that blueprint models are more factual than alternatives which do not resort to planning and allow tighter control of the generation output.
Shashi Narayan, Joshua Maynez, Reinald Kim Amplayo, Kuzman Ganchev, Annie Louis, Fantine Huot, Anders Sandholm 0001, Dipanjan Das 0001, Mirella Lapata
Trans. Assoc. Comput. Linguistics1
2022 A Well-Composed Text is Half Done! Composition Sampling for Diverse Conditional Generation
abstract
Shashi Narayan, Gonçalo Simões, Yao Zhao, Joshua Maynez, Dipanjan Das, Michael Collins, Mirella Lapata. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Shashi Narayan, Gonçalo Simões, Joshua Maynez, Dipanjan Das 0001, Michael Collins 0001, Mirella Lapata
ACL (1)1
2021 Focus Attention: Promoting Faithfulness and Diversity in Summarization
abstract
Rahul Aralikatte, Shashi Narayan, Joshua Maynez, Sascha Rothe, Ryan McDonald. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Rahul Aralikatte, Shashi Narayan, Joshua Maynez, Sascha Rothe, Ryan T. McDonald
ACL/IJCNLP (1)2
2021 A Thorough Evaluation of Task-Specific Pretraining for Summarization
abstract
Task-agnostic pretraining objectives like masked language models or corrupted span prediction are applicable to a wide range of NLP downstream tasks (Raffel et al., 2019), but are outperformed by task-specific pretraining objectives like predicting extracted gap sentences on summarization (Zhang et al., 2020).We compare three summarization specific pretraining objectives with the task agnostic corrupted span prediction pretraining in a controlled study.We also extend our study to a low resource and zero shot setup, to understand how many training examples are needed in order to ablate the task-specific pretraining without quality loss.Our results show that task-agnostic pretraining is sufficient for most cases which hopefully reduces the need for costly task-specific pretraining.We also report new state-of-the-art number for two summarization tasks using a T5 model with 11 billion parameters and an optimal beam search length penalty.
Sascha Rothe, Joshua Maynez, Shashi Narayan
EMNLP (1)3
2021 Planning with Learned Entity Prompts for Abstractive Summarization
abstract
Abstract We introduce a simple but flexible mechanism to learn an intermediate plan to ground the generation of abstractive summaries. Specifically, we prepend (or prompt) target summaries with entity chains—ordered sequences of entities mentioned in the summary. Transformer-based sequence-to-sequence models are then trained to generate the entity chain and then continue generating the summary conditioned on the entity chain and the input. We experimented with both pretraining and finetuning with this content planning objective. When evaluated on CNN/DailyMail, XSum, SAMSum, and BillSum, we demonstrate empirically that the grounded generation with the planning objective improves entity specificity and planning in summaries for all datasets, and achieves state-of-the-art performance on XSum and SAMSum in terms of rouge. Moreover, we demonstrate empirically that planning with entity chains provides a mechanism to control hallucinations in abstractive summaries. By prompting the decoder with a modified content plan that drops hallucinated entities, we outperform state-of-the-art approaches for faithfulness when evaluated automatically and by humans.
Shashi Narayan, Joshua Maynez, Gonçalo Simões, Vitaly Nikolaev, Ryan T. McDonald
Trans. Assoc. Comput. Linguistics1
2020 On Faithfulness and Factuality in Abstractive Summarization
abstract
It is well known that the standard likelihood training and approximate decoding objectives in neural text generation models lead to less human-like responses for open-ended tasks such as language modeling and story generation.In this paper we have analyzed limitations of these models for abstractive document summarization and found that these models are highly prone to hallucinate content that is unfaithful to the input document.We conducted a large scale human evaluation of several neural abstractive summarization systems to better understand the types of hallucinations they produce.Our human annotators found substantial amounts of hallucinated content in all model generated summaries.However, our analysis does show that pretrained models are better summarizers not only in terms of raw metrics, i.e., ROUGE, but also in generating faithful and factual summaries as evaluated by humans.Furthermore, we show that textual entailment measures better correlate with faithfulness than standard metrics, potentially leading the way to automatic evaluation metrics as well as training and decoding criteria.1 * The first two authors contributed equally.
Joshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonald
ACL2
2020 Stepwise Extractive Summarization and Planning with Structured Transformers
abstract
We propose encoder-centric stepwise models for extractive summarization using structured transformers -HiBERT (Zhang et al., 2019) and Extended Transformers (Ainslie et al., 2020).We enable stepwise summarization by injecting the previously generated summary into the structured transformer as an auxiliary sub-structure.Our models are not only efficient in modeling the structure of long inputs, but they also do not rely on task-specific redundancy-aware modeling, making them a general purpose extractive content planner for different tasks.When evaluated on CNN/DailyMail extractive summarization, stepwise models achieve state-of-the-art performance in terms of Rouge without any redundancy aware modeling or sentence filtering.This also holds true for Rotowire tableto-text generation, where our models surpass previously reported metrics for content selection, planning and ordering, highlighting the strength of stepwise modeling.Amongst the two structured transformers we test, stepwise Extended Transformers provides the best performance across both datasets and sets a new standard for these challenges. 1 * Equal contribution.
Shashi Narayan, Joshua Maynez, Jakub Adámek, Daniele Pighin, Blaz Bratanic, Ryan T. McDonald
EMNLP (1)1
2020 Leveraging Pre-trained Checkpoints for Sequence Generation Tasks
abstract
Unsupervised pre-training of large neural models has recently revolutionized Natural Language Processing. By warm-starting from the publicly released checkpoints, NLP practitioners have pushed the state-of-the-art on multiple benchmarks while saving significant amounts of compute time. So far the focus has been mainly on the Natural Language Understanding tasks. In this paper, we demonstrate the efficacy of pre-trained checkpoints for Sequence Generation. We developed a Transformer-based sequence-to-sequence model that is compatible with publicly available pre-trained BERT, GPT-2, and RoBERTa checkpoints and conducted an extensive empirical study on the utility of initializing our model, both encoder and decoder, with these checkpoints. Our models result in new state-of-the-art results on Machine Translation, Text Summarization, Sentence Splitting, and Sentence Fusion.
Sascha Rothe, Shashi Narayan, Aliaksei Severyn
Trans. Assoc. Comput. Linguistics2
2019 HighRES: Highlight-based Reference-less Evaluation of Summarization
abstract
There has been substantial progress in summarization research enabled by the availability of novel, often large-scale, datasets and recent advances on neural network-based approaches.However, manual evaluation of the system generated summaries is inconsistent due to the difficulty the task poses to human non-expert readers.To address this issue, we propose a novel approach for manual evaluation, HIGHlight-based Reference-less Evaluation of Summarization (HIGHRES), in which summaries are assessed by multiple annotators against the source document via manually highlighted salient content in the latter.Thus summary assessment on the source document by human judges is facilitated, while the highlights can be used for evaluating multiple systems.To validate our approach we employ crowd-workers to augment with highlights a recently proposed dataset and compare two state-of-the-art systems.We demonstrate that HIGHRES improves inter-annotator agreement in comparison to using the source document directly, while they help emphasize differences among systems that would be ignored under other evaluation approaches. 1
Hardy, Shashi Narayan, Andreas Vlachos 0001
ACL (1)2
2019 What is this Article about? Extreme Summarization with Topic-aware Convolutional Neural Networks
abstract
We introduce "extreme summarization," a new single-document summarization task which aims at creating a short, one-sentence news summary answering the question "What is the article about?". We argue that extreme summarization, by nature, is not amenable to extractive strategies and requires an abstractive modeling approach. In the hope of driving research on this task further: (a) we collect a real-world, large scale dataset by harvesting online articles from the British Broadcasting Corporation (BBC); and (b) propose a novel abstractive model which is conditioned on the article's topics and based entirely on convolutional neural networks. We demonstrate experimentally that this architecture captures long-range dependencies in a document and recognizes pertinent content, outperforming an oracle extractive system and state-of-the-art abstractive approaches when evaluated automatically and by humans on the extreme summarization dataset.
Shashi Narayan, Shay B. Cohen, Mirella Lapata
J. Artif. Intell. Res.1
2018 Document Modeling with External Attention for Sentence Extraction
abstract
Shashi Narayan, Ronald Cardenas, Nikos Papasarantopoulos, Shay B. Cohen, Mirella Lapata, Jiangsheng Yu, Yi Chang. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Shashi Narayan, Ronald Cardenas, Nikos Papasarantopoulos, Shay B. Cohen, Mirella Lapata, Jiangsheng Yu, Yi Chang 0001
ACL (1)1
2018 Local String Transduction as Sequence Labeling
abstract
We show that the general problem of string transduction can be reduced to the problem of sequence labeling. While character deletion and insertions are allowed in string transduction, they do not exist in sequence labeling. We show how to overcome this difference. Our approach can be used with any sequence labeling algorithm and it works best for problems in which string transduction imposes a strong notion of locality (no long range dependencies). We experiment with spelling correction for social media, OCR correction, and morphological inflection, and we see that it behaves better than seq2seq models and yields state-of-the-art results in several cases.
Joana Ribeiro, Shashi Narayan, Shay B. Cohen, Xavier Carreras
COLING2
2018 Privacy-preserving Neural Representations of Text
abstract
This article deals with adversarial attacks towards deep learning systems for Natural Language Processing (NLP), in the context of privacy protection.We study a specific type of attack: an attacker eavesdrops on the hidden representations of a neural text classifier and tries to recover information about the input text.Such scenario may arise in situations when the computation of a neural network is shared across multiple devices, e.g.some hidden representation is computed by a user's device and sent to a cloud-based model.We measure the privacy of a hidden representation by the ability of an attacker to predict accurately specific private information from it and characterize the tradeoff between the privacy and the utility of neural representations.Finally, we propose several defense methods based on modified training objectives and show that they improve the privacy of neural representations.
Maximin Coavoux, Shashi Narayan, Shay B. Cohen
EMNLP2
2018 Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization
abstract
We introduce extreme summarization, a new single-document summarization task which does not favor extractive strategies and calls for an abstractive modeling approach.The idea is to create a short, one-sentence news summary answering the question "What is the article about?".We collect a real-world, large scale dataset for this task by harvesting online articles from the British Broadcasting Corporation (BBC).We propose a novel abstractive model which is conditioned on the article's topics and based entirely on convolutional neural networks.We demonstrate experimentally that this architecture captures longrange dependencies in a document and recognizes pertinent content, outperforming an oracle extractive system and state-of-the-art abstractive approaches when evaluated automatically and by humans. 1
Shashi Narayan, Shay B. Cohen, Mirella Lapata
EMNLP1
2018 Ranking Sentences for Extractive Summarization with Reinforcement Learning
abstract
Shashi Narayan, Shay B. Cohen, Mirella Lapata. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Shashi Narayan, Shay B. Cohen, Mirella Lapata
NAACL-HLT1
2017 Creating Training Corpora for NLG Micro-Planners
abstract
In this paper, we present a novel framework for semi-automatically creating linguistically challenging microplanning data-to-text corpora from existing Knowledge Bases.Because our method pairs data of varying size and shape with texts ranging from simple clauses to short texts, a dataset created using this framework provides a challenging benchmark for microplanning.Another feature of this framework is that it can be applied to any large scale knowledge base and can therefore be used to train and learn KB verbalisers.We apply our framework to DBpedia data and compare the resulting dataset with Wen et al. (2016)'s.We show that while Wen et al.'s dataset is more than twice larger than ours, it is less diverse both in terms of input and in terms of text.We thus propose our corpus generation framework as a novel method for creating challenging data sets from which NLG models can be learned which are capable of handling the complex interactions occurring during in micro-planning between lexicalisation, aggregation, surface realisation, referring expression generation and sentence segmentation.To encourage researchers to take up this challenge, we recently made available a dataset created using this framework in the context of the WEBNLG shared task.
Claire Gardent, Anastasia Shimorina, Shashi Narayan, Laura Perez-Beltrachini
ACL (1)3
2017 Split and Rephrase
abstract
We propose a new sentence simplification task (Split-and-Rephrase) where the aim is to split a complex sentence into a meaning preserving sequence of shorter sentences.Like sentence simplification, splitting-and-rephrasing has the potential of benefiting both natural language processing and societal applications.Because shorter sentences are generally better processed by NLP systems, it could be used as a preprocessing step which facilitates and improves the performance of parsers, semantic role labelers and machine translation systems.It should also be of use for people with reading disabilities because it allows the conversion of longer sentences into shorter ones.This paper makes two contributions towards this new task.First, we create and make available a benchmark consisting of 1,066,115 tuples mapping a single complex sentence to a sequence of sentences expressing the same meaning.1 Second, we propose five models (vanilla sequence-to-sequence to semantically-motivated models) to understand the difficulty of the proposed task.
Shashi Narayan, Claire Gardent, Shay B. Cohen, Anastasia Shimorina
EMNLP1
2017 The WebNLG Challenge: Generating Text from RDF Data
abstract
The WebNLG challenge consists in mapping sets of RDF triples to text.It provides a common benchmark on which to train, evaluate and compare "microplanners", i.e. generation systems that verbalise a given content by making a range of complex interacting choices including referring expression generation, aggregation, lexicalisation, surface realisation and sentence segmentation.In this paper, we introduce the microplanning task, describe data preparation, introduce our evaluation methodology, analyse participant results and provide a brief description of the participating systems.(3) a. LOCATION-COUNTRY-STARTDATE ⇒ Passive-Apposition-Active
Claire Gardent, Anastasia Shimorina, Shashi Narayan, Laura Perez-Beltrachini
INLG3
2016 Optimizing Spectral Learning for Parsing
abstract
We describe a search algorithm for optimizing the number of latent states when estimating latent-variable PCFGs with spectral methods.Our results show that contrary to the common belief that the number of latent states for each nonterminal in an L-PCFG can be decided in isolation with spectral methods, parsing results significantly improve if the number of latent states for each nonterminal is globally optimized, while taking into account interactions between the different nonterminals.In addition, we contribute an empirical analysis of spectral algorithms on eight morphologically rich languages: Basque, French, German, Hebrew, Hungarian, Korean, Polish and Swedish.Our results show that our estimation consistently performs better or close to coarse-to-fine expectation-maximization techniques for these languages.
Shashi Narayan, Shay B. Cohen
ACL (1)1
2016 The WebNLG Challenge: Generating Text from DBPedia Data
Émilie Colin, Claire Gardent, Yassine Mrabet, Shashi Narayan, Laura Perez-Beltrachini
INLG4
2016 Unsupervised Sentence Simplification Using Deep Semantics
abstract
We present a novel approach to sentence simplification which departs from previous work in two main ways.First, it requires neither hand written rules nor a training corpus of aligned standard and simplified sentences.Second, sentence splitting operates on deep semantic structure.We show (i) that the unsupervised framework we propose is competitive with four state-of-the-art supervised systems and (ii) that our semantic based approach allows for a principled and effective handling of sentence splitting.
Shashi Narayan, Claire Gardent
INLG1
2016 Paraphrase Generation from Latent-Variable PCFGs for Semantic Parsing
abstract
One of the limitations of semantic parsing approaches to open-domain question answering is the lexicosyntactic gap between natural language questions and knowledge base entries -there are many ways to ask a question, all with the same answer.In this paper we propose to bridge this gap by generating paraphrases of the input question with the goal that at least one of them will be correctly mapped to a knowledge-base query.We introduce a novel grammar model for paraphrase generation that does not require any sentence-aligned paraphrase corpus.Our key idea is to leverage the flexibility and scalability of latent-variable probabilistic context-free grammars to sample paraphrases.We do an extrinsic evaluation of our paraphrases by plugging them into a semantic parser for Freebase.Our evaluation experiments on the WebQuestions benchmark dataset show that the performance of the semantic parser improves over strong baselines.
Shashi Narayan, Siva Reddy, Shay B. Cohen
INLG1
2016 Encoding Prior Knowledge with Eigenword Embeddings
abstract
Canonical correlation analysis (CCA) is a method for reducing the dimension of data represented using two views. It has been previously used to derive word embeddings, where one view indicates a word, and the other view indicates its context. We describe a way to incorporate prior knowledge into CCA, give a theoretical justification for it, and test it by deriving word embeddings and evaluating them on a myriad of datasets.
Dominique Osborne, Shashi Narayan, Shay B. Cohen
Trans. Assoc. Comput. Linguistics2
2015 Diversity in Spectral Learning for Natural Language Parsing
abstract
We describe an approach to create a diverse set of predictions with spectral learning of latent-variable PCFGs (L-PCFGs).Our approach works by creating multiple spectral models where noise is added to the underlying features in the training set before the estimation of each model.We describe three ways to decode with multiple models.In addition, we describe a simple variant of the spectral algorithm for L-PCFGs that is fast and leads to compact models.Our experiments for natural language parsing, for English and German, show that we get a significant improvement over baselines comparable to state of the art.For English, we achieve the F 1 score of 90.18, and for German we achieve the F 1 score of 83.38.
Shashi Narayan, Shay B. Cohen
EMNLP1
2015 Multiple Adjunction in Feature-Based Tree-Adjoining Grammar
abstract
In parsing with Tree Adjoining Grammar (TAG), independent derivations have been shown by Schabes and Shieber (1994) to be essential for correctly supporting syntactic analysis, semantic interpretation, and statistical language modeling. However, the parsing algorithm they propose is not directly applicable to Feature-Based TAGs (FB-TAG). We provide a recognition algorithm for FB-TAG that supports both dependent and independent derivations. The resulting algorithm combines the benefits of independent derivations with those of Feature-Based grammars. In particular, we show that it accounts for a range of interactions between dependent vs. independent derivation on the one hand, and syntactic constraints, linear ordering, and scopal vs. nonscopal semantic dependencies on the other hand.
Claire Gardent, Shashi Narayan
Comput. Linguistics2
2014 Hybrid Simplification using Deep Semantics and Machine Translation
abstract
We present a hybrid approach to sentence simplification which combines deep semantics and monolingual machine translation to derive simple sentences from complex ones. The approach differs from previous work in two main ways. First, it is semantic based in that it takes as input a deep semantic representation rather than e.g., a sentence or a parse tree. Second, it combines a simplification model for splitting and deletion with a monolingual translation model for phrase substitution and reordering. When compared against current state of the art methods, our model yields significantly simpler output that is both grammatical and meaning preserving.
Shashi Narayan, Claire Gardent
ACL (1)1
2012 Error Mining on Dependency Trees
Claire Gardent, Shashi Narayan
ACL (1)2
2012 Error Mining with Suspicion Trees: Seeing the Forest for the Trees
Shashi Narayan, Claire Gardent
COLING1
2012 Structure-Driven Lexicalist Generation
Shashi Narayan, Claire Gardent
COLING1
2010 A composite kernel for named entity recognition
Sujan Kumar Saha, Shashi Narayan, Sudeshna Sarkar, Pabitra Mitra
Pattern Recognit. Lett.2