EDBT 2026 Demo / reviewers in the wild / expert
Joshua Maynez
dblp:220/3863
· DBLP profile ↗
17ranked-venue papers
2as first author
14since 2021 · last 2024
0000-0003-4948-2875ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 2 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Learning to Plan and Generate Text with CitationsabstractConstanza Fierro, Reinald Kim Amplayo, Fantine Huot, Nicola De Cao, Joshua Maynez, Shashi Narayan, Mirella Lapata. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Constanza Fierro, Reinald Kim Amplayo, Fantine Huot, Nicola De Cao, Joshua Maynez, Shashi Narayan, Mirella Lapata |
ACL (1) | 5 |
| 2024 | Little Red Riding Hood Goes around the Globe: Crosslingual Story Planning and Generation with Large Language ModelsabstractPrevious work has demonstrated the effectiveness of planning for story generation exclusively in a monolingual setting focusing primarily on English. We consider whether planning brings advantages to automatic story generation across languages. We propose a new task of crosslingual story generation with planning and present a new dataset for this task. We conduct a comprehensive study of different plans and generate stories in several languages, by leveraging the creative and reasoning capabilities of large pretrained language models. Our results demonstrate that plans which structure stories into three acts lead to more coherent and interesting narratives, while allowing to explicitly control their content and structure. Evgeniia Razumovskaia, Joshua Maynez, Annie Louis, Mirella Lapata, Shashi Narayan |
LREC/COLING | 2 |
| 2024 | μPLAN: Summarizing using a Content Plan as Cross-Lingual BridgeabstractFantine Huot, Joshua Maynez, Chris Alberti, Reinald Kim Amplayo, Priyanka Agrawal, Constanza Fierro, Shashi Narayan, Mirella Lapata. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Fantine Huot, Joshua Maynez, Christopher Alberti, Reinald Kim Amplayo, Priyanka Agrawal, Constanza Fierro, Shashi Narayan, Mirella Lapata |
EACL (1) | 2 |
| 2024 | Stratified Prediction-Powered Inference for Effective Hybrid Evaluation of Language ModelsabstractPrediction-powered inference (PPI) is a method that improves statistical estimates based on limited human-labeled data. PPI achieves this by combining small amounts of human-labeled data with larger amounts of data labeled by a reasonably accurate---but potentially biased---automatic system, in a way that results in tighter confidence intervals for certain parameters of interest (e.g., the mean performance of a language model). In this paper, we propose a method called Stratified Prediction-Powered Inference (StratPPI), in which we show that the basic PPI estimates can be considerably improved by employing simple data stratification strategies. Without making any assumptions on the underlying automatic labeling system or data distribution, we derive an algorithm for computing provably valid confidence intervals for parameters of any dimensionality that is based on stratified sampling. In particular, we show both theoretically and empirically that, with appropriate choices of stratification and sample allocation, our approach can provide substantially tighter confidence intervals than unstratified approaches. Specifically, StratPPI is expected to improve in cases where the performance of the autorater varies across different conditional distributions of the target data. Adam Fisch, Joshua Maynez, R. Alex Hofer, Bhuwan Dhingra, Amir Globerson, William W. Cohen |
NeurIPS | 2 |
| 2023 | Benchmarking Large Language Model Capabilities for Conditional GenerationabstractPre-trained large language models (PLMs) underlie most new developments in natural language processing.They have shifted the field from application-specific model pipelines to a single model that is adapted to a wide range of tasks.Autoregressive PLMs like GPT-3 or PaLM, alongside techniques like few-shot learning, have additionally shifted the output modality to generation instead of classification or regression.Despite their ubiquitous use, the generation quality of language models is rarely evaluated when these models are introduced.Additionally, it is unclear how existing generation tasks--while they can be used to compare systems at a high level-relate to the real world use cases for which people have been adopting them.In this work, we discuss how to adapt existing applicationspecific generation benchmarks to PLMs and provide an in-depth, empirical study of the limitations and capabilities of PLMs in natural language generation tasks along dimensions such as scale, architecture, input and output language.Our results show that PLMs differ in their applicability to different data regimes and their generalization to multiple languages and inform which PLMs to use for a given generation task setup.We share best practices to be taken into consideration when benchmarking generation capabilities during the development of upcoming PLMs. Joshua Maynez, Priyanka Agrawal, Sebastian Gehrmann |
ACL (1) | 1 |
| 2023 | SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization EvaluationabstractElizabeth Clark, Shruti Rijhwani, Sebastian Gehrmann, Joshua Maynez, Roee Aharoni, Vitaly Nikolaev, Thibault Sellam, Aditya Siddhant, Dipanjan Das, Ankur Parikh. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Elizabeth Clark, Shruti Rijhwani, Sebastian Gehrmann, Joshua Maynez, Roee Aharoni, Vitaly Nikolaev, Thibault Sellam, Aditya Siddhant, Dipanjan Das 0001, Ankur P. Parikh |
EMNLP | 4 |
| 2023 | PaLM: Scaling Language Modeling with PathwaysabstractLarge language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model (PaLM). We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of-the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state-of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies. Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Adam Roberts, Paul Barham 0001, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du 0002, Ben Hutchinson, Reiner Pope, Jacob Austin, Michael Isard, Guy Gur-Ari, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, William Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang 0002, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeffrey Dean, Slav Petrov, Noah Fiedel |
J. Mach. Learn. Res. | 14 |
| 2023 | QAmeleon: Multilingual QA with Only 5 ExamplesabstractAbstract The availability of large, high-quality datasets has been a major driver of recent progress in question answering (QA). Such annotated datasets, however, are difficult and costly to collect, and rarely exist in languages other than English, rendering QA technology inaccessible to underrepresented languages. An alternative to building large monolingual training datasets is to leverage pre-trained language models (PLMs) under a few-shot learning setting. Our approach, QAmeleon, uses a PLM to automatically generate multilingual data upon which QA models are fine-tuned, thus avoiding costly annotation. Prompt tuning the PLM with only five examples per language delivers accuracy superior to translation-based baselines; it bridges nearly 60% of the gap between an English-only baseline and a fully-supervised upper bound fine-tuned on almost 50,000 hand-labeled examples; and consistently leads to improvements compared to directly fine-tuning a QA model on labeled examples in low resource settings. Experiments on the TyDiqa-GoldP and MLQA benchmarks show that few-shot prompt tuning for data synthesis scales across languages and is a viable alternative to large-scale annotation.1 Priyanka Agrawal, Christopher Alberti, Fantine Huot, Joshua Maynez, Ji Ma 0004, Sebastian Ruder, Kuzman Ganchev, Dipanjan Das 0001, Mirella Lapata |
Trans. Assoc. Comput. Linguistics | 4 |
| 2023 | Conditional Generation with a Question-Answering BlueprintabstractAbstract The ability to convey relevant and faithful information is critical for many tasks in conditional generation and yet remains elusive for neural seq-to-seq models whose outputs often reveal hallucinations and fail to correctly cover important details. In this work, we advocate planning as a useful intermediate representation for rendering conditional generation less opaque and more grounded. We propose a new conceptualization of text plans as a sequence of question-answer (QA) pairs and enhance existing datasets (e.g., for summarization) with a QA blueprint operating as a proxy for content selection (i.e., what to say) and planning (i.e., in what order). We obtain blueprints automatically by exploiting state-of-the-art question generation technology and convert input-output pairs into input-blueprint-output tuples. We develop Transformer-based models, each varying in how they incorporate the blueprint in the generated output (e.g., as a global plan or iteratively). Evaluation across metrics and datasets demonstrates that blueprint models are more factual than alternatives which do not resort to planning and allow tighter control of the generation output. Shashi Narayan, Joshua Maynez, Reinald Kim Amplayo, Kuzman Ganchev, Annie Louis, Fantine Huot, Anders Sandholm 0001, Dipanjan Das 0001, Mirella Lapata |
Trans. Assoc. Comput. Linguistics | 2 |
| 2022 | A Well-Composed Text is Half Done! Composition Sampling for Diverse Conditional GenerationabstractShashi Narayan, Gonçalo Simões, Yao Zhao, Joshua Maynez, Dipanjan Das, Michael Collins, Mirella Lapata. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Shashi Narayan, Gonçalo Simões, Joshua Maynez, Dipanjan Das 0001, Michael Collins 0001, Mirella Lapata |
ACL (1) | 4 |
| 2021 | Focus Attention: Promoting Faithfulness and Diversity in SummarizationabstractRahul Aralikatte, Shashi Narayan, Joshua Maynez, Sascha Rothe, Ryan McDonald. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Rahul Aralikatte, Shashi Narayan, Joshua Maynez, Sascha Rothe, Ryan T. McDonald |
ACL/IJCNLP (1) | 3 |
| 2021 | Action Item Detection in Meetings Using Pretrained TransformersabstractDetecting the action items that were agreed upon during a meeting has important practical applications. But, extremely sparse positive labels, noisy annotations, and small datasets limited the improvement in this task via hand-crafted features techniques. Given the breakthrough performance demonstrated by pretrained transformer-based models in a wide variety of NLP tasks, such as BERT [1] and ETC [2], we revisit this task using these modern techniques. We empirically show how these modelling techniques advance the state-of-the-art on action item detection for the ICSI simulated meeting corpus [3] by 75% and establish a baseline for the action item detection from the AMI [4] meeting corpus. We show that these models are competitive on the related ICSI MRDA classification problem. In order to push the performance even further, we re-evaluate the task definition for action item detection, drawing upon the similarities with the span boundary detection realm. We propose the use of Generalized Hamming Distance ghd as an alternative evaluation. We hope to motivate further interest into the action item detection task by the community. Kishan Sachdeva, Joshua Maynez, Olivier Siohan |
ASRU | 2 |
| 2021 | A Thorough Evaluation of Task-Specific Pretraining for SummarizationabstractTask-agnostic pretraining objectives like masked language models or corrupted span prediction are applicable to a wide range of NLP downstream tasks (Raffel et al., 2019), but are outperformed by task-specific pretraining objectives like predicting extracted gap sentences on summarization (Zhang et al., 2020).We compare three summarization specific pretraining objectives with the task agnostic corrupted span prediction pretraining in a controlled study.We also extend our study to a low resource and zero shot setup, to understand how many training examples are needed in order to ablate the task-specific pretraining without quality loss.Our results show that task-agnostic pretraining is sufficient for most cases which hopefully reduces the need for costly task-specific pretraining.We also report new state-of-the-art number for two summarization tasks using a T5 model with 11 billion parameters and an optimal beam search length penalty. Sascha Rothe, Joshua Maynez, Shashi Narayan |
EMNLP (1) | 2 |
| 2021 | Planning with Learned Entity Prompts for Abstractive SummarizationabstractAbstract We introduce a simple but flexible mechanism to learn an intermediate plan to ground the generation of abstractive summaries. Specifically, we prepend (or prompt) target summaries with entity chains—ordered sequences of entities mentioned in the summary. Transformer-based sequence-to-sequence models are then trained to generate the entity chain and then continue generating the summary conditioned on the entity chain and the input. We experimented with both pretraining and finetuning with this content planning objective. When evaluated on CNN/DailyMail, XSum, SAMSum, and BillSum, we demonstrate empirically that the grounded generation with the planning objective improves entity specificity and planning in summaries for all datasets, and achieves state-of-the-art performance on XSum and SAMSum in terms of rouge. Moreover, we demonstrate empirically that planning with entity chains provides a mechanism to control hallucinations in abstractive summaries. By prompting the decoder with a modified content plan that drops hallucinated entities, we outperform state-of-the-art approaches for faithfulness when evaluated automatically and by humans. Shashi Narayan, Joshua Maynez, Gonçalo Simões, Vitaly Nikolaev, Ryan T. McDonald |
Trans. Assoc. Comput. Linguistics | 3 |
| 2020 | On Faithfulness and Factuality in Abstractive SummarizationabstractIt is well known that the standard likelihood training and approximate decoding objectives in neural text generation models lead to less human-like responses for open-ended tasks such as language modeling and story generation.In this paper we have analyzed limitations of these models for abstractive document summarization and found that these models are highly prone to hallucinate content that is unfaithful to the input document.We conducted a large scale human evaluation of several neural abstractive summarization systems to better understand the types of hallucinations they produce.Our human annotators found substantial amounts of hallucinated content in all model generated summaries.However, our analysis does show that pretrained models are better summarizers not only in terms of raw metrics, i.e., ROUGE, but also in generating faithful and factual summaries as evaluated by humans.Furthermore, we show that textual entailment measures better correlate with faithfulness than standard metrics, potentially leading the way to automatic evaluation metrics as well as training and decoding criteria.1 * The first two authors contributed equally. Joshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonald |
ACL | 1 |
| 2020 | Stepwise Extractive Summarization and Planning with Structured TransformersabstractWe propose encoder-centric stepwise models for extractive summarization using structured transformers -HiBERT (Zhang et al., 2019) and Extended Transformers (Ainslie et al., 2020).We enable stepwise summarization by injecting the previously generated summary into the structured transformer as an auxiliary sub-structure.Our models are not only efficient in modeling the structure of long inputs, but they also do not rely on task-specific redundancy-aware modeling, making them a general purpose extractive content planner for different tasks.When evaluated on CNN/DailyMail extractive summarization, stepwise models achieve state-of-the-art performance in terms of Rouge without any redundancy aware modeling or sentence filtering.This also holds true for Rotowire tableto-text generation, where our models surpass previously reported metrics for content selection, planning and ordering, highlighting the strength of stepwise modeling.Amongst the two structured transformers we test, stepwise Extended Transformers provides the best performance across both datasets and sets a new standard for these challenges. 1 * Equal contribution. Shashi Narayan, Joshua Maynez, Jakub Adámek, Daniele Pighin, Blaz Bratanic, Ryan T. McDonald |
EMNLP (1) | 2 |
| 2018 | Morphosyntactic Tagging with a Meta-BiLSTM Model over Context Sensitive Token EncodingsabstractBernd Bohnet, Ryan McDonald, Gonçalo Simões, Daniel Andor, Emily Pitler, Joshua Maynez. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Bernd Bohnet, Ryan T. McDonald, Gonçalo Simões, Daniel Andor, Emily Pitler, Joshua Maynez |
ACL (1) | 6 |