Wojciech Kryscinski

dblp:225/5389 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
10since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 3 first-author · 10 since 2021
YearPublicationVenuePosition
2024 FOLIO: Natural Language Reasoning with First-Order Logic
abstract
Simeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan, Yixin Liu, Brian Wong, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu, Rui Zhang, Alexander Fabbri, Wojciech Maciej Kryscinski, Semih Yavuz, Ye Liu, Xi Victoria Lin, Shafiq Joty, Yingbo Zhou, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir Radev. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Simeng Han, Hailey Schoelkopf, Yilun Zhao 0001, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan 0001, Yixin Liu 0003, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu 0009, Rui Zhang 0037, Alexander R. Fabbri, Wojciech Kryscinski, Semih Yavuz, Ye Liu 0006, Xi Victoria Lin, Shafiq R. Joty, Yingbo Zhou 0002, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir R. Radev
EMNLP26
2023 SWiPE: A Dataset for Document-Level Simplification of Wikipedia Pages
abstract
Philippe Laban, Jesse Vig, Wojciech Kryscinski, Shafiq Joty, Caiming Xiong, Chien-Sheng Wu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Philippe Laban, Jesse Vig, Wojciech Kryscinski, Shafiq R. Joty, Caiming Xiong, Chien-Sheng Wu
ACL (1)3
2023 Socratic Pretraining: Question-Driven Pretraining for Controllable Summarization
abstract
In long document controllable summarization, where labeled data is scarce, pretrained models struggle to adapt to the task and effectively respond to user queries.In this paper, we introduce SOCRATIC pretraining, a question-driven, unsupervised pretraining objective specifically designed to improve controllability in summarization tasks.By training a model to generate and answer relevant questions in a given context, SOCRATIC pretraining enables the model to more effectively adhere to user-provided queries and identify relevant content to be summarized.We demonstrate the effectiveness of this approach through extensive experimentation on two summarization domains, short stories and dialogue, and multiple control strategies: keywords, questions, and factoid QA pairs.Our pretraining method relies only on unlabeled documents and a question generation system and outperforms pre-finetuning approaches that use additional supervised data.Furthermore, our results show that SOCRATIC pretraining cuts task-specific labeled data requirements in half, is more faithful to userprovided queries, and achieves state-of-the-art performance on QMSum and SQuALITY.Joseph L Fleiss.1971.Measuring nominal scale agreement among many raters.
Artidoro Pagnoni, Alexander R. Fabbri, Wojciech Kryscinski, Chien-Sheng Wu
ACL (1)3
2023 Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error Detectors
abstract
Liyan Tang, Tanya Goyal, Alex Fabbri, Philippe Laban, Jiacheng Xu, Semih Yavuz, Wojciech Kryscinski, Justin Rousseau, Greg Durrett. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Liyan Tang, Tanya Goyal, Alexander R. Fabbri, Philippe Laban, Jiacheng Xu 0001, Semih Yavuz, Wojciech Kryscinski, Justin F. Rousseau, Greg Durrett
ACL (1)7
2023 What's New? Summarizing Contributions in Scientific Literature
abstract
With thousands of academic articles shared on a daily basis, it has become increasingly difficult to keep up with the latest scientific findings.To overcome this problem, we introduce a new task of disentangled paper summarization, which seeks to generate separate summaries for the paper contributions and the context of the work, making it easier to identify the key findings shared in articles.For this purpose, we extend the S2ORC corpus of academic articles, which spans a diverse set of domains ranging from economics to psychology, by adding disentangled "contribution" and "context" reference labels.Together with the dataset, we introduce and analyze three baseline approaches: 1) a unified model controlled by input code prefixes, 2) a model with separate generation heads specialized in generating the disentangled outputs, and 3) a training strategy that guides the model using additional supervision coming from inbound and outbound citations.We also propose a comprehensive automatic evaluation protocol which reports the relevance, novelty, and disentanglement of generated outputs.Through a human study involving expert annotators, we show that in 79%, of cases our new task is considered more helpful than traditional scientific paper summarization.
Hiroaki Hayashi, Wojciech Kryscinski, Bryan McCann, Nazneen Fatema Rajani, Caiming Xiong
EACL2
2023 SummEdits: Measuring LLM Ability at Factual Reasoning Through The Lens of Summarization
abstract
Philippe Laban, Wojciech Kryscinski, Divyansh Agarwal, Alexander Fabbri, Caiming Xiong, Shafiq Joty, Chien-Sheng Wu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Philippe Laban, Wojciech Kryscinski, Divyansh Agarwal, Alexander R. Fabbri, Caiming Xiong, Shafiq R. Joty, Chien-Sheng Wu
EMNLP2
2022 HydraSum: Disentangling Style Features in Text Summarization with Multi-Decoder Models
abstract
Summarization systems make numerous "decisions" about summary properties during inference, e.g.degree of copying, specificity and length of outputs, etc.However, these are implicitly encoded within model parameters and specific styles cannot be enforced.To address this, we introduce HYDRASUM, a new summarization architecture that extends the single decoder framework of current models to a mixture-of-experts version with multiple decoders.We show that HYDRASUM's multiple decoders automatically learn contrasting summary styles when trained under the standard training objective without any extra supervision.Through experiments on three summarization datasets (CNN, NEWSROOM and XSUM), we show that HYDRASUM provides a simple mechanism to obtain stylistically-diverse summaries by sampling from either individual decoders or their mixtures, outperforming baseline models.Finally, we demonstrate that a small modification to the gating strategy during training can enforce an even stricter style partitioning, e.g.high-vs low-abstractiveness or high-vs low-specificity, allowing users to sample from a larger area in the generation space and vary summary styles along multiple dimensions. 1Input Article: Insights into the workings of the human body that Leonardo da Vinci could only obtain by dissecting scores of corpses and recording the results in exquisite drawings will be displayed for the first time beside modern 3D films, CT and MRI scans, which show how close the Renaissance genius got to the truth of what lies under the skin.[…] the Edinburgh show will be the first to compare Leonardo's results with scalpel and pen with the best results of modern technology.[…] The exhibition will show how close Leonardo got in some of his last medical experiments to discovering the role of the beating heart in the circulation of the blood, a century before William Harvey worked it out.[…] Edinburgh show will be first to compare Renaissance genius's results with best results of modern technology.Edinburgh show will be first to compare Renaissance genius's results with the best results of modern technology.Edinburgh show will be first to compare Leonardo's results with the best results of modern technology. Low diversity Baseline BARTEdinburgh show will be first to compare Leonardo's results with best results of modern technology.Modern imaging techniques will be displayed alongside Leonardo da Vinci's anatomical drawings in Edinburgh exhibition.
Tanya Goyal, Nazneen Fatema Rajani, Wenhao Liu 0003, Wojciech Kryscinski
EMNLP4
2022 CTRLsum: Towards Generic Controllable Text Summarization
abstract
Current summarization systems yield generic summaries that are disconnected from users' preferences and expectations.To address this limitation, we present CTRLSUM, a generic framework to control generated summaries through a set of keywords.During training keywords are extracted automatically without requiring additional human annotations.At test time CTRLSUM features a control function to map control signal to keywords; through engineering the control function, the same trained model is able to be applied to control summaries on various dimensions, while neither affecting the model training process nor the pretrained models.We additionally explore the combination of keywords and text prompts for more control tasks.Experiments demonstrate the effectiveness of CTRLSUM on three domains of summarization datasets and five control tasks: (1) entity-centric and (2) length-controllable summarization, (3) contribution summarization on scientific papers, (4) invention purpose summarization on patent filings, and (5) question-guided summarization on news articles.Moreover, when used in a standard, unconstrained summarization setting, CTRLSUM is comparable or better than strong pretrained systems. 1
Junxian He, Wojciech Kryscinski, Bryan McCann, Nazneen Fatema Rajani, Caiming Xiong
EMNLP2
2022 FeTaQA: Free-form Table Question Answering
abstract
Abstract Existing table question answering datasets contain abundant factual questions that primarily evaluate a QA system’s comprehension of query and tabular data. However, restricted by their short-form answers, these datasets fail to include question–answer interactions that represent more advanced and naturally occurring information needs: questions that ask for reasoning and integration of information pieces retrieved from a structured knowledge source. To complement the existing datasets and to reveal the challenging nature of the table-based question answering task, we introduce FeTaQA, a new dataset with 10K Wikipedia-based {table, question, free-form answer, supporting table cells} pairs. FeTaQA is collected from noteworthy descriptions of Wikipedia tables that contain information people tend to seek; generation of these descriptions requires advanced processing that humans perform on a daily basis: Understand the question and table, retrieve, integrate, infer, and conduct text planning and surface realization to generate an answer. We provide two benchmark methods for the proposed task: a pipeline method based on semantic parsing-based QA systems and an end-to-end method based on large pretrained text generation models, and show that FeTaQA poses a challenge for both methods.
Linyong Nan, Chiachun Hsieh, Ziming Mao, Xi Victoria Lin, Neha Verma 0001, Rui Zhang 0037, Wojciech Kryscinski, Hailey Schoelkopf, Riley Kong, Xiangru Tang, Mutethia Mutuma, Ben Rosand, Isabel Trindade, Renusree Bandaru, Jacob Cunningham, Caiming Xiong, Dragomir R. Radev
Trans. Assoc. Comput. Linguistics7
2021 SummEval: Re-evaluating Summarization Evaluation
abstract
Abstract The scarcity of comprehensive up-to-date studies on evaluation metrics for text summarization and the lack of consensus regarding evaluation protocols continue to inhibit progress. We address the existing shortcomings of summarization evaluation methods along five dimensions: 1) we re-evaluate 14 automatic evaluation metrics in a comprehensive and consistent fashion using neural summarization model outputs along with expert and crowd-sourced human annotations; 2) we consistently benchmark 23 recent summarization models using the aforementioned automatic evaluation metrics; 3) we assemble the largest collection of summaries generated by models trained on the CNN/DailyMail news dataset and share it in a unified format; 4) we implement and share a toolkit that provides an extensible and unified API for evaluating summarization models across a broad range of automatic metrics; and 5) we assemble and share the largest and most diverse, in terms of model types, collection of human judgments of model-generated summaries on the CNN/Daily Mail dataset annotated by both expert judges and crowd-source workers. We hope that this work will help promote a more complete evaluation protocol for text summarization as well as advance research in developing evaluation metrics that better correlate with human judgments.
Alexander R. Fabbri, Wojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard Socher, Dragomir R. Radev
Trans. Assoc. Comput. Linguistics2
2020 Evaluating the Factual Consistency of Abstractive Text Summarization
abstract
Currently used metrics for assessing summarization algorithms do not account for whether summaries are factually consistent with source documents.We propose a weakly-supervised, model-based approach for verifying factual consistency and identifying conflicts between source documents and a generated summary.Training data is generated by applying a series of rule-based transformations to the sentences of source documents.The factual consistency model is then trained jointly for three tasks: 1) identify whether sentences remain factually consistent after transformation, 2) extract a span in the source documents to support the consistency prediction, 3) extract a span in the summary sentence that is inconsistent if one exists.Transferring this model to summaries generated by several state-of-the art models reveals that this highly scalable approach substantially outperforms previous models, including those trained with strong supervision using standard datasets for natural language inference and fact checking.Additionally, human evaluation shows that the auxiliary span extraction tasks provide useful assistance in the process of verifying factual consistency.
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard Socher
EMNLP (1)1
2019 Neural Text Summarization: A Critical Evaluation
abstract
Wojciech Kryscinski, Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, Richard Socher. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Wojciech Kryscinski, Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, Richard Socher
EMNLP/IJCNLP (1)1
2018 Improving Abstraction in Text Summarization
abstract
Abstractive text summarization aims to shorten long text documents into a human readable form that contains the most important facts from the original document.However, the level of actual abstraction as measured by novel phrases that do not appear in the source document remains low in existing approaches.We propose two techniques to improve the level of abstraction of generated summaries.First, we decompose the decoder into a contextual network that retrieves relevant parts of the source document, and a pretrained language model that incorporates prior knowledge about language generation.Second, we propose a novelty metric that is optimized directly through policy learning to encourage the generation of novel phrases.Our model achieves results comparable to state-of-the-art models, as determined by ROUGE scores and human evaluations, while achieving a significantly higher level of abstraction as measured by n-gram overlap with the source document.
Wojciech Kryscinski, Romain Paulus, Caiming Xiong, Richard Socher
EMNLP1