Sebastian Gehrmann

dblp:131/1378 · DBLP profile ↗
← Back
27ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-8257-9516ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Domain Generalizable AI Guardrails with Augmented Policy Training
abstract
Minqian Liu, Ioana Baldini, David Rabinowitz, David S Rosenberg, Sebastian Gehrmann, Mark Dredze. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Minqian Liu, Ioana Baldini, David Rabinowitz, David S. Rosenberg, Sebastian Gehrmann, Mark Dredze
ACL (1)5
2024 Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning
abstract
Shivalika Singh, Freddie Vargus, Daniel D’souza, Börje F. Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura O’Mahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Moura, Dominik Krzemiński, Hakimeh Fadaei, Irem Ergun, Ifeoma Okoh, Aisha Alaagib, Oshan Mudannayake, Zaid Alyafeai, Vu Chien, Sebastian Ruder, Surya Guthikonda, Emad Alghamdi, Sebastian Gehrmann, Niklas Muennighoff, Max Bartolo, Julia Kreutzer, Ahmet Üstün, Marzieh Fadaee, Sara Hooker. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Shivalika Singh, Freddie Vargus, Daniel D'souza, Börje Karlsson 0001, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Deividas Mataciunas, Laura O'Mahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Souza Moura, Dominik Krzeminski, Hakimeh Fadaei, Irem Ergün, Ifeoma Okoh, Aisha Alaagib, Oshan Mudannayake, Zaid Alyafeai, Minh Vu Chien, Sebastian Ruder, Surya Guthikonda, Emad A. Alghamdi, Sebastian Gehrmann, Niklas Muennighoff, Max Bartolo, Julia Kreutzer, Ahmet Üstün, Marzieh Fadaee, Sara Hooker
ACL (1)27
2024 Academics Can Contribute to Domain-Specialized Language Models
abstract
Mark Dredze, Genta Indra Winata, Prabhanjan Kambadur, Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, David S Rosenberg, Sebastian Gehrmann. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Mark Dredze, Genta Indra Winata, Prabhanjan Kambadur, Ozan Irsoy, Steven Lu 0003, Vadim Dabravolski, David S. Rosenberg, Sebastian Gehrmann
EMNLP9
2024 Do LLMs Plan Like Human Writers? Comparing Journalist Coverage of Press Releases with LLMs
abstract
Journalists engage in multiple steps in the news writing process that depend on human creativity, like exploring different "angles" (i.e. the specific perspectives a reporter takes).These can potentially be aided by large language models (LLMs).By affecting planning decisions, such interventions can have an outsize impact on creative output.We advocate a careful approach to evaluating these interventions to ensure alignment with human values.In a case study of journalistic coverage of press releases, we assemble a large dataset of 250k press releases 1 and 650k articles covering them. 2 We develop methods to identify news articles that challenge and contextualize press releases.Finally, we evaluate suggestions made by LLMs for these articles and compare these with decisions made by human journalists.Our findings are three-fold: (1) Human-written news articles that challenge and contextualize press releases more take more creative angles and use more informational sources.(2) LLMs align better with humans when recommending angles, compared with informational sources.(3) Both the angles and sources LLMs suggest are significantly less creative than humans.
Alexander Spangher, Nanyun Peng 0001, Sebastian Gehrmann, Mark Dredze
EMNLP3
2023 Benchmarking Large Language Model Capabilities for Conditional Generation
abstract
Pre-trained large language models (PLMs) underlie most new developments in natural language processing.They have shifted the field from application-specific model pipelines to a single model that is adapted to a wide range of tasks.Autoregressive PLMs like GPT-3 or PaLM, alongside techniques like few-shot learning, have additionally shifted the output modality to generation instead of classification or regression.Despite their ubiquitous use, the generation quality of language models is rarely evaluated when these models are introduced.Additionally, it is unclear how existing generation tasks--while they can be used to compare systems at a high level-relate to the real world use cases for which people have been adopting them.In this work, we discuss how to adapt existing applicationspecific generation benchmarks to PLMs and provide an in-depth, empirical study of the limitations and capabilities of PLMs in natural language generation tasks along dimensions such as scale, architecture, input and output language.Our results show that PLMs differ in their applicability to different data regimes and their generalization to multiple languages and inform which PLMs to use for a given generation task setup.We share best practices to be taken into consideration when benchmarking generation capabilities during the development of upcoming PLMs.
Joshua Maynez, Priyanka Agrawal, Sebastian Gehrmann
ACL (1)3
2023 Dialect-robust Evaluation of Generated Text
abstract
Jiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, Sebastian Gehrmann. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, Sebastian Gehrmann
ACL (1)9
2023 A Needle in a Haystack: An Analysis of High-Agreement Workers on MTurk for Summarization
abstract
Lining Zhang, Simon Mille, Yufang Hou, Daniel Deutsch, Elizabeth Clark, Yixin Liu, Saad Mahamood, Sebastian Gehrmann, Miruna Clinciu, Khyathi Raghavi Chandu, João Sedoc. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Lining Zhang, Simon Mille, Yufang Hou 0001, Daniel Deutsch, Elizabeth Clark, Yixin Liu 0003, Saad Mahamood, Sebastian Gehrmann, Miruna-Adriana Clinciu, Khyathi Raghavi Chandu, João Sedoc
ACL (1)8
2023 SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization Evaluation
abstract
Elizabeth Clark, Shruti Rijhwani, Sebastian Gehrmann, Joshua Maynez, Roee Aharoni, Vitaly Nikolaev, Thibault Sellam, Aditya Siddhant, Dipanjan Das, Ankur Parikh. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Elizabeth Clark, Shruti Rijhwani, Sebastian Gehrmann, Joshua Maynez, Roee Aharoni, Vitaly Nikolaev, Thibault Sellam, Aditya Siddhant, Dipanjan Das 0001, Ankur P. Parikh
EMNLP3
2023 Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text
abstract
Evaluation practices in natural language generation (NLG) have many known flaws, but improved evaluation approaches are rarely widely adopted. This issue has become more urgent, since neural generation models have improved to the point where their outputs can often no longer be distinguished based on the surface-level features that older metrics rely on. This paper surveys the issues with human and automatic model evaluations and with commonly used datasets in NLG that have been pointed out over the past 20 years. We summarize, categorize, and discuss how researchers have been addressing these issues and what their findings mean for the current state of model evaluations. Building on those insights, we lay out a long-term vision for evaluation research and propose concrete steps for researchers to improve their evaluation processes. Finally, we analyze 66 generation papers from recent NLP conferences in how well they already follow these suggestions and identify which areas require more drastic changes to the status quo.
Sebastian Gehrmann, Elizabeth Clark, Thibault Sellam
J. Artif. Intell. Res.1
2023 Diagnosing AI Explanation Methods with Folk Concepts of Behavior
Alon Jacovi, Jasmijn Bastings, Sebastian Gehrmann, Yoav Goldberg, Katja Filippova
J. Artif. Intell. Res.3
2023 PaLM: Scaling Language Modeling with Pathways
abstract
Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model (PaLM). We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of-the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state-of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies.
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Adam Roberts, Paul Barham 0001, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du 0002, Ben Hutchinson, Reiner Pope, Jacob Austin, Michael Isard, Guy Gur-Ari, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, William Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang 0002, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeffrey Dean, Slav Petrov, Noah Fiedel
J. Mach. Learn. Res.10
2022 Intriguing Properties of Compression on Multilingual Models
abstract
Multilingual models are often particularly dependent on scaling to generalize to a growing number of languages.Compression techniques are widely relied upon to reconcile the growth in model size with real world resource constraints, but compression can have a disparate effect on model performance for lowresource languages.It is thus crucial to understand the trade-offs between scale, multilingualism, and compression.In this work, we propose an experimental framework to characterize the impact of sparsifying multilingual pre-trained language models during finetuning.Applying this framework to mBERT named entity recognition models across 40 languages, we find that compression confers several intriguing and previously unknown generalization properties.In contrast to prior findings, we find that compression may improve model robustness over dense models.We additionally observe that under certain sparsification regimes compression may aid, rather than disproportionately impact the performance of low-resource languages.
Kelechi Ogueji, Orevaoghene Ahia, Gbemileke Onilude, Sebastian Gehrmann, Sara Hooker, Julia Kreutzer
EMNLP4
2021 Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models
abstract
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen, Yonatan Belinkov. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart M. Shieber, Tal Linzen, Yonatan Belinkov
ACL/IJCNLP (1)3
2021 Learning Compact Metrics for MT
abstract
Recent developments in machine translation and multilingual text generation have led researchers to adopt trained metrics such as COMET or BLEURT, which treat evaluation as a regression problem and use representations from multilingual pre-trained models such as XLM-RoBERTa or mBERT.Yet studies on related tasks suggest that these models are most efficient when they are large, which is costly and impractical for evaluation.We investigate the trade-off between multilinguality and model capacity with RemBERT, a stateof-the-art multilingual language model, using data from the WMT Metrics Shared Task.We present a series of experiments which show that model size is indeed a bottleneck for cross-lingual transfer, then demonstrate how distillation can help addressing this bottleneck, by leveraging synthetic data generation and transferring knowledge from one teacher to multiple students trained on related languages.Our method yields up to 10.5% improvement over vanilla fine-tuning and reaches 92.6% of RemBERT's performance using only a third of its parameters. * Work done during an internship at Google.
Amy Pu, Hyung Won Chung, Ankur P. Parikh, Sebastian Gehrmann, Thibault Sellam
EMNLP (1)4
2020 ToTTo: A Controlled Table-To-Text Generation Dataset
abstract
Ankur Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, Dipanjan Das. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Ankur P. Parikh, Xuezhi Wang 0002, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, Dipanjan Das 0001
EMNLP (1)3
2020 A Corpus for Detecting High-Context Medical Conditions in Intensive Care Patient Notes Focusing on Frequently Readmitted Patients
abstract
A crucial step within secondary analysis of electronic health records (EHRs) is to identify the patient cohort under investigation. While EHRs contain medical billing codes that aim to represent the conditions and treatments patients may have, much of the information is only present in the patient notes. Therefore, it is critical to develop robust algorithms to infer patients’ conditions and treatments from their written notes. In this paper, we introduce a dataset for patient phenotyping, a task that is defined as the identification of whether a patient has a given medical condition (also referred to as clinical indication or phenotype) based on their patient note. Nursing Progress Notes and Discharge Summaries from the Intensive Care Unit of a large tertiary care hospital were manually annotated for the presence of several high-context phenotypes relevant to treatment and risk of re-hospitalization. This dataset contains 1102 Discharge Summaries and 1000 Nursing Progress Notes. Each Discharge Summary and Progress Note has been annotated by at least two expert human annotators (one clinical researcher and one resident physician). Annotated phenotypes include treatment non-adherence, chronic pain, advanced/metastatic cancer, as well as 10 other phenotypes. This dataset can be utilized for academic and industrial research in medicine and computer science, particularly within the field of medical natural language processing.
Edward T. Moseley, Joy T. Wu, Jonathan Welt, John Foote Jr., Patrick D. Tyler, David W. Grant, Eric T. Carlson, Sebastian Gehrmann, Franck Dernoncourt, Leo A. Celi
LREC8
2020 Investigating Gender Bias in Language Models Using Causal Mediation Analysis
abstract
Many interpretation methods for neural models in natural language processing investigate how information is encoded inside hidden representations. However, these methods can only measure whether the information exists, not whether it is actually used by the model. We propose a methodology grounded in the theory of causal mediation analysis for interpreting which parts of a model are causally implicated in its behavior. The approach enables us to analyze the mechanisms that facilitate the flow of information from input to output through various model components, known as mediators. As a case study, we apply this methodology to analyzing gender bias in pre-trained Transformer language models. We study the role of individual neurons and attention heads in mediating gender bias across three datasets designed to gauge a model's sensitivity to gender bias. Our mediation analysis reveals that gender bias effects are concentrated in specific components of the model that may exhibit highly specialized behavior.
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, Stuart M. Shieber
NeurIPS2
2020 Evaluating an automated mediator for joint narratives in a conflict situation
abstract
Joint narratives are often used in the context of reconciliation interventions for people in social conflict situations, which arise, for example, due to ethnic or religious differences. The interventions aim to encourage a change in attitudes of the participants towards each other. Typically, a human mediator is fundamental for achieving a successful intervention. In this work, we present an automated approach to support remote interactions between pairs of participants as they contribute to a shared story in their own language. A key component is an automated cognitive tutor that guides the participants through a controlled escalation/de-escalation process during the development of a joint narrative. We performed a controlled study comparing a trained human mediator to the automated mediator. The results demonstrate that an automated mediator, although simple at this stage, effectively supports interactions and helps to achieve positive outcomes comparable to those attained by the trained human mediator.
Massimo Zancanaro, Oliviero Stock, Gianluca Schiavo, Alessandro Cappelletti, Sebastian Gehrmann, Daphna Canetti, Ohad Shaked, Shani Fachter, Rachel Yifat, Ravit Mimran, Patrice L. (Tamar) Weiss
Behav. Inf. Technol.5
2020 Visual Interaction with Deep Learning Models through Collaborative Semantic Inference
abstract
Automation of tasks can have critical consequences when humans lose agency over decision processes. Deep learning models are particularly susceptible since current black-box approaches lack explainable reasoning. We argue that both the visual interface and model structure of deep learning systems need to take into account interaction design. We propose a framework of collaborative semantic inference (CSI) for the co-design of interactions and models to enable visual collaboration between humans and algorithms. The approach exposes the intermediate reasoning process of models which allows semantic interactions with the visual metaphors of a problem, which means that a user can both understand and control parts of the model reasoning process. We demonstrate the feasibility of CSI with a co-designed case study of a document summarization system.
Sebastian Gehrmann, Hendrik Strobelt, Robert Krüger, Hanspeter Pfister, Alexander M. Rush
IEEE Trans. Vis. Comput. Graph.1
2019 Identifying documented medical non-adherence from clinical notes using natural language processing
Joy T. Wu, David W. Grant, Shrey Lakhotia, Patrick D. Tyler, Daniel Gruhl, Chaitanya P. Shivade, Sebastian Gehrmann, Leo A. Celi
AMIA7
2019 Generating Abstractive Summaries with Finetuned Language Models
abstract
Neural abstractive document summarization is commonly approached by models that exhibit a mostly extractive behavior.This behavior is facilitated by a copy-attention which allows models to copy words from a source document.While models in the mostly extractive news summarization domain benefit from this inductive bias, they commonly fail to paraphrase or compress information from the source document.Recent advances in transferlearning from large pretrained language models give rise to alternative approaches that do not rely on copy-attention and instead learn to generate concise and abstractive summaries.In this paper, as part of the TL;DR challenge, we compare the abstractiveness of summaries from different summarization approaches and show that transfer-learning can be efficiently utilized without any changes to the model architecture.We demonstrate that the approach leads to a higher level of abstraction for a similar performance on the TL;DR challenge tasks, enabling true natural language compression.
Sebastian Gehrmann, Zachary M. Ziegler, Alexander M. Rush
INLG1
2019 Margin Call: an Accessible Web-based Text Viewer with Generated Paragraph Summaries in the Margin
abstract
We present Margin Call, an accessible webbased text viewer that automatically generates short summaries for each paragraph of the text and displays the summaries in the margin of the text next to the corresponding paragraph.On the back-end, the summarizer first identifies the most important sentence for each paragraph in the text file uploaded by the user.The selected sentence is then automatically compressed to produce the short summary.The resulting summary is a few words long.The displayed summaries can help the user understand and retrieve information faster from the text, while increasing the retention of information.
Naba Rizvi, Sebastian Gehrmann, Franck Dernoncourt
INLG2
2019 Seq2seq-Vis: A Visual Debugging Tool for Sequence-to-Sequence Models
abstract
Neural sequence-to-sequence models have proven to be accurate and robust for many sequence prediction tasks, and have become the standard approach for automatic translation of text. The models work with a five-stage blackbox pipeline that begins with encoding a source sequence to a vector space and then decoding out to a new target sequence. This process is now standard, but like many deep learning methods remains quite difficult to understand or debug. In this work, we present a visual analysis tool that allows interaction and "what if"-style exploration of trained sequence-to-sequence models through each stage of the translation process. The aim is to identify which patterns have been learned, to detect model errors, and to probe the model with counterfactual scenario. We demonstrate the utility of our tool through several real-world sequence-to-sequence use cases on large-scale models.
Hendrik Strobelt, Sebastian Gehrmann, Michael Behrisch 0001, Adam Perer, Hanspeter Pfister, Alexander M. Rush
IEEE Trans. Vis. Comput. Graph.2
2018 Bottom-Up Abstractive Summarization
abstract
Neural network-based methods for abstractive summarization produce outputs that are more fluent than other techniques, but which can be poor at content selection.This work proposes a simple technique for addressing this issue: use a data-efficient content selector to overdetermine phrases in a source document that should be part of the summary.We use this selector as a bottom-up attention step to constrain the model to likely phrases.We show that this approach improves the ability to compress text, while still generating fluent summaries.This two-step process is both simpler and higher performing than other end-to-end content selection models, leading to significant improvements on ROUGE for both the CNN-DM and NYT corpus.Furthermore, the content selector can be trained with as little as 1,000 sentences, making it easy to transfer a trained summarizer to a new domain.
Sebastian Gehrmann, Yuntian Deng, Alexander M. Rush
EMNLP1
2018 E2E NLG Challenge Submission: Towards Controllable Generation of Diverse Natural Language
abstract
In natural language generation (NLG), the task is to generate utterances from a more abstract input, such as structured data.An added challenge is to generate utterances that contain an accurate representation of the input, while reflecting the fluency and variety of human-generated text.In this paper, we report experiments with NLG models that can be used in task oriented dialogue systems.We explore the use of additional input to the model to encourage diversity and control of outputs.While our submission does not rank highly using automated metrics, qualitative investigation of generated utterances suggests the use of additional information in neural network NLG systems to be a promising research direction.
Henry Elder, Sebastian Gehrmann, Alexander O'Connor, Qun Liu 0001
INLG2
2018 End-to-End Content and Plan Selection for Data-to-Text Generation
abstract
Learning to generate fluent natural language from structured data with neural networks has become an common approach for NLG.This problem can be challenging when the form of the structured data varies between examples.This paper presents a survey of several extensions to sequence-to-sequence models to account for the latent content selection process, particularly variants of copy attention and coverage decoding.We further propose a training method based on diverse ensembling to encourage models to learn distinct sentence templates during training.An empirical evaluation of these techniques shows an increase in the quality of generated text across five automated metrics, as well as human evaluation.
Sebastian Gehrmann, Falcon Z. Dai, Henry Elder, Alexander M. Rush
INLG1
2018 LSTMVis: A Tool for Visual Analysis of Hidden State Dynamics in Recurrent Neural Networks
abstract
Recurrent neural networks, and in particular long short-term memory (LSTM) networks, are a remarkably effective tool for sequence modeling that learn a dense black-box hidden representation of their sequential input. Researchers interested in better understanding these models have studied the changes in hidden state representations over time and noticed some interpretable patterns but also significant noise. In this work, we present LSTMVis, a visual analysis tool for recurrent neural networks with a focus on understanding these hidden state dynamics. The tool allows users to select a hypothesis input range to focus on local state changes, to match these states changes to similar patterns in a large data set, and to align these results with structural annotations from their domain. We show several use cases of the tool for analyzing specific hidden state properties on dataset containing nesting, phrase structure, and chord progressions, and demonstrate how the tool can be used to isolate patterns for further statistical analysis. We characterize the domain, the different stakeholders, and their goals and tasks. Long-term usage data after putting the tool online revealed great interest in the machine learning community.
Hendrik Strobelt, Sebastian Gehrmann, Hanspeter Pfister, Alexander M. Rush
IEEE Trans. Vis. Comput. Graph.2