Sam Wiseman

dblp:149/1260 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
6since 2021 · last 2024
0000-0003-0923-1086ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 8 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
18 papers
Language models and text generation · 33% Probabilistic and Bayesian machine learning · 18% Information extraction and text analysis · 17%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 30 heaviest of 46, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
coreference resolution
1.332023
Seq2seq is All You Need for Coreference Resolution · EMNLP 2023
Learning to Ignore: Long Document Coreference with Bounded Memory Neural Networks · EMNLP (1) 2020
Learning Anaphoricity and Antecedent Ranking Features for Coreference Resolution · ACL (1) 2015
Machine learning › Representation and self-supervised learning
discrete representation
0.922020
Learning Discrete Structured Representations by Adversarially Maximizing Mutual Information · ICML 2020
Discrete Latent Variable Representations for Low-Resource Text Classification · ACL 2020
Natural language and speech › Language models and text generation › text generation
data-to-text generation
0.822021
Data-to-text Generation by Splicing Together Nearest Neighbors · EMNLP (1) 2021
Challenges in Data-to-Document Generation · EMNLP 2017
Natural language and speech › Language models and text generation
text generation
0.822021
Data-to-text Generation by Splicing Together Nearest Neighbors · EMNLP (1) 2021
Challenges in Data-to-Document Generation · EMNLP 2017
Natural language and speech › Language models and text generation
controllable text generation
0.722019
Controllable Paraphrase Generation with a Syntactic Exemplar · ACL (1) 2019
Learning Neural Templates for Text Generation · EMNLP 2018
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
sequence-to-sequence generation
0.712023
Seq2seq is All You Need for Coreference Resolution · EMNLP 2023
Natural language and speech › Language models and text generation
large language model evaluation
0.612022
Chess as a Testbed for Language Model State Tracking · AAAI 2022
Machine learning › Deep learning architectures and training › sequence modeling
state tracking
0.612022
Chess as a Testbed for Language Model State Tracking · AAAI 2022
Natural language and speech › Language models and text generation
text summarization
0.612022
SummScreen: A Dataset for Abstractive Screenplay Summarization · ACL (1) 2022
Natural language and speech › Language models and text generation › language modeling › language model architecture
transformer language model
0.612022
Chess as a Testbed for Language Model State Tracking · AAAI 2022
Information retrieval
evaluation
0.612022
SummScreen: A Dataset for Abstractive Screenplay Summarization · ACL (1) 2022
Information retrieval › text summarization
summarization evaluation
0.612022
SummScreen: A Dataset for Abstractive Screenplay Summarization · ACL (1) 2022
Machine learning › Generative modeling › diffusion model
controllable generation
0.512021
Data-to-text Generation by Splicing Together Nearest Neighbors · EMNLP (1) 2021
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
discrete latent variable model
0.412020
Discrete Latent Variable Representations for Low-Resource Text Classification · ACL 2020
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.412020
Discrete Latent Variable Representations for Low-Resource Text Classification · ACL 2020
Natural language and speech › Information extraction and text analysis › coreference resolution
long document coreference resolution
0.412020
Learning to Ignore: Long Document Coreference with Bounded Memory Neural Networks · EMNLP (1) 2020
Machine learning › Deep learning architectures and training
memory-augmented neural networks
0.412020
Learning to Ignore: Long Document Coreference with Bounded Memory Neural Networks · EMNLP (1) 2020
Machine learning › Representation and self-supervised learning
mutual information maximization
0.412020
Learning Discrete Structured Representations by Adversarially Maximizing Mutual Information · ICML 2020
Natural language and speech › Machine translation › neural machine translation
non-autoregressive machine translation
0.412020
ENGINE: Energy-Based Inference Networks for Non-Autoregressive Machine Translation · ACL 2020
Natural language and speech › Information extraction and text analysis
text classification
0.412020
Discrete Latent Variable Representations for Low-Resource Text Classification · ACL 2020
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference
0.412019
Amortized Bethe Free Energy Minimization for Learning MRFs · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
bethe free energy
0.412019
Amortized Bethe Free Energy Minimization for Learning MRFs · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › markov random field
markov network structure learning
0.412019
Amortized Bethe Free Energy Minimization for Learning MRFs · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
markov random field
0.412019
Amortized Bethe Free Energy Minimization for Learning MRFs · NeurIPS 2019
Natural language and speech › Language models and text generation › text generation
paraphrase generation
0.412019
Controllable Paraphrase Generation with a Syntactic Exemplar · ACL (1) 2019
Natural language and speech › Information extraction and text analysis
sequence labeling
0.412019
Label-Agnostic Sequence Labeling by Copying Nearest Neighbors · ACL (1) 2019
Natural language and speech › Language models and text generation › controllable text generation
syntactic control
0.412019
Controllable Paraphrase Generation with a Syntactic Exemplar · ACL (1) 2019
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › amortized inference
amortized variational inference
0.312018
Semi-Amortized Variational Autoencoders · ICML 2018
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
cloze-style reading comprehension
0.312018
Entity Tracking Improves Cloze-style Reading Comprehension · EMNLP 2018
Natural language and speech › Information extraction and text analysis
entity tracking
0.312018
Entity Tracking Improves Cloze-style Reading Comprehension · EMNLP 2018

Methods — techniques the papers use, named apart from their topics

neural summarization · 1.1nearest neighbor retrieval · 1.1parsing · 1.0multimodal instruction tuning · 0.8large-scale pretraining · 0.8multi-task learning · 0.7seq2seq transformer · 0.7fine-tuning · 0.7probing · 0.6attention analysis · 0.6policy learning · 0.5
YearPublicationVenuePosition
2024 Sequence Reducible Holdout Loss for Language Model Pretraining
abstract
Data selection techniques, which adaptively select datapoints inside the training loop, have demonstrated empirical benefits in reducing the number of gradient steps to train neural models. However, these techniques have so far largely been applied to classification. In this work, we study their applicability to language model pretraining, a highly time-intensive task. We propose a simple modification to an existing data selection technique (reducible hold-out loss training) in order to adapt it to the sequence losses typical in language modeling. We experiment on both autoregressive and masked language modelling, and show that applying data selection to pretraining offers notable benefits including a 4.3% reduction in total number of steps, a 21.5% steps reduction in average, to an intermediate target perplexity, over the course of pretraining an autoregressive language model. Further, data selection trained language models demonstrate significantly better generalization ability on out of domain datasets - 7.9% reduction in total number of steps and 23.2% average steps reduction to an intermediate target perplexity.
Raghuveer Thirukovalluru, Nicholas Monath, Bhuwan Dhingra, Sam Wiseman
LREC/COLING4
2024 MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang 0002, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Anton Belyi, Haotian Zhang 0005, Karanjeet Singh 0003, Doug Kang, Hongyu Hè, Max Schwarzer, Tom Gunter, Xiang Kong, Aonan Zhang, Nan Du 0002, Tao Lei 0001, Sam Wiseman, Mark Lee 0003, Ruoming Pang, Peter Grasch, Alexander Toshev, Yinfei Yang
ECCV (29)23
2023 Seq2seq is All You Need for Coreference Resolution
abstract
Existing works on coreference resolution suggest that task-specific models are necessary to achieve state-of-the-art performance.In this work, we present compelling evidence that such models are not necessary.We finetune a pretrained seq2seq transformer to map an input document to a tagged sequence encoding the coreference annotation.Despite the extreme simplicity, our model outperforms or closely matches the best coreference systems in the literature on an array of datasets.We also propose an especially simple seq2seq approach that generates only tagged spans rather than the spans interleaved with the original text.Our analysis shows that the model size, the amount of supervision, and the choice of sequence representations are key factors in performance.
Wenzheng Zhang 0003, Sam Wiseman, Karl Stratos
EMNLP2
2022 Chess as a Testbed for Language Model State Tracking
abstract
Transformer language models have made tremendous strides in natural language understanding tasks. However, the complexity of natural language makes it challenging to ascertain how accurately these models are tracking the world state underlying the text. Motivated by this issue, we consider the task of language modeling for the game of chess. Unlike natural language, chess notations describe a simple, constrained, and deterministic domain. Moreover, we observe that the appropriate choice of chess notation allows for directly probing the world state, without requiring any additional probing-related machinery. We find that: (a) With enough training data, transformer language models can learn to track pieces and predict legal moves with high accuracy when trained solely on move sequences. (b) For small training sets providing access to board state information during training can yield significant improvements. (c) The success of transformer language models is dependent on access to the entire game history i.e. “full attention”. Approximating this full attention results in a significant performance drop. We propose this testbed as a benchmark for future work on the development and analysis of transformer language models.
Shubham Toshniwal, Sam Wiseman, Karen Livescu, Kevin Gimpel
AAAI2
2022 SummScreen: A Dataset for Abstractive Screenplay Summarization
abstract
We introduce SUMMSCREEN, a summarization dataset comprised of pairs of TV series transcripts and human written recaps.The dataset provides a challenging testbed for abstractive summarization for several reasons.Plot details are often expressed indirectly in character dialogues and may be scattered across the entirety of the transcript.These details must be found and integrated to form the succinct plot descriptions in the recaps.Also, TV scripts contain content that does not directly pertain to the central plot but rather serves to develop characters or provide comic relief.This information is rarely contained in recaps.Since characters are fundamental to TV series, we also propose two entity-centric evaluation metrics.Empirically, we characterize the dataset by evaluating several methods, including neural models and those based on nearest neighbors.An oracle extractive approach outperforms all benchmarked models according to automatic metrics, showing that the neural models are unable to fully exploit the input transcripts.Human evaluation and qualitative analysis reveal that our nonoracle models are competitive with their oracle counterparts in terms of generating faithful plot events and can benefit from better content selectors.Both oracle and non-oracle models generate unfaithful facts, suggesting future research directions.
Mingda Chen, Zewei Chu, Sam Wiseman, Kevin Gimpel
ACL (1)3
2021 Data-to-text Generation by Splicing Together Nearest Neighbors
abstract
We propose to tackle data-to-text generation tasks by directly splicing together retrieved segments of text from "neighbor" sourcetarget pairs.Unlike recent work that conditions on retrieved neighbors but generates text token-by-token, left-to-right, we learn a policy that directly manipulates segments of neighbor text, by inserting or replacing them in partially constructed generations.Standard techniques for training such a policy require an oracle derivation for each generation, and we prove that finding the shortest such derivation can be reduced to parsing under a particular weighted context-free grammar.We find that policies learned in this way perform on par with strong baselines in terms of automatic and human evaluation, but allow for more interpretable and controllable generation.
Sam Wiseman, Arturs Backurs, Karl Stratos
EMNLP (1)1
2020 Discrete Latent Variable Representations for Low-Resource Text Classification
abstract
While much work on deep latent variable models of text uses continuous latent variables, discrete latent variables are interesting because they are more interpretable and typically more space efficient.We consider several approaches to learning discrete latent variable models for text in the case where exact marginalization over these variables is intractable.We compare the performance of the learned representations as features for lowresource document and sentence classification.Our best models outperform the previous best reported results with continuous representations in these low-resource settings, while learning significantly more compressed representations.Interestingly, we find that an amortized variant of Hard EM performs particularly well in the lowest-resource regimes. 1
Shuning Jin, Sam Wiseman, Karl Stratos, Karen Livescu
ACL2
2020 ENGINE: Energy-Based Inference Networks for Non-Autoregressive Machine Translation
abstract
We propose to train a non-autoregressive machine translation model to minimize the energy defined by a pretrained autoregressive model.In particular, we view our non-autoregressive translation system as an inference network (Tu and Gimpel, 2018) trained to minimize the autoregressive teacher energy.This contrasts with the popular approach of training a non-autoregressive model on a distilled corpus consisting of the beam-searched outputs of such a teacher model.Our approach, which we call ENGINE (ENerGy-based Inference NEtworks), achieves state-of-the-art non-autoregressive results on the IWSLT 2014 DE-EN and WMT 2016 RO-EN datasets, approaching the performance of autoregressive models.
Lifu Tu, Richard Yuanzhe Pang, Sam Wiseman, Kevin Gimpel
ACL3
2020 Learning to Ignore: Long Document Coreference with Bounded Memory Neural Networks
abstract
Long document coreference resolution remains a challenging task due to the large memory and runtime requirements of current models.Recent work doing incremental coreference resolution using just the global representation of entities shows practical benefits but requires keeping all entities in memory, which can be impractical for long documents.We argue that keeping all entities in memory is unnecessary, and we propose a memoryaugmented neural network that tracks only a small bounded number of entities at a time, thus guaranteeing a linear runtime in length of document.We show that (a) the model remains competitive with models with high memory and computational requirements on OntoNotes and LitBank, and (b) the model learns an efficient memory management strategy easily outperforming a rule-based strategy.
Shubham Toshniwal, Sam Wiseman, Allyson Ettinger, Karen Livescu, Kevin Gimpel
EMNLP (1)2
2020 Learning Discrete Structured Representations by Adversarially Maximizing Mutual Information
abstract
We propose learning discrete structured representations from unlabeled data by maximizing the mutual information between a structured latent variable and a target variable. Calculating mutual information is intractable in this setting. Our key technical contribution is an adversarial objective that can be used to tractably estimate mutual information assuming only the feasibility of cross entropy calculation. We develop a concrete realization of this general formulation with Markov distributions over binary encodings. We report critical and unexpected findings on practical aspects of the objective such as the choice of variational priors. We apply our model on document hashing and show that it outperforms current best baselines based on discrete and vector quantized variational autoencoders. It also yields highly compressed interpretable representations.
Karl Stratos, Sam Wiseman
ICML2
2019 Controllable Paraphrase Generation with a Syntactic Exemplar
abstract
Prior work on controllable text generation usually assumes that the controlled attribute can take on one of a small set of values known a priori.In this work, we propose a novel task, where the syntax of a generated sentence is controlled rather by a sentential exemplar.To evaluate quantitatively with standard metrics, we create a novel dataset with human annotations.We also develop a variational model with a neural module specifically designed for capturing syntactic knowledge and several multitask training objectives to promote disentangled representation learning.Empirically, the proposed model is observed to achieve improvements over baselines and learn to capture desirable characteristics.
Mingda Chen, Qingming Tang, Sam Wiseman, Kevin Gimpel
ACL (1)3
2019 Label-Agnostic Sequence Labeling by Copying Nearest Neighbors
abstract
Retrieve-and-edit based approaches to structured prediction, where structures associated with retrieved neighbors are edited to form new structures, have recently attracted increased interest.However, much recent work merely conditions on retrieved structures (e.g., in a sequence-to-sequence framework), rather than explicitly manipulating them.We show we can perform accurate sequence labeling by explicitly (and only) copying labels from retrieved neighbors.Moreover, because this copying is label-agnostic, we can achieve impressive performance in zero-shot sequencelabeling tasks.We additionally consider a dynamic programming approach to sequence labeling in the presence of retrieved neighbors, which allows for controlling the number of distinct (copied) segments used to form a prediction, and leads to both more interpretable and accurate predictions.
Sam Wiseman, Karl Stratos
ACL (1)1
2019 Amortized Bethe Free Energy Minimization for Learning MRFs
abstract
We propose to learn deep undirected graphical models (i.e., MRFs) with a non-ELBO objective for which we can calculate exact gradients. In particular, we optimize a saddle-point objective deriving from the Bethe free energy approximation to the partition function. Unlike much recent work in approximate inference, the derived objective requires no sampling, and can be efficiently computed even for very expressive MRFs. We furthermore amortize this optimization with trained inference networks. Experimentally, we find that the proposed approach compares favorably with loopy belief propagation, but is faster, and it allows for attaining better held out log likelihood than other recent approximate inference schemes.
Sam Wiseman
NeurIPS1
2018 Entity Tracking Improves Cloze-style Reading Comprehension
abstract
Reading comprehension tasks test the ability of models to process long-term context and remember salient information.Recent work has shown that relatively simple neural methods such as the Attention Sum-Reader can perform well on these tasks; however, these systems still significantly trail human performance.Analysis suggests that many of the remaining hard instances are related to the inability to track entity-references throughout documents.This work focuses on these hard entity tracking cases with two extensions: (1) additional entity features, and (2) training with a multi-task tracking objective.We show that these simple modifications improve performance both independently and in combination, and we outperform the previous state of the art on the LAMBADA dataset, particularly on difficult entity examples.
Luong Hoang, Sam Wiseman, Alexander M. Rush
EMNLP2
2018 Learning Neural Templates for Text Generation
abstract
While neural, encoder-decoder models have had significant empirical success in text generation, there remain several unaddressed problems with this style of generation.Encoderdecoder models are largely (a) uninterpretable, and (b) difficult to control in terms of their phrasing or content.This work proposes a neural generation system using a hidden semimarkov model (HSMM) decoder, which learns latent, discrete templates jointly with learning to generate.We show that this model learns useful templates, and that these templates make generation both more interpretable and controllable.Furthermore, we show that this approach scales to real data sets and achieves strong performance nearing that of encoderdecoder text generation models.
Sam Wiseman, Stuart M. Shieber, Alexander M. Rush
EMNLP1
2018 Semi-Amortized Variational Autoencoders
abstract
Amortized variational inference (AVI) replaces instance-specific local inference with a global inference network. While AVI has enabled efficient training of deep generative models such as variational autoencoders (VAE), recent empirical work suggests that inference networks can produce suboptimal variational parameters. We propose a hybrid approach, to use AVI to initialize the variational parameters and run stochastic variational inference (SVI) to refine them. Crucially, the local SVI procedure is itself differentiable, so the inference network and generative model can be trained end-to-end with gradient-based optimization. This semi-amortized approach enables the use of rich generative models without experiencing the posterior-collapse phenomenon common in training VAEs for problems like text generation. Experiments show this approach outperforms strong autoregressive and variational baselines on standard text and image datasets.
Sam Wiseman, Andrew C. Miller, David A. Sontag, Alexander M. Rush
ICML2
2017 Challenges in Data-to-Document Generation
abstract
Recent neural models have shown significant progress on the problem of generating short descriptive texts conditioned on a small number of database records.In this work, we suggest a slightly more difficult data-to-text generation task, and investigate how effective current approaches are on this task.In particular, we introduce a new, large-scale corpus of data records paired with descriptive documents, propose a series of extractive evaluation methods for analyzing performance, and obtain baseline results using current neural generation methods.Experiments show that these models produce fluent text, but fail to convincingly approximate humangenerated documents.Moreover, even templated baselines exceed the performance of these neural models on some metrics, though copy-and reconstructionbased extensions lead to noticeable improvements.
Sam Wiseman, Stuart M. Shieber, Alexander M. Rush
EMNLP1
2016 Sequence-to-Sequence Learning as Beam-Search Optimization
abstract
Sequence-to-Sequence (seq2seq) modeling has rapidly become an important generalpurpose NLP tool that has proven effective for many text-generation and sequence-labeling tasks.Seq2seq builds on deep neural language modeling and inherits its remarkable accuracy in estimating local, next-word distributions.In this work, we introduce a model and beamsearch training scheme, based on the work of Daumé III and Marcu (2005), that extends seq2seq to learn global sequence scores.This structured approach avoids classical biases associated with local training and unifies the training loss with the test-time usage, while preserving the proven model architecture of seq2seq and its efficient training approach.We show that our system outperforms a highlyoptimized attention-based seq2seq system and other baselines on three different sequence to sequence tasks: word ordering, parsing, and machine translation.
Sam Wiseman, Alexander M. Rush
EMNLP1
2016 Learning Global Features for Coreference Resolution
abstract
There is compelling evidence that coreference prediction would benefit from modeling global information about entity-clusters.Yet, state-of-the-art performance can be achieved with systems treating each mention prediction independently, which we attribute to the inherent difficulty of crafting informative clusterlevel features.We instead propose to use recurrent neural networks (RNNs) to learn latent, global representations of entity clusters directly from their mentions.We show that such representations are especially useful for the prediction of pronominal mentions, and can be incorporated into an end-to-end coreference system that outperforms the state of the art without requiring any additional search.
Sam Wiseman, Alexander M. Rush, Stuart M. Shieber
HLT-NAACL1
2015 Learning Anaphoricity and Antecedent Ranking Features for Coreference Resolution
abstract
Sam Wiseman, Alexander M. Rush, Stuart Shieber, Jason Weston. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Sam Wiseman, Alexander M. Rush, Stuart M. Shieber, Jason Weston
ACL (1)1