EDBT 2026 Demo / reviewers in the wild / expert
John Wieting
dblp:156/0158 · also John Frederick Wieting
· DBLP profile ↗
26ranked-venue papers
10as first author
13since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 10 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | PostMark: A Robust Blackbox Watermark for Large Language ModelsabstractThe most effective techniques to detect LLMgenerated text rely on inserting a detectable signature-or watermark-during the model's decoding process.Most existing watermarking methods require access to the underlying LLM's logits, which LLM API providers are loath to share due to fears of model distillation.As such, these watermarks must be implemented independently by each LLM provider.In this paper, we develop POSTMARK, a modular post-hoc watermarking procedure in which an input-dependent set of words (determined via a semantic embedding) is inserted into the text after the decoding process has completed.Critically, POSTMARK does not require logit access, which means it can be implemented by a third party.We also show that POST-MARK is more robust to paraphrasing attacks than existing watermarking methods: our experiments cover eight baseline algorithms, five base LLMs, and three datasets.Finally, we evaluate the impact of POSTMARK on text quality using both automated and human assessments, highlighting the trade-off between quality and robustness to paraphrasing.We release our code, outputs, and annotations at https://github.com/lilakk/PostMark. Yapei Chang, Kalpesh Krishna, Amir Houmansadr, John Wieting, Mohit Iyyer |
EMNLP | 4 |
| 2024 | Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense RetrievalabstractNandan Thakur, Jianmo Ni, Gustavo Hernandez Abrego, John Wieting, Jimmy Lin, Daniel Cer. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Nandan Thakur, Jianmo Ni, Gustavo Hernández Ábrego, John Wieting, Jimmy Lin, Daniel M. Cer |
NAACL-HLT | 4 |
| 2023 | QA Is the New KR: Question-Answer Pairs as Knowledge BasesabstractWe propose a new knowledge representation (KR) based on knowledge bases (KBs) derived from text, based on question generation and entity linking. We argue that the proposed type of KB has many of the key advantages of a traditional symbolic KB: in particular, it consists of small modular components, which can be combined compositionally to answer complex queries, including relational queries and queries involving ``multi-hop'' inferences. However, unlike a traditional KB, this information store is well-aligned with common user information needs. We present one such KB, called a QEDB, and give qualitative evidence that the atomic components are high-quality and meaningful, and that atomic components can be combined in ways similar to the triples in a symbolic KB. We also show experimentally that questions reflective of typical user questions are more easily answered with a QEDB than a symbolic KB. William W. Cohen, Wenhu Chen, Michiel de Jong, Nitish Gupta, Alessandro Presta, Patrick Verga, John Wieting |
AAAI | 7 |
| 2023 | Beyond Contrastive Learning: A Variational Generative Model for Multilingual RetrievalabstractJohn Wieting, Jonathan Clark, William Cohen, Graham Neubig, Taylor Berg-Kirkpatrick. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. John Wieting, Jonathan H. Clark, William W. Cohen, Graham Neubig, Taylor Berg-Kirkpatrick |
ACL (1) | 1 |
| 2023 | Augmenting Pre-trained Language Models with QA-Memory for Open-Domain Question AnsweringabstractExisting state-of-the-art methods for opendomain question-answering (ODQA) use anopen book approach in which information is first retrieved from a large text corpus or knowledge base (KB) and then reasoned over to produce an answer.A recent alternative is to retrieve from a collection of previouslygenerated question-answer pairs; this has several practical advantages including being more memory and compute-efficient.Questionanswer pairs are also appealing in that they can be viewed as an intermediate between text and KB triples: like KB triples, they often concisely express a single relationship, but like text, have much higher coverage than traditional KBs.In this work, we describe a new QA system that augments a text-to-text model with a large memory of question-answer pairs, and a new pre-training task for the latent step of question retrieval.The pre-training task substantially simplifies training and greatly improves performance on smaller QA benchmarks.Unlike prior systems of this sort, our QA system can also answer multi-hop questions that do not explicitly appear in the collection of stored question-answer pairs. Wenhu Chen, Patrick Verga, Michiel de Jong, John Wieting, William W. Cohen |
EACL | 4 |
| 2023 | Evaluating and Modeling Attribution for Cross-Lingual Question AnsweringabstractBenjamin Muller, John Wieting, Jonathan Clark, Tom Kwiatkowski, Sebastian Ruder, Livio Soares, Roee Aharoni, Jonathan Herzig, Xinyi Wang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Benjamin Muller, John Wieting, Jonathan H. Clark, Tom Kwiatkowski, Sebastian Ruder, Livio B. Soares, Roee Aharoni, Jonathan Herzig, Xinyi Wang 0001 |
EMNLP | 2 |
| 2023 | Evaluating Large Language Models on Controlled Generation TasksabstractJiao Sun, Yufei Tian, Wangchunshu Zhou, Nan Xu, Qian Hu, Rahul Gupta, John Wieting, Nanyun Peng, Xuezhe Ma. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Jiao Sun, Yufei Tian, Wangchunshu Zhou, Rahul Gupta 0001, John Wieting, Nanyun Peng 0001, Xuezhe Ma |
EMNLP | 7 |
| 2023 | Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseabstractThe rise in malicious usage of large language models, such as fake content creation and academic plagiarism, has motivated the development of approaches that identify AI-generated text, including those based on watermarking or outlier detection. However, the robustness of these detection algorithms to paraphrases of AI-generated text remains unclear. To stress test these detectors, we build a 11B parameter paraphrase generation model (DIPPER) that can paraphrase paragraphs, condition on surrounding context, and control lexical diversity and content reordering. Paraphrasing text generated by three large language models (including GPT3.5-davinci-003) with DIPPER successfully evades several detectors, including watermarking, GPTZero, DetectGPT, and OpenAI's text classifier. For example, DIPPER drops detection accuracy of DetectGPT from 70.3% to 4.6% (at a constant false positive rate of 1%), without appreciably modifying the input semantics.
To increase the robustness of AI-generated text detection to paraphrase attacks, we introduce a simple defense that relies on retrieving semantically-similar generations and must be maintained by a language model API provider. Given a candidate text, our algorithm searches a database of sequences previously generated by the API, looking for sequences that match the candidate text within a certain threshold. We empirically verify our defense using a database of 15M generations from a fine-tuned T5-XXL model and find that it can detect 80% to 97% of paraphrased generations across different settings while only classifying 1% of human-written sequences as AI-generated. We open-source our models, code and data. Kalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting, Mohit Iyyer |
NeurIPS | 4 |
| 2022 | On The Ingredients of an Effective Zero-shot Semantic ParserabstractSemantic parsers map natural language utterances into meaning representations (e.g.programs).Such models are typically bottlenecked by the paucity of training data due to the laborious annotation efforts.Recent studies have performed zero-shot learning by synthesizing training examples of canonical utterances and programs from a grammar, and further paraphrasing these utterances to improve linguistic diversity.However, such synthetic examples cannot fully capture patterns in real data.In this paper we analyze zero-shot parsers through the lenses of the language and logical gaps (Herzig and Berant, 2019), which quantify the discrepancy of language and programmatic patterns between the synthetic canonical examples and real-world user-issued ones.We propose bridging these gaps using improved grammars, stronger paraphrasers, and efficient learning methods using canonical examples that most likely reflect real user intents.Our model achieves strong results on the SCHOLAR and GEO benchmarks with zero labeled data. 1 John Wieting, Avirup Sil, Graham Neubig |
ACL (1) | 2 |
| 2022 | RankGen: Improving Text Generation with Large Ranking ModelsabstractGiven an input sequence (or prefix), modern language models often assign high probabilities to output sequences that are repetitive, incoherent, or irrelevant to the prefix; as such, model-generated text also contains such artifacts.To address these issues we present RANKGEN, a 1.2B parameter encoder model for English that scores model generations given a prefix.RANKGEN can be flexibly incorporated as a scoring function in beam search and used to decode from any pretrained language model.We train RANKGEN using large-scale contrastive learning to map a prefix close to the ground-truth sequence that follows it and far away from two types of negatives:(1) random sequences from the same document as the prefix, and (2) sequences generated from a large language model conditioned on the prefix.Experiments across four different language models (345M-11B parameters) and two domains show that RANKGEN significantly outperforms decoding algorithms like nucleus, top-k, and typical sampling on both automatic metrics (85.0 vs 77.3 MAUVE) as well as human evaluations with English writers (74.5% human preference over nucleus sampling).Analysis reveals that RANKGEN outputs are more relevant to the prefix and improve continuity and coherence compared to baselines.We release our model checkpoints, code, and human preference data with explanations to facilitate future research.1 Kalpesh Krishna, Yapei Chang, John Wieting, Mohit Iyyer |
EMNLP | 3 |
| 2022 | Exploring Document-Level Literary Machine Translation with Parallel Paragraphs from World LiteratureabstractLiterary translation is a culturally significant task, but it is bottlenecked by the small number of qualified literary translators relative to the many untranslated works published around the world.Machine translation (MT) holds potential to complement the work of human translators by improving both training procedures and their overall efficiency.Literary translation is less constrained than more traditional MT settings since translators must balance meaning equivalence, readability, and critical interpretability in the target language.This property, along with the complex discourse-level context present in literary texts, also makes literary MT more challenging to computationally model and evaluate.To explore this task, we collect a dataset (PAR3) of non-English language novels in the public domain, each aligned at the paragraph level to both human and automatic English translations.Using PAR3, we discover that expert literary translators prefer reference human translations over machinetranslated paragraphs at a rate of 84%, while state-of-the-art automatic MT metrics do not correlate with those preferences.The experts note that MT outputs contain not only mistranslations, but also discourse-disrupting errors and stylistic inconsistencies.To address these problems, we train a post-editing model whose output is preferred over normal MT output at a rate of 69% by experts.We publicly release PAR3 to spur future research into literary MT. 1 Katherine Thai, Marzena Karpinska, Kalpesh Krishna, Bill Ray, Moira Inghilleri, John Wieting, Mohit Iyyer |
EMNLP | 6 |
| 2022 | Canine: Pre-training an Efficient Tokenization-Free Encoder for Language RepresentationabstractAbstract Pipelined NLP systems have largely been superseded by end-to-end neural modeling, yet nearly all commonly used models still require an explicit tokenization step. While recent tokenization approaches based on data-derived subword lexicons are less brittle than manually engineered tokenizers, these techniques are not equally suited to all languages, and the use of any fixed vocabulary may limit a model’s ability to adapt. In this paper, we present Canine, a neural encoder that operates directly on character sequences—without explicit tokenization or vocabulary—and a pre-training strategy that operates either directly on characters or optionally uses subwords as a soft inductive bias. To use its finer-grained input effectively and efficiently, Canine combines downsampling, which reduces the input sequence length, with a deep transformer stack, which encodes context. Canine outperforms a comparable mBert model by 5.7 F1 on TyDi QA, a challenging multilingual benchmark, despite having fewer model parameters. Jonathan H. Clark, Dan Garrette, Iulia Turc, John Wieting |
Trans. Assoc. Comput. Linguistics | 4 |
| 2021 | On Learning Text Style Transfer with Direct RewardsabstractIn most cases, the lack of parallel corpora makes it impossible to directly train supervised models for the text style transfer task.In this paper, we explore training algorithms that instead optimize reward functions that explicitly consider different aspects of the styletransferred outputs.In particular, we leverage semantic similarity metrics originally used for fine-tuning neural machine translation models to explicitly assess the preservation of content between system outputs and input texts.We also investigate the potential weaknesses of the existing automatic metrics and propose efficient strategies of using these metrics for training.The experimental results show that our model provides significant gains in both automatic and human evaluation over strong baselines, indicating the effectiveness of our proposed methods and training strategies. 1 Yixin Liu 0003, Graham Neubig, John Wieting |
NAACL-HLT | 3 |
| 2020 | Reformulating Unsupervised Style Transfer as Paraphrase GenerationabstractModern NLP defines the task of style transfer as modifying the style of a given sentence without appreciably changing its semantics, which implies that the outputs of style transfer systems should be paraphrases of their inputs.However, many existing systems purportedly designed for style transfer inherently warp the input's meaning through attribute transfer, which changes semantic properties such as sentiment.In this paper, we reformulate unsupervised style transfer as a paraphrase generation problem, and present a simple methodology based on fine-tuning pretrained language models on automatically generated paraphrase data.Despite its simplicity, our method significantly outperforms state-of-the-art style transfer systems on both human and automatic evaluations.We also survey 23 style transfer papers and discover that existing automatic metrics can be easily gamed and propose fixed variants.Finally, we pivot to a more real-world style transfer setting by collecting a large dataset of 15M sentences in 11 diverse styles, which we use for an in-depth analysis of our system. Kalpesh Krishna, John Wieting, Mohit Iyyer |
EMNLP (1) | 2 |
| 2020 | A Bilingual Generative Transformer for Semantic Sentence EmbeddingabstractSemantic sentence embedding models encode natural language sentences into vectors, such that closeness in embedding space indicates closeness in the semantics between the sentences.Bilingual data offers a useful signal for learning such embeddings: properties shared by both sentences in a translation pair are likely semantic, while divergent properties are likely stylistic or language-specific.We propose a deep latent variable model that attempts to perform source separation on parallel sentences, isolating what they have in common in a latent semantic vector, and explaining what is left over with language-specific latent vectors.Our proposed approach differs from past work on semantic sentence encoding in two ways.First, by using a variational probabilistic framework, we introduce priors that encourage source separation, and can use our model's posterior to predict sentence embeddings for monolingual data at test time.Second, we use high-capacity transformers as both data generating distributions and inference networkscontrasting with most past work on sentence embeddings.In experiments, our approach substantially outperforms the state-of-the-art on a standard suite of unsupervised semantic similarity evaluations.Further, we demonstrate that our approach yields the largest gains on more difficult subsets of these evaluations where simple word overlap is not a good indicator of similarity. 1 John Wieting, Graham Neubig, Taylor Berg-Kirkpatrick |
EMNLP (1) | 1 |
| 2020 | Improving Candidate Generation for Low-resource Cross-lingual Entity LinkingabstractCross-lingual entity linking (XEL) is the task of finding referents in a target-language knowledge base (KB) for mentions extracted from source-language texts. The first step of (X)EL is candidate generation, which retrieves a list of plausible candidate entities from the target-language KB for each mention. Approaches based on resources from Wikipedia have proven successful in the realm of relatively high-resource languages, but these do not extend well to low-resource languages with few, if any, Wikipedia pages. Recently, transfer learning methods have been shown to reduce the demand for resources in the low-resource languages by utilizing resources in closely related languages, but the performance still lags far behind their high-resource counterparts. In this paper, we first assess the problems faced by current entity candidate generation methods for low-resource XEL, then propose three improvements that (1) reduce the disconnect between entity mentions and KB entries, and (2) improve the robustness of the model to low-resource scenarios. The methods are simple, but effective: We experiment with our approach on seven XEL datasets and find that they yield an average gain of 16.9% in Top-30 gold candidate recall, compared with state-of-the-art baselines. Our improved model also yields an average gain of 7.9% in in-KB accuracy of end-to-end XEL. 1 Shuyan Zhou, Shruti Rijhwani, John Wieting, Jaime G. Carbonell, Graham Neubig |
Trans. Assoc. Comput. Linguistics | 3 |
| 2019 | Beyond BLEU: Training Neural Machine Translation with Semantic SimilarityabstractWhile most neural machine translation (NMT) systems are still trained using maximum likelihood estimation, recent work has demonstrated that optimizing systems to directly improve evaluation metrics such as BLEU can substantially improve final translation accuracy.However, training with BLEU has some limitations: it doesn't assign partial credit, it has a limited range of output values, and it can penalize semantically correct hypotheses if they differ lexically from the reference.In this paper, we introduce an alternative reward function for optimizing NMT systems that is based on recent work in semantic similarity.We evaluate on four disparate languages translated to English, and find that training with our proposed metric results in better translations as evaluated by BLEU, semantic similarity, and human evaluation, and also that the optimization procedure converges faster.Analysis suggests that this is because the proposed metric is more conducive to optimization, assigning partial credit and providing more diversity in scores than BLEU. 1 John Wieting, Taylor Berg-Kirkpatrick, Kevin Gimpel, Graham Neubig |
ACL (1) | 1 |
| 2019 | Simple and Effective Paraphrastic Similarity from Parallel TranslationsabstractWe present a model and methodology for learning paraphrastic sentence embeddings directly from bitext, removing the timeconsuming intermediate step of creating paraphrase corpora.Further, we show that the resulting model can be applied to cross-lingual tasks where it both outperforms and is orders of magnitude faster than more complex stateof-the-art baselines.1 John Wieting, Kevin Gimpel, Graham Neubig, Taylor Berg-Kirkpatrick |
ACL (1) | 1 |
| 2019 | No Training Required: Exploring Random Encoders for Sentence Classification
John Wieting, Douwe Kiela |
ICLR (Poster) | 1 |
| 2018 | ParaNMT-50M: Pushing the Limits of Paraphrastic Sentence Embeddings with Millions of Machine TranslationsabstractWe describe PARANMT-50M, a dataset of more than 50 million English-English sentential paraphrase pairs.We generated the pairs automatically by using neural machine translation to translate the non-English side of a large parallel corpus, following Wieting et al. (2017).Our hope is that PARANMT-50M can be a valuable resource for paraphrase generation and can provide a rich source of semantic knowledge to improve downstream natural language understanding tasks.To show its utility, we use PARANMT-50M to train paraphrastic sentence embeddings that outperform all supervised systems on every SemEval semantic textual similarity competition, in addition to showing how it can be used for paraphrase generation. 1 John Wieting, Kevin Gimpel |
ACL (1) | 1 |
| 2018 | CogCompNLP: Your Swiss Army Knife for NLP
Daniel Khashabi, Mark Sammons, Ben Zhou, Tom Redman, Christos Christodoulopoulos 0001, Vivek Srikumar, Nick Rizzolo, Lev-Arie Ratinov, Guanheng Luo, Quang Do, Chen-Tse Tsai, Subhro Roy, Stephen Mayhew 0001, Zhili Feng, John Wieting, Xiaodong Yu 0003, Yangqiu Song, Shashank Gupta 0007, Shyam Upadhyay, Naveen Arivazhagan, Qiang Ning, Shaoshi Ling, Dan Roth 0001 |
LREC | 15 |
| 2018 | Adversarial Example Generation with Syntactically Controlled Paraphrase NetworksabstractMohit Iyyer, John Wieting, Kevin Gimpel, Luke Zettlemoyer. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Mohit Iyyer, John Wieting, Kevin Gimpel, Luke Zettlemoyer |
NAACL-HLT | 2 |
| 2017 | Revisiting Recurrent Networks for Paraphrastic Sentence EmbeddingsabstractWe consider the problem of learning general-purpose, paraphrastic sentence embeddings, revisiting the setting of Wieting et al. (2016b).While they found LSTM recurrent networks to underperform word averaging, we present several developments that together produce the opposite conclusion.These include training on sentence pairs rather than phrase pairs, averaging states to represent sequences, and regularizing aggressively.These improve LSTMs in both transfer learning and supervised settings.We also introduce a new recurrent architecture, the GATED RECURRENT AVER-AGING NETWORK, that is inspired by averaging and LSTMs while outperforming them both.We analyze our learned models, finding evidence of preferences for particular parts of speech and dependency relations.1 John Wieting, Kevin Gimpel |
ACL (1) | 1 |
| 2017 | Learning Paraphrastic Sentence Embeddings from Back-Translated BitextabstractWe consider the problem of learning general-purpose, paraphrastic sentence embeddings in the setting of Wieting et al. (2016b).We use neural machine translation to generate sentential paraphrases via back-translation of bilingual sentence pairs.We evaluate the paraphrase pairs by their ability to serve as training data for learning paraphrastic sentence embeddings.We find that the data quality is stronger than prior work based on bitext and on par with manually-written English paraphrase pairs, with the advantage that our approach can scale up to generate large training sets for many languages and domains.We experiment with several language pairs and data sources, and develop a variety of data filtering techniques.In the process, we explore how neural machine translation output differs from humanwritten sentences, finding clear differences in length, the amount of repetition, and the use of rare words. 1 1 Generated paraphrases and code are available at http: //ttic.uchicago.edu/˜wieting.R: We understand that has already commenced, but there is a long way to go.T: This situation has already commenced, but much still needs to be done.R: The restaurant is closed on Sundays.No breakfast is available on Sunday mornings John Wieting, Jonathan Mallinson, Kevin Gimpel |
EMNLP | 1 |
| 2016 | Charagram: Embedding Words and Sentences via Character n-gramsabstractWe present CHARAGRAM embeddings, a simple approach for learning character-based compositional models to embed textual sequences.A word or sentence is represented using a character n-gram count vector, followed by a single nonlinear transformation to yield a low-dimensional embedding.We use three tasks for evaluation: word similarity, sentence similarity, and part-of-speech tagging.We demonstrate that CHARAGRAM embeddings outperform more complex architectures based on character-level recurrent and convolutional neural networks, achieving new state-of-the-art performance on several similarity tasks. 1 John Wieting, Mohit Bansal, Kevin Gimpel, Karen Livescu |
EMNLP | 1 |
| 2015 | From Paraphrase Database to Compositional Paraphrase Model and BackabstractThe Paraphrase Database (PPDB; Ganitkevitch et al., 2013) is an extensive semantic resource, consisting of a list of phrase pairs with (heuristic) confidence estimates. However, it is still unclear how it can best be used, due to the heuristic nature of the confidences and its necessarily incomplete coverage. We propose models to leverage the phrase pairs from the PPDB to build parametric paraphrase models that score paraphrase pairs more accurately than the PPDB’s internal scores while simultaneously improving its coverage. They allow for learning phrase embeddings as well as improved word embeddings. Moreover, we introduce two new, manually annotated datasets to evaluate short-phrase paraphrasing models. Using our paraphrase model trained using PPDB, we achieve state-of-the-art results on standard word and bigram similarity tasks and beat strong baselines on our new short phrase paraphrase tasks. John Wieting, Mohit Bansal, Kevin Gimpel, Karen Livescu |
Trans. Assoc. Comput. Linguistics | 1 |