Kartik Goyal

dblp:136/8676 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Language models and text generation · 61% Generative modeling · 23% Representation and self-supervised learning · 10%
Computer graphics and multimedia
2 papers
Multimedia analysis and retrieval · 89% Visualization and visual analytics · 11%
Network and information security
2 papers
Security and privacy of machine learning · 83% Privacy and data protection · 17%

Topics — the 16 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
energy-based model
1.122022
Exposing the Implicit Energy Networks behind Masked Language Models via Metropolis--Hastings · ICLR 2022
Mix and Match: Learning-free Controllable Text Generationusing Energy Language Models · ACL (1) 2022
Multimedia analysis and retrieval › image analysis › image understanding
historical document analysis
1.122023
Contrastive Attention Networks for Attribution of Early Modern Print · AAAI 2023
A Probabilistic Generative Model for Typographical Analysis of Early Modern Printing · ACL 2020
Natural language and speech › Language models and text generation
tokenization
0.912025
Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language Models · EMNLP 2025
Natural language and speech › Language models and text generation
decoding
0.812024
MAP's not dead yet: Uncovering true language model modes by conditioning away degeneracy · ACL (1) 2024
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.712023
Contrastive Attention Networks for Attribution of Early Modern Print · AAAI 2023
Natural language and speech › Language models and text generation
controllable text generation
0.612022
Mix and Match: Learning-free Controllable Text Generationusing Energy Language Models · ACL (1) 2022
Natural language and speech › Language models and text generation
masked language modeling
0.612022
Exposing the Implicit Energy Networks behind Masked Language Models via Metropolis--Hastings · ICLR 2022
Security and privacy of machine learning
membership inference
0.612022
Quantifying Privacy Risks of Masked Language Models Using Membership Inference Attacks · EMNLP 2022
Machine learning › Generative modeling › generative model
probabilistic generative model
0.412020
A Probabilistic Generative Model for Typographical Analysis of Early Modern Printing · ACL 2020
Natural language and speech › Language models and text generation › decoding › decoding strategy
beam search
0.312018
A Continuous Relaxation of Beam Search for End-to-End Training of Neural Sequence Models · AAAI 2018
Natural language and speech › Language models and text generation › decoding
sequence decoding
0.312018
A Continuous Relaxation of Beam Search for End-to-End Training of Neural Sequence Models · AAAI 2018
Machine learning › Deep learning architectures and training
training objective
0.312018
A Continuous Relaxation of Beam Search for End-to-End Training of Neural Sequence Models · AAAI 2018
Security and privacy of machine learning › privacy attack
training data extraction
0.312025
Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language Models · EMNLP 2025
Natural language and speech › Language models and text generation
prompting
0.212022
Mix and Match: Learning-free Controllable Text Generationusing Energy Language Models · ACL (1) 2022
Privacy and data protection › privacy risk assessment
privacy risk estimation
0.212022
Quantifying Privacy Risks of Masked Language Models Using Membership Inference Attacks · EMNLP 2022
Natural language and speech › Information extraction and text analysis
named entity recognition
0.112018
A Continuous Relaxation of Beam Search for End-to-End Training of Neural Sequence Models · AAAI 2018

Methods — techniques the papers use, named apart from their topics

data synthesis · 1.3contrastive attention · 1.3likelihood ratio hypothesis testing · 1.1exact search · 0.8classifier-based approximate search · 0.8reference models · 0.6reference model · 0.6metropolis-hastings sampling · 0.6metropolis-hastings · 0.6energy-based model · 0.6black-box model combination · 0.6variational autoencoder · 0.4neural editor model · 0.4latent variable model · 0.4sub-differentiable surrogate · 0.3continuous relaxation · 0.3
YearPublicationVenuePosition
2025 Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language Models
abstract
Standard Byte-Pair Encoding (BPE) tokenization compresses text by pairing a learned token vocabulary with a detailed merge list.Recent work has shown that this merge list exposes a potential attack surface for extracting information about language model's training data.In this paper, we explore the downstream impact of BPE inference algorithms that do not rely on this merge list at all, and hence differ from the encoding process during BPE training.To address this question, we investigate two broad classes of BPE inference schemes that differ from BPE application during training: a) targeted deviation from merge-lists including random merge orders, and various corruptions of merge list involving deletion/truncation, and b) non-targeted BPE inference algorithms that do not depend on the merge list but focus on compressing the text either greedily or exactly.Extensive experiments across diverse language modeling tasks like accuracy-based QA benchmarks, machine translation, and open-ended generation reveal that while targeted deviation from the merge lists exhibits significant degradation in language model performance, the nontargeted merge-list-free inference algorithms result in minimal impact on downstream performance that is often much smaller than expected.These findings pave way for simpler and potentially more privacy-preserving tokenization schemes that do not catastrophically compromise model performance.
Tomohiro Sawada, Kartik Goyal
EMNLP2
2024 MAP's not dead yet: Uncovering true language model modes by conditioning away degeneracy
abstract
It has been widely observed that exact or approximate MAP (mode-seeking) decoding from natural language generation (NLG) models consistently leads to degenerate outputs (Holtzman et al., 2019;Stahlberg and Byrne, 2019).Prior work has attributed this behavior to either a fundamental and unavoidable inadequacy of modes in probabilistic models or weaknesses in language modeling.Contrastingly, we argue that degenerate modes can even occur in the absence of any modeling error, due to contamination of the training data.Specifically, we argue that mixing even a tiny amount of low-entropy noise with a population text distribution can cause the data distribution's mode to become degenerate.We therefore propose to apply MAP decoding to the model's true conditional distribution where the conditioning variable explicitly avoids specific degenerate behavior.Using exact search, we empirically verify that the length-conditional modes of machine translation models and language models are indeed more fluent and topical than their unconditional modes.For the first time, we also share many examples of exact modal sequences from these models, and from several variants of the LLaMA-7B model.Notably, we observe that various kinds of degenerate modes persist, even at the scale of LLaMA-7B.Although we cannot tractably address these degeneracies with exact search, we perform a classifier-based approximate search on LLaMA-7B, a model which was not trained for instruction following, and find that we are able to elicit reasonable outputs without any finetuning.
Davis Yoshida, Kartik Goyal, Kevin Gimpel
ACL (1)2
2024 Clustering Running Titles to Understand the Printing of Early Modern Books
Nikolai Vogler, Kartik Goyal, Samuel V. Lemley, D. J. Schuldt, Christopher N. Warren, Max G'Sell, Taylor Berg-Kirkpatrick
ICDAR (3)2
2023 Contrastive Attention Networks for Attribution of Early Modern Print
abstract
In this paper, we develop machine learning techniques to identify unknown printers in early modern (c.~1500--1800) English printed books. Specifically, we focus on matching uniquely damaged character type-imprints in anonymously printed books to works with known printers in order to provide evidence of their origins. Until now, this work has been limited to manual investigations by analytical bibliographers. We present a Contrastive Attention-based Metric Learning approach to identify similar damage across character image pairs, which is sensitive to very subtle differences in glyph shapes, yet robust to various confounding sources of noise associated with digitized historical books. To overcome the scarce amount of supervised data, we design a random data synthesis procedure that aims to simulate bends, fractures, and inking variations induced by the early printing process. Our method successfully improves downstream damaged type-imprint matching among printed works from this period, as validated by in-domain human experts. The results of our approach on two important philosophical works from the Early Modern period demonstrate potential to extend the extant historical research about the origins and content of these books.
Nikolai Vogler, Kartik Goyal, Kishore PV Reddy, Elizaveta Pertseva, Samuel V. Lemley, Christopher N. Warren, Max G'Sell, Taylor Berg-Kirkpatrick
AAAI2
2022 Mix and Match: Learning-free Controllable Text Generationusing Energy Language Models
abstract
Recent work on controlled text generation has either required attribute-based fine-tuning of the base language model (LM), or has restricted the parameterization of the attribute discriminator to be compatible with the base autoregressive LM.In this work, we propose Mix and Match LM, a global score-based alternative for controllable text generation that combines arbitrary pre-trained black-box models for achieving the desired attributes in the generated text without involving any fine-tuning or structural assumptions about the black-box models.We interpret the task of controllable generation as drawing samples from an energy-based model whose energy values are a linear combination of scores from black-box models that are separately responsible for fluency, the control attribute, and faithfulness to any conditioning context.We use a Metropolis-Hastings sampling scheme to sample from this energy-based model using bidirectional context and global attribute features.We validate the effectiveness of our approach on various controlled generation and style-based text revision tasks by outperforming recently proposed methods that involve extra training, fine-tuning, or restrictive assumptions over the form of models.
Niloofar Mireshghallah, Kartik Goyal, Taylor Berg-Kirkpatrick
ACL (1)2
2022 Quantifying Privacy Risks of Masked Language Models Using Membership Inference Attacks
abstract
The wide adoption and application of Masked language models (MLMs) on sensitive data (from legal to medical) necessitates a thorough quantitative investigation into their privacy vulnerabilities.Prior attempts at measuring leakage of MLMs via membership inference attacks have been inconclusive, implying potential robustness of MLMs to privacy attacks.In this work, we posit that prior attempts were inconclusive because they based their attack solely on the MLM's model score.We devise a stronger membership inference attack based on likelihood ratio hypothesis testing that involves an additional reference MLM to more accurately quantify the privacy risks of memorization in MLMs.We show that masked language models are indeed susceptible to likelihood ratio membership inference attacks: Our empirical results, on models trained on medical notes, show that our attack improves the AUC of prior membership inference attacks from 0.66 to an alarmingly high 0.90 level.
Niloofar Mireshghallah, Kartik Goyal, Archit Uniyal, Taylor Berg-Kirkpatrick, Reza Shokri
EMNLP2
2022 Exposing the Implicit Energy Networks behind Masked Language Models via Metropolis--Hastings
Kartik Goyal, Chris Dyer, Taylor Berg-Kirkpatrick
ICLR1
2020 A Probabilistic Generative Model for Typographical Analysis of Early Modern Printing
abstract
We propose a deep and interpretable probabilistic generative model to analyze glyph shapes in printed Early Modern documents.We focus on clustering extracted glyph images into underlying templates in the presence of multiple confounding sources of variance.Our approach introduces a neural editor model that first generates well-understood printing phenomena like spatial perturbations from template parameters via interpertable latent variables, and then modifies the result by generating a non-interpretable latent vector responsible for inking variations, jitter, noise from the archiving process, and other unforeseen phenomena associated with Early Modern printing.Critically, by introducing an inference network whose input is restricted to the visual residual between the observation and the interpretably-modified template, we are able to control and isolate what the vector-valued latent variable captures.We show that our approach outperforms rigid interpretable clustering baselines (Ocular) and overly-flexible deep generative models (VAE) alike on the task of completely unsupervised discovery of typefaces in mixed-font documents.
Kartik Goyal, Chris Dyer, Christopher N. Warren, Max G'Sell, Taylor Berg-Kirkpatrick
ACL1
2018 A Continuous Relaxation of Beam Search for End-to-End Training of Neural Sequence Models
abstract
Beam search is a desirable choice of test-time decoding algorithm for neural sequence models because it potentially avoids search errors made by simpler greedy methods. However, typical cross entropy training procedures for these models do not directly consider the behaviour of the final decoding method. As a result, for cross-entropy trained models, beam decoding can sometimes yield reduced test performance when compared with greedy decoding. In order to train models that can more effectively make use of beam search, we propose a new training procedure that focuses on the final loss metric (e.g. Hamming loss) evaluated on the output of beam search. While well-defined, this "direct loss" objective is itself discontinuous and thus difficult to optimize. Hence, in our approach, we form a sub-differentiable surrogate objective by introducing a novel continuous approximation of the beam search decoding procedure.In experiments, we show that optimizing this new training objective yields substantially better results on two sequence tasks (Named Entity Recognition and CCG Supertagging) when compared with both cross entropy trained greedy decoding and cross entropy trained beam decoding baselines.
Kartik Goyal, Graham Neubig, Chris Dyer, Taylor Berg-Kirkpatrick
AAAI1
2016 Named Entity Recognition for Linguistic Rapid Response in Low-Resource Languages: Sorani Kurdish and Tajik
abstract
This paper describes our construction of named-entity recognition (NER) systems in two Western Iranian languages, Sorani Kurdish and Tajik, as a part of a pilot study of “Linguistic Rapid Response” to potential emergency humanitarian relief situations. In the absence of large annotated corpora, parallel corpora, treebanks, bilingual lexica, etc., we found the following to be effective: exploiting distributional regularities in monolingual data, projecting information across closely related languages, and utilizing human linguist judgments. We show promising results on both a four-month exercise in Sorani and a two-day exercise in Tajik, achieved with minimal annotation costs.
Patrick Littell, Kartik Goyal, David R. Mortensen, Alexa Little, Chris Dyer, Lori S. Levin
COLING2
2016 PanPhon: A Resource for Mapping IPA Segments to Articulatory Feature Vectors
abstract
This paper contributes to a growing body of evidence that—when coupled with appropriate machine-learning techniques–linguistically motivated, information-rich representations can outperform one-hot encodings of linguistic data. In particular, we show that phonological features outperform character-based models. PanPhon is a database relating over 5,000 IPA segments to 21 subsegmental articulatory features. We show that this database boosts performance in various NER-related tasks. Phonologically aware, neural CRF models built on PanPhon features are able to perform better on monolingual Spanish and Turkish NER tasks that character-based models. They have also been shown to work well in transfer models (as between Uzbek and Turkish). PanPhon features also contribute measurably to Orthography-to-IPA conversion tasks.
David R. Mortensen, Patrick Littell, Akash Bharadwaj, Kartik Goyal, Chris Dyer, Lori S. Levin
COLING4
2016 Bridge-Language Capitalization Inference in Western Iranian: Sorani, Kurmanji, Zazaki, and Tajik
Patrick Littell, David R. Mortensen, Kartik Goyal, Chris Dyer, Lori S. Levin
LREC3
2014 Unsupervised Word Sense Induction using Distributional Statistics
Kartik Goyal, Eduard H. Hovy
COLING1