EDBT 2026 Demo / reviewers in the wild / expert
Kartik Goyal
dblp:136/8676
· DBLP profile ↗
13ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Language models and text generation · 61% Generative modeling · 23% Representation and self-supervised learning · 10% | |
| Computer graphics and multimedia
2 papers |
Multimedia analysis and retrieval · 89% Visualization and visual analytics · 11% | |
| Network and information security
2 papers |
Security and privacy of machine learning · 83% Privacy and data protection · 17% |
Topics — the 16 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
energy-based model |
1.1 | 2 | 2022 | Exposing the Implicit Energy Networks behind Masked Language Models via Metropolis--Hastings · ICLR 2022 Mix and Match: Learning-free Controllable Text Generationusing Energy Language Models · ACL (1) 2022 |
Multimedia analysis and retrieval › image analysis › image understanding
historical document analysis |
1.1 | 2 | 2023 | Contrastive Attention Networks for Attribution of Early Modern Print · AAAI 2023 A Probabilistic Generative Model for Typographical Analysis of Early Modern Printing · ACL 2020 |
Natural language and speech › Language models and text generation
tokenization |
0.9 | 1 | 2025 | Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language Models · EMNLP 2025 |
Natural language and speech › Language models and text generation
decoding |
0.8 | 1 | 2024 | MAP's not dead yet: Uncovering true language model modes by conditioning away degeneracy · ACL (1) 2024 |
Machine learning › Representation and self-supervised learning › representation learning
metric learning |
0.7 | 1 | 2023 | Contrastive Attention Networks for Attribution of Early Modern Print · AAAI 2023 |
Natural language and speech › Language models and text generation
controllable text generation |
0.6 | 1 | 2022 | Mix and Match: Learning-free Controllable Text Generationusing Energy Language Models · ACL (1) 2022 |
Natural language and speech › Language models and text generation
masked language modeling |
0.6 | 1 | 2022 | Exposing the Implicit Energy Networks behind Masked Language Models via Metropolis--Hastings · ICLR 2022 |
Security and privacy of machine learning
membership inference |
0.6 | 1 | 2022 | Quantifying Privacy Risks of Masked Language Models Using Membership Inference Attacks · EMNLP 2022 |
Machine learning › Generative modeling › generative model
probabilistic generative model |
0.4 | 1 | 2020 | A Probabilistic Generative Model for Typographical Analysis of Early Modern Printing · ACL 2020 |
Natural language and speech › Language models and text generation › decoding › decoding strategy
beam search |
0.3 | 1 | 2018 | A Continuous Relaxation of Beam Search for End-to-End Training of Neural Sequence Models · AAAI 2018 |
Natural language and speech › Language models and text generation › decoding
sequence decoding |
0.3 | 1 | 2018 | A Continuous Relaxation of Beam Search for End-to-End Training of Neural Sequence Models · AAAI 2018 |
Machine learning › Deep learning architectures and training
training objective |
0.3 | 1 | 2018 | A Continuous Relaxation of Beam Search for End-to-End Training of Neural Sequence Models · AAAI 2018 |
Security and privacy of machine learning › privacy attack
training data extraction |
0.3 | 1 | 2025 | Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language Models · EMNLP 2025 |
Natural language and speech › Language models and text generation
prompting |
0.2 | 1 | 2022 | Mix and Match: Learning-free Controllable Text Generationusing Energy Language Models · ACL (1) 2022 |
Privacy and data protection › privacy risk assessment
privacy risk estimation |
0.2 | 1 | 2022 | Quantifying Privacy Risks of Masked Language Models Using Membership Inference Attacks · EMNLP 2022 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.1 | 1 | 2018 | A Continuous Relaxation of Beam Search for End-to-End Training of Neural Sequence Models · AAAI 2018 |
Methods — techniques the papers use, named apart from their topics
data synthesis · 1.3contrastive attention · 1.3likelihood ratio hypothesis testing · 1.1exact search · 0.8classifier-based approximate search · 0.8reference models · 0.6reference model · 0.6metropolis-hastings sampling · 0.6metropolis-hastings · 0.6energy-based model · 0.6black-box model combination · 0.6variational autoencoder · 0.4neural editor model · 0.4latent variable model · 0.4sub-differentiable surrogate · 0.3continuous relaxation · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language ModelsabstractStandard Byte-Pair Encoding (BPE) tokenization compresses text by pairing a learned token vocabulary with a detailed merge list.Recent work has shown that this merge list exposes a potential attack surface for extracting information about language model's training data.In this paper, we explore the downstream impact of BPE inference algorithms that do not rely on this merge list at all, and hence differ from the encoding process during BPE training.To address this question, we investigate two broad classes of BPE inference schemes that differ from BPE application during training: a) targeted deviation from merge-lists including random merge orders, and various corruptions of merge list involving deletion/truncation, and b) non-targeted BPE inference algorithms that do not depend on the merge list but focus on compressing the text either greedily or exactly.Extensive experiments across diverse language modeling tasks like accuracy-based QA benchmarks, machine translation, and open-ended generation reveal that while targeted deviation from the merge lists exhibits significant degradation in language model performance, the nontargeted merge-list-free inference algorithms result in minimal impact on downstream performance that is often much smaller than expected.These findings pave way for simpler and potentially more privacy-preserving tokenization schemes that do not catastrophically compromise model performance. Tomohiro Sawada, Kartik Goyal |
EMNLP | 2 |
| 2024 | MAP's not dead yet: Uncovering true language model modes by conditioning away degeneracyabstractIt has been widely observed that exact or approximate MAP (mode-seeking) decoding from natural language generation (NLG) models consistently leads to degenerate outputs (Holtzman et al., 2019;Stahlberg and Byrne, 2019).Prior work has attributed this behavior to either a fundamental and unavoidable inadequacy of modes in probabilistic models or weaknesses in language modeling.Contrastingly, we argue that degenerate modes can even occur in the absence of any modeling error, due to contamination of the training data.Specifically, we argue that mixing even a tiny amount of low-entropy noise with a population text distribution can cause the data distribution's mode to become degenerate.We therefore propose to apply MAP decoding to the model's true conditional distribution where the conditioning variable explicitly avoids specific degenerate behavior.Using exact search, we empirically verify that the length-conditional modes of machine translation models and language models are indeed more fluent and topical than their unconditional modes.For the first time, we also share many examples of exact modal sequences from these models, and from several variants of the LLaMA-7B model.Notably, we observe that various kinds of degenerate modes persist, even at the scale of LLaMA-7B.Although we cannot tractably address these degeneracies with exact search, we perform a classifier-based approximate search on LLaMA-7B, a model which was not trained for instruction following, and find that we are able to elicit reasonable outputs without any finetuning. Davis Yoshida, Kartik Goyal, Kevin Gimpel |
ACL (1) | 2 |
| 2024 | Clustering Running Titles to Understand the Printing of Early Modern Books
Nikolai Vogler, Kartik Goyal, Samuel V. Lemley, D. J. Schuldt, Christopher N. Warren, Max G'Sell, Taylor Berg-Kirkpatrick |
ICDAR (3) | 2 |
| 2023 | Contrastive Attention Networks for Attribution of Early Modern PrintabstractIn this paper, we develop machine learning techniques to identify unknown printers in early modern (c.~1500--1800) English printed books. Specifically, we focus on matching uniquely damaged character type-imprints in anonymously printed books to works with known printers in order to provide evidence of their origins. Until now, this work has been limited to manual investigations by analytical bibliographers. We present a Contrastive Attention-based Metric Learning approach to identify similar damage across character image pairs, which is sensitive to very subtle differences in glyph shapes, yet robust to various confounding sources of noise associated with digitized historical books. To overcome the scarce amount of supervised data, we design a random data synthesis procedure that aims to simulate bends, fractures, and inking variations induced by the early printing process. Our method successfully improves downstream damaged type-imprint matching among printed works from this period, as validated by in-domain human experts. The results of our approach on two important philosophical works from the Early Modern period demonstrate potential to extend the extant historical research about the origins and content of these books. Nikolai Vogler, Kartik Goyal, Kishore PV Reddy, Elizaveta Pertseva, Samuel V. Lemley, Christopher N. Warren, Max G'Sell, Taylor Berg-Kirkpatrick |
AAAI | 2 |
| 2022 | Mix and Match: Learning-free Controllable Text Generationusing Energy Language ModelsabstractRecent work on controlled text generation has either required attribute-based fine-tuning of the base language model (LM), or has restricted the parameterization of the attribute discriminator to be compatible with the base autoregressive LM.In this work, we propose Mix and Match LM, a global score-based alternative for controllable text generation that combines arbitrary pre-trained black-box models for achieving the desired attributes in the generated text without involving any fine-tuning or structural assumptions about the black-box models.We interpret the task of controllable generation as drawing samples from an energy-based model whose energy values are a linear combination of scores from black-box models that are separately responsible for fluency, the control attribute, and faithfulness to any conditioning context.We use a Metropolis-Hastings sampling scheme to sample from this energy-based model using bidirectional context and global attribute features.We validate the effectiveness of our approach on various controlled generation and style-based text revision tasks by outperforming recently proposed methods that involve extra training, fine-tuning, or restrictive assumptions over the form of models. Niloofar Mireshghallah, Kartik Goyal, Taylor Berg-Kirkpatrick |
ACL (1) | 2 |
| 2022 | Quantifying Privacy Risks of Masked Language Models Using Membership Inference AttacksabstractThe wide adoption and application of Masked language models (MLMs) on sensitive data (from legal to medical) necessitates a thorough quantitative investigation into their privacy vulnerabilities.Prior attempts at measuring leakage of MLMs via membership inference attacks have been inconclusive, implying potential robustness of MLMs to privacy attacks.In this work, we posit that prior attempts were inconclusive because they based their attack solely on the MLM's model score.We devise a stronger membership inference attack based on likelihood ratio hypothesis testing that involves an additional reference MLM to more accurately quantify the privacy risks of memorization in MLMs.We show that masked language models are indeed susceptible to likelihood ratio membership inference attacks: Our empirical results, on models trained on medical notes, show that our attack improves the AUC of prior membership inference attacks from 0.66 to an alarmingly high 0.90 level. Niloofar Mireshghallah, Kartik Goyal, Archit Uniyal, Taylor Berg-Kirkpatrick, Reza Shokri |
EMNLP | 2 |
| 2022 | Exposing the Implicit Energy Networks behind Masked Language Models via Metropolis--Hastings
Kartik Goyal, Chris Dyer, Taylor Berg-Kirkpatrick |
ICLR | 1 |
| 2020 | A Probabilistic Generative Model for Typographical Analysis of Early Modern PrintingabstractWe propose a deep and interpretable probabilistic generative model to analyze glyph shapes in printed Early Modern documents.We focus on clustering extracted glyph images into underlying templates in the presence of multiple confounding sources of variance.Our approach introduces a neural editor model that first generates well-understood printing phenomena like spatial perturbations from template parameters via interpertable latent variables, and then modifies the result by generating a non-interpretable latent vector responsible for inking variations, jitter, noise from the archiving process, and other unforeseen phenomena associated with Early Modern printing.Critically, by introducing an inference network whose input is restricted to the visual residual between the observation and the interpretably-modified template, we are able to control and isolate what the vector-valued latent variable captures.We show that our approach outperforms rigid interpretable clustering baselines (Ocular) and overly-flexible deep generative models (VAE) alike on the task of completely unsupervised discovery of typefaces in mixed-font documents. Kartik Goyal, Chris Dyer, Christopher N. Warren, Max G'Sell, Taylor Berg-Kirkpatrick |
ACL | 1 |
| 2018 | A Continuous Relaxation of Beam Search for End-to-End Training of Neural Sequence ModelsabstractBeam search is a desirable choice of test-time decoding algorithm for neural sequence models because it potentially avoids search errors made by simpler greedy methods. However, typical cross entropy training procedures for these models do not directly consider the behaviour of the final decoding method. As a result, for cross-entropy trained models, beam decoding can sometimes yield reduced test performance when compared with greedy decoding. In order to train models that can more effectively make use of beam search, we propose a new training procedure that focuses on the final loss metric (e.g. Hamming loss) evaluated on the output of beam search. While well-defined, this "direct loss" objective is itself discontinuous and thus difficult to optimize. Hence, in our approach, we form a sub-differentiable surrogate objective by introducing a novel continuous approximation of the beam search decoding procedure.In experiments, we show that optimizing this new training objective yields substantially better results on two sequence tasks (Named Entity Recognition and CCG Supertagging) when compared with both cross entropy trained greedy decoding and cross entropy trained beam decoding baselines. Kartik Goyal, Graham Neubig, Chris Dyer, Taylor Berg-Kirkpatrick |
AAAI | 1 |
| 2016 | Named Entity Recognition for Linguistic Rapid Response in Low-Resource Languages: Sorani Kurdish and TajikabstractThis paper describes our construction of named-entity recognition (NER) systems in two Western Iranian languages, Sorani Kurdish and Tajik, as a part of a pilot study of “Linguistic Rapid Response” to potential emergency humanitarian relief situations. In the absence of large annotated corpora, parallel corpora, treebanks, bilingual lexica, etc., we found the following to be effective: exploiting distributional regularities in monolingual data, projecting information across closely related languages, and utilizing human linguist judgments. We show promising results on both a four-month exercise in Sorani and a two-day exercise in Tajik, achieved with minimal annotation costs. Patrick Littell, Kartik Goyal, David R. Mortensen, Alexa Little, Chris Dyer, Lori S. Levin |
COLING | 2 |
| 2016 | PanPhon: A Resource for Mapping IPA Segments to Articulatory Feature VectorsabstractThis paper contributes to a growing body of evidence that—when coupled with appropriate machine-learning techniques–linguistically motivated, information-rich representations can outperform one-hot encodings of linguistic data. In particular, we show that phonological features outperform character-based models. PanPhon is a database relating over 5,000 IPA segments to 21 subsegmental articulatory features. We show that this database boosts performance in various NER-related tasks. Phonologically aware, neural CRF models built on PanPhon features are able to perform better on monolingual Spanish and Turkish NER tasks that character-based models. They have also been shown to work well in transfer models (as between Uzbek and Turkish). PanPhon features also contribute measurably to Orthography-to-IPA conversion tasks. David R. Mortensen, Patrick Littell, Akash Bharadwaj, Kartik Goyal, Chris Dyer, Lori S. Levin |
COLING | 4 |
| 2016 | Bridge-Language Capitalization Inference in Western Iranian: Sorani, Kurmanji, Zazaki, and Tajik
Patrick Littell, David R. Mortensen, Kartik Goyal, Chris Dyer, Lori S. Levin |
LREC | 3 |
| 2014 | Unsupervised Word Sense Induction using Distributional Statistics
Kartik Goyal, Eduard H. Hovy |
COLING | 1 |