Clara Meister

dblp:245/7485 · DBLP profile ↗
← Back
35ranked-venue papers
11as first author
30since 2021 · last 2026
0000-0002-3775-4426ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 11 first-author · 29 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
abstract
Negar Foroutan, Clara Meister, Debjit Paul, Joel Niklaus, Sina Ahmadi, Antoine Bosselut, Rico Sennrich. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Negar Foroutan Eghlidi, Clara Meister, Debjit Paul, Joel Niklaus, Sina Ahmadi, Antoine Bosselut, Rico Sennrich
ACL (1)2
2025 Causal Estimation of Tokenisation Bias
abstract
Modern language models are typically trained over subword sequences, but ultimately define probabilities over character-strings.Ideally, the choice of the tokeniser-which maps characterstrings to subwords-should not affect the probability assigned to the underlying characterstring; in practice, it does.We define this mismatch as tokenisation bias.In this work, we quantify one particular type of tokenisation bias: the effect of including or not a subword (e.g., ⟨hello⟩) in a tokeniser's vocabulary on the probability a trained model assigns to the corresponding characters (i.e., "hello").Estimating this effect is challenging because each model is trained with only one tokeniser.We address this by framing tokenisation bias as a causal effect and estimating it using the regression discontinuity design.Specifically, we exploit the fact that tokenisation algorithms rank subwords and add the first K to a tokeniser's vocabulary, where K is an arbitrary cutoff point.As such, we can estimate a causal effect by comparing similar subwords around this cutoff.Experimentally, we find that tokenisation consistently affects models' outputs across scales, vocabularies, and tokenisers.Notably, a subword's presence in a small model's vocabulary may increase its characters' probability by up to 17 times, highlighting tokenisation as a key design choice in language modelling.
Pietro Lesci, Clara Meister, Thomas Hofmann 0001, Andreas Vlachos 0001, Tiago Pimentel
ACL (1)2
2025 Information Theory and Cognitive Science
Noga Zaslavsky, Thomas A. Langlois, Nathaniel Imel, Clara Meister, Eleonora Gualdoni, Daniel Polani
CogSci4
2025 Uncertainty-Aware Decoding with Minimum Bayes Risk
abstract
Despite their outstanding performance in the majority of scenarios, contemporary language models still occasionally generate undesirable outputs, for example, hallucinated text. While such behaviors have previously been linked to uncertainty, there is a notable lack of methods that actively consider uncertainty during text generation. In this work, we show how Minimum Bayes Risk (MBR) decoding, which selects model generations according to an expected risk, can be generalized into a principled uncertainty-aware decoding method. In short, we account for model uncertainty during decoding by incorporating a posterior over model parameters into MBR’s computation of expected risk. We show that this modified expected risk is useful for both choosing outputs and deciding when to abstain from generation and can provide improvements without incurring overhead. We benchmark different methods for learning posteriors and show that performance improves with prediction diversity. We release our code publicly.
Nico Daheim, Clara Meister, Thomas Möllenhoff, Iryna Gurevych
ICLR2
2024 Causal Estimation of Memorisation Profiles
abstract
Understanding memorisation in language models has practical and societal implications, e.g., studying models' training dynamics or preventing copyright infringements.Prior work defines memorisation as the causal effect of training with an instance on the model's ability to predict that instance.This definition relies on a counterfactual: the ability to observe what would have happened had the model not seen that instance.Existing methods struggle to provide computationally efficient and accurate estimates of this counterfactual.Further, they often estimate memorisation for a model architecture rather than for a specific model instance.This paper fills an important gap in the literature, proposing a new, principled, and efficient method to estimate memorisation based on the difference-in-differences design from econometrics.Using this method, we characterise a model's memorisation profile-its memorisation trends across training-by only observing its behaviour on a small set of instances throughout training.In experiments with the Pythia model suite, we find that memorisation (i) is stronger and more persistent in larger models, (ii) is determined by data order and learning rate, and (iii) has stable trends across model sizes, thus making memorisation in larger models predictable from smaller ones. pietrolesci/memorisation-profiles
Pietro Lesci, Clara Meister, Thomas Hofmann 0001, Andreas Vlachos 0001, Tiago Pimentel
ACL (1)2
2024 Towards a Similarity-adjusted Surprisal Theory
abstract
Surprisal theory posits that the cognitive effort required to comprehend a word is determined by its contextual predictability, quantified as surprisal.Traditionally, surprisal theory treats words as distinct entities, overlooking any potential similarity between them.Giulianelli et al. (2023) address this limitation by introducing information value, a measure of predictability designed to account for similarities between communicative units.Our work leverages Ricotta and Szeidl's (2006) diversity index to extend surprisal into a metric that we term similarity-adjusted surprisal, exposing a mathematical relationship between surprisal and information value.Similarity-adjusted surprisal aligns with information value when considering graded similarities and reduces to standard surprisal when words are treated as distinct.Experimental results with reading time data indicate that similarity-adjusted surprisal adds predictive power beyond standard surprisal for certain datasets, suggesting it serves as a complementary measure of comprehension effort.
Clara Meister, Mario Giulianelli, Tiago Pimentel
EMNLP1
2024 How to Compute the Probability of a Word
abstract
Language models (LMs) estimate a probability distribution over strings in a natural language; these distributions are crucial for computing perplexity and surprisal in linguistics research.While we are usually concerned with measuring these values for words, most LMs operate over subwords.Despite seemingly straightforward, accurately computing probabilities over one unit given probabilities over the other requires care.Indeed, we show here that many recent linguistic studies have been incorrectly computing these values.This paper derives the correct methods for computing word probabilities, highlighting issues when relying on language models that use beginning-of-word (bow)-marking tokenisers, e.g., the GPT family.Empirically, we show that correcting the widespread bug in probability computations affects measured outcomes in sentence comprehension and lexical optimisation analyses.tpimentelms/probability-of-a-word pip install wordsprobability
Tiago Pimentel, Clara Meister
EMNLP2
2024 The Role of n-gram Smoothing in the Age of Neural Networks
abstract
Luca Malagutti, Andrius Buinovskij, Anej Svete, Clara Meister, Afra Amini, Ryan Cotterell. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Luca Malagutti, Andrius Buinovskij, Anej Svete, Clara Meister, Afra Amini, Ryan Cotterell
NAACL-HLT4
2023 A Measure-Theoretic Characterization of Tight Language Models
abstract
Li Du, Lucas Torroba Hennigen, Tiago Pimentel, Clara Meister, Jason Eisner, Ryan Cotterell. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Lucas Torroba Hennigen, Tiago Pimentel, Clara Meister, Jason Eisner, Ryan Cotterell
ACL (1)4
2023 On the Efficacy of Sampling Adapters
abstract
Sampling is a common strategy for generating text from probabilistic models, yet standard ancestral sampling often results in text that is incoherent or ungrammatical.To alleviate this issue, various modifications to a model's sampling distribution, such as nucleus or top-k sampling, have been introduced and are now ubiquitously used in language generation systems.We propose a unified framework for understanding these techniques, which we term sampling adapters.Sampling adapters often lead to qualitatively better text, which raises the question: From a formal perspective, how are they changing the (sub)word-level distributions of language generation models?And why do these local changes lead to higher-quality text?We argue that the shift they enforce can be viewed as a trade-off between precision and recall: while the model loses its ability to produce certain strings, its precision rate on desirable text increases.While this trade-off is not reflected in standard metrics of distribution quality (such as perplexity), we find that several precision-emphasizing measures indeed indicate that sampling adapters can lead to probability distributions more aligned with the true distribution.Further, these measures correlate with higher sequence-level quality scores, specifically, MAUVE.https://github.com/rycolab/ sampling-adapters
Clara Meister, Tiago Pimentel, Luca Malagutti, Ethan Wilcox, Ryan Cotterell
ACL (1)1
2023 Tokenization and the Noiseless Channel
abstract
Vilém Zouhar, Clara Meister, Juan Gastaldi, Li Du, Mrinmaya Sachan, Ryan Cotterell. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Vilém Zouhar, Clara Meister, Juan Luis Gastaldi, Mrinmaya Sachan, Ryan Cotterell
ACL (1)2
2023 Revisiting the Optimality of Word Lengths
abstract
Zipf (1935) posited that wordforms are optimized to minimize utterances' communicative costs.Under the assumption that cost is given by an utterance's length, he supported this claim by showing that words' lengths are inversely correlated with their frequencies.Communicative cost, however, can be operationalized in different ways.Piantadosi et al. (2011) claim that cost should be measured as the distance between an utterance's information rate and channel capacity, which we dub the channel capacity hypothesis (CCH) here.Following this logic, they then proposed that a word's length should be proportional to the expected value of its surprisal (negative log-probability in context).In this work, we show that Piantadosi et al.'s derivation does not minimize CCH's cost, but rather a lower bound, which we term CCH ↓ .We propose a novel derivation, suggesting an improved way to minimize CCH's cost.Under this method, we find that a language's word lengths should instead be proportional to the surprisal's expectation plus its variance-tomean ratio.Experimentally, we compare these three communicative cost functions: Zipf's, CCH ↓ , and CCH.Across 13 languages and several experimental settings, we find that length is better predicted by frequency than either of the other hypotheses.In fact, when surprisal's expectation, or expectation plus variance-to-mean ratio, is estimated using better language models, it leads to worse word length predictions.We take these results as evidence that Zipf's longstanding hypothesis holds.https://github.com/tpimentelms/ optimality-of-word-lengths
Tiago Pimentel, Clara Meister, Ethan Wilcox, Kyle Mahowald, Ryan Cotterell
EMNLP2
2023 Language Model Quality Correlates with Psychometric Predictive Power in Multiple Languages
abstract
Surprisal theory (Hale, 2001;Levy, 2008) posits that a word's reading time is proportional to its surprisal (i.e., to its negative log probability given the proceeding context).It has been empirically tested using surprisal estimates from language models (LMs).Under the premise that surprisal theory holds, we would expect that higher quality language models, whose predictions are more accurate, provide more powerful predictors of human reading behavior-a conjecture we dub the quality-power (QP) hypothesis.Unfortunately, empirical support for the QP hypothesis is mixed.Some studies in English have found correlations between LM quality and psychometric predictive power, but other studies using Japanese data, as well as using larger English LMs, find no such correlations.In this work, we conduct a systematic crosslinguistic assessment of the QP hypothesis.We train LMs from scratch on small-and medium-sized datasets from 13 languages (across five language families) and assess their ability to predict eye tracking data.We find correlations between LM quality and psychometric predictive power in eleven of these thirteen languages, suggesting that, within the range of model classes and sizes tested, better language models provide better predictors of human language processing behaviors.https://github.com/rycolab/ quality-power-hypothesis
Ethan Wilcox, Clara Meister, Ryan Cotterell, Tiago Pimentel
EMNLP2
2023 On the Usefulness of Embeddings, Clusters and Strings for Text Generation Evaluation
Tiago Pimentel, Clara Meister, Ryan Cotterell
ICLR2
2023 Naturalistic Causal Probing for Morpho-Syntax
abstract
Abstract Probing has become a go-to methodology for interpreting and analyzing deep neural models in natural language processing. However, there is still a lack of understanding of the limitations and weaknesses of various types of probes. In this work, we suggest a strategy for input-level intervention on naturalistic sentences. Using our approach, we intervene on the morpho-syntactic features of a sentence, while keeping the rest of the sentence unchanged. Such an intervention allows us to causally probe pre-trained models. We apply our naturalistic causal probing framework to analyze the effects of grammatical gender and number on contextualized representations extracted from three pre-trained models in Spanish, the multilingual versions of BERT, RoBERTa, and GPT-2. Our experiments suggest that naturalistic interventions lead to stable estimates of the causal effects of various linguistic properties. Moreover, our experiments demonstrate the importance of naturalistic causal probing when analyzing pre-trained models. https://github.com/rycolab/naturalistic-causal-probing
Afra Amini, Tiago Pimentel, Clara Meister, Ryan Cotterell
Trans. Assoc. Comput. Linguistics3
2023 A Cross-Linguistic Pressure for Uniform Information Density in Word Order
abstract
Abstract While natural languages differ widely in both canonical word order and word order flexibility, their word orders still follow shared cross-linguistic statistical patterns, often attributed to functional pressures. In the effort to identify these pressures, prior work has compared real and counterfactual word orders. Yet one functional pressure has been overlooked in such investigations: The uniform information density (UID) hypothesis, which holds that information should be spread evenly throughout an utterance. Here, we ask whether a pressure for UID may have influenced word order patterns cross-linguistically. To this end, we use computational models to test whether real orders lead to greater information uniformity than counterfactual orders. In our empirical study of 10 typologically diverse languages, we find that: (i) among SVO languages, real word orders consistently have greater uniformity than reverse word orders, and (ii) only linguistically implausible counterfactual orders consistently exceed the uniformity of real orders. These findings are compatible with a pressure for information uniformity in the development and usage of natural languages.1
Thomas Hikaru Clark, Clara Meister, Tiago Pimentel, Michael Hahn 0001, Ryan Cotterell, Richard Futrell, Roger Levy
Trans. Assoc. Comput. Linguistics2
2023 Locally Typical Sampling
abstract
Abstract Today’s probabilistic language generators fall short when it comes to producing coherent and fluent text despite the fact that the underlying models perform well under standard metrics (e.g., perplexity). This discrepancy has puzzled the language generation community for the last few years. In this work, we posit that the abstraction of natural language generation as a discrete stochastic process—which allows for an information-theoretic analysis—can provide new insights into the behavior of probabilistic language generators, for example, why high-probability texts can be dull or repetitive. Humans use language as a means of communicating information, aiming to do so in a simultaneously efficient and error-minimizing manner; in fact, psycholinguistics research suggests humans choose each word in a string with this subconscious goal in mind. We formally define the set of strings that meet this criterion: Those for which each word has an information content close to the expected information content, namely, the conditional entropy of our model. We then propose a simple and efficient procedure for enforcing this criterion when generating from probabilistic models, which we call locally typical sampling. Automatic and human evaluations show that, in comparison to nucleus and top-k sampling, locally typical sampling offers competitive performance (in both abstractive summarization and story generation) in terms of quality while consistently reducing degenerate repetitions.
Clara Meister, Tiago Pimentel, Gian Wiher, Ryan Cotterell
Trans. Assoc. Comput. Linguistics1
2023 On the Effect of Anticipation on Reading Times
abstract
Abstract Over the past two decades, numerous studies have demonstrated how less-predictable (i.e., higher surprisal) words take more time to read. In general, these studies have implicitly assumed the reading process is purely responsive: Readers observe a new word and allocate time to process it as required. We argue that prior results are also compatible with a reading process that is at least partially anticipatory: Readers could make predictions about a future word and allocate time to process it based on their expectation. In this work, we operationalize this anticipation as a word’s contextual entropy. We assess the effect of anticipation on reading by comparing how well surprisal and contextual entropy predict reading times on four naturalistic reading datasets: two self-paced and two eye-tracking. Experimentally, across datasets and analyses, we find substantial evidence for effects of contextual entropy over surprisal on a word’s reading time (RT): In fact, entropy is sometimes better than surprisal in predicting a word’s RT. Spillover effects, however, are generally not captured by entropy, but only by surprisal. Further, we hypothesize four cognitive mechanisms through which contextual entropy could impact RTs—three of which we are able to design experiments to analyze. Overall, our results support a view of reading that is not just responsive, but also anticipatory.1
Tiago Pimentel, Clara Meister, Ethan Wilcox, Roger Levy, Ryan Cotterell
Trans. Assoc. Comput. Linguistics2
2023 Testing the Predictions of Surprisal Theory in 11 Languages
abstract
Abstract Surprisal theory posits that less-predictable words should take more time to process, with word predictability quantified as surprisal, i.e., negative log probability in context. While evidence supporting the predictions of surprisal theory has been replicated widely, much of it has focused on a very narrow slice of data: native English speakers reading English texts. Indeed, no comprehensive multilingual analysis exists. We address this gap in the current literature by investigating the relationship between surprisal and reading times in eleven different languages, distributed across five language families. Deriving estimates from language models trained on monolingual and multilingual corpora, we test three predictions associated with surprisal theory: (i) whether surprisal is predictive of reading times, (ii) whether expected surprisal, i.e., contextual entropy, is predictive of reading times, and (iii) whether the linking function between surprisal and reading times is linear. We find that all three predictions are borne out crosslinguistically. By focusing on a more diverse set of languages, we argue that these results offer the most robust link to date between information theory and incremental language processing across languages.
Ethan Wilcox, Tiago Pimentel, Clara Meister, Ryan Cotterell, Roger Levy
Trans. Assoc. Comput. Linguistics3
2022 Mutual Information Alleviates Hallucinations in Abstractive Summarization
abstract
Despite significant progress in the quality of language generated from abstractive summarization models, these models still exhibit the tendency to hallucinate, i.e., output content not supported by the source document.A number of works have tried to fix-or at least uncover the source of-the problem with limited success.In this paper, we identify a simple criterion under which models are significantly more likely to assign more probability to hallucinated content during generation: high model uncertainty.This finding offers a potential explanation for hallucinations: models default to favoring text with high marginal probability, i.e., high-frequency occurrences in the training set, when uncertain about a continuation.It also motivates possible routes for real-time intervention during decoding to prevent such hallucinations.We propose a decoding strategy that switches to optimizing for pointwise mutual information of the source and target token-rather than purely the probability of the target token-when the model exhibits uncertainty.Experiments on the XSUM dataset show that our method decreases the probability of hallucinated tokens while maintaining the ROUGE and BERTS scores of top-performing decoding strategies.
Liam van der Poel, Ryan Cotterell, Clara Meister
EMNLP3
2022 On Decoding Strategies for Neural Text Generators
abstract
Abstract When generating text from probabilistic models, the chosen decoding strategy has a profound effect on the resulting text. Yet the properties elicited by various decoding strategies do not always transfer across natural language generation tasks. For example, while mode-seeking methods like beam search perform remarkably well for machine translation, they have been observed to lead to incoherent and repetitive text in story generation. Despite such observations, the effectiveness of decoding strategies is often assessed on only a single task. This work—in contrast—provides a comprehensive analysis of the interaction between language generation tasks and decoding strategies. Specifically, we measure changes in attributes of generated text as a function of both decoding strategy and task using human and automatic evaluation. Our results reveal both previously observed and novel findings. For example, the nature of the diversity–quality trade-off in language generation is very task-specific; the length bias often attributed to beam search is not constant across tasks. https://github.com/gianwiher/decoding-NLG
Clara Meister, Gian Wiher, Ryan Cotterell
Trans. Assoc. Comput. Linguistics1
2021 Language Model Evaluation Beyond Perplexity
abstract
Clara Meister, Ryan Cotterell. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Clara Meister, Ryan Cotterell
ACL/IJCNLP (1)1
2021 Determinantal Beam Search
abstract
Beam search is a go-to strategy for decoding neural sequence models. The algorithm can naturally be viewed as a subset optimization problem, albeit one where the corresponding set function does not reflect interactions between candidates. Empirically, this leads to sets often exhibiting high overlap, e.g., strings may differ by only a single word. Yet in use-cases that call for multiple solutions, a diverse or representative set is often desired. To address this issue, we propose a reformulation of beam search, which we call determinantal beam search. Determinantal beam search has a natural relationship to determinantal point processes (DPPs), models over sets that inherently encode intra-set interactions. By posing iterations in beam search as a series of subdeterminant maximization problems, we can turn the algorithm into a diverse subset selection process. In a case study, we use the string subsequence kernel to explicitly encourage n-gram coverage in text generated from a sequence model. We observe that our algorithm offers competitive performance against other diverse set generation strategies in the context of language generation, while providing a more general approach to optimizing for diversity.
Clara Meister, Martina Forster, Ryan Cotterell
ACL/IJCNLP (1)1
2021 A Cognitive Regularizer for Language Modeling
abstract
Jason Wei, Clara Meister, Ryan Cotterell. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jason Wei, Clara Meister, Ryan Cotterell
ACL/IJCNLP (1)2
2021 Searching for Search Errors in Neural Morphological Inflection
abstract
Neural sequence-to-sequence models are currently the predominant choice for language generation tasks.Yet, on word-level tasks, exact inference of these models reveals the empty string is often the global optimum.Prior works have speculated this phenomenon is a result of the inadequacy of neural models for language generation.However, in the case of morphological inflection, we find that the empty string is almost never the most probable solution under the model.Further, greedy search often finds the global optimum.These observations suggest that the poor calibration of many neural models may stem from characteristics of a specific subset of tasks rather than general ill-suitedness of such models for language generation.
Martina Forster, Clara Meister, Ryan Cotterell
EACL2
2021 Conditional Poisson Stochastic Beams
abstract
Beam search is the default decoding strategy for many sequence generation tasks in NLP.The set of approximate K-best items returned by the algorithm is a useful summary of the distribution for many applications; however, the candidates typically exhibit high overlap and may give a highly biased estimate for expectations under our model.These problems can be addressed by instead using stochastic decoding strategies.In this work, we propose a new method for turning beam search into a stochastic process: Conditional Poisson stochastic beam search.Rather than taking the maximizing set at each iteration, we sample K candidates without replacement according to the conditional Poisson sampling design.We view this as a more natural alternative to Kool et al. (2019)'s stochastic beam search (SBS).Furthermore, we show how samples generated under the CPSBS design can be used to build consistent estimators and sample diverse sets from sequence models.In our experiments, we observe CPSBS produces lower variance and more efficient estimators than SBS, even showing improvements in high entropy settings.1
Clara Meister, Afra Amini, Tim Vieira, Ryan Cotterell
EMNLP (1)1
2021 Revisiting the Uniform Information Density Hypothesis
abstract
The uniform information density (UID) hypothesis posits a preference among language users for utterances structured such that information is distributed uniformly across a signal.While its implications on language production have been well explored, the hypothesis potentially makes predictions about language comprehension and linguistic acceptability as well.Further, it is unclear how uniformity in a linguistic signal-or lack thereof-should be measured, and over which linguistic unit, e.g., the sentence or language level, this uniformity should hold.Here we investigate these facets of the UID hypothesis using reading time and acceptability data.While our reading time results are generally consistent with previous work, they are also consistent with a weakly super-linear effect of surprisal, which would be compatible with UID's predictions.For acceptability judgments, we find clearer evidence that non-uniformity in information density is predictive of lower acceptability.We then explore multiple operationalizations of UID, motivated by different interpretations of the original hypothesis, and analyze the scope over which the pressure towards uniformity is exerted.The explanatory power of a subset of the proposed operationalizations suggests that the strongest trend may be a regression towards a mean surprisal across the language, rather than the phrase, sentence, or document-a finding that supports a typical interpretation of UID, namely that it is the byproduct of language users maximizing the use of a (hypothetical) communication channel. 1
Clara Meister, Tiago Pimentel, Patrick Haller 0001, Lena A. Jäger, Ryan Cotterell, Roger Levy
EMNLP (1)1
2021 A surprisal-duration trade-off across and within the world's languages
abstract
While there exist scores of natural languages, each with its unique features and idiosyncrasies, they all share a unifying theme: enabling human communication.We may thus reasonably predict that human cognition shapes how these languages evolve and are used.Assuming that the capacity to process information is roughly constant across human populations, we expect a surprisal-duration trade-off to arise both across and within languages.We analyse this trade-off using a corpus of 600 languages and, after controlling for several potential confounds, we find strong supporting evidence in both settings.Specifically, we find that, on average, phones are produced faster in languages where they are less surprising, and vice versa.Further, we confirm that more surprising phones are longer, on average, in 319 languages out of the 600.We thus conclude that there is strong evidence of a surprisal-duration trade-off in operation, both across and within the world's languages.
Tiago Pimentel, Clara Meister, Elizabeth Salesky, Simone Teufel, Damián E. Blasi, Ryan Cotterell
EMNLP (1)2
2021 On Homophony and Rényi Entropy
abstract
Homophony's widespread presence in natural languages is a controversial topic.Recent theories of language optimality have tried to justify its prevalence, despite its negative effects on cognitive processing time; e.g., Piantadosi et al. (2012) argued homophony enables the reuse of efficient wordforms and is thus beneficial for languages.This hypothesis has recently been challenged by Trott and Bergen (2020), who posit that good wordforms are more often homophonous simply because they are more phonotactically probable.In this paper, we join in on the debate.We first propose a new information-theoretic quantification of a language's homophony: the sample Rényi entropy.Then, we use this quantification to revisit Trott and Bergen's claims.While their point is theoretically sound, a specific methodological issue in their experiments raises doubts about their results.After addressing this issue, we find no clear pressure either towards or against homophony-a much more nuanced result than either Piantadosi et al.'s or Trott and Bergen's findings.
Tiago Pimentel, Clara Meister, Simone Teufel, Ryan Cotterell
EMNLP (1)2
2021 Testing Machine Translation via Referential Transparency
abstract
Machine translation software has seen rapid progress in recent years due to the advancement of deep Neural Networks. People routinely use machine translation software in their daily lives for tasks such as ordering food in a foreign restaurant, receiving medical diagnosis and treatment from foreign doctors, and reading international political news online. However, due to the complexity and intractability of the underlying Neural Networks, modern machine translation software is still far from robust and can produce poor or incorrect translations; this can lead to misunderstanding, financial loss, threats to personal safety and health, and political conflicts. To address this problem, we introduce referentially transparent inputs (RTIs), a simple, widely applicable methodology for validating machine translation software. A referentially transparent input is a piece of text that should have similar translations when used in different contexts. Our practical implementation, Purity, detects when this property is broken by a translation. To evaluate RTI, we use Purity to test Google Translate and Bing Microsoft Translator with 200 unlabeled sentences, which detected 123 and 142 erroneous translations with high precision (79.3% and 78.3%). The translation errors are diverse, including examples of under-translation, over-translation, word/phrase mistranslation, incorrect modification, and unclear logic.
Pinjia He, Clara Meister, Zhendong Su 0001
ICSE2
2020 Generalized Entropy Regularization or: There's Nothing Special about Label Smoothing
abstract
Prior work has explored directly regularizing the output distributions of probabilistic models to alleviate peaky (i.e.over-confident) predictions, a common sign of overfitting.This class of techniques, of which label smoothing is one, has a connection to entropy regularization.Despite the consistent success of label smoothing across architectures and datasets in language generation tasks, two problems remain open:(1) there is little understanding of the underlying effects entropy regularizers have on models, and (2) the full space of entropy regularization techniques is largely unexplored.We introduce a parametric family of entropy regularizers, which includes label smoothing as a special case, and use it to gain a better understanding of the relationship between the entropy of a trained model and its performance on language generation tasks.We also find that variance in model performance can be explained largely by the resulting entropy of the model.Lastly, we find that label smoothing provably does not allow for sparse distributions, an undesirable property for language generation models, and therefore advise the use of other entropy regularization methods in its place.Our code is available online at https://github.com/ rycolab/entropyRegularization.2 H(p, q) := -z∈Z p(z) log q(z) is cross-entropy and H(p) := H(p, p) = -z∈Z p(z) log p(z) is the Shannon entropy, for which log = log 2 and Z = supp(p).3 The notation used by Pereyra et al. (2017) is imprecise.
Clara Meister, Elizabeth Salesky, Ryan Cotterell
ACL1
2020 If beam search is the answer, what was the question?
abstract
Quite surprisingly, exact maximum a posteriori (MAP) decoding of neural language generators frequently leads to low-quality results (Stahlberg and Byrne, 2019).Rather, most state-of-the-art results on language generation tasks are attained using beam search despite its overwhelmingly high search error rate.This implies that the MAP objective alone does not express the properties we desire in text, which merits the question: if beam search is the answer, what was the question?We frame beam search as the exact solution to a different decoding objective in order to gain insights into why high probability under a model alone may not indicate adequacy.We find that beam search enforces uniform information density in text, a property motivated by cognitive science.We suggest a set of decoding objectives that explicitly enforce this property and find that exact decoding with these objectives alleviates the problems encountered when decoding poorly calibrated language generation models.Additionally, we analyze the text produced using various decoding strategies and see that, in our neural machine translation experiments, the extent to which this property is adhered to strongly correlates with BLEU.
Clara Meister, Ryan Cotterell, Tim Vieira
EMNLP (1)1
2020 Structure-invariant testing for machine translation
abstract
In recent years, machine translation software has increasingly been integrated into our daily lives. People routinely use machine translation for various applications, such as describing symptoms to a foreign doctor and reading political news in a foreign language. However, the complexity and intractability of neural machine translation (NMT) models that power modern machine translation make the robustness of these systems difficult to even assess, much less guarantee. Machine translation systems can return inferior results that lead to misunderstanding, medical misdiagnoses, threats to personal safety, or political conflicts. Despite its apparent importance, validating the robustness of machine translation systems is very difficult and has, therefore, been much under-explored.
Pinjia He, Clara Meister, Zhendong Su 0001
ICSE2
2020 Machine translation testing via pathological invariance
abstract
Machine translation software has become heavily integrated into our daily lives due to the recent improvement in the performance of deep neural networks. However, machine translation software has been shown to regularly return erroneous translations, which can lead to harmful consequences such as economic loss and political conflicts. Additionally, due to the complexity of the underlying neural models, testing machine translation systems presents new challenges. To address this problem, we introduce a novel methodology called PatInv. The main intuition behind PatInv is that sentences with different meanings should not have the same translation. Under this general idea, we provide two realizations of PatInv that given an arbitrary sentence, generate syntactically similar but semantically different sentences by: (1) replacing one word in the sentence using a masked language model or (2) removing one word or phrase from the sentence based on its constituency structure. We then test whether the returned translations are the same for the original and modified sentences. We have applied PatInv to test Google Translate and Bing Microsoft Translator using 200 English sentences. Two language settings are considered: English-Hindi (En-Hi) and English-Chinese (En-Zh). The results show that PatInv can accurately find 308 erroneous translations in Google Translate and 223 erroneous translations in Bing Microsoft Translator, most of which cannot be found by the state-of-the-art approaches.
Shashij Gupta, Pinjia He, Clara Meister, Zhendong Su 0001
ESEC/SIGSOFT FSE3
2020 Best-First Beam Search
abstract
Decoding for many NLP tasks requires an effective heuristic algorithm for approximating exact search because the problem of searching the full output space is often intractable, or impractical in many settings. The default algorithm for this job is beam search—a pruned version of breadth-first search. Quite surprisingly, beam search often returns better results than exact inference due to beneficial search bias for NLP tasks. In this work, we show that the standard implementation of beam search can be made up to 10x faster in practice. Our method assumes that the scoring function is monotonic in the sequence length, which allows us to safely prune hypotheses that cannot be in the final set of hypotheses early on. We devise effective monotonic approximations to popular nonmonontic scoring functions, including length normalization and mutual information decoding. Lastly, we propose a memory-reduced variant of best-first beam search, which has a similar beneficial search bias in terms of downstream performance, but runs in a fraction of the time.
Clara Meister, Ryan Cotterell, Tim Vieira
Trans. Assoc. Comput. Linguistics1