Tiago Pimentel

dblp:203/8292 · DBLP profile ↗
← Back
49ranked-venue papers
19as first author
37since 2021 · last 2026
0000-0002-5159-4641ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 47 · 18 first-author · 36 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple Channels
abstract
Aditya Yadavalli, Tiago Pimentel, Tamar I Regev, Ethan Gotlieb Wilcox, Alex Warstadt. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Aditya Yadavalli, Tiago Pimentel, Tamar I. Regev, Ethan Wilcox, Alex Warstadt
ACL (1)2
2025 Causal Estimation of Tokenisation Bias
abstract
Modern language models are typically trained over subword sequences, but ultimately define probabilities over character-strings.Ideally, the choice of the tokeniser-which maps characterstrings to subwords-should not affect the probability assigned to the underlying characterstring; in practice, it does.We define this mismatch as tokenisation bias.In this work, we quantify one particular type of tokenisation bias: the effect of including or not a subword (e.g., ⟨hello⟩) in a tokeniser's vocabulary on the probability a trained model assigns to the corresponding characters (i.e., "hello").Estimating this effect is challenging because each model is trained with only one tokeniser.We address this by framing tokenisation bias as a causal effect and estimating it using the regression discontinuity design.Specifically, we exploit the fact that tokenisation algorithms rank subwords and add the first K to a tokeniser's vocabulary, where K is an arbitrary cutoff point.As such, we can estimate a causal effect by comparing similar subwords around this cutoff.Experimentally, we find that tokenisation consistently affects models' outputs across scales, vocabularies, and tokenisers.Notably, a subword's presence in a small model's vocabulary may increase its characters' probability by up to 17 times, highlighting tokenisation as a key design choice in language modelling.
Pietro Lesci, Clara Meister, Thomas Hofmann 0001, Andreas Vlachos 0001, Tiago Pimentel
ACL (1)5
2025 The time scale of redundancy between prosody and linguistic context
abstract
Tamar I Regev, Chiebuka Ohams, Shaylee Xie, Lukas Wolf, Evelina Fedorenko, Alex Warstadt, Ethan Wilcox, Tiago Pimentel. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Tamar I. Regev, Chiebuka Ohams, Shaylee Xie, Lukas Wolf, Evelina Fedorenko, Alex Warstadt, Ethan Wilcox, Tiago Pimentel
ACL (1)8
2025 Tokenisation is NP-Complete
abstract
In this work, we prove the NP-completeness of two variants of tokenisation, defined here as the problem of compressing a dataset to at most δ symbols by either finding a vocabulary directly (direct tokenisation), or selecting a sequence of merge operations (bottom-up tokenisation).
Philip Whittington, Gregor Bachmann, Tiago Pimentel
ACL (1)3
2025 Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-Accent
abstract
Ethan Wilcox, Cui Ding, Giovanni Acampa, Tiago Pimentel, Alex Warstadt, Tamar I Regev. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ethan Wilcox, Cui Ding, Giovanni Acampa, Tiago Pimentel, Alex Warstadt, Tamar I. Regev
ACL (1)4
2025 Convergence and Divergence of Language Models under Different Random Seeds
abstract
In this paper, we investigate the convergence of language models (LMs) trained under different random seeds, measuring convergence as the expected per-token Kullback-Leibler (KL) divergence across seeds.By comparing LM convergence as a function of model size and training checkpoint, we identify a four-phase convergence pattern: (i) an initial uniform phase, (ii) a sharp-convergence phase, (iii) a sharp-divergence phase, and (iv) a slowreconvergence phase.Further, we observe that larger models reconverge faster in later training stages, while smaller models never actually reconverge; these results suggest that a certain model size may be necessary to learn stable distributions.Restricting our analysis to specific token frequencies or part-of-speech (PoS) tags further reveals that convergence is uneven across linguistic categories: frequent tokens and function words converge faster and more reliably than their counterparts (infrequent tokens and content words).Overall, our findings highlight factors that influence the stability of the learned distributions in model training.
Finlay Fehlauer, Kyle Mahowald, Tiago Pimentel
EMNLP3
2025 The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
abstract
The concept of causal abstraction got recently popularised to demystify the opaque decision-making processes of machine learning models; in short, a neural network can be abstracted as a higher-level algorithm if there exists a function which allows us to map between them. Notably, most interpretability papers implement these maps as linear functions, motivated by the linear representation hypothesis: the idea that features are encoded linearly in a model's representations. However, this linearity constraint is not required by the definition of causal abstraction. In this work, we critically examine the concept of causal abstraction by considering arbitrarily powerful alignment maps. In particular, we prove that under reasonable assumptions, any neural network can be mapped to any algorithm, rendering this unrestricted notion of causal abstraction trivial and uninformative. We complement these theoretical findings with empirical evidence, demonstrating that it is possible to perfectly map models to algorithms even when these models are incapable of solving the actual task; e.g., on an experiment using randomly initialised language models, our alignment maps reach 100\% interchange-intervention accuracy on the indirect object identification task. This raises the non-linear representation dilemma: if we lift the linearity constraint imposed to alignment maps in causal abstraction analyses, we are left with no principled way to balance the inherent trade-off between these maps' complexity and accuracy. Together, these results suggest an answer to our title's question: causal abstraction is not enough for mechanistic interpretability, as it becomes vacuous without assumptions about how models encode information. Studying the connection between this information-encoding assumption and causal abstraction should lead to exciting future work.
Denis Sutter, Julian Minder, Thomas Hofmann 0001, Tiago Pimentel
NeurIPS4
2025 Investigating Critical Period Effects in Language Acquisition through Neural Language Models
abstract
Abstract Humans appear to have a critical period (CP) for language acquisition: Second language (L2) acquisition becomes harder after early childhood, and ceasing exposure to a first language (L1) after this period (but not before) typically does not lead to substantial loss of L1 proficiency. It is unknown whether these CP effects result from innately determined brain maturation or as a stabilization of neural connections naturally induced by experience. In this study, we use language models (LMs) to test the extent to which these phenomena are peculiar to humans, or shared by a broader class of language learners. We vary the age of exposure by training LMs on language pairs in various experimental conditions, and find that LMs, which lack any direct analog to innate maturational stages, do not show CP effects when the age of exposure of L2 is delayed. Our results contradict the claim that CP effects are an inevitable result of statistical learning, and they are consistent with an innate mechanism for CP effects. We show that we can reverse-engineer the CP by introducing a regularizer partway through training to simulate a maturational decrease in plasticity. All in all, our results suggest that L1 learning on its own may not be enough to induce a CP, and additional engineering is necessary to make language models more cognitively plausible.
Ionut Constantinescu, Tiago Pimentel, Ryan Cotterell, Alex Warstadt
Trans. Assoc. Comput. Linguistics2
2024 Causal Estimation of Memorisation Profiles
abstract
Understanding memorisation in language models has practical and societal implications, e.g., studying models' training dynamics or preventing copyright infringements.Prior work defines memorisation as the causal effect of training with an instance on the model's ability to predict that instance.This definition relies on a counterfactual: the ability to observe what would have happened had the model not seen that instance.Existing methods struggle to provide computationally efficient and accurate estimates of this counterfactual.Further, they often estimate memorisation for a model architecture rather than for a specific model instance.This paper fills an important gap in the literature, proposing a new, principled, and efficient method to estimate memorisation based on the difference-in-differences design from econometrics.Using this method, we characterise a model's memorisation profile-its memorisation trends across training-by only observing its behaviour on a small set of instances throughout training.In experiments with the Pythia model suite, we find that memorisation (i) is stronger and more persistent in larger models, (ii) is determined by data order and learning rate, and (iii) has stable trends across model sizes, thus making memorisation in larger models predictable from smaller ones. pietrolesci/memorisation-profiles
Pietro Lesci, Clara Meister, Thomas Hofmann 0001, Andreas Vlachos 0001, Tiago Pimentel
ACL (1)5
2024 Towards a Similarity-adjusted Surprisal Theory
abstract
Surprisal theory posits that the cognitive effort required to comprehend a word is determined by its contextual predictability, quantified as surprisal.Traditionally, surprisal theory treats words as distinct entities, overlooking any potential similarity between them.Giulianelli et al. (2023) address this limitation by introducing information value, a measure of predictability designed to account for similarities between communicative units.Our work leverages Ricotta and Szeidl's (2006) diversity index to extend surprisal into a metric that we term similarity-adjusted surprisal, exposing a mathematical relationship between surprisal and information value.Similarity-adjusted surprisal aligns with information value when considering graded similarities and reduces to standard surprisal when words are treated as distinct.Experimental results with reading time data indicate that similarity-adjusted surprisal adds predictive power beyond standard surprisal for certain datasets, suggesting it serves as a complementary measure of comprehension effort.
Clara Meister, Mario Giulianelli, Tiago Pimentel
EMNLP3
2024 How to Compute the Probability of a Word
abstract
Language models (LMs) estimate a probability distribution over strings in a natural language; these distributions are crucial for computing perplexity and surprisal in linguistics research.While we are usually concerned with measuring these values for words, most LMs operate over subwords.Despite seemingly straightforward, accurately computing probabilities over one unit given probabilities over the other requires care.Indeed, we show here that many recent linguistic studies have been incorrectly computing these values.This paper derives the correct methods for computing word probabilities, highlighting issues when relying on language models that use beginning-of-word (bow)-marking tokenisers, e.g., the GPT family.Empirically, we show that correcting the widespread bug in probability computations affects measured outcomes in sentence comprehension and lexical optimisation analyses.tpimentelms/probability-of-a-word pip install wordsprobability
Tiago Pimentel, Clara Meister
EMNLP1
2023 A Measure-Theoretic Characterization of Tight Language Models
abstract
Li Du, Lucas Torroba Hennigen, Tiago Pimentel, Clara Meister, Jason Eisner, Ryan Cotterell. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Lucas Torroba Hennigen, Tiago Pimentel, Clara Meister, Jason Eisner, Ryan Cotterell
ACL (1)3
2023 On the Efficacy of Sampling Adapters
abstract
Sampling is a common strategy for generating text from probabilistic models, yet standard ancestral sampling often results in text that is incoherent or ungrammatical.To alleviate this issue, various modifications to a model's sampling distribution, such as nucleus or top-k sampling, have been introduced and are now ubiquitously used in language generation systems.We propose a unified framework for understanding these techniques, which we term sampling adapters.Sampling adapters often lead to qualitatively better text, which raises the question: From a formal perspective, how are they changing the (sub)word-level distributions of language generation models?And why do these local changes lead to higher-quality text?We argue that the shift they enforce can be viewed as a trade-off between precision and recall: while the model loses its ability to produce certain strings, its precision rate on desirable text increases.While this trade-off is not reflected in standard metrics of distribution quality (such as perplexity), we find that several precision-emphasizing measures indeed indicate that sampling adapters can lead to probability distributions more aligned with the true distribution.Further, these measures correlate with higher sequence-level quality scores, specifically, MAUVE.https://github.com/rycolab/ sampling-adapters
Clara Meister, Tiago Pimentel, Luca Malagutti, Ethan Wilcox, Ryan Cotterell
ACL (1)2
2023 On the Intersection of Context-Free and Regular Languages
abstract
Clemente Pasti, Andreas Opedal, Tiago Pimentel, Tim Vieira, Jason Eisner, Ryan Cotterell. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Clemente Pasti, Andreas Opedal, Tiago Pimentel, Tim Vieira, Jason Eisner, Ryan Cotterell
EACL3
2023 An Exploration of Left-Corner Transformations
abstract
The left-corner transformation (Rosenkrantz and Lewis, 1970) is used to remove left recursion from context-free grammars, which is an important step towards making the grammar parsable top-down with simple techniques.This paper generalizes prior left-corner transformations to support semiring-weighted production rules and to provide finer-grained control over which left corners may be moved.Our generalized left-corner transformation (GLCT) arose from unifying the left-corner transformation and speculation transformation (Eisner and Blatz, 2007), originally for logic programming.Our new transformation and speculation define equivalent weighted languages.Yet, their derivation trees are structurally different in an important way: GLCT replaces left recursion with right recursion, and speculation does not.We also provide several technical results regarding the formal relationships between the outputs of GLCT, speculation, and the original grammar.Lastly, we empirically investigate the efficiency of GLCT for left-recursion elimination from grammars of nine languages.
Andreas Opedal, Eleftheria Tsipidi, Tiago Pimentel, Ryan Cotterell, Tim Vieira
EMNLP3
2023 Revisiting the Optimality of Word Lengths
abstract
Zipf (1935) posited that wordforms are optimized to minimize utterances' communicative costs.Under the assumption that cost is given by an utterance's length, he supported this claim by showing that words' lengths are inversely correlated with their frequencies.Communicative cost, however, can be operationalized in different ways.Piantadosi et al. (2011) claim that cost should be measured as the distance between an utterance's information rate and channel capacity, which we dub the channel capacity hypothesis (CCH) here.Following this logic, they then proposed that a word's length should be proportional to the expected value of its surprisal (negative log-probability in context).In this work, we show that Piantadosi et al.'s derivation does not minimize CCH's cost, but rather a lower bound, which we term CCH ↓ .We propose a novel derivation, suggesting an improved way to minimize CCH's cost.Under this method, we find that a language's word lengths should instead be proportional to the surprisal's expectation plus its variance-tomean ratio.Experimentally, we compare these three communicative cost functions: Zipf's, CCH ↓ , and CCH.Across 13 languages and several experimental settings, we find that length is better predicted by frequency than either of the other hypotheses.In fact, when surprisal's expectation, or expectation plus variance-to-mean ratio, is estimated using better language models, it leads to worse word length predictions.We take these results as evidence that Zipf's longstanding hypothesis holds.https://github.com/tpimentelms/ optimality-of-word-lengths
Tiago Pimentel, Clara Meister, Ethan Wilcox, Kyle Mahowald, Ryan Cotterell
EMNLP1
2023 Language Model Quality Correlates with Psychometric Predictive Power in Multiple Languages
abstract
Surprisal theory (Hale, 2001;Levy, 2008) posits that a word's reading time is proportional to its surprisal (i.e., to its negative log probability given the proceeding context).It has been empirically tested using surprisal estimates from language models (LMs).Under the premise that surprisal theory holds, we would expect that higher quality language models, whose predictions are more accurate, provide more powerful predictors of human reading behavior-a conjecture we dub the quality-power (QP) hypothesis.Unfortunately, empirical support for the QP hypothesis is mixed.Some studies in English have found correlations between LM quality and psychometric predictive power, but other studies using Japanese data, as well as using larger English LMs, find no such correlations.In this work, we conduct a systematic crosslinguistic assessment of the QP hypothesis.We train LMs from scratch on small-and medium-sized datasets from 13 languages (across five language families) and assess their ability to predict eye tracking data.We find correlations between LM quality and psychometric predictive power in eleven of these thirteen languages, suggesting that, within the range of model classes and sizes tested, better language models provide better predictors of human language processing behaviors.https://github.com/rycolab/ quality-power-hypothesis
Ethan Wilcox, Clara Meister, Ryan Cotterell, Tiago Pimentel
EMNLP4
2023 Quantifying the redundancy between prosody and text
abstract
Prosody-the suprasegmental component of speech, including pitch, loudness, and tempocarries critical aspects of meaning.However, the relationship between the information conveyed by prosody vs. by the words themselves remains poorly understood.We use large language models (LLMs) to estimate how much information is redundant between prosody and the words themselves.Using a large spoken corpus of English audiobooks, we extract prosodic features aligned to individual words and test how well they can be predicted from LLM embeddings, compared to non-contextual word embeddings.We find a high degree of redundancy between the information carried by the words and prosodic information across several prosodic features, including intensity, duration, pauses, and pitch contours.Furthermore, a word's prosodic information is redundant with both the word itself and the context preceding as well as following it.Still, we observe that prosodic features can not be fully predicted from text, suggesting that prosody carries information above and beyond the words.Along with this paper, we release a general-purpose data processing pipeline for quantifying the relationship between linguistic information and extra-linguistic features.https://github.com/lu-wo/ quantifying-redundancy
Lukas Wolf, Tiago Pimentel, Evelina Fedorenko, Ryan Cotterell, Alex Warstadt, Ethan Wilcox, Tamar I. Regev
EMNLP2
2023 On the Usefulness of Embeddings, Clusters and Strings for Text Generation Evaluation
Tiago Pimentel, Clara Meister, Ryan Cotterell
ICLR1
2023 Naturalistic Causal Probing for Morpho-Syntax
abstract
Abstract Probing has become a go-to methodology for interpreting and analyzing deep neural models in natural language processing. However, there is still a lack of understanding of the limitations and weaknesses of various types of probes. In this work, we suggest a strategy for input-level intervention on naturalistic sentences. Using our approach, we intervene on the morpho-syntactic features of a sentence, while keeping the rest of the sentence unchanged. Such an intervention allows us to causally probe pre-trained models. We apply our naturalistic causal probing framework to analyze the effects of grammatical gender and number on contextualized representations extracted from three pre-trained models in Spanish, the multilingual versions of BERT, RoBERTa, and GPT-2. Our experiments suggest that naturalistic interventions lead to stable estimates of the causal effects of various linguistic properties. Moreover, our experiments demonstrate the importance of naturalistic causal probing when analyzing pre-trained models. https://github.com/rycolab/naturalistic-causal-probing
Afra Amini, Tiago Pimentel, Clara Meister, Ryan Cotterell
Trans. Assoc. Comput. Linguistics2
2023 A Cross-Linguistic Pressure for Uniform Information Density in Word Order
abstract
Abstract While natural languages differ widely in both canonical word order and word order flexibility, their word orders still follow shared cross-linguistic statistical patterns, often attributed to functional pressures. In the effort to identify these pressures, prior work has compared real and counterfactual word orders. Yet one functional pressure has been overlooked in such investigations: The uniform information density (UID) hypothesis, which holds that information should be spread evenly throughout an utterance. Here, we ask whether a pressure for UID may have influenced word order patterns cross-linguistically. To this end, we use computational models to test whether real orders lead to greater information uniformity than counterfactual orders. In our empirical study of 10 typologically diverse languages, we find that: (i) among SVO languages, real word orders consistently have greater uniformity than reverse word orders, and (ii) only linguistically implausible counterfactual orders consistently exceed the uniformity of real orders. These findings are compatible with a pressure for information uniformity in the development and usage of natural languages.1
Thomas Hikaru Clark, Clara Meister, Tiago Pimentel, Michael Hahn 0001, Ryan Cotterell, Richard Futrell, Roger Levy
Trans. Assoc. Comput. Linguistics3
2023 Locally Typical Sampling
abstract
Abstract Today’s probabilistic language generators fall short when it comes to producing coherent and fluent text despite the fact that the underlying models perform well under standard metrics (e.g., perplexity). This discrepancy has puzzled the language generation community for the last few years. In this work, we posit that the abstraction of natural language generation as a discrete stochastic process—which allows for an information-theoretic analysis—can provide new insights into the behavior of probabilistic language generators, for example, why high-probability texts can be dull or repetitive. Humans use language as a means of communicating information, aiming to do so in a simultaneously efficient and error-minimizing manner; in fact, psycholinguistics research suggests humans choose each word in a string with this subconscious goal in mind. We formally define the set of strings that meet this criterion: Those for which each word has an information content close to the expected information content, namely, the conditional entropy of our model. We then propose a simple and efficient procedure for enforcing this criterion when generating from probabilistic models, which we call locally typical sampling. Automatic and human evaluations show that, in comparison to nucleus and top-k sampling, locally typical sampling offers competitive performance (in both abstractive summarization and story generation) in terms of quality while consistently reducing degenerate repetitions.
Clara Meister, Tiago Pimentel, Gian Wiher, Ryan Cotterell
Trans. Assoc. Comput. Linguistics2
2023 On the Effect of Anticipation on Reading Times
abstract
Abstract Over the past two decades, numerous studies have demonstrated how less-predictable (i.e., higher surprisal) words take more time to read. In general, these studies have implicitly assumed the reading process is purely responsive: Readers observe a new word and allocate time to process it as required. We argue that prior results are also compatible with a reading process that is at least partially anticipatory: Readers could make predictions about a future word and allocate time to process it based on their expectation. In this work, we operationalize this anticipation as a word’s contextual entropy. We assess the effect of anticipation on reading by comparing how well surprisal and contextual entropy predict reading times on four naturalistic reading datasets: two self-paced and two eye-tracking. Experimentally, across datasets and analyses, we find substantial evidence for effects of contextual entropy over surprisal on a word’s reading time (RT): In fact, entropy is sometimes better than surprisal in predicting a word’s RT. Spillover effects, however, are generally not captured by entropy, but only by surprisal. Further, we hypothesize four cognitive mechanisms through which contextual entropy could impact RTs—three of which we are able to design experiments to analyze. Overall, our results support a view of reading that is not just responsive, but also anticipatory.1
Tiago Pimentel, Clara Meister, Ethan Wilcox, Roger Levy, Ryan Cotterell
Trans. Assoc. Comput. Linguistics1
2023 Testing the Predictions of Surprisal Theory in 11 Languages
abstract
Abstract Surprisal theory posits that less-predictable words should take more time to process, with word predictability quantified as surprisal, i.e., negative log probability in context. While evidence supporting the predictions of surprisal theory has been replicated widely, much of it has focused on a very narrow slice of data: native English speakers reading English texts. Indeed, no comprehensive multilingual analysis exists. We address this gap in the current literature by investigating the relationship between surprisal and reading times in eleven different languages, distributed across five language families. Deriving estimates from language models trained on monolingual and multilingual corpora, we test three predictions associated with surprisal theory: (i) whether surprisal is predictive of reading times, (ii) whether expected surprisal, i.e., contextual entropy, is predictive of reading times, and (iii) whether the linking function between surprisal and reading times is linear. We find that all three predictions are borne out crosslinguistically. By focusing on a more diverse set of languages, we argue that these results offer the most robust link to date between information theory and incremental language processing across languages.
Ethan Wilcox, Tiago Pimentel, Clara Meister, Ryan Cotterell, Roger Levy
Trans. Assoc. Comput. Linguistics2
2022 Probing for the Usage of Grammatical Number
abstract
A central quest of probing is to uncover how pre-trained models encode a linguistic property within their representations.An encoding, however, might be spurious-i.e., the model might not rely on it when making predictions.In this paper, we try to find an encoding that the model actually uses, introducing a usage-based probing setup.We first choose a behavioral task which cannot be solved without using the linguistic property.Then, we attempt to remove the property by intervening on the model's representations.We contend that, if an encoding is used by the model, its removal should harm the performance on the chosen behavioral task.As a case study, we focus on how BERT encodes grammatical number, and on how it uses this encoding to solve the number agreement task.Experimentally, we find that BERT relies on a linear encoding of grammatical number to produce the correct behavioral output.We also find that BERT uses a separate encoding of grammatical number for nouns and verbs.Finally, we identify in which layers information about grammatical number is transferred from a noun to its head verb.
Karim Lasri, Tiago Pimentel, Alessandro Lenci, Thierry Poibeau, Ryan Cotterell
ACL (1)2
2022 Attentional Probe: Estimating a Module's Functional Potential
abstract
In this paper, we seek to measure how much information a component in a neural network could extract from the representations fed into it.Our work stands in contrast to prior probing work, most of which investigates how much information a model's representations contain.This shift in perspective leads us to propose a new principle for probing, the architectural bottleneck principle: In order to estimate how much information a given component could extract, a probe should look exactly like the component.Relying on this principle, we estimate how much syntactic information is available to transformers through our attentional probe, a probe that exactly resembles a transformer's self-attention head.Experimentally, we find that, in three models (BERT, ALBERT, and RoBERTa), a sentence's syntax tree is mostly extractable by our probe, suggesting these models have access to syntactic information while composing their contextual representations.Whether this information is actually used by these models, however, remains an open question.
Tiago Pimentel, Josef Valvoda, Niklas Stoehr, Ryan Cotterell
EMNLP1
2022 UniMorph 4.0: Universal Morphology
abstract
The Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation, and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements on several fronts that were made in the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 66 new languages, including 24 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g., missing gender and macrons information. We have amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet.
Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieras, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina J. Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Lane 0002, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóga, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer C. White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo Maria Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar 0002, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Tucker Prud'hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova
LREC15
2022 Rethinking Reinforcement Learning for Recommendation: A Prompt Perspective
abstract
Modern recommender systems aim to improve user experience. As reinforcement learning (RL) naturally fits this objective---maximizing an user's reward per session---it has become an emerging topic in recommender systems. Developing RL-based recommendation methods, however, is not trivial due to the offline training challenge. Specifically, the keystone of traditional RL is to train an agent with large amounts of online exploration making lots of 'errors' in the process. In the recommendation setting, though, we cannot afford the price of making 'errors' online. As a result, the agent needs to be trained through offline historical implicit feedback, collected under different recommendation policies; traditional RL algorithms may lead to sub-optimal policies under these offline training settings.
Xin Xin 0003, Tiago Pimentel, Alexandros Karatzoglou, Pengjie Ren, Konstantina Christakopoulou, Zhaochun Ren
SIGIR2
2021 Disambiguatory Signals are Stronger in Word-initial Positions
abstract
Psycholinguistic studies of human word processing and lexical access provide ample evidence of the preferred nature of word-initial versus word-final segments, e.g., in terms of attention paid by listeners (greater) or the likelihood of reduction by speakers (lower).This has led to the conjecture-as in Wedel et al. (2019b), but common elsewhere-that languages have evolved to provide more information earlier in words than later.Informationtheoretic methods to establish such tendencies in lexicons have suffered from several methodological shortcomings that leave open the question of whether this high word-initial informativeness is actually a property of the lexicon or simply an artefact of the incremental nature of recognition.In this paper, we point out the confounds in existing methods for comparing the informativeness of segments early in the word versus later in the word, and present several new measures that avoid these confounds.When controlling for these confounds, we still find evidence across hundreds of languages that indeed there is a cross-linguistic tendency to front-load information in words. 1
Tiago Pimentel, Ryan Cotterell, Brian Roark
EACL1
2021 Revisiting the Uniform Information Density Hypothesis
abstract
The uniform information density (UID) hypothesis posits a preference among language users for utterances structured such that information is distributed uniformly across a signal.While its implications on language production have been well explored, the hypothesis potentially makes predictions about language comprehension and linguistic acceptability as well.Further, it is unclear how uniformity in a linguistic signal-or lack thereof-should be measured, and over which linguistic unit, e.g., the sentence or language level, this uniformity should hold.Here we investigate these facets of the UID hypothesis using reading time and acceptability data.While our reading time results are generally consistent with previous work, they are also consistent with a weakly super-linear effect of surprisal, which would be compatible with UID's predictions.For acceptability judgments, we find clearer evidence that non-uniformity in information density is predictive of lower acceptability.We then explore multiple operationalizations of UID, motivated by different interpretations of the original hypothesis, and analyze the scope over which the pressure towards uniformity is exerted.The explanatory power of a subset of the proposed operationalizations suggests that the strongest trend may be a regression towards a mean surprisal across the language, rather than the phrase, sentence, or document-a finding that supports a typical interpretation of UID, namely that it is the byproduct of language users maximizing the use of a (hypothetical) communication channel. 1
Clara Meister, Tiago Pimentel, Patrick Haller 0001, Lena A. Jäger, Ryan Cotterell, Roger Levy
EMNLP (1)2
2021 A Bayesian Framework for Information-Theoretic Probing
abstract
Pimentel et al. (2020b) recently analysed probing from an information-theoretic perspective.They argue that probing should be seen as approximating a mutual information.This led to the rather unintuitive conclusion that representations encode exactly the same information about a target task as the original sentences.The mutual information, however, assumes the true probability distribution of a pair of random variables is known, leading to unintuitive results in settings where it is not.This paper proposes a new framework to measure what we term Bayesian mutual information, which analyses information from the perspective of Bayesian agents-allowing for more intuitive findings in scenarios with finite data.For instance, under Bayesian MI we have that data can add information, processing can help, and information can hurt, which makes it more intuitive for machine learning applications.Finally, we apply our framework to probing where we believe Bayesian mutual information naturally operationalises ease of extraction by explicitly limiting the available background knowledge to solve a task.
Tiago Pimentel, Ryan Cotterell
EMNLP (1)1
2021 A surprisal-duration trade-off across and within the world's languages
abstract
While there exist scores of natural languages, each with its unique features and idiosyncrasies, they all share a unifying theme: enabling human communication.We may thus reasonably predict that human cognition shapes how these languages evolve and are used.Assuming that the capacity to process information is roughly constant across human populations, we expect a surprisal-duration trade-off to arise both across and within languages.We analyse this trade-off using a corpus of 600 languages and, after controlling for several potential confounds, we find strong supporting evidence in both settings.Specifically, we find that, on average, phones are produced faster in languages where they are less surprising, and vice versa.Further, we confirm that more surprising phones are longer, on average, in 319 languages out of the 600.We thus conclude that there is strong evidence of a surprisal-duration trade-off in operation, both across and within the world's languages.
Tiago Pimentel, Clara Meister, Elizabeth Salesky, Simone Teufel, Damián E. Blasi, Ryan Cotterell
EMNLP (1)1
2021 On Homophony and Rényi Entropy
abstract
Homophony's widespread presence in natural languages is a controversial topic.Recent theories of language optimality have tried to justify its prevalence, despite its negative effects on cognitive processing time; e.g., Piantadosi et al. (2012) argued homophony enables the reuse of efficient wordforms and is thus beneficial for languages.This hypothesis has recently been challenged by Trott and Bergen (2020), who posit that good wordforms are more often homophonous simply because they are more phonotactically probable.In this paper, we join in on the debate.We first propose a new information-theoretic quantification of a language's homophony: the sample Rényi entropy.Then, we use this quantification to revisit Trott and Bergen's claims.While their point is theoretically sound, a specific methodological issue in their experiments raises doubts about their results.After addressing this issue, we find no clear pressure either towards or against homophony-a much more nuanced result than either Piantadosi et al.'s or Trott and Bergen's findings.
Tiago Pimentel, Clara Meister, Simone Teufel, Ryan Cotterell
EMNLP (1)1
2021 How (Non-)Optimal is the Lexicon?
abstract
Tiago Pimentel, Irene Nikkarinen, Kyle Mahowald, Ryan Cotterell, Damián Blasi. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Tiago Pimentel, Irene Nikkarinen, Kyle Mahowald, Ryan Cotterell, Damián E. Blasi
NAACL-HLT1
2021 Finding Concept-specific Biases in Form-Meaning Associations
abstract
Tiago Pimentel, Brian Roark, Søren Wichmann, Ryan Cotterell, Damián Blasi. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Tiago Pimentel, Brian Roark, Søren Wichmann, Ryan Cotterell, Damián E. Blasi
NAACL-HLT1
2021 What About the Precedent: An Information-Theoretic Analysis of Common Law
abstract
Josef Valvoda, Tiago Pimentel, Niklas Stoehr, Ryan Cotterell, Simone Teufel. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Josef Valvoda, Tiago Pimentel, Niklas Stoehr, Ryan Cotterell, Simone Teufel
NAACL-HLT2
2021 A Non-Linear Structural Probe
abstract
Probes are models devised to investigate the encoding of knowledge—e.g. syntactic structure—in contextual representations. Probes are often designed for simplicity, which has led to restrictions on probe design that may not allow for the full exploitation of the structure of encoded information; one such restriction is linearity. We examine the case of a structural probe (Hewitt and Manning, 2019), which aims to investigate the encoding of syntactic structure in contextual representations through learning only linear transformations. By observing that the structural probe learns a metric, we are able to kernelize it and develop a novel non-linear variant with an identical number of parameters. We test on 6 languages and find that the radial-basis function (RBF) kernel, in conjunction with regularization, achieves a statistically significant improvement over the baseline in all languages—implying that at least part of the syntactic knowledge is encoded non-linearly. We conclude by discussing how the RBF kernel resembles BERT’s self-attention layers and speculate that this resemblance leads to the RBF-based probe’s stronger performance.
Jennifer C. White, Tiago Pimentel, Naomi Saphra, Ryan Cotterell
NAACL-HLT2
2020 A Tale of a Probe and a Parser
abstract
Measuring what linguistic information is encoded in neural models of language has become popular in NLP.Researchers approach this enterprise by training "probes"supervised models designed to extract linguistic structure from another model's output.One such probe is the structural probe (Hewitt and Manning, 2019), designed to quantify the extent to which syntactic information is encoded in contextualised word representations.The structural probe has a novel design, unattested in the parsing literature, the precise benefit of which is not immediately obvious.To explore whether syntactic probes would do better to make use of existing techniques, we compare the structural probe to a more traditional parser with an identical lightweight parameterisation.The parser outperforms structural probe on UUAS in seven of nine analysed languages, often by a substantial amount (e.g. by 11.1 points in English).Under a second less common metric, however, there is the opposite trend-the structural probe outperforms the parser.This begs the question: which metric should we prefer?
Rowan Hall Maudslay, Josef Valvoda, Tiago Pimentel, Adina Williams, Ryan Cotterell
ACL3
2020 Information-Theoretic Probing for Linguistic Structure
abstract
The success of neural networks on a diverse set of NLP tasks has led researchers to question how much these networks actually "know" about natural language.Probes are a natural way of assessing this.When probing, a researcher chooses a linguistic task and trains a supervised model to predict annotations in that linguistic task from the network's learned representations.If the probe does well, the researcher may conclude that the representations encode knowledge related to the task.A commonly held belief is that using simpler models as probes is better; the logic is that simpler models will identify linguistic structure, but not learn the task itself.We propose an information-theoretic operationalization of probing as estimating mutual information that contradicts this received wisdom: one should always select the highest performing probe one can, even if it is more complex, since it will result in a tighter estimate, and thus reveal more of the linguistic information inherent in the representation.The experimental portion of our paper focuses on empirically estimating the mutual information between a linguistic property and BERT, comparing these estimates to several baselines.We evaluate on a set of ten typologically diverse languages often underrepresented in NLP research-plus Englishtotalling eleven languages.Our implementation is available in https://github.com/ rycolab/info-theoretic-probing.
Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, Ryan Cotterell
ACL1
2020 A Corpus for Large-Scale Phonetic Typology
abstract
A major hurdle in data-driven research on typology is having sufficient data in many languages to draw meaningful conclusions. We present VoxClamantis v1.0, the first large-scale corpus for phonetic typology, with aligned segments and estimated phoneme-level labels in 690 readings spanning 635 languages, along with acoustic-phonetic measures of vowels and sibilants. Access to such data can greatly facilitate investigation of phonetic typology at a large scale and across many languages. However, it is non-trivial and computationally intensive to obtain such alignments for hundreds of languages, many of which have few to no resources presently available. We describe the methodology to create our corpus, discuss caveats with current methods and their impact on the utility of this data, and illustrate possible research directions through a series of case studies on the 48 highest-quality readings. Our corpus and scripts are publicly available for non-commercial use at https://voxclamantisproject.github.io.
Elizabeth Salesky, Eleanor Chodroff, Tiago Pimentel, Matthew Wiesner, Ryan Cotterell, Alan W. Black, Jason Eisner
ACL3
2020 Predicting Declension Class from Form and Meaning
abstract
The noun lexica of many natural languages are divided into several declension classes with characteristic morphological properties.Class membership is far from deterministic, but the phonological form of a noun and its meaning can often provide imperfect clues.Here, we investigate the strength of those clues.More specifically, we operationalize "strength" as measuring how much information, in bits, we can glean about declension class from knowing the form and meaning of nouns.We know that form and meaning are often also indicative of grammatical gender-which, as we quantitatively verify, can itself share information with declension class-so we also control for gender.We find for two Indo-European languages (Czech and German) that form and meaning share a significant amount of information with class (and contribute additional information beyond gender).The three-way interaction between class, form, and meaning (given gender) is also significant.Our study is important for two reasons: First, we introduce a new method that provides additional quantitative support for a classic linguistic finding that form and meaning are relevant for the classification of nouns into declensions.Second, we show not only that individual declension classes vary in the strength of their clues within a language, but also that the variations between classes vary across languages.The code is publicly available at https://github.com/ rycolab/declension-mi.
Adina Williams, Tiago Pimentel, Hagen Blix, Arya McCarthy, Eleanor Chodroff, Ryan Cotterell
ACL2
2020 Speakers Fill Lexical Semantic Gaps with Context
abstract
Lexical ambiguity is widespread in language, allowing for the reuse of economical word forms and therefore making language more efficient.If ambiguous words cannot be disambiguated from context, however, this gain in efficiency might make language less clearresulting in frequent miscommunication.For a language to be clear and efficiently encoded, we posit that the lexical ambiguity of a word type should correlate with how much information context provides about it, on average.To investigate whether this is the case, we operationalise the lexical ambiguity of a word as the entropy of meanings it can take, and provide two ways to estimate this-one which requires human annotation (using WordNet), and one which does not (using BERT), making it readily applicable to a large number of languages.We validate these measures by showing that, on six high-resource languages, there are significant Pearson correlations between our BERT-based estimate of ambiguity and the number of synonyms a word has in Word-Net (e.g.ρ = 0.40 in English).We then test our main hypothesis-that a word's lexical ambiguity should negatively correlate with its contextual uncertainty-and find significant correlations on all 18 typologically diverse languages we analyse.This suggests that, in the presence of ambiguity, speakers compensate by making contexts more informative.
Tiago Pimentel, Rowan Hall Maudslay, Damián E. Blasi, Ryan Cotterell
EMNLP (1)1
2020 Pareto Probing: Trading Off Accuracy for Complexity
abstract
The question of how to probe contextual word representations for linguistic structure in a way that is both principled and useful has seen significant attention recently in the NLP literature.In our contribution to this discussion, we argue for a probe metric that reflects the fundamental trade-off between probe complexity and performance: the Pareto hypervolume.To measure complexity, we present a number of parametric and non-parametric metrics.Our experiments using Pareto hypervolume as an evaluation metric show that probes often do not conform to our expectations-e.g., why should the non-contextual fastText representations encode more morpho-syntactic information than the contextual BERT representations?These results suggest that common, simplistic probing tasks, such as part-of-speech labeling and dependency arc labeling, are inadequate to evaluate the linguistic structure encoded in contextual word representations.This leads us to propose full dependency parsing as a probing task.In support of our suggestion that harder probing tasks are necessary, our experiments with dependency parsing reveal a wide gap in syntactic knowledge between contextual and non-contextual representations.Our code can be found at https://github. com/rycolab/pareto-probing.
Tiago Pimentel, Naomi Saphra, Adina Williams, Ryan Cotterell
EMNLP (1)1
2020 Deep Active Learning for Anomaly Detection
abstract
Anomalies are intuitively easy for human experts to understand, but they are hard to define mathematically. Therefore, in order to have performance guarantees in unsupervised anomaly detection, priors need to be assumed on what the anomalies are. By contrast, active learning provides the necessary priors through appropriate expert feedback. Thus, in this work we present an active learning method that can be built upon existing deep learning solutions for unsupervised anomaly detection, so that outliers can be separated from normal data effectively. We introduce a new layer that can be easily attached to any deep learning model designed for unsupervised anomaly detection to transform it into an active method. We report results on both synthetic and real anomaly detection datasets, using multi-layer perceptrons and autoencoder architectures empowered with the proposed active layer, and we discuss their performance on finding clustered and low density anomalies.
Tiago Pimentel, Marianne Monteiro, Adriano Veloso, Nivio Ziviani
IJCNN1
2020 Assessing the Reliability of Visual Explanations of Deep Models with Adversarial Perturbations
abstract
The interest in complex deep neural networks for computer vision applications is increasing. This leads to the need for improving the interpretable capabilities of these models. Recent explanation methods present visualizations of the relevance of pixels from input images, thus enabling the direct interpretation of properties of the input that lead to a specific output. These methods produce maps of pixel importance, which are commonly evaluated by visual inspection. This means that the effectiveness of an explanation method is assessed based on human expectation instead of actual feature importance. Thus, in this work we propose an objective measure to evaluate the reliability of explanations of deep models. Specifically, our approach is based on changes in the network's outcome resulting from the perturbation of input images in an adversarial way. We present a comparison between widely-known explanation methods using our proposed approach. Finally, we also propose a straightforward application of our approach to clean relevance maps, creating more interpretable maps without any loss in essential explanation (as per our proposed measure).
Dan Valle, Tiago Pimentel, Adriano Veloso
IJCNN2
2020 Phonotactic Complexity and its Trade-offs
abstract
We present methods for calculating a measure of phonotactic complexity—bits per phoneme— that permits a straightforward cross-linguistic comparison. When given a word, represented as a sequence of phonemic segments such as symbols in the international phonetic alphabet, and a statistical model trained on a sample of word types from the language, we can approximately measure bits per phoneme using the negative log-probability of that word under the model. This simple measure allows us to compare the entropy across languages, giving insight into how complex a language’s phonotactics is. Using a collection of 1016 basic concept words across 106 languages, we demonstrate a very strong negative correlation of − 0.74 between bits per phoneme and the average length of words.
Tiago Pimentel, Brian Roark, Ryan Cotterell
Trans. Assoc. Comput. Linguistics1
2019 Meaning to Form: Measuring Systematicity as Information
abstract
A longstanding debate in semiotics centers on the relationship between linguistic signs and their corresponding semantics: is there an arbitrary relationship between a word form and its meaning, or does some systematic phenomenon pervade?For instance, does the character bigram gl have any systematic relationship to the meaning of words like glisten, gleam and glow?In this work, we offer a holistic quantification of the systematicity of the sign using mutual information and recurrent neural networks.We employ these in a data-driven and massively multilingual approach to the question, examining 106 languages.We find a statistically significant reduction in entropy when modeling a word form conditioned on its semantic representation.Encouragingly, we also recover wellattested English examples of systematic affixes.We conclude with the meta-point: Our approximate effect size (measured in bits) is quite small-despite some amount of systematicity between form and meaning, an arbitrary relationship and its resulting benefits dominate human language.
Tiago Pimentel, Arya McCarthy, Damián E. Blasi, Brian Roark, Ryan Cotterell
ACL (1)1
2019 Efficient Estimation of Node Representations in Large Graphs using Linear Contexts
abstract
Learning distributed representations in graphs has a rising interest in the neural network community. Recent works have proposed new methods for learning low dimensional embeddings of nodes and edges in graphs and networks. Several of these methods rely on the SkipGram algorithm to learn distributed representations, and they usually process a large number of multi-hop neighbors in order to produce the context from which node representations are learned. This is a limiting factor for these methods as graphs and networks keep growing in size. In this paper, we propose a simple alternate method which is as effective as previous methods, but being much faster at learning node representations. Our proposed method employs a restricted number of permutations over the immediate neighborhood of a node as context to generate its representation, thus avoiding long walks and large contexts while learning the representations. We present a thorough evaluation showing that our method outperforms state-of-the-art methods in six different datasets related to the problems of link prediction and node classification, being one to three orders of magnitude faster than baselines when generating node embeddings for very large graphs.
Tiago Pimentel, Rafael Castro, Adriano Veloso, Nivio Ziviani
IJCNN1
2018 Fast and Effective Neural Networks for Translating Natural Language into Denotations
Tiago Pimentel, Juliano Viana, Adriano Veloso, Nivio Ziviani
SPIRE1