Lucas Torroba Hennigen

dblp:267/9755 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0002-8197-9008ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 On the Duality between Gradient Transformations and Adapters
abstract
We study memory-efficient optimization of neural networks (in particular language models) with *linear gradient transformations*, where the gradients are linearly mapped to a lower dimensional space than the full parameter space, thus saving memory required for gradient accumulation and optimizer state persistence. The model parameters are updated by first performing an optimization step in the lower dimensional space and then going back into the original parameter space via the linear map's transpose. We show that optimizing the model in this transformed space is equivalent to reparameterizing the original model through a *linear adapter* that additively modifies the model parameters, and then only optimizing the adapter's parameters. When the transformation is Kronecker-factored, this establishes an equivalence between GaLore and one-sided LoRA. We show that this duality between gradient transformations and adapter-based reparameterizations unifies existing approaches to memory-efficient training and suggests new techniques for improving training efficiency and memory use.
Lucas Torroba Hennigen, Hunter Lang
ICML1
2024 Principled Gradient-Based MCMC for Conditional Sampling of Text
abstract
We consider the problem of sampling text from an energy-based model. This arises, for example, when sampling text from a neural language model subject to soft constraints. Although the target distribution is discrete, the internal computations of the energy function (given by the language model) are differentiable, so one would like to exploit gradient information within a method such as MCMC. Alas, all previous attempts to generalize gradient-based MCMC to text sampling fail to sample correctly from the target distribution. We propose a solution, along with variants, and study its theoretical properties. Through experiments on various forms of text generation, we demonstrate that our unbiased samplers are able to generate more fluent text while better adhering to the control objectives. The same methods could be used to sample from discrete energy-based models unrelated to text.
Afra Amini, Lucas Torroba Hennigen, Xinyan Yu 0001, Holden Lee, Jason Eisner, Ryan Cotterell
ICML3
2023 A Latent-Variable Model for Intrinsic Probing
abstract
The success of pre-trained contextualized representations has prompted researchers to analyze them for the presence of linguistic information. Indeed, it is natural to assume that these pre-trained representations do encode some level of linguistic knowledge as they have brought about large empirical improvements on a wide variety of NLP tasks, which suggests they are learning true linguistic generalization. In this work, we focus on intrinsic probing, an analysis technique where the goal is not only to identify whether a representation encodes a linguistic attribute but also to pinpoint where this attribute is encoded. We propose a novel latent-variable formulation for constructing intrinsic probes and derive a tractable variational approximation to the log-likelihood. Our results show that our model is versatile and yields tighter mutual information estimates than two intrinsic probes previously proposed in the literature. Finally, we find empirical evidence that pre-trained representations develop a cross-lingually entangled notion of morphosyntax.
Karolina Stanczak, Lucas Torroba Hennigen, Adina Williams, Ryan Cotterell, Isabelle Augenstein
AAAI2
2023 A Measure-Theoretic Characterization of Tight Language Models
abstract
Li Du, Lucas Torroba Hennigen, Tiago Pimentel, Clara Meister, Jason Eisner, Ryan Cotterell. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Lucas Torroba Hennigen, Tiago Pimentel, Clara Meister, Jason Eisner, Ryan Cotterell
ACL (1)2
2023 Generalizing Backpropagation for Gradient-Based Interpretability
abstract
Many popular feature-attribution methods for interpreting deep neural networks rely on computing the gradients of a model's output with respect to its inputs.While these methods can indicate which input features may be important for the model's prediction, they reveal little about the inner workings of the model itself.In this paper, we observe that the gradient computation of a model is a special case of a more general formulation using semirings.This observation allows us to generalize the backpropagation algorithm to efficiently compute other interpretable statistics about the gradient graph of a neural network, such as the highest-weighted path and entropy.We implement this generalized algorithm, evaluate it on synthetic datasets to better understand the statistics it computes, and apply it to study BERT's behavior on the subject-verb number agreement task (SVA).With this method, we (a) validate that the amount of gradient flow through a component of a model reflects its importance to a prediction and (b) for SVA, identify which pathways of the self-attention mechanism are most important.
Kevin Du, Lucas Torroba Hennigen, Niklas Stoehr, Alex Warstadt, Ryan Cotterell
ACL (1)2
2023 An Ordinal Latent Variable Model of Conflict Intensity
abstract
Niklas Stoehr, Lucas Torroba Hennigen, Josef Valvoda, Robert West, Ryan Cotterell, Aaron Schein. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Niklas Stoehr, Lucas Torroba Hennigen, Josef Valvoda, Robert West 0001, Ryan Cotterell, Aaron Schein
ACL (1)2
2023 Learning to Grow Pretrained Models for Efficient Transformer Training
Peihao Wang, Rameswar Panda, Lucas Torroba Hennigen, Philip Greengard, Leonid Karlinsky, Rogério Feris, David D. Cox, Zhangyang Wang
ICLR3
2022 Probing as Quantifying Inductive Bias
abstract
Pre-trained contextual representations have led to dramatic performance improvements on a range of downstream tasks.Such performance improvements have motivated researchers to quantify and understand the linguistic information encoded in these representations.In general, researchers quantify the amount of linguistic information through probing, an endeavor which consists of training a supervised model to predict a linguistic property directly from the contextual representations.Unfortunately, this definition of probing has been subject to extensive criticism in the literature, and has been observed to lead to paradoxical and counterintuitive results.In the theoretical portion of this paper, we take the position that the goal of probing ought to be measuring the amount of inductive bias that the representations encode on a specific task.We further describe a Bayesian framework that operationalizes this goal and allows us to quantify the representations' inductive bias.In the empirical portion of the paper, we apply our framework to a variety of NLP tasks.Our results suggest that our proposed framework alleviates many previous problems found in probing.Moreover, we are able to offer concrete evidence that-for some tasks-fastText can offer a better inductive bias than BERT. 1 * Equal contribution. 1 Our code is available at https://github.com/ rycolab/evidence-probing.
Alexander Immer, Lucas Torroba Hennigen, Vincent Fortuin, Ryan Cotterell
ACL (1)2
2022 UniMorph 4.0: Universal Morphology
abstract
The Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation, and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements on several fronts that were made in the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 66 new languages, including 24 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g., missing gender and macrons information. We have amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet.
Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieras, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina J. Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Lane 0002, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóga, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer C. White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo Maria Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar 0002, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Tucker Prud'hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova
LREC27
2022 Same Neurons, Different Languages: Probing Morphosyntax in Multilingual Pre-trained Models
abstract
Karolina Stanczak, Edoardo Ponti, Lucas Torroba Hennigen, Ryan Cotterell, Isabelle Augenstein. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Karolina Stanczak, Edoardo Maria Ponti, Lucas Torroba Hennigen, Ryan Cotterell, Isabelle Augenstein
NAACL-HLT3
2021 Classifying Dyads for Militarized Conflict Analysis
abstract
Understanding the origins of militarized conflict is a complex, yet important undertaking.Existing research seeks to build this understanding by considering bi-lateral relationships between entity pairs (dyadic causes) and multi-lateral relationships among multiple entities (systemic causes).The aim of this work is to compare these two causes in terms of how they correlate with conflict between two entities.We do this by devising a set of textual and graph-based features which represent each of the causes.The features are extracted from Wikipedia and modeled as a large graph.Nodes in this graph represent entities connected by labeled edges representing ally or enemy-relationships.This allows casting the problem as an edge classification task, which we term dyad classification.We propose and evaluate classifiers to determine if a particular pair of entities are allies or enemies.Our results suggest that our systemic features might be slightly better correlates of conflict.Further, we find that Wikipedia articles of allies are semantically more similar than enemies. 1
Niklas Stoehr, Lucas Torroba Hennigen, Samin Ahbab, Robert West 0001, Ryan Cotterell
EMNLP (1)2
2020 Machine Reading of Historical Events
abstract
Machine reading is an ambitious goal in NLP that subsumes a wide range of text understanding capabilities.Within this broad framework, we address the task of machine reading the time of historical events, compile datasets for the task, and develop a model for tackling it.Given a brief textual description of an event, we show that good performance can be achieved by extracting relevant sentences from Wikipedia, and applying a combination of taskspecific and general-purpose feature embeddings for the classification.Furthermore, we establish a link between the historical event ordering task and the event focus time task from the information retrieval literature, showing they also provide a challenging test case for machine reading algorithms. 1 1 Code and data are available at https://github.com/ltorroba/ machine-reading-historical-events.* Equal contribution.Year Event text OTD 2005 107 die in Amagasaki rail crash in Japan.1939 BMI (Broadcast Music Incorporated) formed.1864 General Sherman's armies reach Savannah & 12 day siege begins.WOTD 1887 Buffalo Bill Cody's Wild West Show opens in London.1399 Henry IV is proclaimed King of England.1943 First Flight of the Gloster Meteor, Britain's first combat jet aircraft.
Or Honovich, Lucas Torroba Hennigen, Omri Abend, Shay B. Cohen
ACL2
2020 Intrinsic Probing through Dimension Selection
abstract
Most modern NLP systems make use of pretrained contextual representations that attain astonishingly high performance on a variety of tasks.Such high performance should not be possible unless some form of linguistic structure inheres in these representations, and a wealth of research has sprung up on probing for it.In this paper, we draw a distinction between intrinsic probing, which examines how linguistic information is structured within a representation, and the extrinsic probing popular in prior work, which only argues for the presence of such information by showing that it can be successfully extracted.To enable intrinsic probing, we propose a novel framework based on a decomposable multivariate Gaussian probe that allows us to determine whether the linguistic information in word embeddings is dispersed or focal.We then probe fastText and BERT for various morphosyntactic attributes across 36 languages.We find that most attributes are reliably encoded by only a few neurons, with fastText concentrating its linguistic structure more than BERT. 1
Lucas Torroba Hennigen, Adina Williams, Ryan Cotterell
EMNLP (1)1